Most requests to collect job listings are really recruiting-lead requests, and those two need different answers. What the terms of use decide, why candidate data is the sharpest line, and what market data actually tells you.
Key takeaways
- Separate two requests that sound alike: market intelligence about roles and salaries, and a list of companies to sell recruiting services to. They carry different risks and different answers.
- Candidate data is the line you do not cross. Job listings are company-published; CVs and profiles are individuals' personal data, and that is a different legal category entirely.
- Job posts are marketing copy, not a record. A listing tells you what a company advertised, not what it pays or whether the role was ever filled.
- The aggregate signal is real and useful: which roles are being advertised, in what volume, and how that changes. Individual listings are perishable.
The request arrives in two forms that sound alike and are not: "I want to understand what is happening in the market" and "I want a list of companies hiring so I can approach them". They need different answers, and only one is genuinely a data question.
First: which of the two
| The need | What it really is |
|---|---|
| Which roles are in demand, how it changes | Market research |
| A list of companies hiring now | Sales leads |
The distinction matters because the second has cheaper and safer alternatives that involve no automated collection at all - and because it enters the territory of marketing outreach, which carries its own rules.
The gate: terms of use
Job boards have terms of use, and they usually address automated collection explicitly. Read them first, including robots.txt - they are the source, not an article.
And for commercial use, that is a question for a lawyer. The difference between research and a product that is sold is exactly the difference that matters, and I cannot settle it here.
I am not providing circumvention techniques. What is useful is what determines whether this works, and what the alternatives are.
The line you do not cross: candidate data
This is the most important point in the article, because here the difference is one of kind rather than degree.
| A job listing | A candidate profile |
|---|---|
| Published by a company | A person's personal data |
| Intended to be read | Given for a specific context |
| A terms-of-use question | A privacy-law question |
Collecting CVs, profiles or candidate contact details is not "another kind of data". It is private individuals' personal information, in a particularly sensitive context - looking for work.
That is a question for a lawyer before touching it, not after. And in Israeli recruiting it is not theoretical: it comes with real obligations.
The practical consequence: if the need is market intelligence, collect only what the company published about the role, and nothing identifying a person.
What the data actually says - and what it does not
This is what prevents wrong conclusions.
A job listing is marketing copy. A company writes it to attract candidates. What it is not:
- Not a salary record. A range appearing in a listing - where one appears at all - is what was advertised, not what was paid.
- Not evidence the role exists. Listings stay open after being filled, and are sometimes posted to build a pipeline.
- Not an indication of headcount. One listing can be five positions; five listings can be one position posted across several boards.
The conclusion: any inference treating listing counts as job counts will be wrong. What is reliable is direction - whether there is more or less advertising in a category over time.
What is worth it, and to whom
- Which skills are in demand. Text analysis of listings in a field gives a good picture of what is being asked for - useful to someone building a training programme or considering a specialisation.
- Volume trend by category. A rise or fall in advertising within a sector is a genuine market signal.
- Who advertises in a sector. Useful for market mapping - but see below on the cheaper route.
- Differences in how companies write. Indicative of culture and of what they emphasise.
In all of them, the value is in the aggregate and over time, not in an individual listing.
The cost
The same reality as any collection from a website:
- It is an ongoing cost. Boards change structure. Maintenance is part of the cost.
- It breaks silently. Zero listings is not "no hiring in the market" - it is a broken collector. An alert on a coverage collapse is mandatory, or the failure masquerades as an insight.
- The data is dirty. The same role under five names, a location given as a region or a street, salary in free text.
- Duplicates across boards. The same vacancy posted in several places - and counting without deduplication inflates every conclusion.
The alternatives - and here they are particularly strong
Before building anything:
- Official statistics. Employment and pay have government sources published for public use - see the data.gov.il guide. For a salary question they are a far better source than listings, because a listing is advertising and an official figure is measurement.
- Industry salary surveys. They exist, and are usually more accurate than anything extracted from listings.
- Just look. If the need is a list of companies hiring - you can usually see that manually in half an hour, and it produces a shorter, better list than a thousand rows.
The strong point here: if the question is salary, extracting it from listings is almost always the worst route. The figure exists somewhere better.
When not to build
If the need is sales leads - almost always not worth it. Twenty relevant companies chosen by hand are worth more than a thousand collected rows, both in conversion and in your time.
If the need is a one-off question - look manually.
Building pays off only when a time series is needed: tracking a sector trend over months. And even then, first check whether that trend is already published by an official source.
Checklist
- Decide what the need is - research or leads. They are not the same project.
- Read the terms of use and
robots.txt; commercial use goes to a lawyer. - Do not touch candidate data. That is a line, not a consideration.
- Check whether the official source answers it - especially for salary.
- Deduplicate across boards before counting.
- Do not infer headcount from listing counts.
- Alert on a coverage collapse.
Frequently asked questions
Can I collect candidate CVs or profiles from job boards?
That is a line rather than a consideration. A job listing is published by a company to be read; a CV or profile is a private individual's personal data given in a specific context, and in a particularly sensitive one - looking for work. It is a privacy-law question for a lawyer before touching it, and in Israeli recruiting it comes with real obligations.
Can I learn salary levels from job listings?
Poorly. A listing is marketing copy, so a range appearing in one - where any appears - is what was advertised rather than what was paid. For a salary question, official employment statistics published as government open data and industry salary surveys are far better sources, because a listing is advertising while an official figure is a measurement.
Does the number of listings tell me how many jobs there are?
No. One listing can represent five positions, and five listings across different boards can be one position. Listings also stay open after being filled and are sometimes posted to build a pipeline. What is reliable is direction - whether advertising in a category is rising or falling over time - provided you deduplicate across boards before counting.
Is collecting job listings worth it for finding sales leads?
Almost never. Twenty relevant companies chosen by hand outperform a thousand collected rows on both conversion and your own time, and the manual route avoids the terms-of-use gate entirely. Automated collection earns its cost when a time series is needed - tracking a sector trend over months - not when the output is a prospect list.
How do you know a job-listing collector broke?
By alerting on a collapse in coverage, because the failure presents as information. A run returning zero or a handful of listings looks like "hiring stopped in this sector" on a chart, when it is a board that changed its structure. Validate the expected fields on every run and stop loudly when something changed rather than importing empty results.
Keep reading
Related service
Web Scraping
Reliable web scraping and data pipelines that deliver clean data.
About the author
Yehonatan Saadia
Freelance automation, web & MVP engineer
I'm Yehonatan Saadia, a senior engineer who builds business automation, custom websites, and MVPs for small and mid-sized companies across the US, Europe, and Israel. These guides come from real client work, not theory.
Work with meHave a project like this?
Tell me what you're trying to automate or build and I'll tell you the fastest reliable way to ship it.
