Hand over several PDFs and get back a spreadsheet of what they contain. That is the useful shape. The one thing that decides whether it works: whether the PDF holds real text or is a picture of text.
Key takeaways
- A PDF with a real text layer and a scanned PDF are different problems. The scan is a picture, and reading it means recognising characters from an image - which is where Hebrew degrades most.
- A PDF has no structure. What looks like a table is really lines of text positioned on a page, which is why column extraction goes wrong on documents that look perfectly clear.
- The highest-value use is many documents at once - several invoices or quotes into one comparison sheet - not reading one document you could read yourself.
- For anything with legal or financial consequence the output is a draft. Extraction from a document is a reading, and a reading needs checking against the page.
The useful shape is not "summarise this document" but hand over several PDFs and get back a spreadsheet of what they contain. And whether that works depends on one thing: whether the PDF holds real text, or is a picture of text.
Two kinds of PDF, and why it decides everything
The same extension hides two entirely different problems:
| PDF with a text layer | Scanned PDF |
|---|---|
| Produced by software - Word, an invoicing system | Scanned or photographed |
| You can select text with the mouse | You cannot - it is an image |
| The text exists in the document | Characters must be recognised from a picture |
| Accurate reading | Quality depends on the scan |
How to check in two seconds: open the file and try to select a line. If the selection grabs words, there is a text layer. If it draws a rectangle over a picture, it is a scan.
Why this matters especially in Hebrew: character recognition from an image is weaker in Hebrew than in English, and a letter misread inside a name or a number does not necessarily stand out in the output. A low-quality scan, a document photographed at an angle on a phone, a fax - all of these raise the risk.
The practical conclusion: verify more on scans, less on documents produced by software.
The second problem: a PDF has no structure
This is a point that surprises people.
What looks like a table in a document is not a table. A PDF is instructions to draw text at positions on a page. The eye sees columns; the document contains lines of text whose positions create the impression of columns.
Two common consequences follow:
- Columns get mixed where spacing is uneven - and that happens in documents that look perfectly clear.
- A table breaking across pages loses its headers, and rows on the second page can be read wrongly.
What that means: do not expect table extraction to be perfect first time, and check the rows at the edges specifically - the first, the last, and the first after a page break.
Where it delivers real value
1. Many documents at once
This is the case that justifies the tool. One document you can read yourself; twenty you cannot.
- Twenty supplier invoices into a sheet with supplier, date, amount, invoice number.
- Three quotes into a comparison table by line item.
- Monthly reports into a trend over time.
There is an automated, ongoing version of this too, and the difference matters: if it repeats every month, that is automation rather than a conversation.
2. Comparing two versions
A contract before and after the other side's comments, an updated proposal, a changed specification. Ask for a list of differences rather than a summary - the gap between "what changed" and "what this is about" is the whole point.
3. Finding something in a long document
"What does it say about cancellation?", "What are the payment terms?", "Is there a warranty clause?" - asking for the passage to be quoted rather than rephrased. A quote can be verified; a rephrasing is already an interpretation.
4. Turning a document into data
A catalogue, a price list, a specification table into a spreadsheet you can work with.
How to ask
The rule: describe the output, not the task.
| Weak | Works |
|---|---|
| "Extract the details" | "A sheet: supplier, invoice number, date, amount before tax, total. One row per invoice." |
| "Compare the contracts" | "A list of differences by clause. For each: clause number, what it was, what it is now." |
| "What does it say about cancellation" | "Quote the clauses dealing with cancellation, with the clause number." |
And say what to do with what is missing. A field not present in the document - leave it blank and flag it, do not guess. Without an instruction, a missing row and an inferred row look identical.
Four things to check every time
- The count. You supplied 20 documents - did 20 rows come out?
- Amounts. A number copied wrongly looks exactly like a correct one. Check two or three against the file.
- Dates. An Israeli document writes day/month. 03/09 can be read as the 9th of March, and on any day from the 1st to the 12th both readings look valid.
- Hebrew names. A name read wrongly will not match a record in a system later.
Where it is not enough
Worth being straight about the boundary:
- A legal document. A contract summary is a starting point for reading it, not a substitute. A question about what a clause means is a question for a lawyer.
- Numbers entering the books. Extraction is interpretation. Before an amount enters a system, a person approves it.
- A poor scan. If you cannot read the number by eye, there is no reason to assume automatic recognition managed it.
- A sensitive document. An uploaded file leaves your computer. If it contains personal details the question does not need, remove them, or work on one page rather than the whole file.
And in Hebrew
Reading and understanding Hebrew works. What needs attention is the mechanical side:
- A Hebrew scan degrades more than an English one. Verify accordingly.
- Hebrew mixed with numbers - an invoice number or ID beside Hebrew text - is where character order is most likely to go wrong.
- A bilingual document is worth splitting: ask for the extraction once per language.
Frequently asked questions
Why does AI read some PDFs perfectly and others badly?
Because a PDF produced by software contains real text, while a scanned or photographed one is a picture and its characters must be recognised from an image. Check in two seconds by trying to select a line: if the selection grabs words there is a text layer, if it draws a rectangle over a picture it is a scan. Verify far more on scans.
Why does table extraction from a PDF come out wrong?
Because a PDF has no structure - it is instructions to draw text at positions on a page, so what looks like a table is really lines of text whose positions create the impression of columns. Columns get mixed where spacing is uneven, and a table breaking across pages loses its headers. Check the rows at the edges: the first, the last, and the first after a page break.
Is Hebrew OCR reliable enough for invoice data?
Character recognition from an image is weaker in Hebrew than in English, and a letter misread inside a name or a number does not necessarily stand out in the output. Low-quality scans, phone photos at an angle and faxes all raise the risk. Treat the extraction as a reading that needs checking, and have a person approve any amount before it enters a system.
What is the best use of AI with PDFs?
Many documents at once - twenty supplier invoices into one sheet, three quotes into a comparison table by line item, monthly reports into a trend. One document you can read yourself; twenty you cannot, and that is what justifies the tool. If the same job repeats every month, though, it is automation rather than a conversation.
Can I rely on an AI summary of a contract?
As a starting point for reading it, not a substitute - and a question about what a clause means is a question for a lawyer. Where you do use it, ask for the relevant clauses to be quoted with their numbers rather than rephrased: a quote can be verified against the page, while a rephrasing is already an interpretation.
Keep reading
Related service
Business Automation
I build custom automations that remove repetitive work end to end.
About the author
Yehonatan Saadia
Freelance automation, web & MVP engineer
I'm Yehonatan Saadia, a senior engineer who builds business automation, custom websites, and MVPs for small and mid-sized companies across the US, Europe, and Israel. These guides come from real client work, not theory.
Work with meHave a project like this?
Tell me what you're trying to automate or build and I'll tell you the fastest reliable way to ship it.
