PDF guide
FrançaisNative PDF text vs OCR: how to edit a scanned PDF
A PDF can look like text while containing only a page image. The editor checks for meaningful native text first, then offers optional local OCR for pages that need recognition.
Open the PDF editorNative text is already inside the PDF
A text-based PDF has characters and positions that a viewer can select or search. Editing a line in this kind of document starts with the existing text layer.
- 1Open the PDF and wait for the text lines to load.
- 2Choose Edit text and click a line that is already selectable.
- 3Correct the line and review the visual replacement before export.
A scan is a picture of a page
A scanned PDF usually has no useful character layer. The page can still look sharp, but there are no words for a text editor to select.
- 1Open the document and watch for pages reported without meaningful native text.
- 2Keep the native text that is present on hybrid pages.
- 3Run OCR only for the pages containing lines you need to edit.
Run OCR when the document needs it
OCR recognition runs locally and supports English, French or both. Recognition assists editing; it does not add an invisible searchable text layer to an unedited scan on export. Inspect the recognized lines and correct uncertain results before exporting.
- 1Select the OCR language in the editor.
- 2Start OCR for the missing pages or use the optional supplement.
- 3Cancel if the document is too large for the current device, then continue with manual annotations.
Know the boundary
OCR does not recreate the original typography perfectly. Font matching is estimated, and some complex scripts or combining sequences may not reproduce safely.
- 1Compare each edited line with the scan.
- 2Choose a replacement font when the suggested one is not suitable.
- 3Keep the original PDF alongside the exported copy.
Getting the best OCR accuracy
Resolution matters more than any setting. Aim for about 300 DPI: characters are sharp enough for reliable recognition without ballooning the file. Below roughly 200 DPI, small letters blur into each other and error rates climb; much above 400 DPI rarely improves text results while making the scan heavy to handle.
Light the page evenly and keep it flat. Shadows, creases, and pages curving near a book spine all confuse recognition, as does a photo taken at an angle. Hold the camera parallel to the page, or better, use a scanner with the lid closed so the background is a clean white.
Choose the language that matches the document. The editor offers English, French, or both — and this choice genuinely matters. French text recognised with English selected comes back with mangled accents and wrong word guesses; mixed-language documents are what the combined option is for.
Keep the page tidy. Coffee stains, handwritten margin notes, stamps, and dense graphics sitting close to text all become "characters" the recogniser must interpret. None of this needs perfection — a clean, straight, well-lit scan at 300 DPI in the right language gets the large majority of pages right on the first pass.
For phone photos of documents, the same rules apply doubly. Clean the lens, flatten the page under a book if it curls, and shoot from directly above. A photo taken at an angle forces the recogniser to read stretched characters, and no language setting compensates for that — five seconds of care with the camera saves ten minutes of correcting.
If the only copy you have is low-resolution, don't upscale it in an image editor and expect better results — upscaling invents pixels, not detail. Work from the best original you can obtain; when a rescan isn't possible, correct the output carefully rather than hoping software recovers what was never captured.
Hybrid PDFs: treat each page on its own
Many real documents are hybrids: a contract drafted on a computer with a scanned signature page appended, or a report where someone inserted phone photos of receipts between typed pages. Some pages carry native text; others are pure images. Treating the whole file one way wastes effort or destroys quality.
The strategy is simple: leave native-text pages alone and run OCR only on the image pages. Native lines edit directly and keep their original typography; running OCR over them gains nothing and can actually produce worse text than what was already there.
Telling the two apart takes seconds: try selecting a sentence with the mouse. If the text highlights, the page is native. If nothing selects, it's an image and a candidate for OCR. Work through the document page by page rather than deciding once for the whole file.
Finish by reviewing each touched page at full-page scale. On hybrids especially, a line that looked fine in isolation can sit oddly next to untouched native text — a quick visual pass catches it.
One more efficiency rule: resist re-running OCR on a page just because the first pass was imperfect. Fixing a handful of wrong words by hand is faster than re-running recognition and proofreading the whole page a second time. OCR gets you most of the way there; your eyes finish the job.
Troubleshooting garbled recognition
When OCR output looks wrong, the cause is almost always the input, not the recogniser. Match the symptom to the fix:
| Symptom | Likely cause | Fix |
|---|---|---|
| Random characters and symbols | Resolution too low | Rescan at about 300 DPI. |
| "rn" read as "m", blurred pairs | Small or out-of-focus type | Rescan sharper; enlarge small originals first. |
| Accents missing or wrong | Wrong language selected | Re-run with French (or both languages) selected. |
| Two columns merged into one line | Complex multi-column layout | Correct the order manually; keep edits to short lines. |
| 0 read as O, 1 as l | Low-contrast or stylised font | Rescan with better contrast; verify every digit by hand. |
| Entire page is garbage | Skewed or shadowed photo | Straighten the page, light it evenly, rescan. |
| Handwriting comes back as gibberish | Recognition tuned for printed text | Type handwritten passages manually instead. |
| Words split across line breaks | Narrow columns, hyphenated endings | Join split words while correcting. |
One habit beats every fix in that table: proofread the numbers. A misread letter is embarrassing; a misread digit in an invoice total or a contract date is expensive. Read every figure against the original before the file leaves your hands — totals, dates, and reference numbers deserve a second look even when the rest of the page is perfect.
Real-world scenarios
Invoices are the classic case. The line items and totals are what you need, and they're usually printed cleanly — ideal input for recognition. Run OCR on the pages you need, then verify every amount and date character by character. If you only need the figures as text, PDF to text extracts the page content for copying into a spreadsheet.
Contracts demand more caution. OCR the pages you must edit or quote, but never trust recognised legal wording without proofreading each line against the scan. A single misread "not" or a swapped figure changes the meaning, and contracts are exactly where such errors like to hide.
Archives are a question of scale. A hundred-page scanned report doesn't need a hundred pages of OCR — recognise only the pages you'll edit or search, since large documents take noticeably longer to process. Keep the untouched original archived; the OCR-assisted copy is a working version, not a replacement.
Expense receipts are the messiest common input: crumpled thermal paper, faded ink, tiny fonts. Photograph them flattened on a dark background for contrast, run OCR only on the ones you actually need to edit, and cross-check the totals with extra care — faded print is exactly where digit errors cluster.
A final comfort: because recognition runs in your browser, sensitive invoices and contracts never travel to a server. That makes local OCR a reasonable choice even for documents you'd never upload anywhere.
Whatever the scenario, run the recognition pass before you start editing. Correcting text first and running OCR afterwards means redoing work the recogniser could have handed you cleanly.