Recognition is only as good as the scan: the five habits that raise every score

Better OCR starts at the scanner: 200–300 DPI, straight pages, even light, real contrast, clean crops — five habits that measurably raise recognition.

UnboundPDF is a free suite of 43 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.

Diagram comparing in-browser local processing, where a document never leaves the device, to an upload-based tool that sends the document to a server and back

The five habits. 200–300 DPI, pages straight (or deskewed after), even light with no shadows, real black-on-white contrast, and crops that keep the full text block. Then OCR — and watch the per-page confidence agree with you.

Resolution: the floor matters, the ceiling doesn't

Below 200 DPI, the strokes of small characters start merging — 8-point footnotes become guesses. The 200–300 band is where ordinary print resolves cleanly; beyond it, accuracy gains flatten while file sizes climb. If a document has genuinely tiny print, scan THAT one higher rather than everything.

Geometry: straight pages read better

Tilted lines make every character a slightly rotated puzzle, and it shows in the scores. Feed pages straight, and when the world hands you a crooked scan anyway, deskew before recognising — the order is the point; recognition first and the damage is baked into the text layer.

OCR PDF with a scanned Chinese-language document loaded for recognition
A Chinese-language scan loaded in OCR PDF — one of 126 languages, recognised on the device.

Light and contrast: the phone-photo clause

Scanner users get even light for free; phone photographers don't. Daylight from the side, no shadow of your own hand, page flat — the photo-to-searchable walkthrough turns this into a routine. Faded originals (thermal receipts, old faxes) are the hard class: photograph them well NOW; they only fade further.

Then trust, but verify by the score

The per-page confidence tells you where your habits paid off and which page deserves a second look — treat a low-scoring page's numbers as read-from-the-image. Language settings and the 126-language reality live in the full OCR guide; the batch version of this hygiene is the paperwork cleanup pass.

Troubleshooting

Good scan, bad score. Check the language setting first — right script, right pack. Only the footnotes fail. Rescan that document at higher DPI; tiny print is a resolution problem, not an engine one. Handwriting scores terribly. Honest limit: OCR here is for print; handwriting recognition is a different problem class.

Frequently asked questions

What resolution should I scan at?

200–300 DPI for ordinary print. Below 200, small characters merge; far above 300 mostly costs file size, not accuracy.

My scan is already crooked — rescan or fix?

Deskew it — straightening before OCR measurably helps recognition, and the tool is one step in the chain.

Colour or grayscale?

Grayscale for black-on-white text: cleaner contrast, smaller files. Keep colour only when colour carries meaning.

Try it yourself

Free, private, no account. Runs entirely in your browser.

Related articles