• +91 44 400 500 39
  • info@greenbooks.in
  • Chennai, INDIA

OCR a PDF to Make Scans Searchable

OCR reads the words in a scanned page and adds them as an invisible text layer behind the image. The page looks identical, but it becomes searchable, and that single change is what turns a pile of scans into a usable archive.

How to OCR a PDF to Make Scans Searchable

  1. Upload the scanned PDF.
  2. Choose the language of the document.
  3. Run OCR, then download the searchable file.

When you would use this

  • Making a scanned register searchable by name or number.
  • Preparing scans for indexing in a document management system.
  • Enabling text extraction from a historical archive.

Things to watch for

Scan quality decides accuracy, not the software

300 DPI is the practical minimum. Below 200 DPI accuracy falls sharply, and no OCR engine recovers detail that was never captured.

Straighten and clean the page first

Skewed or faint pages read badly. Running Adjust Colours / Contrast before OCR often lifts accuracy more than any OCR setting.

Pick the right language

OCR uses a language model to resolve ambiguous shapes. Running an English model over a Tamil document produces nonsense.

Handwriting is a different problem

Standard OCR is trained on printed text. Handwritten material needs specialist recognition and results vary widely, even at high scan quality.

Choosing the right tool

OCR is the step that makes every other text tool possible. PDF to Word, PDF to Text, PDF to CSV, Compare PDFs and Auto Rename all need a text layer, and on a scan none of them will produce anything until OCR has run. If a text tool returns an empty result on a scanned document, OCR is almost always the missing step.

Questions

How accurate is OCR?

Clean 300 DPI typescript reaches the high nineties. Faint, skewed or handwritten material is far lower. Scan quality matters more than any other factor.

Does the page look different afterwards?

No. The text layer sits invisibly behind the image, so the scan appears unchanged.

Can it read handwriting?

Ordinary OCR is built for printed text. Handwriting needs specialist recognition and results vary greatly.

What resolution should I scan at?

300 DPI is the practical minimum for reliable OCR. Below 200 DPI accuracy drops sharply.

What DPI should I scan at for OCR?

300 DPI for ordinary printed text. Go higher for small type, faded originals or documents with fine detail.

Does OCR change how the page looks?

No. The recognised text is added as an invisible layer behind the existing image, so the scan appears exactly as before.

Can I OCR a PDF that already has some text?

Yes, though pages that already carry real text gain nothing. OCR is for the scanned pages.

Can OCR read Tamil or Hindi?

Recognition depends on the language models installed on the instance. Select the correct language before running it.

Need this at scale?

These tools handle one document at a time. When the job is thousands of files, Greenbooks handles document digitization as a managed service, including bulk scanning, OCR and metadata capture, with the output loaded into DocuVenta DMS. Talk to us about volume work.

Chat with Greenbooks on WhatsApp