ForgemintTools

Make a scanned PDF searchable

A scanned document is a stack of pictures — you can't select a line, copy an address, or find a name in it. This adds an invisible text layer underneath the original page, so the PDF looks exactly the same but is searchable and copyable. Free, in 22 languages, built on OCRmyPDF and Tesseract.

Sign in to run OCR

A free account keeps the OCR engine from being open to the whole internet. It takes a few seconds, costs nothing, and there is no paywall behind it.

How it works

Upload your scan

A scanned PDF, or a photo of a document. Nothing is installed and nothing leaves your control — the file goes straight to the OCR engine.

Pick the languages

Choose every language that appears on the page. Mixed English and Hindi on one document is normal and handled — select both.

Get a searchable file

The page looks identical, but an invisible text layer sits underneath it. Ctrl+F works, text selects, and copy-paste behaves.

Questions

What does OCR actually change in my PDF?
Nothing you can see. A scan is a stack of pictures, so there is no text to select or search. OCR reads the words in those pictures and stores them as an invisible layer positioned underneath the original page. The document looks pixel-for-pixel the same, but you can now search it, select a line, and copy an address out of it.
Which languages are supported?
Twenty-two, with Indian languages first: Hindi, Punjabi, Gujarati, Marathi, Bengali, Tamil, Telugu, Kannada, Malayalam and Urdu, alongside English, French, German, Spanish, Portuguese, Italian, Dutch, Russian, Arabic, Chinese, Japanese and Korean. You can select several at once for a document that mixes them.
Is my document stored anywhere?
No. The file is processed in a temporary directory that is deleted as soon as the result is sent back. It is never written to a database, never logged, and never used for anything else.
My scan came out badly recognised. What helps?
Scan at 300 DPI if you can — it is the resolution the recognition engine is tuned for. Turn on Deskew if the page went through the scanner at an angle, and Clean if there is speckle or scanner noise. If the PDF already has a bad text layer from previous software, choose Force re-OCR to discard it and start again.
What is PDF/A, and should I pick it?
PDF/A is the archival version of PDF, required by many government, legal and institutional filings. It embeds everything needed to render the document identically decades from now. Choose it if you are filing the document somewhere that asks for it; standard PDF is smaller and fine for everything else.
Does it cost anything?
No. The tool is free, with no page limits and no watermark. You need a free ForgeMint account so the engine is not open to the whole internet, and that is the only requirement.

What happens to your file

Your document is processed in a temporary directory that is deleted the moment the finished PDF is sent back to you. It is never stored, never logged, and never used to train anything.