gizmobench

Searchable PDF OCR

Searchable PDF OCR reads the printed English on the pages of a scanned PDF and makes a new PDF you can search and copy from: each page is redrawn as an image at up to 200 DPI with the recognised words laid invisibly over it, then put back at its original size in points, in the order you chose and turned the way you set it. Before you convert, it lists what redrawing loses in your file, such as links, form fields and signatures. Afterwards it shows each page's text with the words Tesseract was least sure of marked, and gives you the text of every page as a plain text file too. The OCR engine and its English model are files this site serves and run in your browser, so the PDF is never uploaded; it takes PDFs up to 25 MB and 30 pages.

Drop a scanned PDF here

, up to 25 MB and 30 pages, or

Never uploadedNo accountEngine served by this siteOriginal left untouched
Render
Turn page

Pages are converted in the order you type them, like 1-3, 5 or 4, 1. Turn a sideways page upright before converting; the arrows above the preview move between pages. PDFs of up to 30 pages are read.

  • Searchable PDFeach page an image with hidden text
    nothing to download yet
  • Plain text.txt, page by page, UTF-8
    nothing to download yet
  • Clipboardthe same text as the .txt
    nothing to copy yet
  • Word boxesoutlined on the preview
    after the first conversion
Read in this tab, with files this site serves. The engine is tesseract.js 7.0.0 with tesseract.js-core 7.0.0, LSTM build without SIMD, and the model is tessdata_fast English (eng.traineddata), SHA-256 beginning 7d4322bd2a77, retrieved 2026-09-21. The model is checked against that pinned checksum before any page is read, and a mismatch stops the tool instead of guessing. Pages are drawn by PDF.js and the new file is put together by pdf-lib, both in this tab. The PDF is never uploaded; only your DPI and word box settings are remembered on this device. Credit: Tesseract.js and tessdata_fast, both under the Apache License 2.0 (engine licence, model licence, bundled notices).

Common questions

How do I make a scanned PDF searchable?
Choose or drop the PDF, check the page list, and turn any sideways page upright with the Turn buttons, using the arrows above the preview to move between pages. Pick a render DPI (200 is the default and the highest) and press Convert. Each page is drawn, read and added to a new PDF; when every page is done, download the searchable PDF or the text as a .txt file, or copy the text.
Is my PDF uploaded anywhere?
No. The PDF is opened, drawn and read in your browser tab: PDF.js draws the pages, Tesseract reads them in a Web Worker, and pdf-lib writes the new file, all from code and a model this site serves. Nothing about the file or its text is sent to this site or to anyone else, and the only things remembered, on your device, are your DPI and word box settings.
What does converting lose?
Every page becomes a picture with invisible text over it, so links stop working, form fields can no longer be filled in, a digital signature is not carried over, and vector text and drawings become pixels. The tool lists this before you convert, and names the pages in your file that have links or already hold selectable text. Your original PDF is not changed: a new file is made.
Does it keep the page size and order?
Yes. Tesseract sizes its own page from a resolution it guesses, so each page it returns is scaled back to the original page's size in points before it is added: a US Letter page stays 612 x 792 points, and one you turned a quarter becomes 792 x 612. Pages come out in the order you type in the page list, so 3, 1 puts page 3 first.
Which PDFs does it refuse?
Files over 25 MB or 30 pages, encrypted or password protected PDFs, and PDFs whose scanned images are compressed as JBIG2, CCITT fax or JPEG 2000 are refused with the reason before any page is read. Those three image formats cannot be decoded by this page in the browser, so their pages would come out blank; scanning again in grayscale or colour usually gives JPEG images, which work. A damaged file, or a page that fails during conversion, stops with the reason, and no partial PDF is offered.
Why are some words marked?
After converting, the review shows each page's text, and words Tesseract scored below 60 out of 100 are marked, with the lowest-scoring word named, so you know where to check first. With word boxes on, the same words are outlined in red on the page preview. A high score is not a promise that a word is right, so read the rest too.
Can it read handwriting, other languages or tables?
No. The model is Tesseract's fast English model for printed text. Handwriting and other languages come out wrong or not at all, and a table is read as lines of words rather than rebuilt as a table.

Printed English is read by a pinned Tesseract model in your browser, and OCR can misread characters, so check the flagged low-scoring words against the page. Each page is redrawn as an image at up to 200 DPI with the text laid over it: page order, size and your chosen rotation are kept, but links, form fields, signatures and sharp vector text are lost, and this is shown before you convert. It does not read handwriting or other languages or rebuild tables; PDFs up to 25 MB and 30 pages are processed on your device and never uploaded.