Skip to content
Main Site News Console

Convert PDFs to Markdown in 20ms: Open-Source OCR Tool Nearly 300× Faster

· 量子位
国内AI

The miserable days of manually extracting text from PDFs, only to end up with a mess of garbled characters that even LLMs couldn’t understand—

are finally coming to an end!

Image

The team behind Firecrawl, a star project incubated by Y Combinator, recently pulled off something big:

Its co-founder and CTO, Nicolas Camara, announced that their latest creation, OCR It, can extract all the text from a non-copyable PDF in just 20 milliseconds and convert it into clean Markdown that AI can readily understand.

Image

What does 20 milliseconds actually mean?

It means that right after you click “Confirm,” before you even have time to blink, a clean, processed Markdown file is already in your hands.

Best of all, it doesn’t require an internet connection. It can run 100% offline using the built-in Tesseract.

Not long after its release as an open-source project, this OCR tool attracted widespread attention on GitHub.

Nicolas Camara proudly stated that while its processing quality is on par with Docling, OCR It is nearly 300 times faster!

Image

After testing it, users also took to X to share their thoughts:

Fast—really, really fast! 👍

Image

Image

Image

How to Use OCR It

First, OCR It is a free browser extension for Chrome and Firefox.

You only need to select an area once. After that, each time you press the shortcut—or enable auto mode—it can convert an entire book or document into complete, editable plain text, making it easy to hand directly to an AI for summarization or questions.

During the process, OCR It automatically stops when it detects two identical pages in a row, can no longer turn the page, OCR fails, or it reaches the 300-page limit, preventing it from repeatedly capturing the final page indefinitely.

Image

Specifically, manual mode and auto mode work as follows:

Note: On Windows and Linux, replace the Option key with the Alt key.

1. Manual Mode

Select and Recognize

First, press Option+Shift+R and select the area containing the text you want to recognize. Once you’ve confirmed the selection is correct, press Enter to save it.

After selecting the text area, press Option+Shift+S once. OCR It will recognize the current page and automatically move to the next page when finished.

Then simply repeat the process until the entire document has been recognized.

To save even more time, you can turn the pages as you go and press Option+Shift+S once on each page. The system will immediately take a screenshot and automatically queue the recognition tasks in the background.

Image

Review

Once recognition is complete, you can review and edit the text in the extension panel, or re-recognize an individual page. Finally, select “Copy all” or “Download .txt” to export the full text, and you’re done.

2. Auto Mode

Auto mode is even simpler: just press Option+Shift+A or click Start auto-run, and OCR It will automatically repeat the cycle of “screenshot—OCR—turn the page” until it reaches the end of the document.

To stop the task midway, press the Esc key. The review process is essentially the same as in manual mode.

In addition, the OCR It repository interface shows that it requires neither an API key nor an internet connection. It requests no website permissions during installation, applying for permission to access the current site only when needed for automatic execution or for interacting with cross-origin iframes.

One more thing: browsers’ built-in PDF viewers currently do not support automatic page turning, so manual mode is still recommended when processing this type of PDF.

Still Unstable When Handling Complex Tasks

That said, while OCR It looks fast and convenient, it is still far from the point where you can hand every PDF over to it and rest easy.

Image

After hands-on testing, users reached a consensus:

OCR It currently handles PDFs with simple layouts and clear text fairly well, but it still struggles with complex pages that combine headings, footnotes, tables, mathematical formulas, and body text, so there is still plenty of room for improvement.

Image

References:

[1] https://x.com/nickscamara_/status/2083295265793212827

[2] https://github.com/thiagotigaz/ocr-it

[3] https://www.firecrawl.dev/about