Tesseract OCR PDF to Text
Upload a scanned PDF, extract searchable text with Tesseract OCR, and return page-level results.
The full canvas, before you import it
Click any node to see its config.
Click a node to select it — same as the Heym editor; the panel shows its settings.
3 nodes · Free & source-available
Tesseract OCR PDF to Text
Turn a scanned PDF into searchable plain text without sending the document to an external OCR service. The Converter node reads the file from Heym Drive, runs Tesseract on each page, and returns both the complete text and page-level results.
What this workflow does
- ScannedPDF creates a single-use upload link for a PDF
- ExtractText rasterizes the PDF and runs Tesseract OCR with automatic language detection
- OCRResult returns the extracted text, detected language, encoding, page count, and per-page text
Use cases
- Make scanned contracts, invoices, receipts, and forms searchable
- Prepare image-only PDFs for an LLM, Agent, RAG, or structured extraction step
- Keep OCR inside a self-hosted Heym deployment
- Preserve page boundaries for downstream review or indexing
Setup
Run the workflow once to create a single-use upload command, then upload a PDF with the generated curl command. The default settings use automatic language detection, UTF-8 output, Unicode normalization, 300 DPI, and Tesseract page segmentation mode 3.
Notes
To process only part of a long PDF, set Page range on ExtractText to a page such as 3 or a range such as 2-5. If you know the document language, selecting a specific Tesseract model can improve accuracy. For screenshots or image files, change the conversion to imageToText and allow image/* uploads.
How to import this template
- 1Click Import → Copy JSON on this page.
- 2Open your Heym and navigate to a workflow canvas.
- 3PressCmd+V/Ctrl+V— nodes appear instantly.
- 4Add your API keys in the node config panels and click Run.
Discover more automations
- Document OpsFile Upload Intake MetadataMint a single-use upload link, receive a multipart file, and return clean metadata plus the Drive download URL.
- Document OpsWebsite LoaderFetch clean, LLM-ready website text from any URL using the Jina Reader API.
- Document OpsPDF / DOCX Translation AgentTranslate the full text of any uploaded document using an AI agent.
- Document OpsBatch URL FetcherIterate over a JSON array of URLs with the Loop node, fetch each via HTTP, and merge all responses into one payload.
- Document OpsRAG Document IngestChunk and embed a document into a Qdrant vector store so it can be retrieved later by the RAG Search node.
- Document OpsRAG Q&A AgentSearch your Qdrant vector store for relevant context, then answer with an LLM — grounded in your own documents.
- Document OpsDrive Share Link MailerFetch a remote file into Drive, return the download link immediately, and email the same link asynchronously.
- Document OpsChat with Drive DocsAsk questions about your Google Drive PDF files. An agent lists your Drive, reads the file you need, extracts its text, and answers in plain language.
- Document OpsHTML Table to Markdown (Agent Skill)An agent runs a bundled Python skill that converts HTML tables into clean Markdown tables using only the standard library.
- Document OpsGoogle Drive to Heym Drive BackupMirror a Google Drive folder into Heym Drive on a schedule, exporting Google Docs automatically and posting a summary to Slack.
- Customer SupportHITL Support Reply AgentDraft a customer-facing support response, pause for human approval, then continue with the reviewed reply.
- Dev & IT OpsCodex PR Fix AgentSend a coding task to Codex, open a draft PR when it succeeds, and notify Slack when Codex needs follow-up input.