Surya OCR with layout detection
OCR + layout + tables in 90+ languages — scanned PDFs
- Video memory
- 6 GB
- Setup
- ~4 min
- Access
- port 8501
- Billing
- hourly, in BRL
Surya reads scanned documents, understands the reading order and returns the layout — columns, blocks and tables — instead of scrambled text. It covers more than 90 languages.
Surya is next-gen OCR: detects layout (text, figures, tables), table recognition, supports 90+ languages. Ideal for scanned PDFs that Marker misses (invoices, forms, legal documents).
What it is for
- Digitise old contracts, invoices and case files
- Extract tables from scans without retyping
- Prepare a document archive for AI search
- Run in batch, with documents never leaving your machine
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the Surya OCR template and a machine — the console hides the ones that do not meet the requirement.
- In about 4 minutes the setup finishes and the access address shows up in the panel, on port 8501.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
Does it work on invoices and receipts?
Yes; layout detection is exactly what helps on documents full of blocks and tables.
How is it different from Marker?
Surya is the OCR and layout engine; Marker uses that kind of engine to deliver a whole document as Markdown. For batches of scans, Surya gives you more control.
How much video memory do I need?
6 GB is plenty. Page volume matters more than card size.