Marker: PDF into clean Markdown
Convert PDFs to structured Markdown, 10x better than pdftotext
- Video memory
- 6 GB
- Setup
- ~3 min
- Access
- port 8501
- Billing
- hourly, in BRL
Marker turns PDFs into Markdown while keeping structure: headings, lists, tables and formulas stay readable. It is the missing step in nearly every RAG project, where a badly extracted PDF quietly ruins the model's answers.
Marker preserves tables, equations, code and formatting. REST API + web interface for drag-and-drop PDFs. Supports batch processing.
What it is for
- Prepare a PDF archive for AI-powered search
- Turn manuals and technical standards into searchable text
- Pull tables out of reports without retyping
- Process thousands of pages in batch on the same machine
How to deploy
- Create your account and add balance (card or Pix, no subscription).
- In the console, pick the Marker template and a machine — the console hides the ones that do not meet the requirement.
- In about 3 minutes the setup finishes and the access address shows up in the panel, on port 8501.
Done? Just destroy the machine and billing stops with it. No contract, no minimum commitment.
FAQ
Does it handle scanned PDFs?
Yes, with built-in OCR. For badly degraded documents, pair it with Surya.
Do I need a large GPU?
No. From 6 GB of video memory it runs; bigger cards mainly speed up batches.
Does it preserve tables?
Yes, tables come out as Markdown, which makes them easy to load into a database or a search index.