Convert PDF to Markdown from your code
Plug pdf2markdown into your backend, RAG pipeline or automation. Same converter and limits as the web app.
Quick start
- 1Create an API key on the API keys page.
- 2Upload a PDF: POST /convert/upload returns a conversion id.
- 3Poll GET /convert/status/{id} every 2–5 seconds until the status is done or failed.
- 4Download the result: GET /convert/download/{id} returns .md (or .zip when the PDF has images).
Examples
# 1. Upload
curl -s -X POST https://pdf2markdown.ru/api/v1/convert/upload \
-H "X-API-Key: $PDF2MARKDOWN_API_KEY" \
-F "file=@report.pdf;type=application/pdf" -F "format=md"
# → {"id":"<id>","status":"queued"}
# 2. Poll until "done"
curl -s https://pdf2markdown.ru/api/v1/convert/status/<id> -H "X-API-Key: $PDF2MARKDOWN_API_KEY"
# 3. Download (.md, or .zip when the PDF has images)
curl -s -OJ https://pdf2markdown.ru/api/v1/convert/download/<id> -H "X-API-Key: $PDF2MARKDOWN_API_KEY"Endpoints
Base URL: https://pdf2markdown.ru/api/v1
| Method | Path | Description |
|---|---|---|
POST | /convert/upload | Upload a PDF (multipart: file, format=md). Returns {id, status}. |
GET | /convert/status/{id} | Conversion status: queued, processing, done or failed (with error). |
GET | /convert/download/{id} | Result file: text/markdown or application/zip. |
GET | /convert/history | Your conversions, ?page=&size=&status=. |
GET | /convert/limits | Your plan, quota used and limits. |
GET | /convert/batch/{batch_id} | Batch status and all results as one .zip. |
DELETE | /convert/{id} | Delete a conversion and its files. |
Authentication
Send the key in the X-API-Key header (or Authorization: Bearer mk_…). A key only works for /convert/* endpoints: it cannot change the account, billing or other keys.
Limits
API calls use your plan: the monthly quota and file size are shared with the web app. Rate limits are counted per key: 30 requests/min for upload, download and delete, 120/min for status.
Formats: Markdown and JSON for RAG
Every conversion is available as Markdown and as structured JSON: blocks with type, page and heading path (section), tables as rows, images with captions, formulas in LaTeX, plus ready-made sections for chunking. Download with ?format=json.
JSON is available on Lite and Pro. Markdown is on every plan.
{
"schema_version": 1,
"source": { "filename": "report.pdf", "pages": 12 },
"stats": { "blocks": 148, "tables": 3, "images": 2, "formulas": 5, "ocr_pages": [] },
"blocks": [
{ "type": "heading", "level": 1, "page": 1, "section": ["Annual report"],
"text": "Annual report", "markdown": "# Annual report" },
{ "type": "table", "page": 4, "section": ["Annual report", "Revenue"],
"rows": [["Region", "Q1"], ["Moscow", "120"]], "caption": "Table 2. Revenue",
"markdown": "| Region | Q1 |\n|---|---|\n| Moscow | 120 |" },
{ "type": "formula", "page": 7, "latex": "E = m c^{2}", "markdown": "$\nE = m c^{2}\n$" },
{ "type": "image", "page": 9, "image": "assets/page9-img1.png",
"caption": "Figure 3. Architecture" }
],
"sections": [
{ "title": "Revenue", "level": 2, "path": ["Annual report"],
"page_start": 4, "page_end": 5, "markdown": "## Revenue\n\n| Region | Q1 | ..." }
]
}Webhooks (Pro)
Pass callback_url on upload and we POST to it when the conversion finishes. For batches also pass batch_id and batch_size: you get a single batch.completed event when all files are done. Every request is signed: verify X-Pdf2Markdown-Signature with your webhook secret (API keys page).
# 1. Upload with a callback (Pro). For batches add batch_id and batch_size.
curl -X POST https://pdf2markdown.ru/api/v1/convert/upload \
-H "X-API-Key: $PDF2MARKDOWN_API_KEY" \
-F "file=@report.pdf;type=application/pdf" \
-F "callback_url=https://your-app.example.com/pdf2markdown-webhook"
# 2. We POST: {"event": "conversion.completed", "conversion": {"id", "status", "download_url", ...}}
# Headers: X-Pdf2Markdown-Event, X-Pdf2Markdown-Timestamp, X-Pdf2Markdown-Signature: sha256=<hex>
# 3. Verify the signature (Python):
import hashlib, hmac, time
def verify(secret: str, timestamp: str, body: bytes, signature: str) -> bool:
if abs(time.time() - int(timestamp)) > 300: # reject replays
return False
expected = hmac.new(secret.encode(), timestamp.encode() + b"." + body, hashlib.sha256)
return hmac.compare_digest("sha256=" + expected.hexdigest(), signature)Scanned PDFs (OCR, Pro)
On Pro, pages without a text layer are recognized automatically (Russian and English). On other plans such files fail with scan_needs_ocr.
Errors
Errors always look like {"detail": "…", "code": "…"}. Branch on code, not on detail.
401 unauthorized 402 conversion_limit_reached · file_too_large 404 not_found 409 not_ready 410 expired 402 webhooks_not_allowed · batch_not_allowed 415 invalid_file_type 429 rate_limited (Retry-After)
Keep your key safe
- • Call the API from your server only: never put the key in frontend code or a mobile app.
- • Store it in an environment variable or a secret manager, not in the repository.
- • Use one key per integration and revoke a key right away if it may have leaked.
All PDF tools
API for developers
Convert PDFs from your backend, bot or RAG pipeline.