FastAPI service wrapping the Surya-OCR-2 model (datalab-to) served through vLLM: legacy /v1/api/ai/* endpoints, an OpenAI-compatible /v1/chat/completions endpoint, a coalescing request batcher, a local OCR CLI, Docker packaging, multilingual example outputs, and quantization/concurrency benchmarks. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
15 KiB
Suya OCR API
A FastAPI service that wraps the Surya-OCR-2
document-OCR model and serves it through vLLM for production-grade throughput and latency.
A single 650M vision-language model (qwen3_5 architecture) performs both layout detection
and text recognition — there is no separate torch detection model in the hot path.
HTTP client ──► FastAPI app (surya/endpoint/, served via api.py shim :5002)
├─ legacy.py /v1/api/ai/* envelope endpoints
├─ openai.py /v1/chat/completions (OpenAI-compatible)
├─ request batcher (vllm_batcher.py) ← coalesces concurrent requests
│ ▼
└─ vLLM OpenAI server (:8000) ← Surya-OCR-2 VLM
Inbound OpenAI: clients call
/v1/chat/completions. Outbound OpenAI: the service itself calls the vLLM server viasurya/inference/backends/openai_client.py.
What was improved
This service started as a torch-spawn pipeline and was re-engineered for serving, with the measured results summarized below:
| Area | Change | Result |
|---|---|---|
| Serving engine | Recognition + layout now run through a vLLM OpenAI-compatible server instead of in-process torch | Continuous batching, CUDA graphs, prefix caching |
| vLLM tuning | CUDA graph on, prefix caching on, chunked prefill on (pinned in the launch script) | CUDA graph is the biggest lever — ~9.5× vs eager; prefix cache ~15% |
| Concurrency fix | Matched the batcher's HTTP fan-out (SURYA_INFERENCE_MAX_INFLIGHT) to vLLM's 16 sequence slots |
Closed a 33% mean / 32% p95 regression |
| Request batching | Added a coalescing batcher so concurrent page requests share vLLM batches | Fills all 16 slots under load |
| Quantization study | Benchmarked fp8/int8/awq/gptq/bnb4/bnb8 vs BF16 on A100 | No method beats BF16 on latency; int8 (GPTQ W8A8) near-lossless at ~40% smaller |
| MTP speculative decode | Tested --speculative-config mtp |
Rejected: +27–36% slower at 16-wide block concurrency |
Bottom line: on A100 the latency levers are serving config + batching, not weight quantization or speculative decoding. Keep BF16 with CUDA graphs and prefix caching on.
Examples — multilingual OCR
Real outputs from this service (block mode, single A100) on public-domain newspaper/magazine
scans: the original page, the same page with detected text-block boxes (coloured by block type /
reading order; orange = non-text region skipped), and an excerpt of the recognised text. Full
text for each sample lives in assets/.
Excerpts are the model's raw output on hard, dense historical scans — ornate mastheads (e.g. the gothic “THE NATION”) are the failure cases. Reproduce any row with, e.g.,
python -m surya.scripts.ocr_text assets/russian_original.jpg, or via the API/CLI examples below.
API
FastAPI app served on port 5002. Interactive docs: http://<host>:5002/v1/api/ai/swagger.
All OCR endpoints accept the same JSON body and return the same envelope.
Request body:
{
"file": "<base64-encoded image bytes>",
"type": "png", // one of: png | jpg | jpeg | gif
"skip_text_detection": false, // true ⇒ full-page OCR, skip layout
"skip_table_detection": false, // table endpoint only
"recognize_math": false,
"ocr_with_boxes": true
}
Response envelope:
{ "data": { "ocr_text_json": {...}, "text_lines": "line1\nline2\n...", "elapsed_seconds": 4.64 },
"message": "success",
"code": 200 }
| Method | Path | Purpose | Backend |
|---|---|---|---|
| POST | /v1/api/ai/suya_ocr_vllm/ |
Primary OCR — layout + recognition, request-batched | vLLM |
| GET | /v1/api/ai/suya_ocr_vllm/health |
Backend info (model, base URL, mode, parallelism) | — |
| POST | /v1/api/ai/suya_ocr/ |
OCR (routes to the same vLLM path) | vLLM |
| POST | /v1/api/ai/suya_layout_det/ |
Layout detection only | vLLM |
| POST | /v1/api/ai/suya_table_rec/ |
Table structure recognition | vLLM |
| POST | /image2text |
OCR a multipart file upload (no base64) | vLLM |
The legacy
/v1/api/ai/suya_text_det/endpoint targets the standalone torch detection model, which is not loaded in the vLLM deployment (the VLM does detection internally). It will return an error envelope — use/suya_ocr_vllm/instead.
OpenAI-compatible endpoint (POST /v1/chat/completions)
Point any OpenAI client at http://<host>:5002/v1. Attach the page as an image_url
content part; the joined OCR text comes back in choices[0].message.content, and the
structured per-line JSON + timing come back in a non-standard surya field.
OCR options are passed through extra_body (the OpenAI SDK delivers these at the top
level of the request, which is where the server reads them):
| Field | Default | Purpose |
|---|---|---|
mode |
block |
block (layout→per-block), full_page (one call), or table |
skip_text_detection |
false |
Full-page OCR, skip layout (same as mode:full_page) |
recognize_math |
false |
Math-aware recognition |
skip_table_detection |
false |
mode:table only — treat whole image as one table |
ocr_with_boxes |
true |
When false, omit surya.ocr_text_json (text only) |
Streaming (stream:true) is not supported and returns HTTP 400 — OCR completes
all at once, so there is no token stream.
OpenAI-compatible backend
Recognition and layout are issued as OpenAI Chat Completions calls to the vLLM server
(POST {SURYA_INFERENCE_URL}/chat/completions, default http://127.0.0.1:8000/v1), one per
image crop, with the page image passed as a base64 image_url content part. This means the
inference tier is a standard vLLM OpenAI server — you can point SURYA_INFERENCE_URL at any
OpenAI-compatible vLLM endpoint (local, remote, or a shared cluster).
Concurrency support
Two cooperating layers keep the GPU saturated without overcommitting:
- Request batcher (
vllm_batcher.py): a background thread coalesces incoming page requests that share the same options into one batch — up toSUYA_MAX_BATCH_SIZE(8), waiting at mostSUYA_BATCH_WAIT_MS(25 ms). Requests queue up toSUYA_MAX_QUEUE_SIZE(128); callers block on a future until their page completes. - vLLM sequence slots: within a page, the detected blocks are fanned out over HTTP up to
SURYA_INFERENCE_MAX_INFLIGHT(16) concurrent workers, matching vLLM's--max-num-seqs(16) so every sequence slot is filled. (Overcommitting past 16 only queues — see the diagnosis.)
Measured behaviour (single A100, benchmarks/concurrency_sweep_results.csv) — 1→50 concurrent
clients, 0 failures at every level. Throughput is GPU-bound and plateaus around
0.20 req/s on one model replica:
| Concurrency | Success / Failed | Throughput (req/s) | Mean latency (s) | p95 latency (s) |
|---|---|---|---|---|
| 1 | 1 / 0 | 0.13 | 7.66 | 7.66 |
| 4 | 4 / 0 | 0.19 | 20.95 | 20.96 |
| 8 | 8 / 0 | 0.19 | 34.05 | 41.46 |
| 12 | 12 / 0 | 0.20 | 45.55 | 59.96 |
| 20 | 20 / 0 | 0.20 | 66.68 | 92.91 |
| 30 | 30 / 0 | 0.20 | 90.50 | 147.47 |
| 50 | 50 / 0 | 0.21 | 149.29 | 240.87 |
Scale throughput by adding GPU replicas behind a load balancer, each with its own vLLM server.
How to use
Run with Docker (recommended)
docker build -t suya-ocr-api .
docker run --rm --gpus all -p 5002:5002 suya-ocr-api
The container entrypoint (scripts/start_single_container.sh) starts the vLLM server with the
tuned flags, waits for it to become healthy, then launches the API on :5002.
Deployment target: the Dockerfile defaults target T4 (
VLLM_DTYPE=float16,VLLM_GPU_TYPE=t4). On bf16-capable GPUs (A100/L40/4090) setVLLM_DTYPE=bfloat16.
API mode — call the HTTP service (curl)
# encode an image and OCR it
B64=$(base64 -w0 page.png)
curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/ \
-H 'Content-Type: application/json' \
-d "{\"file\":\"$B64\",\"type\":\"png\",\"ocr_with_boxes\":true}" \
| python3 -c "import sys,json; d=json.load(sys.stdin)['data']; print(d['text_lines']); print('elapsed:', d['elapsed_seconds'])"
# full-page OCR (skip layout) — good for single-column pages
curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/ \
-H 'Content-Type: application/json' \
-d "{\"file\":\"$B64\",\"type\":\"png\",\"skip_text_detection\":true}"
# health / backend info
curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/health
Demo: OCR one image via the OpenAI API
OpenAI Python SDK (the headline "it's really OpenAI-compatible" path):
import base64
from openai import OpenAI
client = OpenAI(base_url="http://localhost:5002/v1", api_key="not-needed")
with open("page.png", "rb") as f:
b64 = base64.b64encode(f.read()).decode()
resp = client.chat.completions.create(
model="surya-ocr",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
{"type": "text", "text": "Extract the text"},
],
}],
extra_body={"mode": "block", "ocr_with_boxes": True},
)
print(resp.choices[0].message.content) # plain OCR text
print(resp.model_dump()["surya"]["ocr_text_json"]) # structured boxes
curl (data-URL base64 image):
B64=$(base64 -w0 page.png)
curl -s http://localhost:5002/v1/chat/completions \
-H 'Content-Type: application/json' \
-d "{\"model\":\"surya-ocr\",\"messages\":[{\"role\":\"user\",\"content\":[{\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/png;base64,$B64\"}},{\"type\":\"text\",\"text\":\"ocr\"}]}]}" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])"
CLI mode — local OCR (no HTTP server)
Run OCR straight from the command line. The surya package spins up its own vLLM-backed
inference backend, so the FastAPI server does not need to be running. Input can be a single
image, a PDF, or a folder of them; results are written as results.json.
# OCR an image or PDF → writes results/surya/<name>/results.json
python -m surya.scripts.ocr_text page.png
# pick an output dir and OCR only some PDF pages (0-indexed; ranges allowed)
python -m surya.scripts.ocr_text doc.pdf --output_dir out --page_range 0,2-4
# also save annotated bbox images, and print per-stage timing
python -m surya.scripts.ocr_text page.png --images --debug
By default the CLI autostarts a vLLM server (SURYA_INFERENCE_AUTOSTART=true) and tears it
down on exit. To reuse a vLLM server that is already running (for example the one started by the
Docker container or scripts/start_single_container.sh), point the CLI at it instead of spawning
a new one:
SURYA_INFERENCE_URL=http://127.0.0.1:8000/v1 python -m surya.scripts.ocr_text page.png
Pass --keep_server to leave a spawned server up so back-to-back commands reuse it. Sibling
commands share the same options (run with --help):
surya.scripts.detect_layout, surya.scripts.detect_text, surya.scripts.table_recognition.
Run without Docker
Requires a CUDA GPU with the matching vLLM. Start the two processes (vLLM first, then the API)
exactly as scripts/start_single_container.sh does, or run that script directly inside an
environment that already has the dependencies from requirements.txt.
Configuration
Set via environment variables (defaults shown; see Dockerfile and surya/settings.py):
| Variable | Default | Purpose |
|---|---|---|
SURYA_INFERENCE_URL |
http://127.0.0.1:8000/v1 |
vLLM OpenAI endpoint |
SURYA_INFERENCE_PARALLEL |
8 |
Recognition worker hint |
SURYA_INFERENCE_MAX_INFLIGHT |
16 |
Max concurrent HTTP calls to vLLM — keep = VLLM_MAX_NUM_SEQS |
SUYA_OCR_MODE |
block |
block (layout→per-block) or full_page (single call) |
SUYA_MAX_BATCH_SIZE |
8 |
Request-batcher coalesce size |
SUYA_BATCH_WAIT_MS |
25 |
Request-batcher wait window |
SUYA_MAX_QUEUE_SIZE |
128 |
Request-batcher backlog cap |
VLLM_DTYPE |
float16 |
bfloat16 on A100/L40/4090; float16 on T4 |
VLLM_MAX_MODEL_LEN |
18000 |
vLLM context length |
VLLM_MAX_NUM_SEQS |
16 |
vLLM sequence slots |
VLLM_GPU_MEMORY_UTILIZATION |
0.85 |
vLLM GPU memory fraction |
Benchmarks
benchmarks/— concurrency sweep scripts and results (see the table above)scripts/quant/— the quantization benchmarking toolchain (build, serve, capture, score, plot)
Acknowledgements
This project is a serving wrapper around Surya OCR by
Datalab. All layout detection and text recognition is done by their
datalab-to/surya-ocr-2 model — this repo only
adds the vLLM serving, request batching, and API/CLI around it. Huge thanks to the Surya team for
building and open-sourcing such a capable multilingual OCR model.
- Official site: https://www.datalab.to/
- Surya GitHub: https://github.com/datalab-to/surya
- Model: https://huggingface.co/datalab-to/surya-ocr-2










