# Suya OCR API A FastAPI service that wraps the [Surya-OCR-2](https://huggingface.co/datalab-to/surya-ocr-2) document-OCR model and serves it through **vLLM** for production-grade throughput and latency. A single 650M vision-language model (`qwen3_5` architecture) performs both **layout detection** and **text recognition** — there is no separate torch detection model in the hot path. ``` HTTP client ──► FastAPI app (surya/endpoint/, served via api.py shim :5002) ├─ legacy.py /v1/api/ai/* envelope endpoints ├─ openai.py /v1/chat/completions (OpenAI-compatible) ├─ request batcher (vllm_batcher.py) ← coalesces concurrent requests │ ▼ └─ vLLM OpenAI server (:8000) ← Surya-OCR-2 VLM ``` > **Inbound OpenAI:** clients call `/v1/chat/completions`. **Outbound OpenAI:** the service itself calls the vLLM server via `surya/inference/backends/openai_client.py`. --- ## What was improved This service started as a torch-spawn pipeline and was re-engineered for serving, with the measured results summarized below: | Area | Change | Result | | --- | --- | --- | | **Serving engine** | Recognition + layout now run through a vLLM **OpenAI-compatible** server instead of in-process torch | Continuous batching, CUDA graphs, prefix caching | | **vLLM tuning** | CUDA graph on, prefix caching on, chunked prefill on (pinned in the launch script) | CUDA graph is the biggest lever — **~9.5× vs eager**; prefix cache ~15% | | **Concurrency fix** | Matched the batcher's HTTP fan-out (`SURYA_INFERENCE_MAX_INFLIGHT`) to vLLM's 16 sequence slots | Closed a **33% mean / 32% p95** regression | | **Request batching** | Added a coalescing batcher so concurrent page requests share vLLM batches | Fills all 16 slots under load | | **Quantization study** | Benchmarked fp8/int8/awq/gptq/bnb4/bnb8 vs BF16 on A100 | **No method beats BF16 on latency**; int8 (GPTQ W8A8) near-lossless at ~40% smaller | | **MTP speculative decode** | Tested `--speculative-config mtp` | **Rejected: +27–36% slower** at 16-wide block concurrency | **Bottom line:** on A100 the latency levers are *serving config + batching*, not weight quantization or speculative decoding. Keep BF16 with CUDA graphs and prefix caching on. --- ## Examples — multilingual OCR Real outputs from **this service** (block mode, single A100) on public-domain newspaper/magazine scans: the original page, the same page with detected text-block boxes (coloured by block type / reading order; orange = non-text region skipped), and an excerpt of the recognised text. Full text for each sample lives in [`assets/`](assets/). | Language | Original | Detected boxes | OCR text (excerpt) | | --- | --- | --- | --- | | **Chinese**
叻報, 1890 | | | 大清光緒十六年
本館新聞除禮拜外日出一張
諸君賜閱本報者無論本埠外埠…
工務局告示 | | **English**
The Nation, 1846 | | | VOL. IV. No. 181.
DUBLIN, SATURDAY, MARCH 28, 1846.
PRICE 6 D.
DAVIS TESTIMONIAL. | | **Arabic**
Al-Ahram, 1981 | | | رئيس مجلس الادارة — عبد الله عبد البخاري
المنطقات المنتصرة
السنة ١٠٠٧ — العدد ٣٩٦٣٨ | | **Russian**
Виттова Пляска, 1905 | | | Виттова Пляска
ОДНОДНЕВНАЯ ГАЗЕТА. ПАЛИ—ТИКО—ФИ—НАНСОВАЯ
Цѣна 5 копѣекъ. | | **French**
Le Miroir des Sports, 1937 | | | LE MIROIR DES SPORTS
Le plus fort tirage des hebdomadaires sportifs
Mardi 10 Août 1937 | > Excerpts are the model's **raw** output on hard, dense historical scans — ornate mastheads > (e.g. the gothic “THE NATION”) are the failure cases. Reproduce any row with, e.g., > `python -m surya.scripts.ocr_text assets/russian_original.jpg`, or via the API/CLI examples below. --- ## API FastAPI app served on port **5002**. Interactive docs: `http://:5002/v1/api/ai/swagger`. All OCR endpoints accept the same JSON body and return the same envelope. **Request body:** ```json { "file": "", "type": "png", // one of: png | jpg | jpeg | gif "skip_text_detection": false, // true ⇒ full-page OCR, skip layout "skip_table_detection": false, // table endpoint only "recognize_math": false, "ocr_with_boxes": true } ``` **Response envelope:** ```json { "data": { "ocr_text_json": {...}, "text_lines": "line1\nline2\n...", "elapsed_seconds": 4.64 }, "message": "success", "code": 200 } ``` | Method | Path | Purpose | Backend | | --- | --- | --- | --- | | POST | `/v1/api/ai/suya_ocr_vllm/` | **Primary OCR** — layout + recognition, request-batched | vLLM | | GET | `/v1/api/ai/suya_ocr_vllm/health` | Backend info (model, base URL, mode, parallelism) | — | | POST | `/v1/api/ai/suya_ocr/` | OCR (routes to the same vLLM path) | vLLM | | POST | `/v1/api/ai/suya_layout_det/` | Layout detection only | vLLM | | POST | `/v1/api/ai/suya_table_rec/` | Table structure recognition | vLLM | | POST | `/image2text` | OCR a multipart file upload (no base64) | vLLM | > The legacy `/v1/api/ai/suya_text_det/` endpoint targets the standalone torch detection model, > which is **not loaded** in the vLLM deployment (the VLM does detection internally). It will > return an error envelope — use `/suya_ocr_vllm/` instead. ### OpenAI-compatible endpoint (`POST /v1/chat/completions`) Point any OpenAI client at `http://:5002/v1`. Attach the page as an `image_url` content part; the joined OCR text comes back in `choices[0].message.content`, and the structured per-line JSON + timing come back in a non-standard `surya` field. OCR options are passed through `extra_body` (the OpenAI SDK delivers these at the top level of the request, which is where the server reads them): | Field | Default | Purpose | | --- | --- | --- | | `mode` | `block` | `block` (layout→per-block), `full_page` (one call), or `table` | | `skip_text_detection` | `false` | Full-page OCR, skip layout (same as `mode:full_page`) | | `recognize_math` | `false` | Math-aware recognition | | `skip_table_detection` | `false` | `mode:table` only — treat whole image as one table | | `ocr_with_boxes` | `true` | When `false`, omit `surya.ocr_text_json` (text only) | Streaming (`stream:true`) is **not supported** and returns HTTP 400 — OCR completes all at once, so there is no token stream. ### OpenAI-compatible backend Recognition and layout are issued as **OpenAI Chat Completions** calls to the vLLM server (`POST {SURYA_INFERENCE_URL}/chat/completions`, default `http://127.0.0.1:8000/v1`), one per image crop, with the page image passed as a base64 `image_url` content part. This means the inference tier is a standard vLLM OpenAI server — you can point `SURYA_INFERENCE_URL` at any OpenAI-compatible vLLM endpoint (local, remote, or a shared cluster). --- ## Concurrency support Two cooperating layers keep the GPU saturated without overcommitting: 1. **Request batcher** (`vllm_batcher.py`): a background thread coalesces incoming page requests that share the same options into one batch — up to `SUYA_MAX_BATCH_SIZE` (8), waiting at most `SUYA_BATCH_WAIT_MS` (25 ms). Requests queue up to `SUYA_MAX_QUEUE_SIZE` (128); callers block on a future until their page completes. 2. **vLLM sequence slots**: within a page, the detected blocks are fanned out over HTTP up to `SURYA_INFERENCE_MAX_INFLIGHT` (16) concurrent workers, matching vLLM's `--max-num-seqs` (16) so every sequence slot is filled. (Overcommitting past 16 only queues — see the diagnosis.) **Measured behaviour** (single A100, `benchmarks/concurrency_sweep_results.csv`) — 1→50 concurrent clients, **0 failures** at every level. Throughput is GPU-bound and plateaus around **0.20 req/s** on one model replica: | Concurrency | Success / Failed | Throughput (req/s) | Mean latency (s) | p95 latency (s) | | --- | --- | --- | --- | --- | | 1 | 1 / 0 | 0.13 | 7.66 | 7.66 | | 4 | 4 / 0 | 0.19 | 20.95 | 20.96 | | 8 | 8 / 0 | 0.19 | 34.05 | 41.46 | | 12 | 12 / 0 | 0.20 | 45.55 | 59.96 | | 20 | 20 / 0 | 0.20 | 66.68 | 92.91 | | 30 | 30 / 0 | 0.20 | 90.50 | 147.47 | | 50 | 50 / 0 | 0.21 | 149.29 | 240.87 | ![Latency vs concurrency](benchmarks/concurrency_sweep_latency.png) Scale throughput by adding GPU replicas behind a load balancer, each with its own vLLM server. --- ## How to use ### Run with Docker (recommended) ```bash docker build -t suya-ocr-api . docker run --rm --gpus all -p 5002:5002 suya-ocr-api ``` The container entrypoint (`scripts/start_single_container.sh`) starts the vLLM server with the tuned flags, waits for it to become healthy, then launches the API on `:5002`. > **Deployment target:** the Dockerfile defaults target **T4** (`VLLM_DTYPE=float16`, > `VLLM_GPU_TYPE=t4`). On bf16-capable GPUs (A100/L40/4090) set `VLLM_DTYPE=bfloat16`. ### API mode — call the HTTP service (curl) ```bash # encode an image and OCR it B64=$(base64 -w0 page.png) curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/ \ -H 'Content-Type: application/json' \ -d "{\"file\":\"$B64\",\"type\":\"png\",\"ocr_with_boxes\":true}" \ | python3 -c "import sys,json; d=json.load(sys.stdin)['data']; print(d['text_lines']); print('elapsed:', d['elapsed_seconds'])" # full-page OCR (skip layout) — good for single-column pages curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/ \ -H 'Content-Type: application/json' \ -d "{\"file\":\"$B64\",\"type\":\"png\",\"skip_text_detection\":true}" # health / backend info curl -s http://localhost:5002/v1/api/ai/suya_ocr_vllm/health ``` ### Demo: OCR one image via the OpenAI API **OpenAI Python SDK** (the headline "it's really OpenAI-compatible" path): ```python import base64 from openai import OpenAI client = OpenAI(base_url="http://localhost:5002/v1", api_key="not-needed") with open("page.png", "rb") as f: b64 = base64.b64encode(f.read()).decode() resp = client.chat.completions.create( model="surya-ocr", messages=[{ "role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}}, {"type": "text", "text": "Extract the text"}, ], }], extra_body={"mode": "block", "ocr_with_boxes": True}, ) print(resp.choices[0].message.content) # plain OCR text print(resp.model_dump()["surya"]["ocr_text_json"]) # structured boxes ``` **curl** (data-URL base64 image): ```bash B64=$(base64 -w0 page.png) curl -s http://localhost:5002/v1/chat/completions \ -H 'Content-Type: application/json' \ -d "{\"model\":\"surya-ocr\",\"messages\":[{\"role\":\"user\",\"content\":[{\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/png;base64,$B64\"}},{\"type\":\"text\",\"text\":\"ocr\"}]}]}" \ | python3 -c "import sys,json; print(json.load(sys.stdin)['choices'][0]['message']['content'])" ``` ### CLI mode — local OCR (no HTTP server) Run OCR straight from the command line. The `surya` package spins up its own vLLM-backed inference backend, so the FastAPI server does **not** need to be running. Input can be a single image, a PDF, or a folder of them; results are written as `results.json`. ```bash # OCR an image or PDF → writes results/surya//results.json python -m surya.scripts.ocr_text page.png # pick an output dir and OCR only some PDF pages (0-indexed; ranges allowed) python -m surya.scripts.ocr_text doc.pdf --output_dir out --page_range 0,2-4 # also save annotated bbox images, and print per-stage timing python -m surya.scripts.ocr_text page.png --images --debug ``` By default the CLI **autostarts a vLLM server** (`SURYA_INFERENCE_AUTOSTART=true`) and tears it down on exit. To reuse a vLLM server that is already running (for example the one started by the Docker container or `scripts/start_single_container.sh`), point the CLI at it instead of spawning a new one: ```bash SURYA_INFERENCE_URL=http://127.0.0.1:8000/v1 python -m surya.scripts.ocr_text page.png ``` Pass `--keep_server` to leave a spawned server up so back-to-back commands reuse it. Sibling commands share the same options (run with `--help`): `surya.scripts.detect_layout`, `surya.scripts.detect_text`, `surya.scripts.table_recognition`. ### Run without Docker Requires a CUDA GPU with the matching vLLM. Start the two processes (vLLM first, then the API) exactly as `scripts/start_single_container.sh` does, or run that script directly inside an environment that already has the dependencies from `requirements.txt`. --- ## Configuration Set via environment variables (defaults shown; see `Dockerfile` and `surya/settings.py`): | Variable | Default | Purpose | | --- | --- | --- | | `SURYA_INFERENCE_URL` | `http://127.0.0.1:8000/v1` | vLLM OpenAI endpoint | | `SURYA_INFERENCE_PARALLEL` | `8` | Recognition worker hint | | `SURYA_INFERENCE_MAX_INFLIGHT` | `16` | Max concurrent HTTP calls to vLLM — **keep = `VLLM_MAX_NUM_SEQS`** | | `SUYA_OCR_MODE` | `block` | `block` (layout→per-block) or `full_page` (single call) | | `SUYA_MAX_BATCH_SIZE` | `8` | Request-batcher coalesce size | | `SUYA_BATCH_WAIT_MS` | `25` | Request-batcher wait window | | `SUYA_MAX_QUEUE_SIZE` | `128` | Request-batcher backlog cap | | `VLLM_DTYPE` | `float16` | `bfloat16` on A100/L40/4090; `float16` on T4 | | `VLLM_MAX_MODEL_LEN` | `18000` | vLLM context length | | `VLLM_MAX_NUM_SEQS` | `16` | vLLM sequence slots | | `VLLM_GPU_MEMORY_UTILIZATION` | `0.85` | vLLM GPU memory fraction | --- ## Benchmarks - `benchmarks/` — concurrency sweep scripts and results (see the table above) - `scripts/quant/` — the quantization benchmarking toolchain (build, serve, capture, score, plot) --- ## Acknowledgements This project is a serving wrapper around **[Surya OCR](https://github.com/datalab-to/surya)** by **[Datalab](https://www.datalab.to/)**. All layout detection and text recognition is done by their [`datalab-to/surya-ocr-2`](https://huggingface.co/datalab-to/surya-ocr-2) model — this repo only adds the vLLM serving, request batching, and API/CLI around it. Huge thanks to the Surya team for building and open-sourcing such a capable multilingual OCR model. - Official site: - Surya GitHub: - Model: