- _generate_one_stream: streams responses, optional live repeat-loop
detection (abort) that closes the stream early and flags the partial
result as repeat=True so the existing retry loop re-runs with higher
temperature
- _should_retry honors result.repeat; stream_mode threaded through both
first attempt and retry calls in chat_completions_batch
- GenerationResult.repeat field; SURYA_STREAM_MODE setting (default off)
- diagnostic retry logs (reason=repeat|error|detected, temp)
- docker-compose: SURYA_STREAM_MODE=off, 4090 GPU profile, logs volume
- logger.yaml: TimedRotatingFileHandler for error handler (fixes startup
crash without maxBytes)
FastAPI service wrapping the Surya-OCR-2 model (datalab-to) served through vLLM:
legacy /v1/api/ai/* endpoints, an OpenAI-compatible /v1/chat/completions endpoint,
a coalescing request batcher, a local OCR CLI, Docker packaging, multilingual
example outputs, and quantization/concurrency benchmarks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>