Add SURYA_STREAM_MODE streaming (off/on/abort) for chat completions
- _generate_one_stream: streams responses, optional live repeat-loop detection (abort) that closes the stream early and flags the partial result as repeat=True so the existing retry loop re-runs with higher temperature - _should_retry honors result.repeat; stream_mode threaded through both first attempt and retry calls in chat_completions_batch - GenerationResult.repeat field; SURYA_STREAM_MODE setting (default off) - diagnostic retry logs (reason=repeat|error|detected, temp) - docker-compose: SURYA_STREAM_MODE=off, 4090 GPU profile, logs volume - logger.yaml: TimedRotatingFileHandler for error handler (fixes startup crash without maxBytes)
This commit is contained in:
@@ -29,6 +29,9 @@ class GenerationResult:
|
||||
raw: str
|
||||
token_count: int
|
||||
error: bool = False
|
||||
# True when streaming mode aborted generation early due to a detected
|
||||
# token repetition loop (see SURYA_STREAM_MODE=abort)
|
||||
repeat: bool = False
|
||||
# Mean of exp(logprob) across response tokens, if logprobs requested
|
||||
mean_token_prob: Optional[float] = None
|
||||
# Per-token logprobs (raw OpenAI-style content list), if requested - phase 2 use
|
||||
|
||||
Reference in New Issue
Block a user