fix(llm): resolve context window via max_input_tokens with 128k fallback for unmapped models

- get_context_window() now reads litellm.get_model_info().max_input_tokens
  instead of get_max_tokens(), which returns the output cap and severely
  under-reported the window for every mapped model (e.g. deepseek 8k vs 1M)
- models litellm cannot map fall back to 128000 with a single warning per
  model instead of warning on every call and returning None
- raise LLMConfig.max_tokens default 8192 -> 32768 to match the template
- README: configure the API key directly in llm_config.yaml; document
  max_tokens / context_window in the example

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
LZH-YS1998
2026-07-04 17:53:12 +08:00
parent 08e48c2f9c
commit 6c8d3f3dc9
5 changed files with 82 additions and 38 deletions
+5 -5
View File
@@ -11,11 +11,11 @@ llm:
fallback: {}
# Optional: total input context window, used as the denominator for the
# context-usage ring and to trigger auto-summarization. You normally do NOT
# need this — for models litellm can map (most OpenAI / Anthropic models) the
# window is resolved automatically. Only set it when your model is served
# through a proxy / self-hosted endpoint that litellm cannot map, so the ring
# has no denominator. Scalar applies to the default model:
# context-usage ring and to trigger auto-summarization. Auto-detected via
# litellm for most OpenAI / Anthropic models; models litellm cannot map
# (e.g. deepseek/doubao/glm or proxy / self-hosted endpoints) fall back to
# 128000. Set this only when the fallback is wrong for your model.
# Scalar applies to the default model:
# context_window: 200000
# Or per-model (wins over the scalar); keys are matched by model name:
# context_window_overrides: