fix(llm): resolve context window via max_input_tokens with 128k fallback for unmapped models
- get_context_window() now reads litellm.get_model_info().max_input_tokens instead of get_max_tokens(), which returns the output cap and severely under-reported the window for every mapped model (e.g. deepseek 8k vs 1M) - models litellm cannot map fall back to 128000 with a single warning per model instead of warning on every call and returning None - raise LLMConfig.max_tokens default 8192 -> 32768 to match the template - README: configure the API key directly in llm_config.yaml; document max_tokens / context_window in the example Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -11,11 +11,11 @@ llm:
|
||||
fallback: {}
|
||||
|
||||
# Optional: total input context window, used as the denominator for the
|
||||
# context-usage ring and to trigger auto-summarization. You normally do NOT
|
||||
# need this — for models litellm can map (most OpenAI / Anthropic models) the
|
||||
# window is resolved automatically. Only set it when your model is served
|
||||
# through a proxy / self-hosted endpoint that litellm cannot map, so the ring
|
||||
# has no denominator. Scalar applies to the default model:
|
||||
# context-usage ring and to trigger auto-summarization. Auto-detected via
|
||||
# litellm for most OpenAI / Anthropic models; models litellm cannot map
|
||||
# (e.g. deepseek/doubao/glm or proxy / self-hosted endpoints) fall back to
|
||||
# 128000. Set this only when the fallback is wrong for your model.
|
||||
# Scalar applies to the default model:
|
||||
# context_window: 200000
|
||||
# Or per-model (wins over the scalar); keys are matched by model name:
|
||||
# context_window_overrides:
|
||||
|
||||
Reference in New Issue
Block a user