fix(company): stop/resume identity truth, failure-path closure, quota park
OBS-11 — stop/resume killed pure-native runs over a phantom external pin. Role templates' preferred_external_agent leaked into execution identity even when the user requested native and execution actually ran native; on resume the availability gate trusted the pin and failed every non-terminal item. Root fixes across the whole chain: - Staffing card per-role defaults are now the RESOLVED backend (explicit session agent choice > runnable template preference > native), never a hardcoded external default; seat enrichment and the dispatch selector's locked branch downgrade provably unavailable externals to native and record the wish in execution_agent_unavailable. - The resume availability gate fails closed only when a resumable external session actually exists; a bare pin heals to native (snapshot AND task durable identity) and the run resumes — mirroring dispatch fallback. - Suspend-checkpoint replies: force_resume (chat/headless spelling) is recognized alongside ui_force_resume, and bare continuation tokens (English and Chinese spellings) take the plain-resume path instead of being routed to the final decider as content, which reopened the already-approved intake card. OBS-5 — failed runs never closed and dropped new input. The dispatcher's convergence exit now settles terminally-failed runs (status=failed, lifecycle=closed_failed, run_failure metadata) and emits a company_run_failure_review card whose replies never swallow messages: dismiss acknowledges, content falls through so normal routing starts a fresh run. _maybe_resume_existing_company_runtime no longer re-executes a terminally-failed tree: control replies get an honest closed status, content-bearing input starts a new run. OBS-6 — provider quota exhaustion terminally failed work items. Rate-limit rejections are classified (LLMProvider.is_rate_limit_error, covering status codes, exception types, and English/Chinese provider error text), the agent runtime raises typed ProviderQuotaExhaustedError instead of burning conversation-feedback retries, and the company dispatcher parks: the item returns to READY (attempt interrupted, no terminal failure), the member session idles, and claiming backs off exponentially (60s doubling to a 900s cap; a quiet 30min resets the streak) before resuming automatically. Verified end-to-end on the real minimax-m3 campaign: same goal, same 300s stop point, same run shape that previously killed the whole tree within 90s now resumes cleanly and completes with all items approved; staffing defaults native for all 11 roles. Tests: test_stop_resume_native_pin (10), test_run_failure_settlement (6), test_provider_quota_park (9); attempt-ledger, recruiter, and suspend-resume suites updated to the new contracts (their old assertions pinned the defective behaviors). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -204,6 +204,16 @@ def _parse_tool_arguments(tool_name: str, arguments: Any) -> tuple[Any, str | No
|
||||
return raw, raw, error
|
||||
|
||||
|
||||
class ProviderQuotaExhaustedError(RuntimeError):
|
||||
"""The provider rejected the request for quota/rate-limit reasons.
|
||||
|
||||
Raised by the agent runtime instead of retrying in place: replaying the
|
||||
same payload against an exhausted quota can only fail, so the company
|
||||
dispatcher parks the work (returns the item to READY and backs off)
|
||||
rather than failing it terminally (OBS-6).
|
||||
"""
|
||||
|
||||
|
||||
class LLMProvider:
|
||||
"""Unified LLM interface via LiteLLM supporting tool calls."""
|
||||
|
||||
@@ -401,6 +411,44 @@ class LLMProvider:
|
||||
raw = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str)
|
||||
return hashlib.sha256(raw.encode("utf-8")).hexdigest()
|
||||
|
||||
def is_rate_limit_error(self, error: Exception) -> bool:
|
||||
"""Classify provider quota / rate-limit rejections.
|
||||
|
||||
These never produce model output and never get better by replaying
|
||||
the identical payload, so callers must park/back off instead of
|
||||
burning conversation-feedback retries (OBS-6). Classification is by
|
||||
exception type when available and by error text otherwise — the
|
||||
streaming path re-raises provider errors as plain RuntimeError with
|
||||
only the message preserved.
|
||||
"""
|
||||
if isinstance(error, litellm.exceptions.RateLimitError):
|
||||
return True
|
||||
if "ratelimit" in type(error).__name__.lower():
|
||||
return True
|
||||
if getattr(error, "status_code", None) == 429:
|
||||
return True
|
||||
message = str(error).lower()
|
||||
keywords = (
|
||||
"rate limit",
|
||||
"rate_limit",
|
||||
"ratelimit",
|
||||
"too many requests",
|
||||
"insufficient_quota",
|
||||
"quota exceeded",
|
||||
"exceeded your quota",
|
||||
"quota exhausted",
|
||||
"error code: 429",
|
||||
"status code: 429",
|
||||
"http 429",
|
||||
# Chinese-provider spellings of the same rejection (Volces/DeepSeek
|
||||
# and other domestic endpoints return localized error text).
|
||||
"请求过于频繁",
|
||||
"配额已用完",
|
||||
"配额耗尽",
|
||||
"触发限流",
|
||||
)
|
||||
return any(keyword in message for keyword in keywords)
|
||||
|
||||
def is_context_overflow_error(self, error: Exception) -> bool:
|
||||
if isinstance(error, litellm.exceptions.ContextWindowExceededError):
|
||||
return True
|
||||
|
||||
Reference in New Issue
Block a user