fix(company): stop/resume identity truth, failure-path closure, quota park

OBS-11 — stop/resume killed pure-native runs over a phantom external pin.
Role templates' preferred_external_agent leaked into execution identity even
when the user requested native and execution actually ran native; on resume
the availability gate trusted the pin and failed every non-terminal item.
Root fixes across the whole chain:
- Staffing card per-role defaults are now the RESOLVED backend (explicit
  session agent choice > runnable template preference > native), never a
  hardcoded external default; seat enrichment and the dispatch selector's
  locked branch downgrade provably unavailable externals to native and
  record the wish in execution_agent_unavailable.
- The resume availability gate fails closed only when a resumable external
  session actually exists; a bare pin heals to native (snapshot AND task
  durable identity) and the run resumes — mirroring dispatch fallback.
- Suspend-checkpoint replies: force_resume (chat/headless spelling) is
  recognized alongside ui_force_resume, and bare continuation tokens
  (English and Chinese spellings) take the plain-resume path instead of
  being routed to the final decider as content, which reopened the
  already-approved intake card.

OBS-5 — failed runs never closed and dropped new input. The dispatcher's
convergence exit now settles terminally-failed runs (status=failed,
lifecycle=closed_failed, run_failure metadata) and emits a
company_run_failure_review card whose replies never swallow messages:
dismiss acknowledges, content falls through so normal routing starts a
fresh run. _maybe_resume_existing_company_runtime no longer re-executes a
terminally-failed tree: control replies get an honest closed status,
content-bearing input starts a new run.

OBS-6 — provider quota exhaustion terminally failed work items. Rate-limit
rejections are classified (LLMProvider.is_rate_limit_error, covering
status codes, exception types, and English/Chinese provider error text),
the agent runtime raises typed ProviderQuotaExhaustedError instead of
burning conversation-feedback retries, and the company dispatcher parks:
the item returns to READY (attempt interrupted, no terminal failure), the
member session idles, and claiming backs off exponentially (60s doubling
to a 900s cap; a quiet 30min resets the streak) before resuming
automatically.

Verified end-to-end on the real minimax-m3 campaign: same goal, same 300s
stop point, same run shape that previously killed the whole tree within
90s now resumes cleanly and completes with all items approved; staffing
defaults native for all 11 roles.

Tests: test_stop_resume_native_pin (10), test_run_failure_settlement (6),
test_provider_quota_park (9); attempt-ledger, recruiter, and
suspend-resume suites updated to the new contracts (their old assertions
pinned the defective behaviors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
LZH-YS1998
2026-07-28 16:59:56 +08:00
parent 14ee8806de
commit d14f3920e0
10 changed files with 1185 additions and 39 deletions
+48
View File
@@ -204,6 +204,16 @@ def _parse_tool_arguments(tool_name: str, arguments: Any) -> tuple[Any, str | No
return raw, raw, error
class ProviderQuotaExhaustedError(RuntimeError):
"""The provider rejected the request for quota/rate-limit reasons.
Raised by the agent runtime instead of retrying in place: replaying the
same payload against an exhausted quota can only fail, so the company
dispatcher parks the work (returns the item to READY and backs off)
rather than failing it terminally (OBS-6).
"""
class LLMProvider:
"""Unified LLM interface via LiteLLM supporting tool calls."""
@@ -401,6 +411,44 @@ class LLMProvider:
raw = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str)
return hashlib.sha256(raw.encode("utf-8")).hexdigest()
def is_rate_limit_error(self, error: Exception) -> bool:
"""Classify provider quota / rate-limit rejections.
These never produce model output and never get better by replaying
the identical payload, so callers must park/back off instead of
burning conversation-feedback retries (OBS-6). Classification is by
exception type when available and by error text otherwise — the
streaming path re-raises provider errors as plain RuntimeError with
only the message preserved.
"""
if isinstance(error, litellm.exceptions.RateLimitError):
return True
if "ratelimit" in type(error).__name__.lower():
return True
if getattr(error, "status_code", None) == 429:
return True
message = str(error).lower()
keywords = (
"rate limit",
"rate_limit",
"ratelimit",
"too many requests",
"insufficient_quota",
"quota exceeded",
"exceeded your quota",
"quota exhausted",
"error code: 429",
"status code: 429",
"http 429",
# Chinese-provider spellings of the same rejection (Volces/DeepSeek
# and other domestic endpoints return localized error text).
"请求过于频繁",
"配额已用完",
"配额耗尽",
"触发限流",
)
return any(keyword in message for keyword in keywords)
def is_context_overflow_error(self, error: Exception) -> bool:
if isinstance(error, litellm.exceptions.ContextWindowExceededError):
return True