fix(company): stop/resume identity truth, failure-path closure, quota park

OBS-11 — stop/resume killed pure-native runs over a phantom external pin.
Role templates' preferred_external_agent leaked into execution identity even
when the user requested native and execution actually ran native; on resume
the availability gate trusted the pin and failed every non-terminal item.
Root fixes across the whole chain:
- Staffing card per-role defaults are now the RESOLVED backend (explicit
  session agent choice > runnable template preference > native), never a
  hardcoded external default; seat enrichment and the dispatch selector's
  locked branch downgrade provably unavailable externals to native and
  record the wish in execution_agent_unavailable.
- The resume availability gate fails closed only when a resumable external
  session actually exists; a bare pin heals to native (snapshot AND task
  durable identity) and the run resumes — mirroring dispatch fallback.
- Suspend-checkpoint replies: force_resume (chat/headless spelling) is
  recognized alongside ui_force_resume, and bare continuation tokens
  (English and Chinese spellings) take the plain-resume path instead of
  being routed to the final decider as content, which reopened the
  already-approved intake card.

OBS-5 — failed runs never closed and dropped new input. The dispatcher's
convergence exit now settles terminally-failed runs (status=failed,
lifecycle=closed_failed, run_failure metadata) and emits a
company_run_failure_review card whose replies never swallow messages:
dismiss acknowledges, content falls through so normal routing starts a
fresh run. _maybe_resume_existing_company_runtime no longer re-executes a
terminally-failed tree: control replies get an honest closed status,
content-bearing input starts a new run.

OBS-6 — provider quota exhaustion terminally failed work items. Rate-limit
rejections are classified (LLMProvider.is_rate_limit_error, covering
status codes, exception types, and English/Chinese provider error text),
the agent runtime raises typed ProviderQuotaExhaustedError instead of
burning conversation-feedback retries, and the company dispatcher parks:
the item returns to READY (attempt interrupted, no terminal failure), the
member session idles, and claiming backs off exponentially (60s doubling
to a 900s cap; a quiet 30min resets the streak) before resuming
automatically.

Verified end-to-end on the real minimax-m3 campaign: same goal, same 300s
stop point, same run shape that previously killed the whole tree within
90s now resumes cleanly and completes with all items approved; staffing
defaults native for all 11 roles.

Tests: test_stop_resume_native_pin (10), test_run_failure_settlement (6),
test_provider_quota_park (9); attempt-ledger, recruiter, and
suspend-resume suites updated to the new contracts (their old assertions
pinned the defective behaviors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
LZH-YS1998
2026-07-28 16:59:56 +08:00
parent 14ee8806de
commit d14f3920e0
10 changed files with 1185 additions and 39 deletions
+18 -1
View File
@@ -39,7 +39,7 @@ from opc.layer4_tools.output_budget import clip_text
from opc.layer4_tools.registry import ToolDefinition
from opc.layer4_tools.registry import ToolRegistry
from opc.layer6_observability.cost_tracker import CostEntry
from opc.llm.provider import LLMProvider
from opc.llm.provider import LLMProvider, ProviderQuotaExhaustedError
ApprovalCallback = Callable[[ToolDefinition, dict[str, Any], Optional[Task], Any], Awaitable[tuple[bool, Any]]]
@@ -537,6 +537,23 @@ class NativeRuntimeV2:
},
)
except Exception as exc:
rate_limit_checker = getattr(self.llm, "is_rate_limit_error", None)
if callable(rate_limit_checker) and rate_limit_checker(exc):
# Quota/rate-limit rejections never reach the model, so
# conversation-feedback retries cannot help — surface a
# typed error for the dispatcher to park on instead of
# burning retries and failing the work item (OBS-6).
await self._cancel_early_tool_runs(early_tool_runs)
await self._emit_runtime_event(
runtime_session_id,
task,
"provider_quota_exhausted",
{
"iteration": iteration + 1,
"message": str(exc)[:600],
},
)
raise ProviderQuotaExhaustedError(str(exc)) from exc
if self.llm.is_context_overflow_error(exc) and overflow_retries < max_overflow_retries:
overflow_retries += 1
messages = await self._apply_context_pipeline(