fix(company): stop/resume identity truth, failure-path closure, quota park
OBS-11 — stop/resume killed pure-native runs over a phantom external pin. Role templates' preferred_external_agent leaked into execution identity even when the user requested native and execution actually ran native; on resume the availability gate trusted the pin and failed every non-terminal item. Root fixes across the whole chain: - Staffing card per-role defaults are now the RESOLVED backend (explicit session agent choice > runnable template preference > native), never a hardcoded external default; seat enrichment and the dispatch selector's locked branch downgrade provably unavailable externals to native and record the wish in execution_agent_unavailable. - The resume availability gate fails closed only when a resumable external session actually exists; a bare pin heals to native (snapshot AND task durable identity) and the run resumes — mirroring dispatch fallback. - Suspend-checkpoint replies: force_resume (chat/headless spelling) is recognized alongside ui_force_resume, and bare continuation tokens (English and Chinese spellings) take the plain-resume path instead of being routed to the final decider as content, which reopened the already-approved intake card. OBS-5 — failed runs never closed and dropped new input. The dispatcher's convergence exit now settles terminally-failed runs (status=failed, lifecycle=closed_failed, run_failure metadata) and emits a company_run_failure_review card whose replies never swallow messages: dismiss acknowledges, content falls through so normal routing starts a fresh run. _maybe_resume_existing_company_runtime no longer re-executes a terminally-failed tree: control replies get an honest closed status, content-bearing input starts a new run. OBS-6 — provider quota exhaustion terminally failed work items. Rate-limit rejections are classified (LLMProvider.is_rate_limit_error, covering status codes, exception types, and English/Chinese provider error text), the agent runtime raises typed ProviderQuotaExhaustedError instead of burning conversation-feedback retries, and the company dispatcher parks: the item returns to READY (attempt interrupted, no terminal failure), the member session idles, and claiming backs off exponentially (60s doubling to a 900s cap; a quiet 30min resets the streak) before resuming automatically. Verified end-to-end on the real minimax-m3 campaign: same goal, same 300s stop point, same run shape that previously killed the whole tree within 90s now resumes cleanly and completes with all items approved; staffing defaults native for all 11 roles. Tests: test_stop_resume_native_pin (10), test_run_failure_settlement (6), test_provider_quota_park (9); attempt-ledger, recruiter, and suspend-resume suites updated to the new contracts (their old assertions pinned the defective behaviors). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -420,7 +420,12 @@ class ResumeAvailabilityGateTests(unittest.IsolatedAsyncioTestCase):
|
||||
)
|
||||
return engine
|
||||
|
||||
async def test_resume_fails_closed_when_pinned_agent_unavailable(self) -> None:
|
||||
async def test_resume_heals_pin_to_native_when_agent_unavailable(self) -> None:
|
||||
"""A pin to a disabled external agent with NO resumable external
|
||||
session heals to native and resumes (OBS-11): dispatch would fall
|
||||
back to native anyway, so failing the item punished runs — including
|
||||
fully native ones whose metadata inherited a template preference —
|
||||
for an availability gap that does not block execution."""
|
||||
store = await self._store()
|
||||
task = await self._seed(store, external_agent="codex")
|
||||
engine = self._engine(store, available=["opencode"]) # codex disabled
|
||||
@@ -449,19 +454,25 @@ class ResumeAvailabilityGateTests(unittest.IsolatedAsyncioTestCase):
|
||||
self.assertIsNotNone(response)
|
||||
refreshed_item = await store.get_delegation_work_item("work-item-1")
|
||||
assert refreshed_item is not None
|
||||
self.assertEqual(refreshed_item.phase, Phase.FAILED)
|
||||
self.assertIn("codex", str(refreshed_item.blocked_reason or ""))
|
||||
refreshed_task = await store.get_task(task.id)
|
||||
assert refreshed_task is not None
|
||||
self.assertEqual(refreshed_task.status, TaskStatus.FAILED)
|
||||
self.assertEqual(refreshed_item.phase, Phase.RUNNING)
|
||||
self.assertIn("tasks", executed)
|
||||
resumed_task = executed["tasks"][0]
|
||||
self.assertIsNone(resumed_task.assigned_external_agent)
|
||||
self.assertEqual(
|
||||
refreshed_task.metadata.get("resume_unavailable_external_agent"),
|
||||
resumed_task.metadata.get("selected_execution_agent"), "native"
|
||||
)
|
||||
self.assertEqual(
|
||||
resumed_task.metadata.get("resume_execution_agent_healed_from"),
|
||||
"codex",
|
||||
)
|
||||
# The runtime still executed (the rest of the org resumes normally).
|
||||
self.assertIn("tasks", executed)
|
||||
pin = dict(
|
||||
resumed_task.metadata.get(
|
||||
"_company_runtime_resume_execution_agent_pin", {}
|
||||
)
|
||||
)
|
||||
self.assertEqual(pin.get("selected_execution_agent"), "native")
|
||||
|
||||
async def test_plain_message_after_gate_failure_converges_without_revival(self) -> None:
|
||||
async def test_plain_message_after_terminal_failure_converges_without_revival(self) -> None:
|
||||
"""A plain text follow-up (final-decider routing path) on a run whose
|
||||
decider card failed terminally must drain the checkpoint and must not
|
||||
clobber the FAILED task back to PENDING (InvalidPhaseTransition crash
|
||||
@@ -480,12 +491,14 @@ class ResumeAvailabilityGateTests(unittest.IsolatedAsyncioTestCase):
|
||||
return "runtime resumed"
|
||||
|
||||
engine.company_executor = DummyCompanyExecutor()
|
||||
# First resume: gate fails the codex-pinned decider card closed.
|
||||
await engine._maybe_resume_checkpoint(
|
||||
"continue",
|
||||
"sess-parent",
|
||||
reply_metadata={"ui_force_resume": True},
|
||||
)
|
||||
# The decider card failed terminally after the suspend (the resume
|
||||
# gate no longer fails pins without external sessions — OBS-11 — so
|
||||
# the terminal failure is seeded directly).
|
||||
await store.update_delegation_work_item("work-item-1", phase=Phase.FAILED)
|
||||
failed_task = await store.get_task(task.id)
|
||||
assert failed_task is not None
|
||||
failed_task.status = TaskStatus.FAILED
|
||||
await store.save_task(failed_task)
|
||||
refreshed_item = await store.get_delegation_work_item("work-item-1")
|
||||
assert refreshed_item is not None
|
||||
self.assertEqual(refreshed_item.phase, Phase.FAILED)
|
||||
|
||||
Reference in New Issue
Block a user