Skip to content

Background process jobs

Background process jobs let the existing Pi-native Exec and Bash tools return immediately while mono-agent continues to own the spawned POSIX process group. There is no separate job tool. When the host has an available process-job controller, both schemas gain the optional background: true field. Without that controller, background starts are unavailable. A configured host still discloses request lineage diagnostics in the tool descriptions, including chainDepth, maxChainDepth, remainingStarts, and unavailableReason=chain_depth_exhausted when the budget is spent.

Use this for a command that should outlive the current model turn but still report back to the exact Slack thread, Telegram chat, or web-console thread that started it. It is independent from durable continuations: a process job owns a local Exec or Bash child, while a continuation accepts a later result from a selected external MCP service.

The feature is opt-in and every key is JSON-only. Unknown keys are rejected. stateDir must be a relative child of the agent root. Its canonical path must be disjoint from every root removed by restart --clear-sessions: Pi provider sessions, durable message/tool history, and ACP session authorizations. The check covers every root retained in the durable registry, not only the current configuration. Startup and clear-sessions preflight reject equality or containment in either direction, including lexical and canonical aliases, so clearing conversation state cannot delete process-job records or output.

{
"processJobs": {
"enabled": true,
"unsafeAllowUnprotectedState": false,
"stateDir": ".mono-agent/process-jobs",
"maxConcurrent": 4,
"maxActivePerConversation": 2,
"maxQueued": 8,
"maxRuntimeMs": 1800000,
"maxQueueAgeMs": 300000,
"maxOutputBytes": 1048576,
"previewChars": 2000,
"maxChainDepth": 4,
"retention": {
"maxRecords": 1000,
"maxAgeMs": 604800000,
"artifactMaxBytes": 268435456
}
}
}
SettingDefaultCompiled maximum
unsafeAllowUnprotectedStatefalse
maxConcurrent432
maxActivePerConversation28
maxQueued864
maxRuntimeMs30 minutes24 hours
maxQueueAgeMs5 minutes1 hour
maxOutputBytes1 MiB8 MiB
previewChars2,0008,000
maxChainDepth464
retention.maxRecords1,00010,000
retention.maxAgeMs7 days30 days
retention.artifactMaxBytes256 MiB1 GiB

The compiled maximum bounds configuration, configuration bounds each job, and a tool call may narrow only timeout_ms and max_output_chars. Queue age starts at admission. The runtime deadline starts when the detached launch gate is spawned immediately before ownership is persisted, so waiting in a busy queue does not consume the runtime budget.

Before mono-agent creates an enabled stateDir, opens its store, or writes a store secret, @mono-agent/agent-app securely publishes that root to .mono-agent/process-jobs-roots-v1/registry.json. Opening the service requires the exact registration proof. The v1 manifest is absent exactly while no root has ever been registered for the agent. It stores sorted agent-root-relative segments and a fresh generation id, with these fixed bounds:

Registry boundMaximum
Retained roots64
Segments per root64
UTF-8 bytes per segment255
UTF-8 bytes per relative root2 KiB
Encoded manifest256 KiB

The manifest is an owner-only, no-follow, regular single-link file. Updates use an owner-only mutation lock, atomic secure replacement, directory fsync, and an identity-and-content reread. Unsafe, malformed, or over-bound state fails closed with the path-free Process-job private-state protection is unavailable. error. The registry directory remains strict: its only permitted entry is registry.json. Replacement artifacts instead use the owner-only sibling .mono-agent/process-jobs-roots-v1.recovery/, which is mode 0700, must share the registry filesystem, and permits at most these three mode-0600 regular files:

  • registry.staging.json
  • registry.previous.json
  • registry.failed.json

Cross-directory publication fsyncs each affected directory at the namespace mutation, destination before source. Ordinary request loading performs only a bounded inspection of the recovery directory and fails closed if any artifact is present; it never repairs or removes one. Only root registration and clear-sessions preflight may recover while holding the existing registry mutation lock. Recovery artifacts are single-link in steady processing except for one exact crash-transient: after a previous manifest is linked back to registry.json and before registry.previous.json is unlinked, those two fixed names may be the same proven inode with identical bytes and nlink=2. Requests still fail closed and leave that pair untouched. Locked recovery alone may reprove the exact pair, make the target link durable, unlink the previous name, fsync the recovery directory, and reprove registry.json at nlink=1; every other hard-linked shape remains untouched and fail-closed.

A valid current manifest wins and proven artifacts are removed; an absent current manifest is restored from a valid previous manifest (recreating only a proven-absent registry directory when necessary) before proven staging/failed cleanup; staging alone is discarded so a fresh first registration can proceed. Failed-only state, a corrupt or ambiguous current manifest, unknown or fourth entries, and unsafe directories, links, ownership, modes, or artifact contents remain fail-closed. The recovery directory is empty in steady state.

It is never auto-pruned: disabling or removing processJobs, changing A to B, or restarting retains A, and an A-to-B change protects both roots. The registry directory, recovery directory, and every retained root’s lexical and canonical aliases are native protected roots and reply-artifact private roots. A degraded or failed store open therefore leaves all earlier roots sealed even though the background controller itself remains optional.

Every official local request captures and re-attests the current registry generation, then acquires its generation lease before request resource extensions or provider invocation. A newly registered generation becomes the current generation before its store may open; opening waits for older leases that did not cover the new root. The bounded drain timeout fails closed and does not create the directory, store, or secret. A request releases its lease only from settleCleanup, after runtime.run truly settles. Earlier cleanup, abort, or harness disposal cannot release it beneath a late provider result.

Once the registry is non-empty, every reachable primary, fallback, accepted request override, and named Agent child route must be Pi-native — which every route now is. The configured app rejects the whole incompatible route plan before provider invocation; it does not wait to discover the unsafe fallback after another route fails. An empty registry preserves legitimate non-Pi routes.

Tool-less direct configured memory LLM and embedding-provider calls are the one exception, in every posture. SRT confines the model’s tool loop; these surfaces run no tool loop and touch no filesystem — an embedding provider is a single embed(texts) HTTP call with nothing the model can steer — so confinement has nothing to protect there. They still take the canonical owner and the registry-generation lease and hold both until the provider promise truly settles, which is the control that does apply. Rejecting them instead made processJobs.enabled mutually exclusive with the bujo and journal memory tiers, which require an embedding provider. Eligible Pi-native turns receive the real SRT policy for the registry and every retained root independently of whether they receive the background controller. Model Read, Write, Edit, Glob, Grep, Bash, and Exec cannot read, replace, rename, search, or use those paths as a workdir. Host filesystem tools perform their actual file operation through the native sandbox, closing symlink swaps after path authorization. SRT also denies a rename of any ancestor that would move a protected leaf.

Only the registry and retained state directories are protected: workspace siblings such as .mono-agent/artifacts/attachments remain readable, including when the workspace itself is nested under .mono-agent/. When the configured sandbox is absent or off, this filesystem-only protection preserves unrestricted network behavior for commands, WebFetch, and WebSearch. A configured native network policy remains unchanged. Provider-owned non-Pi tool loops cannot enforce this host policy, so they are rejected while any registered private root exists.

This agent-root-aware coverage belongs to the full app, configured harness/responder, local TUI, the lazy-run wrapper returned by createConfiguredAgentRuntime, configured named children, and direct configured memory LLM/embedding calls. Remote TUI and ACP bridges are thin clients of an already owned host and do not invoke a provider themselves. Lower-level @mono-agent/runtime-adapter and @mono-agent/agent-runtime factories are root-agnostic unless an app-owned caller supplies the protection policy and route gate.

processJobs.unsafeAllowUnprotectedState: true is a JSON-only escape hatch for an operator who intentionally runs trusted same-user host tools. It is accepted only when all of these conditions hold:

  • sandbox.mode is present and exactly "off";
  • ProcessJobs is enabled, or the attested durable registry still retains at least one root; and
  • every configured primary, fallback, named Agent child, and agent-host memory route is Pi-native. Accepted request overrides are checked again before any provider is invoked.

The escape hatch changes only SRT policy injection. Registry load and attestation, including the second request-boundary attestation, owner and generation leases through true settlement, root disjointness, retention, store/service/controller lifecycle, and reply-artifact privateRoots remain in force. A failed or unavailable registry always wins and remains provider-zero. The Pi-only gates are independent from sandbox.protectedRoots, so mixed primary/fallback chains, non-Pi request overrides, non-Pi named children, and non-Pi agent-host memory remain rejected before provider work.

Tool-less direct Ollama memory LLM and embedding calls run in this posture as they do in the safe one, still holding the canonical owner and registry-generation lease until the provider promise truly settles. Agent-host memory remains tool-less and Pi-native. Public package-root runtime, harness, responder, and memory factories do not accept this authority and keep their safe behavior.

Changing this posture takes effect only when the app rebuilds its owned runtime surfaces: a managed configuration apply performs that teardown/rebuild, and a process restart does the same. Existing in-flight runs are not mutated. No state migration or reset occurs, and retained registry roots remain registered when the flag is enabled, disabled, or removed. validate, foreground/background status, trace metadata, and the local TUI summary show the path-free warning UNSAFE: ProcessJobs state and operator secret are model-accessible.

Official local hosts also take one cooperative lifetime lease keyed by the canonical realpath of the agent root and stored under the effective account home. In one process, repeated configured app/harness/responder/local-TUI owners share a reentrant reference count. Physical release waits for all owner references and all true-settlement request leases; a release failure makes every later in-process acquisition fail deterministically. A stale official-process lease is recoverable after a crash. The lease path and random owner token are host-only coordination data, but their permissions and hash-derived pathname are not a secrecy or tamper-resistance claim.

This is cooperative serialization, not a security boundary against actively hostile code running as the same OS user. Such code can signal or SIGKILL the host, rewrite same-UID control state, and attack another registered agent root under the same account. Alternate same-UID coordination mechanisms do not change that fact. The local persistence guarantee covers exactly the roots durably registered for this agent; it is not an account-global provider-zero rule. If a provider is in that threat model, run it under a distinct UID or another real privilege-separation boundary.

The host injects the controller only when all of these are true at call time:

  • processJobs.enabled is true and the owner-private store lock is ready;
  • the platform is POSIX; Windows is unsupported;
  • the selected request route is Pi-native;
  • the ordinary tool policy permits Exec or Bash; and
  • the turn originated in an exact addressable conversation whose channel driver opts into the ProcessJobs capability. The built-ins are Slack, Telegram, the web console, and the WhatsApp plugin; future plugins may claim one unique conversation-id scheme and publish the same running capability.

Direct TUI turns, cron, webhook, OpenAI API, A2A, and plugins without the explicit capability do not get background schemas. Duplicate or malformed scheme claims fail during app startup. A controller also rejects an invalid origin as background_unsupported_channel. Foreground calls continue to use their normal route and policy.

Use the existing tool exactly as before and add background: true:

{
"executable": "pnpm",
"args": ["test"],
"workdir": ".",
"description": "Running the full repository test suite",
"timeout_ms": 900000,
"max_output_chars": 4000,
"background": true
}

Bash preserves its clean non-interactive shell semantics; Exec preserves literal argv semantics. Both reuse the exact sandbox-prepared command that a foreground call would launch. The immediate result contains only an opaque job_id, state, and started_at; it does not expose argv, environment, process ids, paths, or secrets.

description is a short statement of the work’s purpose. For background jobs, the host bounds and redacts it before showing the same safe summary throughout the lifecycle surface and durable terminal record. The raw description is not persisted; argv, command content, environment values, and host-specific workspace/home paths remain hidden.

The agent is told when to reach for this and what not to do afterwards in three places, all gated on the same availability check as the schema itself: the background field description states that foreground is the default, that a background job costs an extra turn and defers the answer, that only work expected to exceed the foreground ceiling or to keep running after the reply belongs there, and that a restart of the agent interrupts every job (so backgrounding is not a way to run something “while replying”); the start result leads with a line saying the conversation is woken on completion, so the agent must not poll, sleep, or re-run the command to check; and the session block of the system prompt repeats both alongside the daemonize prohibition and the fact that job output arrives as untrusted evidence. The same availability check also relaxes the prompt’s continuation rule, which otherwise has the agent announce that background delivery was not scheduled for a job it just started. No operator command is named in any model-facing copy — the agent has a shell, and naming a status command invites the polling this is meant to prevent.

A chain starts at depth zero; a completion wake inherits its parent’s depth plus one. With maxChainDepth: 32, depth 31 permits one further background stage and its depth-32 wake permits none. Steering, queued follow-ups and retries retain host-owned lineage. The default remains 4; the ceiling is 64, and concurrency, queue, output and runtime limits remain independent.

wake_on_completion is an optional boolean on background Exec/Bash calls. It defaults to true. Set it to false explicitly for a helper that should update its terminal lifecycle card without scheduling a completion turn. Using it without background: true is invalid. This preference survives restart; older records retain the default wake behavior. Cancellation never restarts a command.

A genuine completion wake can answer with exactly NOTHING_TO_REPORT to suppress delivery. Narration and rich reply parts remain visible. The host matches the exact active delivery key; unrelated and stale keys cannot silence another turn. Web removes the sentinel from the settled reply and emits no response push for a reply without visible content.

A timeout before the destination confirms steering or durably admits the exact follow-up may leave delivery uncertain. The terminal wake state is unknown, with an explicit outcome-unknown error; automatic replay is suppressed, including after restart. The web destination receipts a durably admitted follow-up without waiting for its model turn to finish; that turn’s later success, failure, cancellation, or interruption remains independently visible. A definite refusal remains failed or follows its existing bounded safe-retry policy.

The process owns its sandbox settings until every process remaining in its owned POSIX process group exits. On POSIX a command-agnostic detached group leader starts first. Mono-agent persists its PID, equal PGID, and process-incarnation evidence before releasing the exact command and environment over an anonymous pipe. A crash before that commit cannot spawn the target; a crash after release leaves a recoverable owned group. Timeout, cancellation, and matched restart recovery therefore wait for inherited-group descendants before sandbox cleanup. Commands that deliberately daemonize into another POSIX process group or session are not contained by this contract and must not use background: true.

Live timeout, cancellation, and shutdown use a bounded SIGTERM then SIGKILL sequence. While the exact self-led ChildProcess leader is live and unreaped, its negative PGID remains authoritative even if the owner’s event loop stalls; the kernel cannot recycle that live identity. The host begins bounded-frequency group observation when the leader reports exit and signals the recorded negative PGID only while that post-exit proof remains continuous. A post-exit over-limit or indeterminate observation gap permanently revokes signalling authority. If termination or final group absence cannot be proved, the job settles with an explicit degraded error and leaves sandbox settings intact for operator investigation instead of hanging or cleaning beneath a possibly live descendant.

Jobs move through this durable state machine:

queued -> starting -> running -> succeeded | failed | timed_out | cancelled
| |
| +----------------> spawn_failed | cancelled
+---------------------------> queue_expired | cancelled
any nonterminal at restart -> interrupted

Wake delivery is orthogonal: pending, delivered, failed, unknown, or suppressed, with a stable delivery key and attempt count. A terminal transition is lock-idempotent and schedules one wake unless explicitly opted out. An adapter result that explicitly proves retry is safe gets at most three attempts with the same delivery key, including across restart. Ambiguous wake attempts are not replayed automatically, because a second post could duplicate a real first delivery.

For an active turn in the exact originating conversation, the adapter first offers the completion as live input targeted to that run id. It reserves the normal follow-up position before making the offer. Confirmed exact transcript consumption keeps the completion in that turn; only an explicit unavailable or requeued settlement runs the reserved normal follow-up turn. Discarded, uncertain, rejected, or unknown settlement is ambiguous and never triggers an automatic duplicate. Consumption does not prove provider receipt or adherence. Slack, Telegram, and WhatsApp use their ordinary visible thinking/tool/final stream for the fallback. The web console creates an assistant-only turn, emits the same NDJSON activity/tool frames, and never invents a user message.

Every steer or fallback carries the stable delivery key out of band. The web console durably records accepted and completed delivery claims. A steered completion means the exact active run accepted the live input; a follow_up completion means the agent accepted that exact assistant turn request, not that its model work succeeded. Acceptance is also when the wake’s host-owned capability binds to the request, so the follow-up keeps the parent’s chain depth and its remaining background starts; a receipt that returned earlier would end the wake and leave the turn it raised unable to start the next job. After restart a completed claim returns its prior receipt while the associated turn can independently recover as interrupted; an accepted but unsettled claim fails closed as ambiguous. A wake is a genuine tool-capable turn, not continuation synthesis. The host raises the active controller to the parent job’s chain depth plus one before any steered tool call can start; a non-consumed offer rolls that provisional depth back, and the configured maximum remains authoritative.

A Slack or Telegram conversation that is already at its pre-turn admission cap does not spend that three-attempt budget. The wake stays durably pending and is re-armed on a separate, longer timer; a restart can deliver it later. Busy refusals do not themselves exhaust the wake: the durable bound is five minutes from its first refusal, so an ordinary busy turn can clear and receive the same delivery identity. Age exhaustion settles the wake as failed and immediately runs retention. Once a turn is admitted, any ambiguous failure is nonretryable and exactly-once wins over automatic replay.

An absent or disabled destination channel is also a proven pre-dispatch refusal, but it has a separate durable bound of three checks. Those checks do not change the delivery attempt count or timestamp and reuse the same stable delivery key. If the channel returns before exhaustion, delivery continues normally; otherwise the wake settles as failed so retention can reclaim its record and artifacts. Conversation-cap busy admission remains distinct and does not spend this absent-channel bound.

Admission counts every pending wake obligation, including a queued or running job whose terminal wake is not due yet. It rejects process_job_capacity at retention.maxRecords + maxConcurrent + maxQueued obligations (1,012 by default; compiled maximum 10,096). This bounds the number of separate busy-wake rearm timers without silently evicting an obligation. Recovery keeps the oldest obligations when repairing legacy overflow and records explicit failed-wake outcomes for the excess before ordinary retention can reclaim them.

The store also has a compiled open ceiling of 20,096 record entries: the 10,000 retained-record maximum plus the 10,096 pending-obligation maximum. Startup streams and bounds the record directory before materializing records. A legacy or externally modified directory above that ceiling fails closed with a stable, path-free storage error; it does not delete pending wakes, artifacts, or active ownership records. The same check prevents transaction replay from growing the store past the ceiling. An operator must inspect and remediate that owner-private state before restart can proceed.

Pending-wake records and their referenced artifacts remain live and are exempt from age, count, and artifact-byte pruning until delivery settles. Other terminal records and artifacts are pruned oldest-first with job-id tie-breaking. Retention runs at startup, after every terminal completion, and after each wake settles. Every 64 retention applications also reconcile orphan artifact directories, so long-running agents reach crash-cleanup work without restarting.

Stdout and stderr are stored separately under the configured output budget. While a process is running, its operator projection exposes a memory-only tail of the newest 100 logical stdout/stderr lines. It is refreshed at most every 250 milliseconds, carries stream headings and an omission marker, and remains under the job’s previewChars bound. Chunk callbacks never mutate the durable store, write artifacts, or schedule channel lifecycle edits. The final redacted tail is persisted when the job settles; full bounded stdout/stderr remain available through their separate artifact references.

The model, CLI, operator API, and web card receive only bounded redacted previews plus agent-root-relative artifact references. Redaction occurs before line and character cuts, including for known secrets split across UTF-8 chunks or physical lines, credential-shaped values, and PEM blocks. Treat every preview as untrusted process output. Records retain a redacted command summary and only environment key names; raw argv and environment values are never projected to operator clients. Distinctive effective environment values and values from sensitive environment names are also scrubbed from previews and artifacts, including a retained secret prefix at the process-runner truncation boundary. Public lastError values and immediate background-tool failures use one stable generic message per error code. Ambient spawn, artifact, cleanup, and store exception text—including absolute paths—never enters a durable public error, operator projection, wake prompt, or model tool result. Older v1 records with free-form error text remain readable; projection replaces that text and the next mutation rewrites it to the stable public form. process_job_cleanup_incomplete remains distinct from a safe process_job_agent_restarted interruption and from an artifact-only process_job_store_error; the terminal lifecycle state still records whether the process was cancelled, timed out, or otherwise failed. Environment-key inventories are bounded independently, and the exact serialized record size is checked before a recovery transaction marker is published. A legacy transaction that can never fit or validate is moved intact into the owner-only quarantine-v1/ directory; the store opens in degraded health and mono-agent validate reports the incident for operator review. Other unsafe or transient failures from every store read, mutation, artifact, wake, recovery, retention, and shutdown boundary remain fail-closed. The controller closes new admission, publishes degraded health to the live TUI and trace source, persists a bounded secret-free health marker for mono-agent validate, and serves its bounded last-known in-memory record view where that is safe. If a live process completes but its terminal record cannot be committed, the controller exposes a failed in-memory projection with wake delivery withheld and still releases the process’s active slot. It preserves the durable nonterminal record so the next owner restart can reconcile it to interrupted and deliver the recovery wake. A clean restart clears the marker only after recovery, retention, and durable readback succeed. Because process jobs are opt-in, an unavailable store disables only the background controller instead of aborting the whole agent. The registry and every retained stateDir remain protected on every model turn; mixed or non-Pi routes and unavailable native protection fail before provider invocation.

Slack and Telegram start one host-owned lifecycle message before the tool call returns, without waiting indefinitely on chat API latency. Updates are serialized per job, edit that same exact-origin thread/chat message when possible, and never let a late running update overwrite a terminal state. Each adapter retains the shared compiled maximum of 10,096 outstanding lifecycle identities, refuses unsafe overflow, and evicts only a terminal identity whose wake has settled. After a restart or an uneditable/missing message reference, the adapter may publish one bounded self-contained terminal fallback in that same origin only. Adapter identity state is deliberately instance-local rather than durable across restart. Empty host lifecycle updates never enter the ordinary responder/model path. The terminal wake itself still uses the channel’s normal proactive turn path. Web wakes through the operator driver without requiring a live browser or HTTP turn, commits one normal agent-history entry, and updates one durable job card in the exact originating thread. Ordinary web:<id> notifications are rejected; only a process-job lifecycle wake carrying that matching web origin routes to the TUI driver.

At startup, mono-agent reuses process-incarnation evidence. Only a stored leader whose PID still matches its incarnation and equals its PGID can authorize a signal. Recovery sends SIGTERM to that owned process group and waits one second. If the group remains, it re-attests the leader immediately before SIGKILL; if the leader vanished or its PID changed during the grace window, it does not signal the group again or clean settings while descendants may remain. Recovery polls for actual group absence after an accepted signal and only then removes the validated one-use sandbox settings directory. A queued or pre-attestation record never crossed the target-release fence, so recovery can remove that directory without signalling. Every path marks the job interrupted and schedules one recovered wake, with conservative wording when owned-group termination could not be proven. Process jobs never claim to survive an agent restart.

mono-agent restart --clear-sessions does not delete process-job records or artifacts. Stop the agent and remove the configured stateDir only when you explicitly intend to discard that audit/output state; there is no process-job purge command.

Use the local discovered-agent CLI:

Terminal window
mono-agent jobs list
mono-agent jobs get JOB_ID
mono-agent jobs cancel JOB_ID
mono-agent jobs list --agent AGENT_LABEL --json

The command refuses remote endpoints, derives an independent owner capability from the selected agent’s private store, and exits 1 with agent_unreachable when the agent cannot be reached. Misuse exits 2.

Successful background Exec/Bash and persistent Agent/AgentSend launches carry a bounded versioned start receipt in their machine-readable tool result. The receipt records the exact job id, tool, admission state, and real start stamp when one exists; the human result text is not an identity source. The web console uses that causal receipt only when both the launching response and its card are loaded. Its response Activity then shows real start and terminal evidence, including failed, timed-out, cancelled, spawn-failed, queue-expired, and interrupted outcomes. Queued or starting jobs without a real start stamp do not gain a start row, and missing timestamps are never synthesized. Older launches without the receipt remain in the separate job stack only.

An enabled local operator endpoint exposes bearer-protected GET /gui/v1/jobs, GET /gui/v1/jobs/:jobId, and POST /gui/v1/jobs/:jobId/cancel. Its info response advertises jobs: true only while the controller and its owner bearer are present. List responses keep every queued, starting, and running projection and add a deterministic newest-terminal prefix within the 16 MiB response ceiling. The web console collects running and terminal jobs from the loaded transcript window into one stack after the conversation (tool, purpose, state and elapsed time on each card; output tail, wake state and the wake’s response behind it, without the host-local artifact paths an operator cannot open from a browser). Queued, starting, and running work stays visible by default. Every terminal outcome remains mounted but hidden until the operator expands history; that choice is remembered per conversation for the browser session. The stack labels its active and history counts as loaded and points to Load earlier messages whenever older history is available. A running card polls its exact job once per second and opens when its first output arrives. Operators may collapse that card; later output and settlement preserve the choice, and the tail follows the bottom only until the operator scrolls upward. Queued/starting jobs and failed reads retain bounded backoff. Each nonterminal card polls only its exact authenticated, source- and thread-bound GET /api/v1/threads/:id/jobs/:jobId proxy with bounded backoff; it does not clone or serialize the retained job list on every refresh.

The stack remains the only live card and poller. Response Activity rows do not poll, show output/artifacts/wake details, offer cancellation, or duplicate the completion response.

mono-agent validate / doctor reports whether the feature is disabled or unsupported on Windows, then inspects only bounded local record counts and owner-only modes, including any quarantined transaction count and the bounded runtime health marker. For valid internal child records it also reports path-free retained-ownership, unresolved-ownership, and owner-unavailable counts. The owner-unavailable count is the conservative subset whose persisted owner is unknown (including legacy busy records without structured ownership); it is not a live process probe. Doctor does not expose job ids, registry roots, or command paths, does not mutate the live controller, and never creates a missing store.

When ProcessJobs is enabled and healthy on an exact-conversation Pi-native route, Agent({persist:true, background:true, prompt:"Review the change"}) returns a durable started receipt. AgentSend({id:"helper", message:"Continue", background:true}) continues the same child transcript. Both require the ordinary Agent/AgentSend policy. Close-only calls remain synchronous. Bare runtime hosts need a supplied background controller; unsupported calls fail clearly.

In the web console, detached launches keep their receipt in the parent’s Activity, which shows Agent job started / Agent job succeeded (or the actual terminal state); AgentSend uses the corresponding label. The child no longer streams foreground-style subagent rows into the parent response. Its Background jobs card uses the subagent glyph and a height-bounded scroll region with clustered tool calls, running/complete/failed status, durations, and a plain-text terminal report. State, Wake, and terminal facts remain on the card. Scrolling upward holds the reading position through later progress and report arrival.

Progress is separate from stdout and from the parent’s completion-wake output. The host retains at most 50 recent calls plus total/failed counts, redacts short argument summaries before persistence, and coalesces progress writes every 250 ms. No prompts or tool result bodies enter progress. Identity/name fields and argument summaries are byte-bounded; the report is a redacted head of at most 8,000 UTF-8 bytes, explicitly marked when truncated. The original terminal output JSON and wake behavior are unchanged.

Old records without progress remain readable and show that progress is unavailable. Upgrade host and console together: the optional internal-only subagentProgress field is strictly validated, and older binaries can reject populated records or projections. This adds no new state directory or web SQLite migration; ordinary process-job retention still owns the data.

Detached children can run long foreground Bash/Exec commands: timeout_ms is capped at the smaller of the owning job’s remaining runtime at child-run setup and subagents.commandTimeoutMs (positive integer milliseconds; default 1800000, or 30 minutes). Tool descriptions report this effective ceiling; the job’s abort signal still stops commands when its deadline arrives, including commands started later in the turn. Raising the command ceiling does not extend subagents.timeoutMs, profile timeouts, or processJobs.maxRuntimeMs. Interactive turns and foreground children retain the 120-second cap, and NodeRepl retains its fixed 120-second timer. Child-owned background commands remain unsupported and are explicitly out of scope.

The child runs inside the owning host, using the existing process-job admission, queue, runtime/output limits, lineage, lifecycle card and exact-origin wake. A queued child is reserved before the receipt returns, so another send or close reports busy. Its prompt and raw parameters are never stored in job metadata. Completion, failure and AskParent deliver one terminal wake; AskParent preserves awaiting_reply and its structured question for a later AgentSend. Do not poll or replay a started job. Message plus close closes only after a successful answer.

Use AgentSend({id, stop: true}) to cooperatively stop a queued or running managed detached child. An optional description string of at most 80 characters is accepted and ignored: stop creates no job to label. Stop is exclusive with message, close, background, inspect and ack (even explicitly false values). It invokes no new provider turn (executed:false) and accepts no job id. Foreground children are not stoppable through this operation. Stop never force-kills an in-process provider or rolls back filesystem/network effects.

Invalid requests return a JSON error receipt with a human-readable message, stopRequested:false and executed:false, before instance lookup:

  • subagent_stop_not_requested: stop must be exactly true.
  • subagent_stop_invalid_id: id must be a string of 1–40 lowercase letters, digits or hyphens, starting with a letter or digit.
  • subagent_stop_invalid_request: description must be a string of at most 80 characters.
  • subagent_stop_unexpected_parameters: only id, stop and description are accepted; the message lists unexpected keys in sorted order.

The operation waits at most six seconds, including storage work:

  • stopped, resumable:true: the matched job, provider, owned commands and registry publication have settled, and native session continuity is certified.
  • already_idle, resumable:true: no stop was needed; ordinary continuation is admissible.
  • stop_requested, childStillBusy:true, resumable:false: cancellation was accepted, but settlement remains unproven. Ownership and capacity remain held. Ordinary messages and close remain blocked; do not poll or replay the job.

Only after a resumable receipt, use AgentSend({id, message: "Continue"}) (optionally background:true) to resume the same warm instance and prior context, or AgentSend({id, close:true}) to retire it. A queued stop charges no turn; a begun stopped turn charges one. Completion winning the race keeps its actual disposition, and pending AskParent questions survive stopping.

Lost or unproven continuity returns subagent_stop_recovery_required, not a resumable receipt. Unsupported ownership/storage returns subagent_stop_unavailable; uncertain cancellation acceptance is reported as stopRequested:"unknown". A stale captured turn is refused. These errors never authorize bypassing recovery fences. Intentional, certified parent stops do not require a failure acknowledgement; unrelated timeout, cancellation and failure recovery rules below are unchanged.

Timeout/cancellation requests abort and wait through the Agent grace period. If execution remains unresolved, the terminal job reports childStillBusy:true. The instance stays busy and retains its turn lock and independent runtime protection lease until actual settlement. Process death alone does not prove that command descendants exited. A late settlement never sends another wake or changes the terminal outcome, and leaves a recovery fence when session continuity is unknown. Do not send a new message merely because the running lock disappeared. The retained card describes the terminal observation; Session context shows the current instance state. Service shutdown does not wait indefinitely for an abandoned child. Restart interrupts stored work and wakes its origin without replaying it; pending questions survive recovery.

Persistent registry failures retain a minimal typed reason and turn identity. Lost or unknown continuity requires explicit close/create after ownership is resolved, not an implicit retry of unretained prose. Foreground persistent turns retain the ordinary command timeout and do not gain durable command ownership. Unresolved ownership or registry publication pins terminal records and their artifacts independently of wake delivery, age, and admission limits. Legacy childStillBusy:true is conservative unknown evidence, not proof of cleanup. A disabled service or an unavailable retained-root index cannot authorize new persistent instances around forgotten work. Do not delete ownership records to bypass this fence. Older runtimes reject records containing the new ownership or registry-intent fields; stripping those fields is not a safe downgrade. Existing legacy command cleanup cannot be retroactively proven from an interrupted job.

For managed detached turns, Bash/Exec still returns its ordinary awaited tool result. The host persists preparation and PID/group incarnation before releasing the gated target, and borrows the existing child job slot rather than scheduling a second job. Overlapping commands and repeated host call identities reject without execution; no rejected owner falls back to an untracked command. Registry reservation identity is verified before admission. The provider’s actual promise settlement is observed before the reporting race; reporting timeout cannot erase an unresolved lease or command. Clean instance release waits for terminal job publication, not merely provider return. A lost registry-confirmation receipt keeps continuation/close/reuse fenced until the registered owner confirms the same publication sequence. Private owner roots and publication receipts are not included in instance handle results.

The private store retains at most 32 command receipts (12 KiB aggregate), with an omitted count under pressure. These contain actual tool/cwd/budget/exit/signal, timeout/cancel/truncation and cleanup measurements, never raw argv, environment, stdout, provider answer or a fabricated “checks passed” verdict. Optional facts are trimmed before they can exceed the job’s real serialized-record budget; mandatory ownership and the non-evicting call ledger are never trimmed for them. Crash recovery can add positive group-cleanup evidence while leaving the exit unobserved. It cannot turn OS cleanup or a matching command label into successful verification. These private receipts do not widen ProcessJob/wake projections.

Explicit child recovery inspection and acknowledgement

Section titled “Explicit child recovery inspection and acknowledgement”

AgentSend({id, inspect: true}) is separate from message/close/background/ack requests and invokes no provider. It can perform one bounded owner reconciliation pass, then returns held/unavailable or current-policy-authorized recovery facts. After independently verifying the work, a retained-only acknowledgement may be submitted with a message; recovery of a detached job requires background: true. The same token and request semantics return subagent_recovery_already_consumed without execution, even while the first continuation is busy. Changed semantics return conflict. Consumption and the new reservation are durable before admission. A requested close:true retires the child only after that acknowledged continuation succeeds. If it fails, any pending AskParent question and the child instance remain available for explicit recovery; repeating the consumed request does not execute it again. A proven rejected admission retains a not-started disposition; ambiguous absence never authorizes retry. Lost/unknown continuity cannot be acknowledged back into retained context: resolve ownership, then explicitly close/create instead.

The configured foreground persistent path currently classifies failures as lost (a native response outside the selected session) or unknown (including late timeout settlement). It does not establish a retained failure epoch eligible for acknowledgement. While the original runtime is unresolved, inspection is held and close/continuation remain blocked. After settlement, authorized inspection reports structured_job_recovery_unavailable with the minimal registry fence, not a ProcessJob, command checkpoint or acknowledgement token. Explicitly close the instance and create another with the necessary context. A late answer or existing JSONL file does not upgrade unknown continuity. Ordinary successful foreground continuation and AskParent replies still resume their retained session without a recovery acknowledgement.

For a managed detached failure explicitly classified as retained, an accepted acknowledgement resumes the same durable session; it does not replay the failed request or authorize a new provider epoch. Inspection and duplicate consumed acknowledgements invoke no provider and append no continuation message.

Persistent Agent’s optional verification: {workdir, reportPath?} declares only an observation target. It changes neither command cwd nor policy/approval. Current readable/protected roots are checked at admission, capture and disclosure; linked common Git metadata outside those roots is not implicitly authorized. Fixed Git probes require an available read-only sandbox, disable external helper paths and never fall back to host execution. The host selects root-installed native Git: /Library/Developer/CommandLineTools/usr/bin/git on macOS and /usr/bin/git on Linux. It checks canonical secure ancestry, ownership, executable format and identity before/after probes; it never consults model-supplied paths or PATH, executes macOS’s /usr/bin/git bootstrap shim, installs tooling or changes the active developer directory. Missing/redirected/untrusted tooling or an unsupported platform yields observation_unavailable. Existing sandbox runtime-read allowances for the executable do not grant repository/private-root read authority; explicit protection of the selected executable denies observation before preparation. The prepared native command must preserve the selected executable and exact declared observation cwd. Repository config includes are unsupported and fail closed before status, preventing an included file from changing helper policy between probes. Status ignores all initialized submodules, so nested repository changes are outside the observation and nested clean filters are not executed. Unsupported alternates, denied paths, replaced roots, changed Git metadata/executable identity/HEAD and unavailable sandboxing produce typed gaps. Report metadata records only a relative path and presence, never content or a hash. Command facts and observations remain bounded; foreground-only recovery is marked structured_job_recovery_unavailable, not presented as a durable command job.