Background process jobs
Background process jobs let the existing Pi-native Exec and Bash tools
return immediately while mono-agent continues to own the spawned POSIX process
group.
There is no separate job tool. When the host has an available process-job
controller, both schemas gain the optional background: true field. Without
that controller, background starts are unavailable. A configured host still
discloses request lineage diagnostics in the tool descriptions, including
chainDepth, maxChainDepth, remainingStarts, and
unavailableReason=chain_depth_exhausted when the budget is spent.
Use this for a command that should outlive the current model turn but still
report back to the exact Slack thread, Telegram chat, or web-console thread that
started it. It is independent from durable continuations:
a process job owns a local Exec or Bash child, while a continuation accepts a
later result from a selected external MCP service.
Configuration
Section titled “Configuration”The feature is opt-in and every key is JSON-only. Unknown keys are rejected.
stateDir must be a relative child of the agent root. Its canonical path must
be disjoint from every root removed by restart --clear-sessions: Pi provider
sessions, durable message/tool history, and ACP session authorizations. The
check covers every root retained in the durable registry, not only the current
configuration. Startup and clear-sessions preflight reject equality or
containment in either direction, including lexical and canonical aliases, so
clearing conversation state cannot delete process-job records or output.
{ "processJobs": { "enabled": true, "unsafeAllowUnprotectedState": false, "stateDir": ".mono-agent/process-jobs", "maxConcurrent": 4, "maxActivePerConversation": 2, "maxQueued": 8, "maxRuntimeMs": 1800000, "maxQueueAgeMs": 300000, "maxOutputBytes": 1048576, "previewChars": 2000, "maxChainDepth": 4, "retention": { "maxRecords": 1000, "maxAgeMs": 604800000, "artifactMaxBytes": 268435456 } }}| Setting | Default | Compiled maximum |
|---|---|---|
unsafeAllowUnprotectedState | false | — |
maxConcurrent | 4 | 32 |
maxActivePerConversation | 2 | 8 |
maxQueued | 8 | 64 |
maxRuntimeMs | 30 minutes | 24 hours |
maxQueueAgeMs | 5 minutes | 1 hour |
maxOutputBytes | 1 MiB | 8 MiB |
previewChars | 2,000 | 8,000 |
maxChainDepth | 4 | 64 |
retention.maxRecords | 1,000 | 10,000 |
retention.maxAgeMs | 7 days | 30 days |
retention.artifactMaxBytes | 256 MiB | 1 GiB |
The compiled maximum bounds configuration, configuration bounds each job, and
a tool call may narrow only timeout_ms and max_output_chars. Queue age starts
at admission. The runtime deadline starts when the detached launch gate is
spawned immediately before ownership is persisted, so waiting in a busy queue
does not consume the runtime budget.
Durable root registry and request barrier
Section titled “Durable root registry and request barrier”Before mono-agent creates an enabled stateDir, opens its store, or writes a
store secret, @mono-agent/agent-app securely publishes that root to
.mono-agent/process-jobs-roots-v1/registry.json. Opening the service requires
the exact registration proof. The v1 manifest is absent exactly while no root
has ever been registered for the agent. It stores sorted agent-root-relative
segments and a fresh generation id, with these fixed bounds:
| Registry bound | Maximum |
|---|---|
| Retained roots | 64 |
| Segments per root | 64 |
| UTF-8 bytes per segment | 255 |
| UTF-8 bytes per relative root | 2 KiB |
| Encoded manifest | 256 KiB |
The manifest is an owner-only, no-follow, regular single-link file. Updates use
an owner-only mutation lock, atomic secure replacement, directory fsync, and an
identity-and-content reread. Unsafe, malformed, or over-bound state fails closed
with the path-free Process-job private-state protection is unavailable. error.
The registry directory remains strict: its only permitted entry is
registry.json. Replacement artifacts instead use the owner-only sibling
.mono-agent/process-jobs-roots-v1.recovery/, which is mode 0700, must share
the registry filesystem, and permits at most these three mode-0600 regular
files:
registry.staging.jsonregistry.previous.jsonregistry.failed.json
Cross-directory publication fsyncs each affected directory at the namespace
mutation, destination before source. Ordinary request loading performs only a
bounded inspection of the recovery directory and fails closed if any artifact
is present; it never repairs or removes one. Only root registration and
clear-sessions preflight may recover while holding the existing registry
mutation lock. Recovery artifacts are single-link in steady processing except
for one exact crash-transient: after a previous manifest is linked back to
registry.json and before registry.previous.json is unlinked, those two fixed
names may be the same proven inode with identical bytes and nlink=2. Requests
still fail closed and leave that pair untouched. Locked recovery alone may
reprove the exact pair, make the target link durable, unlink the previous name,
fsync the recovery directory, and reprove registry.json at nlink=1; every
other hard-linked shape remains untouched and fail-closed.
A valid current manifest wins and proven artifacts are removed; an absent current manifest is restored from a valid previous manifest (recreating only a proven-absent registry directory when necessary) before proven staging/failed cleanup; staging alone is discarded so a fresh first registration can proceed. Failed-only state, a corrupt or ambiguous current manifest, unknown or fourth entries, and unsafe directories, links, ownership, modes, or artifact contents remain fail-closed. The recovery directory is empty in steady state.
It is never auto-pruned: disabling or removing processJobs, changing A to B,
or restarting retains A, and an A-to-B change protects both roots. The registry
directory, recovery directory, and every retained root’s lexical and canonical
aliases are native protected roots and reply-artifact private roots. A degraded
or failed store open therefore leaves all earlier roots sealed even though the
background controller itself remains optional.
Every official local request captures and re-attests the current registry
generation, then acquires its generation lease before request resource
extensions or provider invocation. A newly registered generation becomes the
current generation before its store may open; opening waits for older leases
that did not cover the new root. The bounded drain timeout fails closed and does
not create the directory, store, or secret. A request releases its lease only
from settleCleanup, after runtime.run truly settles. Earlier cleanup, abort,
or harness disposal cannot release it beneath a late provider result.
Once the registry is non-empty, every reachable primary, fallback, accepted
request override, and named Agent child route must be Pi-native — which every
route now is. The configured app rejects the whole
incompatible route plan before provider invocation; it does not wait to discover
the unsafe fallback after another route fails. An empty registry preserves
legitimate non-Pi routes.
Tool-less direct configured memory LLM and embedding-provider calls are the one
exception, in every posture. SRT confines the model’s tool loop; these surfaces
run no tool loop and touch no filesystem — an embedding provider is a single
embed(texts) HTTP call with nothing the model can steer — so confinement has
nothing to protect there. They still take the canonical owner and the
registry-generation lease and hold both until the provider promise truly
settles, which is the control that does apply. Rejecting them instead made
processJobs.enabled mutually exclusive with the bujo and journal memory
tiers, which require an embedding provider.
Eligible Pi-native turns receive the real SRT policy for the registry and every
retained root independently of whether they receive the background controller.
Model Read, Write, Edit, Glob, Grep, Bash, and Exec cannot read,
replace, rename, search, or use those paths as a workdir. Host filesystem tools
perform their actual file operation through the native sandbox, closing symlink
swaps after path authorization. SRT also denies a rename of any ancestor that
would move a protected leaf.
Only the registry and retained state directories are protected: workspace siblings such as
.mono-agent/artifacts/attachments remain readable, including when the
workspace itself is nested under .mono-agent/. When the configured sandbox
is absent or off, this filesystem-only protection preserves unrestricted
network behavior for commands, WebFetch, and WebSearch. A configured native
network policy remains unchanged. Provider-owned non-Pi tool loops cannot
enforce this host policy, so they are rejected while any registered private
root exists.
This agent-root-aware coverage belongs to the full app, configured
harness/responder, local TUI, the lazy-run wrapper returned by
createConfiguredAgentRuntime, configured named children, and direct configured
memory LLM/embedding calls. Remote TUI and ACP bridges are thin clients of an
already owned host and do not invoke a provider themselves. Lower-level
@mono-agent/runtime-adapter and @mono-agent/agent-runtime factories are
root-agnostic unless an app-owned caller supplies the protection policy and
route gate.
Unsafe trusted-host posture
Section titled “Unsafe trusted-host posture”processJobs.unsafeAllowUnprotectedState: true is a JSON-only escape hatch for
an operator who intentionally runs trusted same-user host tools. It is accepted
only when all of these conditions hold:
sandbox.modeis present and exactly"off";- ProcessJobs is enabled, or the attested durable registry still retains at least one root; and
- every configured primary, fallback, named
Agentchild, and agent-host memory route is Pi-native. Accepted request overrides are checked again before any provider is invoked.
The escape hatch changes only SRT policy injection. Registry load and
attestation, including the second request-boundary attestation, owner and
generation leases through true settlement, root disjointness, retention,
store/service/controller lifecycle, and reply-artifact privateRoots remain in
force. A failed or unavailable registry always wins and remains provider-zero.
The Pi-only gates are independent from sandbox.protectedRoots, so mixed
primary/fallback chains, non-Pi request overrides, non-Pi named children, and
non-Pi agent-host memory remain rejected before provider work.
Tool-less direct Ollama memory LLM and embedding calls run in this posture as they do in the safe one, still holding the canonical owner and registry-generation lease until the provider promise truly settles. Agent-host memory remains tool-less and Pi-native. Public package-root runtime, harness, responder, and memory factories do not accept this authority and keep their safe behavior.
Changing this posture takes effect only when the app rebuilds its owned
runtime surfaces: a managed configuration apply performs that teardown/rebuild,
and a process restart does the same. Existing in-flight runs are not mutated.
No state migration or reset occurs, and retained registry roots remain
registered when the flag is enabled, disabled, or removed. validate,
foreground/background status, trace metadata, and the local TUI summary show
the path-free warning UNSAFE: ProcessJobs state and operator secret are model-accessible.
Cooperative ownership and threat boundary
Section titled “Cooperative ownership and threat boundary”Official local hosts also take one cooperative lifetime lease keyed by the canonical realpath of the agent root and stored under the effective account home. In one process, repeated configured app/harness/responder/local-TUI owners share a reentrant reference count. Physical release waits for all owner references and all true-settlement request leases; a release failure makes every later in-process acquisition fail deterministically. A stale official-process lease is recoverable after a crash. The lease path and random owner token are host-only coordination data, but their permissions and hash-derived pathname are not a secrecy or tamper-resistance claim.
This is cooperative serialization, not a security boundary against actively
hostile code running as the same OS user. Such code can signal or SIGKILL the
host, rewrite same-UID control state, and attack another registered agent root
under the same account. Alternate same-UID coordination mechanisms do not
change that fact. The local persistence guarantee covers exactly the roots
durably registered for this agent; it is not an account-global provider-zero
rule. If a provider is in that threat model, run it under a distinct UID or
another real privilege-separation boundary.
Availability and origins
Section titled “Availability and origins”The host injects the controller only when all of these are true at call time:
processJobs.enabledistrueand the owner-private store lock is ready;- the platform is POSIX; Windows is unsupported;
- the selected request route is Pi-native;
- the ordinary tool policy permits
ExecorBash; and - the turn originated in an exact addressable conversation whose channel driver opts into the ProcessJobs capability. The built-ins are Slack, Telegram, the web console, and the WhatsApp plugin; future plugins may claim one unique conversation-id scheme and publish the same running capability.
Direct TUI turns, cron, webhook, OpenAI API, A2A, and plugins without the
explicit capability do not get background schemas. Duplicate or malformed
scheme claims fail during app startup. A controller also rejects an invalid origin as
background_unsupported_channel. Foreground calls continue to use their normal
route and policy.
Starting a job
Section titled “Starting a job”Use the existing tool exactly as before and add background: true:
{ "executable": "pnpm", "args": ["test"], "workdir": ".", "description": "Running the full repository test suite", "timeout_ms": 900000, "max_output_chars": 4000, "background": true}Bash preserves its clean non-interactive shell semantics; Exec preserves
literal argv semantics. Both reuse the exact sandbox-prepared command that a
foreground call would launch. The immediate result contains only an opaque
job_id, state, and started_at; it does not expose argv, environment,
process ids, paths, or secrets.
description is a short statement of the work’s purpose. For background jobs,
the host bounds and redacts it before showing the same safe summary throughout
the lifecycle surface and durable terminal record. The raw description is not
persisted; argv, command content, environment values, and host-specific
workspace/home paths remain hidden.
The agent is told when to reach for this and what not to do afterwards in three
places, all gated on the same availability check as the schema itself: the
background field description states that foreground is the default, that a
background job costs an extra turn and defers the answer, that only work
expected to exceed the foreground ceiling or to keep running after the reply
belongs there, and that a restart of the agent interrupts every job (so
backgrounding is not a way to run something “while replying”); the start result leads
with a line saying the conversation is woken on completion, so the agent must
not poll, sleep, or re-run the command to check; and the session block of the
system prompt repeats both alongside the daemonize prohibition and the fact that
job output arrives as untrusted evidence. The same availability check also
relaxes the prompt’s continuation rule, which otherwise has the agent announce
that background delivery was not scheduled for a job it just started. No
operator command is named in any model-facing copy — the agent has a shell, and
naming a status command invites the polling this is meant to prevent.
A chain starts at depth zero; a completion wake inherits its parent’s depth
plus one. With maxChainDepth: 32, depth 31 permits one further background
stage and its depth-32 wake permits none. Steering, queued follow-ups and
retries retain host-owned lineage. The default remains 4; the ceiling is 64,
and concurrency, queue, output and runtime limits remain independent.
wake_on_completion is an optional boolean on background Exec/Bash calls.
It defaults to true. Set it to false explicitly for a helper that should update
its terminal lifecycle card without scheduling a completion turn. Using it
without background: true is invalid. This preference survives restart;
older records retain the default wake behavior. Cancellation never restarts a
command.
A genuine completion wake can answer with exactly NOTHING_TO_REPORT to
suppress delivery. Narration and rich reply parts remain visible. The host
matches the exact active delivery key; unrelated and stale keys cannot silence
another turn. Web removes the sentinel from the settled reply and emits no
response push for a reply without visible content.
A timeout before the destination confirms steering or durably admits the exact
follow-up may leave delivery uncertain. The terminal wake state is unknown,
with an explicit outcome-unknown error; automatic replay is suppressed,
including after restart. The web destination receipts a durably admitted
follow-up without waiting for its model turn to finish; that turn’s later
success, failure, cancellation, or interruption remains independently visible.
A definite refusal remains failed or follows its existing bounded safe-retry
policy.
The process owns its sandbox settings until every process remaining in its
owned POSIX process group exits.
On POSIX a command-agnostic detached group leader starts first. Mono-agent
persists its PID, equal PGID, and process-incarnation evidence before releasing
the exact command and environment over an anonymous pipe. A crash before that
commit cannot spawn the target; a crash after release leaves a recoverable
owned group. Timeout, cancellation, and matched restart recovery therefore wait
for inherited-group descendants before sandbox cleanup. Commands that
deliberately daemonize into another POSIX process group or session are not
contained by this contract and must not use background: true.
Live timeout, cancellation, and shutdown use a bounded SIGTERM then SIGKILL
sequence. While the exact self-led ChildProcess leader is live and unreaped,
its negative PGID remains authoritative even if the owner’s event loop stalls;
the kernel cannot recycle that live identity. The host begins bounded-frequency
group observation when the leader reports exit and signals the recorded
negative PGID only while that post-exit proof remains continuous. A post-exit
over-limit or indeterminate observation gap permanently revokes signalling
authority. If termination or final group absence cannot be proved, the job
settles with an explicit degraded error and leaves sandbox settings intact for
operator investigation instead of hanging or cleaning beneath a possibly live
descendant.
Lifecycle, output, and wake delivery
Section titled “Lifecycle, output, and wake delivery”Jobs move through this durable state machine:
queued -> starting -> running -> succeeded | failed | timed_out | cancelled | | | +----------------> spawn_failed | cancelled +---------------------------> queue_expired | cancelledany nonterminal at restart -> interruptedWake delivery is orthogonal: pending, delivered, failed, unknown, or
suppressed, with a stable delivery key and attempt count. A terminal transition
is lock-idempotent and schedules one wake unless explicitly opted out.
An adapter result that explicitly proves retry is safe gets
at most three attempts with the same delivery key, including across restart.
Ambiguous wake attempts are not replayed automatically, because a second post
could duplicate a real first delivery.
For an active turn in the exact originating conversation, the adapter first offers the completion as live input targeted to that run id. It reserves the normal follow-up position before making the offer. Confirmed exact transcript consumption keeps the completion in that turn; only an explicit unavailable or requeued settlement runs the reserved normal follow-up turn. Discarded, uncertain, rejected, or unknown settlement is ambiguous and never triggers an automatic duplicate. Consumption does not prove provider receipt or adherence. Slack, Telegram, and WhatsApp use their ordinary visible thinking/tool/final stream for the fallback. The web console creates an assistant-only turn, emits the same NDJSON activity/tool frames, and never invents a user message.
Every steer or fallback carries the stable delivery key out of band. The web
console durably records accepted and completed delivery claims. A steered
completion means the exact active run accepted the live input; a follow_up
completion means the agent accepted that exact assistant turn request, not that
its model work succeeded. Acceptance is also when the wake’s host-owned
capability binds to the request, so the follow-up keeps the parent’s chain depth
and its remaining background starts; a receipt that returned earlier would end
the wake and leave the turn it raised unable to start the next job. After
restart a completed claim returns its prior receipt
while the associated turn can independently recover as interrupted; an accepted
but unsettled claim fails closed as ambiguous. A wake is a genuine tool-capable
turn, not continuation synthesis. The host raises the active controller to the
parent job’s chain depth plus one before any steered tool call can start; a
non-consumed offer rolls that provisional depth back, and the configured maximum
remains authoritative.
A Slack or Telegram conversation that is already at its pre-turn admission cap
does not spend that three-attempt budget. The wake stays durably pending and is
re-armed on a separate, longer timer; a restart can deliver it later. Busy
refusals do not themselves exhaust the wake: the durable bound is five minutes
from its first refusal, so an ordinary busy turn can clear and receive the same
delivery identity. Age exhaustion settles the wake as failed and immediately
runs retention. Once a turn is admitted, any ambiguous failure is nonretryable
and exactly-once wins over automatic replay.
An absent or disabled destination channel is also a proven pre-dispatch refusal,
but it has a separate durable bound of three checks. Those checks do not change
the delivery attempt count or timestamp and reuse the same stable delivery key.
If the channel returns before exhaustion, delivery continues normally; otherwise
the wake settles as failed so retention can reclaim its record and artifacts.
Conversation-cap busy admission remains distinct and does not spend this
absent-channel bound.
Admission counts every pending wake obligation, including a queued or running
job whose terminal wake is not due yet. It rejects process_job_capacity at
retention.maxRecords + maxConcurrent + maxQueued obligations (1,012 by
default; compiled maximum 10,096). This bounds the number of separate busy-wake
rearm timers without silently evicting an obligation. Recovery keeps the oldest
obligations when repairing legacy overflow and records explicit failed-wake
outcomes for the excess before ordinary retention can reclaim them.
The store also has a compiled open ceiling of 20,096 record entries: the 10,000 retained-record maximum plus the 10,096 pending-obligation maximum. Startup streams and bounds the record directory before materializing records. A legacy or externally modified directory above that ceiling fails closed with a stable, path-free storage error; it does not delete pending wakes, artifacts, or active ownership records. The same check prevents transaction replay from growing the store past the ceiling. An operator must inspect and remediate that owner-private state before restart can proceed.
Pending-wake records and their referenced artifacts remain live and are exempt from age, count, and artifact-byte pruning until delivery settles. Other terminal records and artifacts are pruned oldest-first with job-id tie-breaking. Retention runs at startup, after every terminal completion, and after each wake settles. Every 64 retention applications also reconcile orphan artifact directories, so long-running agents reach crash-cleanup work without restarting.
Stdout and stderr are stored separately under the configured output budget.
While a process is running, its operator projection exposes a memory-only tail
of the newest 100 logical stdout/stderr lines. It is refreshed at most every 250
milliseconds, carries stream headings and an omission marker, and remains under
the job’s previewChars bound. Chunk callbacks never mutate the durable store,
write artifacts, or schedule channel lifecycle edits. The final redacted tail
is persisted when the job settles; full bounded stdout/stderr remain available
through their separate artifact references.
The model, CLI, operator API, and web card receive only bounded redacted
previews plus agent-root-relative artifact references. Redaction occurs before
line and character cuts, including for known secrets split across UTF-8 chunks
or physical lines, credential-shaped values, and PEM blocks. Treat every
preview as untrusted process output. Records retain a redacted command summary
and only environment key names; raw argv and environment values are never
projected to operator clients. Distinctive effective environment values and values from
sensitive environment names are also scrubbed from previews and artifacts,
including a retained secret prefix at the process-runner truncation boundary.
Public lastError values and immediate background-tool failures use one stable
generic message per error code. Ambient spawn, artifact, cleanup, and store
exception text—including absolute paths—never enters a durable public error,
operator projection, wake prompt, or model tool result. Older v1 records with
free-form error text remain readable; projection replaces that text and the
next mutation rewrites it to the stable public form.
process_job_cleanup_incomplete remains distinct from a safe
process_job_agent_restarted interruption and from an artifact-only
process_job_store_error; the terminal lifecycle state still records whether
the process was cancelled, timed out, or otherwise failed.
Environment-key inventories are bounded independently, and the exact serialized
record size is checked before a recovery transaction marker is published. A
legacy transaction that can never fit or validate is moved intact into the
owner-only quarantine-v1/ directory; the store opens in degraded health and
mono-agent validate reports the incident for operator review. Other unsafe or
transient failures from every store read, mutation, artifact, wake, recovery,
retention, and shutdown boundary remain fail-closed. The controller closes new
admission, publishes degraded health to the live TUI and trace source, persists
a bounded secret-free health marker for mono-agent validate, and serves its
bounded last-known in-memory record view where that is safe. If a live process
completes but its terminal record cannot be committed, the controller exposes a
failed in-memory projection with wake delivery withheld and still releases the
process’s active slot. It preserves the durable nonterminal record so the next
owner restart can reconcile it to interrupted and deliver the recovery wake.
A clean restart clears the marker only after recovery, retention, and durable
readback succeed.
Because process jobs are opt-in, an unavailable store disables only the
background controller instead of aborting the whole agent. The registry and
every retained stateDir remain protected on every model turn; mixed or non-Pi
routes and unavailable native protection fail before provider invocation.
Slack and Telegram start one host-owned lifecycle message before the tool call
returns, without waiting indefinitely on chat API latency. Updates are
serialized per job, edit that same exact-origin thread/chat message when
possible, and never let a late running update overwrite a terminal state. Each
adapter retains the shared compiled maximum of 10,096 outstanding lifecycle
identities, refuses unsafe overflow, and evicts only a terminal identity whose
wake has settled. After a restart or an uneditable/missing message reference,
the adapter may publish one bounded self-contained terminal fallback in that
same origin only. Adapter identity state is deliberately instance-local rather
than durable across restart. Empty host lifecycle updates never enter the
ordinary responder/model path. The terminal wake itself still uses the
channel’s normal proactive turn path. Web
wakes through the operator driver without requiring a live browser or HTTP
turn, commits one normal agent-history entry, and updates one durable job card
in the exact originating thread. Ordinary web:<id> notifications are rejected;
only a process-job lifecycle wake carrying that matching web origin routes to
the TUI driver.
Restart and cancellation
Section titled “Restart and cancellation”At startup, mono-agent reuses process-incarnation evidence. Only a stored leader
whose PID still matches its incarnation and equals its PGID can authorize a
signal. Recovery sends SIGTERM to that owned process group and waits one
second. If the group remains, it re-attests the leader immediately before
SIGKILL; if the leader vanished or its PID changed during the grace window,
it does not signal the group again or clean settings while descendants may
remain. Recovery polls for actual group absence after an accepted signal and
only then removes the validated one-use sandbox settings directory. A queued or
pre-attestation record never crossed the target-release fence, so recovery can
remove that directory without signalling. Every path marks the job
interrupted and schedules one recovered wake, with conservative wording when
owned-group termination could not be proven. Process jobs never claim to
survive an agent restart.
mono-agent restart --clear-sessions does not delete process-job records or
artifacts. Stop the agent and remove the configured stateDir only when you
explicitly intend to discard that audit/output state; there is no process-job
purge command.
Operator surfaces
Section titled “Operator surfaces”Use the local discovered-agent CLI:
mono-agent jobs listmono-agent jobs get JOB_IDmono-agent jobs cancel JOB_IDmono-agent jobs list --agent AGENT_LABEL --jsonThe command refuses remote endpoints, derives an independent owner capability
from the selected agent’s private store, and exits 1 with
agent_unreachable when the agent cannot be reached. Misuse exits 2.
Successful background Exec/Bash and persistent Agent/AgentSend launches carry a bounded versioned start
receipt in their machine-readable tool result. The receipt records the exact job
id, tool, admission state, and real start stamp when one exists; the human result
text is not an identity source. The web console uses that causal receipt only
when both the launching response and its card are loaded. Its response Activity
then shows real start and terminal evidence, including failed, timed-out,
cancelled, spawn-failed, queue-expired, and interrupted outcomes. Queued or
starting jobs without a real start stamp do not gain a start row, and missing
timestamps are never synthesized. Older launches without the receipt remain in
the separate job stack only.
An enabled local operator endpoint exposes bearer-protected
GET /gui/v1/jobs, GET /gui/v1/jobs/:jobId, and
POST /gui/v1/jobs/:jobId/cancel. Its info response advertises jobs: true
only while the controller and its owner bearer are present. List responses keep
every queued, starting, and running projection and add a deterministic
newest-terminal prefix within the 16 MiB response ceiling. The web console
collects running and terminal jobs from the loaded transcript window into one
stack after the conversation (tool, purpose, state and elapsed time on each
card; output tail, wake state and the wake’s response behind it, without the
host-local artifact paths an operator cannot open from a browser). Queued,
starting, and running work stays visible by default. Every terminal
outcome remains mounted but hidden until the operator expands history; that
choice is remembered per conversation for the browser session. The stack labels
its active and history counts as loaded and points to Load earlier messages
whenever older history is available. A running card polls its exact job once per
second and opens when its first output arrives. Operators may collapse that
card; later output and settlement preserve the choice, and the tail follows the
bottom only until the operator scrolls upward. Queued/starting jobs and failed
reads retain bounded backoff. Each nonterminal card polls only its
exact authenticated, source- and thread-bound
GET /api/v1/threads/:id/jobs/:jobId proxy with bounded backoff; it does not
clone or serialize the retained job list on every refresh.
The stack remains the only live card and poller. Response Activity rows do not poll, show output/artifacts/wake details, offer cancellation, or duplicate the completion response.
mono-agent validate / doctor reports whether the feature is disabled or
unsupported on Windows, then inspects only bounded local record counts and
owner-only modes, including any quarantined transaction count and the bounded
runtime health marker. For valid internal child records it also reports path-free
retained-ownership, unresolved-ownership, and owner-unavailable counts. The
owner-unavailable count is the conservative subset whose persisted owner is
unknown (including legacy busy records without structured ownership); it is not
a live process probe. Doctor does not expose job ids, registry roots, or command
paths, does not mutate the live controller, and never creates a missing store.
Detached persistent children
Section titled “Detached persistent children”When ProcessJobs is enabled and healthy on an exact-conversation Pi-native route,
Agent({persist:true, background:true, prompt:"Review the change"}) returns a
durable started receipt. AgentSend({id:"helper", message:"Continue", background:true})
continues the same child transcript. Both require the ordinary Agent/AgentSend
policy. Close-only calls remain synchronous. Bare runtime hosts need a supplied
background controller; unsupported calls fail clearly.
In the web console, detached launches keep their receipt in the parent’s Activity,
which shows Agent job started / Agent job succeeded (or the actual terminal
state); AgentSend uses the corresponding label. The child no longer streams
foreground-style subagent rows into the parent response. Its Background jobs
card uses the subagent glyph and a height-bounded scroll region with clustered
tool calls, running/complete/failed status, durations, and a plain-text terminal
report. State, Wake, and terminal facts remain on the card. Scrolling upward
holds the reading position through later progress and report arrival.
Progress is separate from stdout and from the parent’s completion-wake output. The host retains at most 50 recent calls plus total/failed counts, redacts short argument summaries before persistence, and coalesces progress writes every 250 ms. No prompts or tool result bodies enter progress. Identity/name fields and argument summaries are byte-bounded; the report is a redacted head of at most 8,000 UTF-8 bytes, explicitly marked when truncated. The original terminal output JSON and wake behavior are unchanged.
Old records without progress remain readable and show that progress is
unavailable. Upgrade host and console together: the optional internal-only
subagentProgress field is strictly validated, and older binaries can reject
populated records or projections. This adds no new state directory or web
SQLite migration; ordinary process-job retention still owns the data.
Detached children can run long foreground Bash/Exec commands: timeout_ms
is capped at the smaller of the owning job’s remaining runtime at child-run
setup and subagents.commandTimeoutMs (positive integer milliseconds; default
1800000, or 30 minutes). Tool descriptions report this effective ceiling;
the job’s abort signal still stops commands when its deadline arrives, including
commands started later in the turn. Raising the command ceiling does not extend
subagents.timeoutMs, profile timeouts, or processJobs.maxRuntimeMs.
Interactive turns and foreground children retain the 120-second cap, and
NodeRepl retains its fixed 120-second timer. Child-owned background commands
remain unsupported and are explicitly out of scope.
The child runs inside the owning host, using the existing process-job admission,
queue, runtime/output limits, lineage, lifecycle card and exact-origin wake.
A queued child is reserved before the receipt returns, so another send or close
reports busy. Its prompt and raw parameters are never stored in job metadata.
Completion, failure and AskParent deliver one terminal wake; AskParent preserves
awaiting_reply and its structured question for a later AgentSend. Do not poll or
replay a started job. Message plus close closes only after a successful answer.
Parent stop and resume
Section titled “Parent stop and resume”Use AgentSend({id, stop: true}) to cooperatively stop a queued or running
managed detached child. An optional description string of at most 80 characters
is accepted and ignored: stop creates no job to label. Stop is exclusive with
message, close, background, inspect and ack (even explicitly false values).
It invokes no new provider turn (executed:false) and accepts no job id. Foreground children are
not stoppable through this operation. Stop never force-kills an in-process
provider or rolls back filesystem/network effects.
Invalid requests return a JSON error receipt with a human-readable message,
stopRequested:false and executed:false, before instance lookup:
subagent_stop_not_requested:stopmust be exactlytrue.subagent_stop_invalid_id:idmust be a string of 1–40 lowercase letters, digits or hyphens, starting with a letter or digit.subagent_stop_invalid_request:descriptionmust be a string of at most 80 characters.subagent_stop_unexpected_parameters: onlyid,stopanddescriptionare accepted; the message lists unexpected keys in sorted order.
The operation waits at most six seconds, including storage work:
stopped,resumable:true: the matched job, provider, owned commands and registry publication have settled, and native session continuity is certified.already_idle,resumable:true: no stop was needed; ordinary continuation is admissible.stop_requested,childStillBusy:true,resumable:false: cancellation was accepted, but settlement remains unproven. Ownership and capacity remain held. Ordinary messages and close remain blocked; do not poll or replay the job.
Only after a resumable receipt, use AgentSend({id, message: "Continue"})
(optionally background:true) to resume the same warm instance and prior
context, or AgentSend({id, close:true}) to retire it. A queued stop charges no
turn; a begun stopped turn charges one. Completion winning the race keeps its
actual disposition, and pending AskParent questions survive stopping.
Lost or unproven continuity returns subagent_stop_recovery_required, not a
resumable receipt. Unsupported ownership/storage returns
subagent_stop_unavailable; uncertain cancellation acceptance is reported as
stopRequested:"unknown". A stale captured turn is refused. These errors never
authorize bypassing recovery fences. Intentional, certified parent stops do not
require a failure acknowledgement; unrelated timeout, cancellation and failure
recovery rules below are unchanged.
Timeout/cancellation requests abort and wait through the Agent grace period.
If execution remains unresolved, the terminal job reports childStillBusy:true.
The instance stays busy and retains its turn lock and independent runtime
protection lease until actual settlement. Process death alone does not prove
that command descendants exited. A late settlement never sends another wake or
changes the terminal outcome, and leaves a recovery fence when session continuity
is unknown. Do not send a new message merely because the running lock disappeared. The retained card describes the terminal observation; Session context shows
the current instance state. Service shutdown does not wait indefinitely for an
abandoned child. Restart interrupts stored work and wakes its origin without
replaying it; pending questions survive recovery.
Persistent registry failures retain a minimal typed reason and turn identity.
Lost or unknown continuity requires explicit close/create after ownership is
resolved, not an implicit retry of unretained prose. Foreground persistent turns
retain the ordinary command timeout and do not gain durable command ownership.
Unresolved ownership or registry publication pins terminal records and their
artifacts independently of wake delivery, age, and admission limits. Legacy
childStillBusy:true is conservative unknown evidence, not proof of cleanup.
A disabled service or an unavailable retained-root index cannot authorize new
persistent instances around forgotten work. Do not delete ownership records to
bypass this fence. Older runtimes reject records containing the new ownership or
registry-intent fields; stripping those fields is not a safe downgrade. Existing
legacy command cleanup cannot be retroactively proven from an interrupted job.
For managed detached turns, Bash/Exec still returns its ordinary awaited tool result. The host persists preparation and PID/group incarnation before releasing the gated target, and borrows the existing child job slot rather than scheduling a second job. Overlapping commands and repeated host call identities reject without execution; no rejected owner falls back to an untracked command. Registry reservation identity is verified before admission. The provider’s actual promise settlement is observed before the reporting race; reporting timeout cannot erase an unresolved lease or command. Clean instance release waits for terminal job publication, not merely provider return. A lost registry-confirmation receipt keeps continuation/close/reuse fenced until the registered owner confirms the same publication sequence. Private owner roots and publication receipts are not included in instance handle results.
The private store retains at most 32 command receipts (12 KiB aggregate), with an omitted count under pressure. These contain actual tool/cwd/budget/exit/signal, timeout/cancel/truncation and cleanup measurements, never raw argv, environment, stdout, provider answer or a fabricated “checks passed” verdict. Optional facts are trimmed before they can exceed the job’s real serialized-record budget; mandatory ownership and the non-evicting call ledger are never trimmed for them. Crash recovery can add positive group-cleanup evidence while leaving the exit unobserved. It cannot turn OS cleanup or a matching command label into successful verification. These private receipts do not widen ProcessJob/wake projections.
Explicit child recovery inspection and acknowledgement
Section titled “Explicit child recovery inspection and acknowledgement”AgentSend({id, inspect: true}) is separate from message/close/background/ack
requests and invokes no provider. It can perform one bounded owner reconciliation
pass, then returns held/unavailable or current-policy-authorized recovery facts.
After independently verifying the work, a retained-only acknowledgement may be
submitted with a message; recovery of a detached job requires background: true.
The same token and request semantics return subagent_recovery_already_consumed
without execution, even while the first continuation is busy. Changed semantics
return conflict. Consumption and the new reservation are durable before admission.
A requested close:true retires the child only after that acknowledged
continuation succeeds. If it fails, any pending AskParent question and the child
instance remain available for explicit recovery; repeating the consumed request
does not execute it again.
A proven rejected admission retains a not-started disposition; ambiguous absence
never authorizes retry. Lost/unknown continuity cannot be acknowledged back into
retained context: resolve ownership, then explicitly close/create instead.
The configured foreground persistent path currently classifies failures as
lost (a native response outside the selected session) or unknown (including
late timeout settlement). It does not establish a retained failure epoch eligible
for acknowledgement. While the original runtime is unresolved, inspection is
held and close/continuation remain blocked. After settlement, authorized
inspection reports structured_job_recovery_unavailable with the minimal registry
fence, not a ProcessJob, command checkpoint or acknowledgement token. Explicitly
close the instance and create another with the necessary context. A late answer
or existing JSONL file does not upgrade unknown continuity. Ordinary successful
foreground continuation and AskParent replies still resume their retained session
without a recovery acknowledgement.
For a managed detached failure explicitly classified as retained, an accepted acknowledgement resumes the same durable session; it does not replay the failed request or authorize a new provider epoch. Inspection and duplicate consumed acknowledgements invoke no provider and append no continuation message.
Persistent Agent’s optional verification: {workdir, reportPath?} declares only
an observation target. It changes neither command cwd nor policy/approval.
Current readable/protected roots are checked at admission, capture and disclosure;
linked common Git metadata outside those roots is not implicitly authorized.
Fixed Git probes require an available read-only sandbox, disable external helper
paths and never fall back to host execution. The host selects root-installed native
Git: /Library/Developer/CommandLineTools/usr/bin/git on macOS and /usr/bin/git
on Linux. It checks canonical secure ancestry, ownership, executable format and
identity before/after probes; it never consults model-supplied paths or PATH,
executes macOS’s /usr/bin/git bootstrap shim, installs tooling or changes the
active developer directory. Missing/redirected/untrusted tooling or an unsupported
platform yields observation_unavailable. Existing sandbox runtime-read allowances
for the executable do not grant repository/private-root read authority; explicit
protection of the selected executable denies observation before preparation.
The prepared native command must preserve the selected executable and exact
declared observation cwd. Repository config includes are unsupported and fail
closed before status, preventing an included file from changing helper policy
between probes. Status ignores all initialized submodules, so nested repository
changes are outside the observation and nested clean filters are not executed.
Unsupported alternates, denied paths, replaced roots,
changed Git metadata/executable identity/HEAD and unavailable sandboxing produce
typed gaps. Report metadata records only a relative path and presence, never
content or a hash.
Command facts and observations remain bounded; foreground-only recovery is marked
structured_job_recovery_unavailable, not presented as a durable command job.