10 KiB
RFC: Subagent capability seam
Status: implemented
The full seam is shipped: the
dsh-subagentinterface, thedsh-subagent-mocktest backend, and thedsh-tool-subagentconsumer; the two in-process backends (dsh-subagent-spawn,dsh-subagent-fork); the nested-agent snapshot infrastructure (per-session snapshot replay); and the out-of-processdsh-subagent-acpbackend (its RFC).
Problem
The harness has a long-deferred seam for subagents — an agent delegating work to another agent. The intent was sketched in the Agent/AgentLoop interfaces (packages/core/agent/src/types.ts, packages/core/agent-loop/src/index.ts): a creation option referencing a parent agent (fork = seed the child session with the parent's event log; spawn = fresh session), with the child returned as an Agent handle so steering and event subscription work uniformly. This RFC realizes that seam; the banner above lists what shipped.
The distinctive requirement — the one that shapes the whole design — is that multiple subagent implementations must coexist at runtime. A parent may want a cheap in-process child for a scoped subtask AND an isolated out-of-process child (over ACP) in the same session. The transports we foresee:
- in-process — a child
ReactLoopAgenton the sameContext(the cheapest, and nearly free given the existing agent factory); - ACP — act as an ACP client driving another agent process (which can be another instance of ourselves);
- later: A2A, the Codex app-server, and the Claude Code Agent SDK — each the same out-of-process "start a child, prompt it, stream updates, cancel" shape as the ACP backend.
Alternatives considered
Why not the bash seam shape
The bash seam (capability seams) registers exactly one BashExecutor per context; loading a second throws. That is correct for bash (one machine, one way to run a command) but wrong here: coexistence is the requirement. So the subagent service is a named-provider registry — each implementation registers under a unique name and a caller picks one by name — mirroring the LLM adapter registry (LlmService.registerAdapter), not the single-service bash executor. The seam is still three-package (interface / implementation / consumer); only the "one vs. many implementations" axis differs.
Decision
The three-package seam
A new package group packages/subagent/:
| Package | Role |
|---|---|
@deepseek-ai/dsh-subagent |
interface: SubagentService (ctx.subagents), SubagentProvider, SubagentRun, the request/result/capability vocabulary, the subagent/* events |
@deepseek-ai/dsh-subagent-spawn |
implementation: a fresh in-process child via ctx.agents.create |
@deepseek-ai/dsh-subagent-fork |
implementation: an in-process child seeded with a snapshot of the parent's log |
@deepseek-ai/dsh-subagent-acp |
implementation: an ACP client driving a configured child process |
@deepseek-ai/dsh-subagent-mock |
support: a scripted provider for testing the seam through the real load path |
@deepseek-ai/dsh-tool-subagent |
consumer: the model-facing subagent tool over ctx.subagents |
The primitive: async start → SubagentRun
A provider exposes start(request) → Promise<SubagentRun>. Promise fulfillment is the publication/readiness and provider-to-caller ownership boundary: for an in-process backend the child is already published in ctx.agents, and for ACP the remote session already exists. SubagentStartRequest.signal is the single cancellation channel before and after readiness; SubagentRun carries the terminal result and a dispose() method that cancels remaining work and awaits quiescence. The transport-neutral verb is start; "spawn" is reserved for the in-process dsh-subagent-spawn backend's identity, not the service verb. A rejected start cleans provider-owned partial resources and emits neither subagent lifecycle event.
Two kinds of optional capability, discovered two ways
- Start-time features (
outputSchema,depthLimit,toolFilter,persona) ride on a staticprovider.capabilitiesdescriptor. The service checks every requested one BEFORE delegating and rejects loud (SubagentError('UNSUPPORTED_CAPABILITY')) if the provider lacks it — never accepted-then-ignored. They must be checked before a run exists, which is why they cannot be runtime methods. - Runtime features (steering via
sendMessage, follow-up viaresume) are optional methods onSubagentRun. The method's presence IS the capability, and TypeScript narrowing is the discovery mechanism: a consumer cannot call an absent method without narrowing first, so there is no silent-degradation path and no separate flags object to keep in sync.
Fork vs. fresh are separate backends, not a flag
Rather than a context: 'fresh' | 'fork' request field, the distinction is the provider's identity: dsh-subagent-spawn (fresh, isolated, own system prompt) and dsh-subagent-fork (seeded from the parent's log) are two registered providers. You pick behavior by picking a provider — consistent with the registry being the selection mechanism. The fork backend seeds only a balanced, completed-turn prefix of the parent log: at tool-execute time the parent's turn is open (it holds the assistant/message and the dangling spawn tool/call with no tool/result), and seeding that raw prefix would give the child an unbalanced turn that the invariants trace replay rejects.
Child isolation and the parent log
Each subagent runs in its own Session (own id, parentSession lineage), persisted independently. The parent's log records only the spawn tool/call and its tool/result (the child's final output) — the child's internal steps and tool calls stay in the child's own session, never injected into the parent log. This is the only design that is identical across transports: an ACP child's internal events physically cannot be injected into our parent log, so making in-process behave the same keeps the seam transport-agnostic.
Synchronous collect (first cut)
The dsh-tool-subagent consumer passes its execution signal into the start request, awaits the ready run's result, and returns the child's final output as the tool result, blocking the parent's turn until the child finishes. A try/finally always dispose()s the run, so no success, failure, or cancellation path leaks an idle child/session. A non-completed stop reason maps to an isError result rather than returning partial output as success. Steering (sendMessage) is part of the contract but intentionally unused in this consumer.
Provider selection is config, not model-facing
dsh-tool-subagent binds to exactly one provider name (Config.provider); the model sees only { description, prompt }. To expose more than one transport, load the tool plugin more than once, each bound to a different provider and a distinct toolName (the tool registry rejects a duplicate name). The service holds the multi-provider registry; the tool picks one — no provider/type parameter in the schema this cut.
Testing
The seam is tested through the real cordis Loader / export path, not a hand-built ctx.plugin mount (which bypasses unwrapExports and cannot catch a broken export shape — postmortem 0001); the registry pins HMR-safety, duplicate-name rejection, and start-time capability rejection; the nested-agent snapshot scenarios replay keyless in the default gate (per-session snapshot replay); in-process backends carry real-loop unit tests plus a with-key e2e.
Consequences
- Recursion. Without a bound, an in-process child can see the delegation tool and recurse. The in-process backends implement the optional absolute depth limit and scoped live-global
toolFilter; ACP advertises both capabilities off and rejects such a request. The subagent composition-controls RFC owns their exact semantics and security limits. - Blocking the parent turn. Synchronous collect holds the parent's
runStepopen for the child's full duration. This is acceptable for the first cut; background / poll / spill semantics are deferred to a future redesign that unifies long-running-tool handling across subagents AND bash (a sub-agent and a longbashbackground task pose the same "the model started something slow, how does it collect later" problem, and should share one mechanism rather than each inventing its own). - Live progress. This cut surfaces only lifecycle + final result; a per-chunk child→parent update stream is deferred with the background redesign.
- ACP client surface. Proxying
fs/terminalfrom the ACP child back to the parent (a shared-workspace mode) is future work; the first cut advertises neither, so the child self-serves in its own process. - Snapshot coverage of nested agents. The snapshot tier (
pnpm run test:snapshot) replays a recorded session throughdsh-llm-replay. It was built single-session: a single GLOBAL positional cursor (the Nthllm/streamcall serves the Nth recorded entry) and a harness that harvested a single session log file. A subagent runs as a second agent with its own session log, so a parent→child scenario needed per-session-keyed replay plus harvest-all-logs and plural-session-id plumbing — self-contained infrastructure orthogonal to the backends, scheduled as a dedicated stacked follow-up rather than folded into the in-process-backends PR. That follow-up has landed: see Per-session snapshot replay for nested agents. Replay now keys each call by its calling session (GenerateOptions.sessionId) and binds live sessions to recorded scripts by first-call order; the harness harvests every log; and two nested scenarios (subagent-spawn,subagent-multi) replay keyless in the default gate. In-process subagents remain covered by real-loop unit tests and a with-key e2e in addition to the snapshot tier.