The bridge now holds each session's `AgentHandle` disposer in its `SessionRecord` and runs it on teardown (client disconnect or fiber dispose) instead of the old `abort()` + `whenIdle()` drain that left agents registered. A bare client disconnect now leaves NO registered agent and NO session-store entry — not an idled-but-still-registered one. The queue-aware `cancel()` inside the disposer also closes the former pre-step best-effort window (a turn about to start is dropped), so teardown reaches true quiescence. The `session/load`-races-teardown leak is fixed: if the bridge closed while `resume()` was pending, the just-resumed handle is disposed before throwing, so it leaves no orphan (it has no SessionRecord, so quiesce() never sees it). Tests: the disconnect test now asserts (through the SAME memoized teardown) that the agent is unregistered AND its session removed; a durability test re-loads the persisted log after dispose and asserts the closing turn/end is on disk (guards the teardown-order contract); a sibling-isolation test proves one handle's dispose() leaves other agents untouched. Docs: agent / agent-loop / acp READMEs, architecture.md, and the stale in-code quiesce() ownership comment updated to the per-agent disposal model; the now-resolved TODO(rfc010-agent-disposal) / TODO(rfc010-cancel-prestep) teardown notes removed.
5.6 KiB
RFC: Multiplex concurrent ACP sessions over one connection
Status: proposed
Implementation status: the multi-session bridge (steps 1, 3, 4) and the bash task-ownership isolation are implemented in
packages/acp+packages/tool-bash. Per-session permission ownership is deferred — it depends on the ACP support permission gate (TODO(rfc010-permission-gate)), which is itself deferred; theagent→sessionIdreverse map the gate will route through is in place. Step 2's per-session disposer scope is now implemented (see agent lifecycle & ownership seams): the factory returns a per-agentAgentHandlewhosedispose()stops the loop, awaits quiescence, unregisters the agent, and removes its session, so a bare client disconnect leaves no registered agent or session-store entry. Status staysproposeduntil per-session permission ownership lands.
Problem
ACP support ships with a single active session per connection: a second session/new is rejected. Editors expect to run several conversations over one agent subprocess — a user opens multiple threads, or a client pre-warms sessions. The single-session guard is a deliberate MVP scope cut, not an architectural limit; this RFC lifts it.
This paragraph is historical: the multi-session bridge has landed. The remaining proposed work is per-session permission ownership plus the lifecycle seams now tracked in agent lifecycle and ownership seams.
Proposal
The harness core already supports many agents (AgentRegistry.list() and AgentLoop.create impose no count limit), so multiplexing is a bridge-layer change in @deepseek-ai/dsh-acp, not a loop or core change.
- Lift the single-session guard in
session/new; allow N live sessions, each mapped to its ownReactLoopAgent. - The bridge's
sessionId→agentandSession→sessionIdmaps (introduced single-entry by the ACP support RFC) become true multi-entry, plus a thirdagent→sessionIdreverse map: thetools/executepermission gate receives onlyexec.agent(no sessionId), so it needs an O(1) reverse lookup to find the owning session. Everyagent/*event and everysession/eventis demuxed strictly by id, so two sessions streaming at once never interleave theirsession/updatenotifications. - Per-session prompt queues: the ACP support RFC's single-entry in-flight-prompt state becomes multi-entry — one in-flight prompt per session, tracked per
sessionId. - Per-session cancel routing:
session/cancelaborts only its own session's agent and settles only that session's in-flight prompt.agent.abort()drives a per-agentAbortController, so the per-sessionexec.signalis the natural isolation fence. - Per-session permission ownership: a
session/request_permissionand its outcome are bound to the originating session via the reverse map, so a permission prompt or a cancel in one session can never resolve another session's pending permission.
Plan
- Generalize the two id maps to multi-entry and add the
agent→sessionIdreverse map; add a per-session record holding the agent, the in-flight-prompt state, the pending-permission registry, and the session's disposer scope (see step 2). - Give each session a real per-session disposer scope, NOT
ctx.extend()— in Cordisctx.extend()only creates a child context/prototype, butctx.on()registered on it is still owned by the current plugin fiber, so disposing it would not remove that session's listeners. Use a genuine child fiber (load a per-session sub-plugin, e.g.ctx.plugin(...)returning a fork, or collect each session'sctx.ondisposers in its session record and call them on teardown). Demux everyagent/*andsession/eventby id into the right session record. Note the single globaltools/executelistener stays on the bridge root (it must see all agents) and routes via the reverse map. - Lift the
session/newguard; keepsession/load(from ACP support) working per session. - Tests for cross-session isolation: two sessions streaming and permission-prompting concurrently never interleave; a cancel/abort in one session leaves the other's stream and pending permission untouched; per-session in-flight-prompt enforcement holds independently; disposing one session leaves the others running.
Risks
Listener fan-out cost: each session adds listeners; ensure disposal of one session removes exactly its own and the connection teardown (from ACP support) still reaches quiescence across all sessions.
The subtle correctness trap is cross-session leakage — a cancel or abort on one session settling another session's pending permission. The per-session permission ownership rule (routed via the agent→sessionId reverse map) and its isolation test are the guard.
Shared background-task state: the bash executor's task ids are global and predictable (bash-1, bash-2, …), and bash_output/bash_kill look up by id without checking the caller. Under one session this is benign; under N sessions one session's agent could read or kill another's background task. This is a pre-existing tool-bash gap that multi-session turns into a real isolation hole — fixing it (validate the caller against the task owner) belongs with this RFC or a companion tool-bash change.