context/message previously defaulted to a <context source="…">…</context> wrapper. No model is trained on a <context> tag either, and message framing does not belong on the session surface: the surface projects the durable log, and a caller that wants a frame formats its own content — which the one heavy producer (workspace-context) already does with its own <system-reminder> frame, opting out via 'raw'. The tag only added machinery — ContextEnvelope plus an envelope field threaded through InjectOptions, HookContext, the context/message event, and the agent-loop inject/additionalContexts plumbing. context/message now projects its content verbatim as a user-role message, sharing one deriveEventMessage case with user/message and steering/message. ContextEnvelope and every envelope field are removed; context/message.meta still carries durable, model-hidden JSON state. Regenerated catalogs and website API; refreshed the three affected keyless snapshots (envelope field only; timestamps unchanged). Broadens and renames the steering Agent Note to cover both envelope removals as one decision. Agent Note: .agents/notes/implemented/simplification/2026-07-20-unwrap-injected-content-envelopes.md
4.9 KiB
Agent Note: Project injected content verbatim, dropping the XML envelopes
Status: implemented
English | 中文
Problem
Two families of injected session content rendered into the model transcript wrapped in XML envelopes: steering/message as <steering source="…">…</steering> and context/message as <context source="…">…</context> (the latter with a 'raw' opt-out that skipped the wrapper). The envelopes aimed to tell the model "this is injected, not the user speaking."
Two problems:
- No model is trained on these tags.
<steering>and<context>are arbitrary markup no model was taught to read, so the framing adds tokens without a reliable effect and can actively mislead — recorded transcripts show a model treating a<steering>instruction as third-party metadata and refusing it while answering only the original prompt. - The session surface is the wrong layer for framing. The surface projects the durable log into the model transcript; deciding how content is worded is not its job. A caller that wants a particular frame formats its own content before injecting it — which the one heavy producer (
workspace-context) already does, owning its complete<system-reminder>frame and opting out of the<context>wrapper withenvelope: 'raw'. The remaining tag machinery (ContextEnvelope, anenvelopefield threaded throughInjectOptions,HookContext, thecontext/messageevent, and the loop) served a distinction that belongs to the caller.
Decision
Injected session content projects verbatim; the caller owns any framing. deriveEventMessage renders user/message, context/message, and steering/message through one shared case returning { role: 'user', content: event.data.content }; their content blocks reach the model unchanged. context/message's source/meta and steering/message's turn stay in the durable event log but do not render.
The ContextEnvelope type and every envelope field are removed — context/message in SessionEventMap, InjectOptions, HookContext, and the inject()/additionalContexts plumbing in dsh-agent-loop. workspace-context no longer requests 'raw'; its self-framed content renders as before. The renderTagged/renderContextEnvelope helpers are deleted. context/message.meta still carries durable, model-hidden JSON state.
The source attribution the envelopes carried is not lost — it remains on the durable events; it simply no longer renders into the transcript.
Alternatives considered
- Keep the
<context>envelope, unwrap only steering — leaves theContextEnvelope/envelopemachinery alive for a framing bit no model reads, and keeps the inconsistency that the main producer already opts out of. - Keep the envelope field for plugin-sourced content only — splits one projection into two on
source.kindfor no observed benefit; a plugin steering the agent (hook-bridge continuation reasons) also wants the instruction followed, not labeled. - Move the unwrapping into adapters — the canonical projection is the model-visible contract ("model-visible ⟺ logged"); per-adapter divergence on framing would make the derived transcript adapter-dependent. Framing that a caller genuinely wants belongs in the caller's content, not in an adapter.
Consequences
- Mid-turn steering and injected context reach the model with the same weight as an ordinary user prompt.
- The transcript no longer distinguishes injected content from a user message; consumers that need the distinction read the durable event log, which keeps the event types,
source, andmetaintact. - The
hook-{cc,codex}-stop-continueACP snapshots were re-recorded: the old recordings captured the model refusing steering as third-party metadata, the fix's exact failure mode. - The content-block-vocabulary Agent Note's tagged-envelope clause is amended to point here.
Deferred
workspace-context already frames its own content: it emits a complete <system-reminder>…</system-reminder> block as the message content instead of leaning on a surface-level wrapper. That caller-owned pattern is the one to keep — the surface passes content through verbatim, and any framing lives in the producer's own content.
Two framing paths existed — caller-baked framing (workspace-context's <system-reminder>) and surface-level wrapping (<context>/<steering> added by deriveEventMessage). This change removes the second, leaving only caller-owned framing. If labeled framing is wanted again, unify it through the event's meta map — the producer-attached, model-hidden metadata field — consumed by a dedicated renderer or adapter, rather than re-hardcoding a tag in deriveEventMessage. A producer declares the frame it wants in meta; one renderer applies it; the session-surface projection stays a verbatim pass-through.