The sandbox execute wrapper JSON round-tripped the return and blindly cast it
to ToolExecuteReturn. A JSON-valid but wrong-shape return — a bare string,
{ content: 'ok' }, blocks without a type tag — sailed through: the registry
spreads result.content, so { content: 'ok' } became ['o','k'], passed the
session log's isJsonValue gate, and the DeepSeek serializer then flattened it
to '(no output)' — silent corruption of the next model request and every
replay, instead of a contained tool error.
The round-tripped value is now shape-checked against the two ToolExecuteReturn
forms (array of content blocks, or { content: blocks, meta? }); block checks
are structural only (plain object + string type tag) because the ContentBlock
union is merge-extensible. A wrong shape — and the formerly cryptic
forgot-return/bare-string cases — fails that one call with a teaching error
echoing a truncated preview of what was returned and the two valid forms.
New specs pin the object-form pass-through (meta included), six rejection
shapes, and the preview truncation; per-file 100% coverage holds.
The sandbox docs overclaimed a containment contract the design never makes:
"capability access is routed through cordis services, never Node built-ins,
so everything a mounted plugin does stays inspectable and disposable". The
host-realm helpers on the sandbox global (harness, console, btoa) are
reachable functions, so mount code that goes looking can reach the host realm
through one of them — accepted under the trust stance, because the ctx a
mount ultimately receives is fully privileged anyway. Reword the sandbox
module doc, the README trust stance, and the RFC sandbox-semantics section to
say exactly that: the traps and small global surface STEER honest code onto
the cordis services; they are not a security boundary.
Two review findings (#220) on the sandbox context façade:
- Undeclared services were reachable: the façade resolved any live global via
ctx.get(name), so ctx.bash worked without inject: ['bash']. A cross-mount
consumer could then depend on a provider cordis never saw — unmounting the
provider would neither park the consumer nor unwind its registered tools,
leaving a model-visible tool that fails only at execution. The façade now
reads ctx.fiber.inject and refuses any service the mount did not declare
(with a teaching error naming the inject fix), so the dependency is always
visible to cordis and its activation/unload semantics bind.
- ctx.tools.get returned the live ToolDefinition, including execute — mount
code could call another tool directly and bypass ToolRegistry.execute and
its pre/post-execute hooks and accounting. get now returns the same
read-only name/description/parameters view as schemas(), never an invocable.
Adds inject-gate and schema-view regression cases to sandbox-context.spec.ts
(undeclared property/get denied, declared allowed, the cross-mount zombie-tool
scenario refused at call time, get exposes no execute). Package stays at
per-file 100% coverage. RFC, mount description, and tool-catalog updated.
Review finding (#220): the guarded proxy only special-cased ctx.tools, so
mount code could reach an UNGUARDED context through ctx.root, ctx.extend(), or
a service instance's .ctx, then ctx.root.tools.register({…}) to bypass the
marker check and host-realm normalization — a raw vm-realm result would later
error a real agent turn at the session-log plainness check.
The sandbox ctx is now a whitelist façade, not a pass-through proxy: it exposes
only what a mount needs — tools.register (marker-guarded), on/once, provide, the
timer helpers, and injected services resolved through a guarded get — and denies
every framework-plumbing member (root, parent, fiber, reflect, registry, extend,
isolate, intercept, plugin, set, mixin, …) with a teaching error. Injected
services are wrapped so a method returning a Context is rejected on the way back
(the .ctx escape), closing the one indirect leak. There is no context-valued
member left to reach; cross-mount provide/inject is untouched (the plugin's own
inject and the fiber's pending/active gating are unchanged). ctx.plugin (child
plugins) and ctx.set are denied by design; ctx.effect is deferred (FIXME).
Adds tests/sandbox-context.spec.ts covering the escape class (root/extend/fiber/
plugin/set/… denied, the classic root.tools.register bypass, the .ctx escape,
read-only writes) plus the async-service and symbol/in-operator paths for 100%
coverage. RFC/README/tool-catalog/config-catalog updated; api-catalog.ts
regenerated (also picks up the codeRuntime service that entered on the master
merge and was left stale).
The tree SHAPE was the wrong surface for the model: what it needs from
cordis_inspect is what services, plugins, and capabilities are loaded, not the
fiber hierarchy. The plugins section is now a flat name + lifecycle-state list
from ctx.registry (deterministically sorted, one line per instance); the ASCII
tree renderer, the parent→child rebuild, and the dyn-id tree annotation are
deleted — dynamic mounts keep their own richer dynamic section (id, state,
provides, waits). Net -49 lines; RFC and READMEs state the flat-list contract.
Field sessions showed models writing tool schemas in the JSON-Schema dialect
by strong prior — type: 'integer', required: false, then the full
{ type:'object', properties, required: [...] } wrapper — and the rejection
text itself pushed a nearly-correct DSL attempt BACK to raw JSON Schema: one
stats tool cost three consecutive schema errors before mounting. The boundary
now normalizes wherever the input has exactly one meaning (wrapper unwrapped
with the required array becoming per-property flags at any nesting level,
integer → number, required: false → optional, all rebuilt as fresh host-realm
objects) and rejects only genuinely meaningless input, enumerating the valid
vocabulary in the error. Re-running the failing session mounts first-try.
The mount description documents both accepted forms.
The design record for tool-cordis: the three-tool contract, the vm sandbox
trust stance and boundary mechanisms, the dynamic-group lifecycle, cross-mount
provide/inject composition, the generated runtime API catalog, and the
alternatives weighed (per-capability registration tools, hand-maintained API
tables, a mount provenance event, a hardened sandbox).