Add bash execution: dsh-bash seam, dsh-bash-local impl, dsh-tool-bash tools

Three packages following the new capability-seam pattern (interface /
implementation / consumer, now documented in docs/architecture.md):

- dsh-bash: abstract BashExecutor service (ctx.bash) + vocabulary types.
- dsh-bash-local: local subprocesses — bash -c per call in a detached
  process group, SIGTERM→SIGKILL group kills, tail-keep truncation with
  full-stream spill files, model-friendly env, background task registry.
- dsh-tool-bash: the bash / bash_output / bash_kill tool schemas with
  runtime arg validation and background completion notices via
  agent.inject(). Non-zero exits are reported, not errored.

Design surveyed against the bash tools of Claude Code, OpenCode, Codex,
and pi (notes in the package READMEs). Permissions/sandbox stay TODO on
the tools/execute waterfall seam; stateful-shell alternatives recorded
in run.ts.
This commit is contained in:
Tianyi Cui
2026-06-12 23:28:44 +08:00
parent 353b0170c7
commit 8b5a3ef730
27 changed files with 2714 additions and 7 deletions
+62
View File
@@ -0,0 +1,62 @@
# @deepseek-ai/dsh-tool-bash
The model-facing bash tools — `bash`, `bash_output`, `bash_kill` — registered
over the `ctx.bash` executor seam (`@deepseek-ai/dsh-bash`). Pure schema +
text shaping; every process concern lives behind the seam, so sandboxed or
remote executor implementations swap in without changing what the model sees.
Requires a loaded executor implementation (e.g.
`@deepseek-ai/dsh-bash-local`); the plugin stays pending until `ctx.bash`
exists (`inject: ['tools', 'bash']`).
## Tools
### `bash`
| Arg | Type | Notes |
|---|---|---|
| `command` | string (required) | Run via `bash -c`. No state persists between calls — use `workdir`, not `cd`. |
| `description` | string (required) | One-line, active-voice summary of the command (5-10 words), for UI/log display only — no effect on execution. |
| `timeoutMs` | number | Default/max from executor config (120s/600s for bash-local). |
| `workdir` | string | Working directory for this call. |
| `run_in_background` | boolean | Return a task id immediately; no timeout applies. |
`command`, `workdir`, and `timeoutMs` are resolved against the executor's
config defaults via `ctx.bash.resolve()` before execution, so the executor
seam (`BashExecSpec`) receives explicit `workdir`/`timeoutMs` values.
Result text: stdout, then a `[stderr]` section, then status markers —
`[timed out after Nms]` whenever the executor's timer fired (reported
independently of how the process ended, so a command that traps SIGTERM and
exits 0 still shows it), `[killed by signal: …]` for a signal death,
`[exit code: N]` for a non-zero exit (reported, **not** `isError`: the model
decides how to react), and `[output truncated; full output: <path>]` when the
tail was kept. Only infrastructure failures (spawn errors, aborts) surface as
`isError` results.
### `bash_output`
`task_id` → output produced **since the previous `bash_output` call** plus a
status line (`running` / `completed, exit code: N` / `killed`). Reads that
lost data to buffer bounds say so and point at the full-output spill file.
### `bash_kill`
`task_id` → SIGTERM→SIGKILL on the task's process group. Killing an
already-finished task is a reported no-op; unknown ids are errors.
## Background completion notices
When a background task finishes, a short notice is injected into the owning
agent's session (`agent.inject()`, source `{kind: 'plugin', plugin:
'tool-bash'}`). Injection is **durable context for the next model request,
not a wake-up** — an idle agent stays idle until something sends a message.
That's why the tool descriptions tell the model to poll with `bash_output`.
## Permissions
`TODO(permissions)`: commands run with the executor's full authority. The
permission/sandbox seam is the `tools/execute` waterfall (veto or ask) plus
sandboxing `BashExecutor` implementations — see docs/architecture.md.
`@cordisjs/plugin-capability` (a named-permission service with a session
`test()`) is a candidate building block for that work.
+40
View File
@@ -0,0 +1,40 @@
{
"name": "@deepseek-ai/dsh-tool-bash",
"description": "Model-facing bash tools (bash, bash_output, bash_kill) over the DeepSeek Harness bash executor seam",
"version": "0.0.1",
"private": true,
"type": "module",
"main": "lib/index.js",
"types": "lib/index.d.ts",
"exports": {
".": {
"types": "./lib/index.d.ts",
"default": "./lib/index.js"
},
"./src/*": "./src/*",
"./package.json": "./package.json"
},
"files": [
"lib",
"src"
],
"license": "BSD-3-Clause",
"peerDependencies": {
"@deepseek-ai/dsh-agent": "^0.0.1",
"@deepseek-ai/dsh-bash": "^0.0.1",
"@deepseek-ai/dsh-llm": "^0.0.1",
"@deepseek-ai/dsh-tools": "^0.0.1",
"cordis": "^4.0.0-rc.6"
},
"devDependencies": {
"@deepseek-ai/dsh-agent": "^0.0.1",
"@deepseek-ai/dsh-agent-loop": "^0.0.1",
"@deepseek-ai/dsh-bash": "^0.0.1",
"@deepseek-ai/dsh-bash-local": "^0.0.1",
"@deepseek-ai/dsh-llm": "^0.0.1",
"@deepseek-ai/dsh-session": "^0.0.1",
"@deepseek-ai/dsh-system-prompt": "^0.0.1",
"@deepseek-ai/dsh-tools": "^0.0.1",
"cordis": "^4.0.0-rc.6"
}
}
+226
View File
@@ -0,0 +1,226 @@
/**
* The model-facing bash tools: `bash`, `bash_output`, `bash_kill`. Pure
* schema + text shaping — every process concern lives behind the `ctx.bash`
* executor seam (`@deepseek-ai/dsh-bash`), so sandbox/permission/remote
* executor implementations swap in without touching what the model sees.
*
* Background notifications: when a background task completes, a short notice
* is injected into the owning agent's session (`agent.inject()` — the
* documented context seam). Injection is durable context for the NEXT model
* request, not a wake-up: an idle agent stays idle until something sends a
* message, which is why the tool descriptions tell the model to poll with
* `bash_output`.
*
* TODO(permissions): commands run with the executor's full authority. The
* permission/sandbox seam is the `tools/execute` waterfall (veto/ask) plus
* sandboxing `BashExecutor` implementations — see docs/architecture.md
* § plugin checklist.
*
* @module @deepseek-ai/dsh-tool-bash
*/
import type { Context } from 'cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'
import type { Agent } from '@deepseek-ai/dsh-agent'
import type { BashRunResult, BashTask, CollectedOutput } from '@deepseek-ai/dsh-bash'
export const name = 'tool-bash'
export const inject = ['tools', 'bash']
/**
* Validate model-produced arguments. `defineTool`'s `InferArgs` typing is
* compile-time only — at runtime `arguments` is whatever JSON the model
* emitted, so every field is checked before it reaches the executor.
*
* TODO(RFC 005): this hand-rolled validation is the per-tool stopgap until
* `defineTool` validates parsed args against the SchemaSpec itself (the
* converter already encodes the structure). When that lands, delete this and
* let the registry reject malformed calls — see docs/rfc/005.
*/
function validateBashArgs(args: {
command: string
description: string
timeoutMs?: number
workdir?: string
run_in_background?: boolean
}): void {
if (typeof args.command !== 'string' || args.command.trim().length === 0) {
throw new Error('invalid command: expected a non-empty string')
}
if (typeof args.description !== 'string' || args.description.trim().length === 0) {
throw new Error('invalid description: expected a non-empty string')
}
if (args.timeoutMs !== undefined
&& (typeof args.timeoutMs !== 'number' || !Number.isFinite(args.timeoutMs) || args.timeoutMs <= 0)) {
throw new Error(`invalid timeoutMs: expected a positive number, got ${JSON.stringify(args.timeoutMs)}`)
}
if (args.workdir !== undefined && typeof args.workdir !== 'string') {
throw new Error(`invalid workdir: expected a string, got ${JSON.stringify(args.workdir)}`)
}
if (args.run_in_background !== undefined && typeof args.run_in_background !== 'boolean') {
throw new Error(`invalid run_in_background: expected a boolean, got ${JSON.stringify(args.run_in_background)}`)
}
}
/** Require a string `task_id` (model-produced, so runtime-checked). */
function validateTaskId(value: unknown): string {
if (typeof value !== 'string' || value.length === 0) {
throw new Error(`invalid task_id: expected a string, got ${JSON.stringify(value)}`)
}
return value
}
/** Append the truncation notice (with the full-output spill path) to a stream's text. */
function streamText(output: CollectedOutput): string {
if (!output.truncated) return output.text
return `${output.text}\n[output truncated; full output: ${output.spillPath ?? '(unavailable)'}]`
}
/**
* Shape one finished run into the text the model sees: stdout, then a marked
* stderr section, then exit-status markers. Non-zero exits are REPORTED, not
* errored — the model decides how to react; only infrastructure failures
* (spawn errors, aborts) surface as isError results.
*/
export function renderResult(result: BashRunResult): string {
const out = streamText(result.stdout)
const err = streamText(result.stderr)
let body = out
if (err.length > 0) {
// Single newline between sections (stdout usually ends with one already).
if (body.length > 0 && !body.endsWith('\n')) body += '\n'
body += `[stderr]\n${err}`
}
if (body.length === 0) body = '(no output)'
const markers: string[] = []
// Timeout is reported independently of how the process actually ended: a
// command can trap SIGTERM and exit 0 after our timer fired (e.g.
// `trap "exit 0" TERM; sleep 60`), giving timedOut:true / exitCode:0 /
// signal:null — the model must still see that the command was cut short.
if (result.timedOut) markers.push(`[timed out after ${result.timeoutMs}ms]`)
if (result.signal !== null) {
markers.push(`[killed by signal: ${result.signal}]`)
} else if (result.exitCode !== 0) {
markers.push(`[exit code: ${result.exitCode}]`)
}
if (markers.length === 0) return body
if (!body.endsWith('\n')) body += '\n'
return body + markers.join('\n')
}
/** Status line for background task reads. */
function statusLine(task: BashTask): string {
switch (task.status) {
case 'running': return '[status: running]'
case 'killed': return `[status: killed${task.signal !== null ? ` by ${task.signal}` : ''}]`
case 'completed': return `[status: completed, exit code: ${task.exitCode ?? 0}]`
}
}
export function apply(ctx: Context): void {
// Background completion → inject a notice into the owning agent's session.
// Tracks the agent per task id; entries drop once notified.
const owners = new Map<string, Agent>()
ctx.bash.onTaskDone((task) => {
const agent = owners.get(task.id)
owners.delete(task.id)
if (!agent) return
try {
agent.inject(
[{ type: 'text', text: `background bash task ${task.id} finished ${statusLine(task)}. Read its output with bash_output.` }],
{ source: { kind: 'plugin', plugin: 'tool-bash' } },
)
} catch (error: unknown) {
// The ONE expected failure: the agent was disposed between task
// completion and this injection (LoopAgent.inject throws
// `agent "<id>" is disposed`). That race is benign — drop the notice.
// Anything else is a real bug and must surface, not be swallowed.
if (error instanceof Error && error.message.includes('is disposed')) return
throw error
}
})
ctx.tools.register(defineTool({
name: 'bash',
description: 'Execute a bash command (`bash -c`) and return its stdout/stderr. '
+ 'Each call runs in a fresh shell: no state (cwd, variables, functions) persists between calls — '
+ 'pass `workdir` instead of using `cd`. Non-zero exits are reported as `[exit code: N]`. '
+ 'Long output is truncated to its tail; the full output is saved to a file whose path is reported. '
+ 'Set `run_in_background: true` for long-running commands: the call returns a task id immediately; '
+ 'poll it with `bash_output` and stop it with `bash_kill`.',
parameters: {
command: { type: 'string', required: true, description: 'The bash command to execute.' },
description: {
type: 'string',
required: true,
description: 'Clear, concise description of what this command does in active voice, '
+ '5-10 words (shown in the UI). Examples: "ls" → "List files in current directory"; '
+ '"git status" → "Show working tree status"; "npm install" → "Install package dependencies".',
},
timeoutMs: { type: 'number', description: 'Timeout in milliseconds (default 120000, max 600000). The command is killed on expiry.' },
workdir: { type: 'string', description: 'Working directory for this command.' },
run_in_background: { type: 'boolean', description: 'Run in the background and return a task id immediately. No timeout applies.' },
},
async execute(args, exec) {
validateBashArgs(args)
// `description` is display/logging metadata only (surfaced to UIs via
// the tool/call session event); it is intentionally NOT forwarded to
// ctx.bash and has no effect on execution.
const request = {
command: args.command,
...args.workdir !== undefined ? { workdir: args.workdir } : {},
...args.timeoutMs !== undefined ? { timeoutMs: args.timeoutMs } : {},
...exec.signal ? { signal: exec.signal } : {},
}
if (args.run_in_background === true) {
const task = ctx.bash.start(ctx.bash.resolve(request))
if (exec.agent) owners.set(task.id, exec.agent)
return [{ type: 'text', text: `started background task ${task.id}` }]
}
const result = await ctx.bash.run(ctx.bash.resolve(request))
if (result.aborted) throw new Error('command aborted')
return [{ type: 'text', text: renderResult(result) }]
},
}))
ctx.tools.register(defineTool({
name: 'bash_output',
description: 'Read new output from a background bash task started with `bash` + `run_in_background`. '
+ 'Returns only output produced since the previous bash_output call, plus the task status. '
+ 'Tasks keep running while you do other work; poll again later for more output.',
parameters: {
task_id: { type: 'string', required: true, description: 'Task id returned by the bash tool.' },
},
// execute is synchronous (registry reads + string shaping) but the
// ToolDefinition contract wants a Promise — hence resolve(), not async.
execute(args) {
const read = ctx.bash.readOutput(validateTaskId(args.task_id))
let text = read.delta.length > 0 ? read.delta : '(no new output)'
if (read.lossy) {
const paths = [read.stdoutSpillPath, read.stderrSpillPath].filter((p): p is string => p !== undefined)
text += `\n[some output was dropped from memory; full output: ${paths.join(', ')}]`
}
text += `\n${statusLine(read.task)}`
return Promise.resolve([{ type: 'text', text }])
},
}))
ctx.tools.register(defineTool({
name: 'bash_kill',
description: 'Kill a running background bash task (SIGTERM, then SIGKILL) by task id.',
parameters: {
task_id: { type: 'string', required: true, description: 'Task id returned by the bash tool.' },
},
execute(args) {
const id = validateTaskId(args.task_id)
const killed = ctx.bash.kill(id)
return Promise.resolve([{
type: 'text',
text: killed ? `killed background task ${id}` : `task ${id} had already finished`,
}])
},
}))
}
@@ -0,0 +1,164 @@
import { describe, expect, it } from 'vitest'
import { Context } from 'cordis'
import LlmService from '@deepseek-ai/dsh-llm'
import SessionStore from '@deepseek-ai/dsh-session'
import type { SessionEvent } from '@deepseek-ai/dsh-session'
import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
import ToolRegistry from '@deepseek-ai/dsh-tools'
import AgentRegistry from '@deepseek-ai/dsh-agent'
import AgentLoop, { LoopAgent } from '@deepseek-ai/dsh-agent-loop'
import { LocalBashExecutor } from '@deepseek-ai/dsh-bash-local'
import * as ToolBash from '@deepseek-ai/dsh-tool-bash'
import { MockAdapter, textResponse, toolCallResponse } from '../../agent-loop/tests/mock-adapter.ts'
/**
* Full-loop integration: a scripted mock model drives the REAL bash tool
* through the agent loop, exercising the same seams a live model would
* (tool/call + tool/result session events, agent.inject notifications).
*/
async function harness(adapter: MockAdapter) {
const ctx = new Context()
await ctx.plugin(LlmService)
await ctx.plugin(SessionStore)
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
await ctx.plugin(AgentRegistry)
await ctx.plugin(AgentLoop, { agents: [] })
await ctx.plugin(LocalBashExecutor, { timeoutMs: 10_000 })
await ctx.plugin(ToolBash)
ctx.llm.registerAdapter(['mock'], adapter)
return ctx
}
function waitForIdle(ctx: Context, agent: LoopAgent): Promise<void> {
return new Promise((resolve) => {
const dispose = ctx.on('agent/status', (subject, status) => {
if (subject === agent && status === 'idle') {
dispose()
resolve()
}
})
})
}
function events(agent: LoopAgent): SessionEvent[] {
return [...agent.session.events]
}
/** Find a session event by type, narrowed; throws when absent. */
function findEvent<T extends SessionEvent['type']>(
log: SessionEvent[],
type: T,
position: 'first' | 'last' = 'first',
): Extract<SessionEvent, { type: T }> {
const found = position === 'first'
? log.find(event => event.type === type)
: log.findLast(event => event.type === type)
if (!found) throw new Error(`no ${type} event in the session log`)
return found as Extract<SessionEvent, { type: T }>
}
function resultText(event: SessionEvent): string {
if (event.type !== 'tool/result') return ''
return event.data.content
.filter(block => block.type === 'text')
.map(block => block.text)
.join('')
}
describe('bash tool through the agent loop', () => {
it('foreground: model calls bash, sees the result, replies', async () => {
const adapter = new MockAdapter([
toolCallResponse('call-1', 'bash', { command: 'echo integration-ok', description: 'test command' }, 'Running it.'),
textResponse('The command printed integration-ok.'),
])
const ctx = await harness(adapter)
const agent = ctx.agentLoop.create('it-fg', { model: 'mock' })
agent.send([{ type: 'text', text: 'run echo integration-ok' }])
await waitForIdle(ctx, agent)
const log = events(agent)
const toolCall = findEvent(log, 'tool/call')
expect(toolCall.data.name).toBe('bash')
const toolResult = findEvent(log, 'tool/result')
expect(toolResult.data.isError).toBe(false)
expect(resultText(toolResult)).toBe('integration-ok\n')
// The second model call saw the tool result in its derived history.
const lastRequest = adapter.requests.at(-1)
const toolResultBlocks = (lastRequest?.messages ?? [])
.flatMap(message => message.content)
.filter(block => block.type === 'tool-result')
expect(toolResultBlocks).toHaveLength(1)
const finalMessage = findEvent(log, 'assistant/message', 'last')
expect(finalMessage.data.content.some(
block => block.type === 'text' && block.text.includes('integration-ok'),
)).toBe(true)
})
it('foreground: non-zero exit is reported in the result text, not as isError', async () => {
const adapter = new MockAdapter([
toolCallResponse('call-1', 'bash', { command: 'exit 9', description: 'test command' }),
textResponse('It failed with code 9.'),
])
const ctx = await harness(adapter)
const agent = ctx.agentLoop.create('it-exit', { model: 'mock' })
agent.send([{ type: 'text', text: 'run exit 9' }])
await waitForIdle(ctx, agent)
const toolResult = findEvent(events(agent), 'tool/result')
expect(toolResult.data.isError).toBe(false)
expect(resultText(toolResult)).toContain('[exit code: 9]')
})
it('background: start → poll → completion notice lands as context/message', async () => {
const adapter = new MockAdapter([
toolCallResponse('call-1', 'bash', { command: 'echo bg-ok', description: 'test command', run_in_background: true }),
toolCallResponse('call-2', 'bash_output', {}, undefined),
textResponse('Background task finished.'),
])
// The second tool call needs the REAL task id from the first result;
// a tools/execute waterfall listener rewrites the scripted arguments.
let taskId = ''
const ctx = await harness(adapter)
const agent = ctx.agentLoop.create('it-bg', { model: 'mock' })
// Intercept the first tool result to capture the generated task id, then
// rewrite the second scripted call's arguments to use it.
ctx.on('session/event', (_session, event) => {
if (event.type === 'tool/result' && taskId === '') {
const match = /task (bash-\d+)/.exec(resultText(event))
if (match) taskId = match[1]!
}
})
ctx.on('tools/execute', async (exec, next) => {
if (exec.name === 'bash_output') {
exec.arguments = { task_id: taskId }
}
return next()
})
agent.send([{ type: 'text', text: 'run echo bg-ok in the background' }])
await waitForIdle(ctx, agent)
// Wait for the background task itself (completion may race turn end).
const task = ctx.bash.get(taskId)
if (!task) throw new Error(`task ${taskId} not registered`)
await task.done
const log = events(agent)
const firstResult = findEvent(log, 'tool/result')
expect(resultText(firstResult)).toBe(`started background task ${taskId}`)
const notice = findEvent(log, 'context/message')
expect(notice.data.content.some(
block => block.type === 'text' && block.text.includes(`background bash task ${taskId} finished`),
)).toBe(true)
expect(notice.data.source).toEqual({ kind: 'plugin', plugin: 'tool-bash' })
})
})
+407
View File
@@ -0,0 +1,407 @@
import { mkdtempSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { describe, expect, it, vi } from 'vitest'
import { Context } from 'cordis'
import { CallId } from '@deepseek-ai/dsh-llm'
import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
import ToolRegistry from '@deepseek-ai/dsh-tools'
import { LocalBashExecutor } from '@deepseek-ai/dsh-bash-local'
import * as ToolBash from '@deepseek-ai/dsh-tool-bash'
import { renderResult } from '@deepseek-ai/dsh-tool-bash'
const spillDir = mkdtempSync(join(tmpdir(), 'dsh-tool-bash-spec-'))
async function setup() {
const ctx = new Context()
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
await ctx.plugin(LocalBashExecutor, { timeoutMs: 10_000 })
;(ctx.bash as LocalBashExecutor).internals = { spillDir, graceMs: 200 }
await ctx.plugin(ToolBash)
return ctx
}
let callCounter = 0
function call(ctx: Context, name: string, args: unknown) {
return ctx.tools.execute({ callId: CallId(`call-${++callCounter}`), name, arguments: args })
}
function text(result: { content: { type: string; text?: string }[] }): string {
return result.content.filter(block => block.type === 'text').map(block => block.text).join('')
}
describe('bash tool', () => {
it('returns stdout for a successful command', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'echo hello', description: 'test command' })
expect(result.isError).toBe(false)
expect(text(result)).toBe('hello\n')
})
it('reports (no output) for silent commands', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'true', description: 'test command' })
expect(text(result)).toBe('(no output)')
})
it('marks stderr sections', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'echo out; echo err >&2', description: 'test command' })
expect(text(result)).toBe('out\n[stderr]\nerr\n')
expect(result.isError).toBe(false)
})
it('reports non-zero exits without isError', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'echo failing; exit 3', description: 'test command' })
expect(result.isError).toBe(false)
expect(text(result)).toBe('failing\n[exit code: 3]')
})
it('reports timeout kills with both markers (timeout first)', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'sleep 60', description: 'test command', timeoutMs: 100 })
expect(result.isError).toBe(false)
expect(text(result)).toBe('(no output)\n[timed out after 100ms]\n[killed by signal: SIGTERM]')
})
it('reports a timeout even when the command traps the signal and exits 0', async () => {
// The signal-independent timeout marker: a trapped SIGTERM that exits 0
// after our timer fired must NOT look like a clean success. (bash may
// print "Terminated" to stderr for the killed sleep — environment
// dependent — so assert the marker, not the exact body.)
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'trap "exit 0" TERM; sleep 60', description: 'test command', timeoutMs: 100 })
expect(result.isError).toBe(false)
expect(text(result)).toContain('[timed out after 100ms]')
expect(text(result)).not.toContain('[exit code:')
})
it('reports truncation with the spill path', async () => {
const ctx = new Context()
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
await ctx.plugin(LocalBashExecutor, { maxOutputBytes: 100 })
;(ctx.bash as LocalBashExecutor).internals = { spillDir, graceMs: 200 }
await ctx.plugin(ToolBash)
const result = await call(ctx, 'bash', { command: 'for i in $(seq 1 100); do printf "line-%04d\\n" $i; done', description: 'test command' })
expect(text(result)).toContain('[output truncated; full output: ')
expect(text(result)).toContain('line-0100')
})
it('honors workdir', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'pwd', description: 'test command', workdir: '/tmp' })
expect(text(result).trim()).toMatch(/\/tmp$/)
})
it('surfaces spawn failures as isError', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'true', description: 'test command', workdir: '/nonexistent-dsh' })
expect(result.isError).toBe(true)
expect(text(result)).toMatch(/ENOENT/)
})
it('surfaces aborts as isError', async () => {
const ctx = await setup()
const controller = new AbortController()
const pending = ctx.tools.execute({
callId: CallId('call-abort'),
name: 'bash',
arguments: { command: 'sleep 60', description: 'test command' },
signal: controller.signal,
})
setTimeout(() => { controller.abort() }, 50)
const result = await pending
expect(result.isError).toBe(true)
expect(text(result)).toMatch(/aborted/)
})
it.each([
[{}, /invalid command/],
[{ command: 42 }, /invalid command/],
[{ command: ' ' }, /invalid command/],
[{ command: 'x' }, /invalid description/],
[{ command: 'x', description: '' }, /invalid description/],
[{ command: 'x', description: 7 }, /invalid description/],
[{ command: 'x', description: 'd', timeoutMs: 'soon' }, /invalid timeoutMs/],
[{ command: 'x', description: 'd', timeoutMs: -1 }, /invalid timeoutMs/],
[{ command: 'x', description: 'd', timeoutMs: Number.NaN }, /invalid timeoutMs/],
[{ command: 'x', description: 'd', workdir: 7 }, /invalid workdir/],
[{ command: 'x', description: 'd', run_in_background: 'yes' }, /invalid run_in_background/],
])('rejects invalid args %j', async (args, pattern) => {
const ctx = await setup()
const result = await call(ctx, 'bash', args)
expect(result.isError).toBe(true)
expect(text(result)).toMatch(pattern)
})
it('registers all three schemas in the system prompt assembly', async () => {
const ctx = await setup()
const names = ctx.tools.schemas().map(schema => schema.name)
expect(names).toEqual(['bash', 'bash_output', 'bash_kill'])
const bashSchema = ctx.tools.schemas()[0]!
expect(bashSchema.parameters).toMatchObject({
type: 'object',
required: ['command', 'description'],
})
})
it('unregisters everything when the plugin fiber is disposed (HMR safety)', async () => {
const ctx = new Context()
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
await ctx.plugin(LocalBashExecutor, {})
const fiber = await ctx.plugin(ToolBash)
expect(ctx.tools.schemas()).toHaveLength(3)
await fiber.dispose()
expect(ctx.tools.schemas()).toHaveLength(0)
})
it('tools depend on the executor: no registration without ctx.bash', async () => {
const ctx = new Context()
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
// inject: ['tools', 'bash'] keeps the plugin pending until bash exists.
await ctx.plugin(ToolBash)
expect(ctx.tools.schemas()).toHaveLength(0)
await ctx.plugin(LocalBashExecutor, {})
await new Promise(resolve => setTimeout(resolve, 0))
expect(ctx.tools.schemas()).toHaveLength(3)
})
})
describe('background tools', () => {
it('bash with run_in_background returns a task id immediately', async () => {
const ctx = await setup()
const result = await call(ctx, 'bash', { command: 'sleep 0.2; echo bg-done', description: 'test command', run_in_background: true })
expect(result.isError).toBe(false)
expect(text(result)).toMatch(/^started background task bash-\d+$/)
})
it('bash_output polls incrementally and reports status', async () => {
const ctx = await setup()
const started = await call(ctx, 'bash', { command: 'echo first; sleep 0.3; echo second', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
await new Promise(resolve => setTimeout(resolve, 150))
const first = await call(ctx, 'bash_output', { task_id: id })
expect(text(first)).toContain('first')
expect(text(first)).toContain('[status: running]')
await ctx.bash.get(id)!.done
const second = await call(ctx, 'bash_output', { task_id: id })
expect(text(second)).toContain('second')
expect(text(second)).not.toContain('first')
expect(text(second)).toContain('[status: completed, exit code: 0]')
const third = await call(ctx, 'bash_output', { task_id: id })
expect(text(third)).toContain('(no new output)')
})
it('bash_output flags lossy reads with spill paths', async () => {
const ctx = new Context()
await ctx.plugin(SystemPrompt)
await ctx.plugin(ToolRegistry)
await ctx.plugin(LocalBashExecutor, { maxOutputBytes: 100 })
;(ctx.bash as LocalBashExecutor).internals = { spillDir, graceMs: 200 }
await ctx.plugin(ToolBash)
const started = await call(ctx, 'bash', { command: 'for i in $(seq 1 200); do printf "line-%04d\\n" $i; done', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
await ctx.bash.get(id)!.done
const read = await call(ctx, 'bash_output', { task_id: id })
expect(text(read)).toContain('[some output was dropped from memory; full output: ')
})
it('bash_kill stops a running task; repeat reports already-finished', async () => {
const ctx = await setup()
const started = await call(ctx, 'bash', { command: 'sleep 60', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
const killed = await call(ctx, 'bash_kill', { task_id: id })
expect(text(killed)).toBe(`killed background task ${id}`)
await ctx.bash.get(id)!.done
const again = await call(ctx, 'bash_kill', { task_id: id })
expect(text(again)).toBe(`task ${id} had already finished`)
const status = await call(ctx, 'bash_output', { task_id: id })
expect(text(status)).toContain('[status: killed by SIGTERM]')
})
it('unknown task ids are isError for both tools', async () => {
const ctx = await setup()
const read = await call(ctx, 'bash_output', { task_id: 'bash-999' })
expect(read.isError).toBe(true)
expect(text(read)).toMatch(/unknown bash task/)
const kill = await call(ctx, 'bash_kill', { task_id: 'bash-999' })
expect(kill.isError).toBe(true)
})
it.each([
['bash_output', {}],
['bash_output', { task_id: 9 }],
['bash_kill', { task_id: '' }],
])('%s rejects invalid task_id %j', async (tool, args) => {
const ctx = await setup()
const result = await call(ctx, tool, args)
expect(result.isError).toBe(true)
expect(text(result)).toMatch(/invalid task_id/)
})
it('injects a completion notice into the owning agent', async () => {
const ctx = await setup()
const inject = vi.fn()
const agent = { inject } as unknown as import('@deepseek-ai/dsh-agent').Agent
const started = await ctx.tools.execute({
callId: CallId('call-bg'),
name: 'bash',
arguments: { command: 'true', description: 'test command', run_in_background: true },
agent,
})
const id = /task (bash-\d+)/.exec(text(started))![1]!
await ctx.bash.get(id)!.done
expect(inject).toHaveBeenCalledTimes(1)
const [content, options] = inject.mock.calls[0] as [
{ type: string; text: string }[],
{ source: { kind: string; plugin: string } },
]
expect(content[0]!.text).toContain(`background bash task ${id} finished`)
expect(content[0]!.text).toContain('bash_output')
expect(options.source).toEqual({ kind: 'plugin', plugin: 'tool-bash' })
})
it('swallows ONLY the disposed-agent inject error', async () => {
const ctx = await setup()
const agent = {
inject: () => { throw new Error('agent "x" is disposed') },
} as unknown as import('@deepseek-ai/dsh-agent').Agent
const started = await ctx.tools.execute({
callId: CallId('call-bg2'),
name: 'bash',
arguments: { command: 'true', description: 'test command', run_in_background: true },
agent,
})
const id = /task (bash-\d+)/.exec(text(started))![1]!
await expect(ctx.bash.get(id)!.done).resolves.toBeUndefined()
})
it('rethrows a non-disposed inject failure (not blindly swallowed)', async () => {
const ctx = await setup()
// A real bug in inject (not the benign disposed race) must surface — the
// base-class notifier contains it (logs, does not reject task.done), but
// the listener itself must have thrown rather than silently eaten it.
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => undefined)
try {
const agent = {
inject: () => { throw new Error('unexpected inject bug') },
} as unknown as import('@deepseek-ai/dsh-agent').Agent
const started = await ctx.tools.execute({
callId: CallId('call-bg3'),
name: 'bash',
arguments: { command: 'true', description: 'test command', run_in_background: true },
agent,
})
const id = /task (bash-\d+)/.exec(text(started))![1]!
await ctx.bash.get(id)!.done
// notifyTaskDone caught and logged the rethrown error.
expect(errorSpy).toHaveBeenCalled()
const logged = errorSpy.mock.calls.flat().some(arg => arg instanceof Error && arg.message === 'unexpected inject bug')
expect(logged).toBe(true)
} finally {
errorSpy.mockRestore()
}
})
it('does not notify when no agent owned the task', async () => {
const ctx = await setup()
const started = await call(ctx, 'bash', { command: 'true', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
await expect(ctx.bash.get(id)!.done).resolves.toBeUndefined()
})
})
describe('renderResult', () => {
const base = {
exitCode: 0 as number | null,
signal: null as NodeJS.Signals | null,
timedOut: false,
aborted: false,
timeoutMs: 1000,
stdout: { text: '', truncated: false },
stderr: { text: '', truncated: false },
}
it('renders stderr-only output without a stdout prefix', () => {
expect(renderResult({ ...base, stderr: { text: 'err\n', truncated: false } }))
.toBe('[stderr]\nerr\n')
})
it('adds a separator when stdout does not end with a newline', () => {
expect(renderResult({
...base,
stdout: { text: 'out', truncated: false },
stderr: { text: 'err', truncated: false },
})).toBe('out\n[stderr]\nerr')
})
it('appends exit-code markers after a newline for unterminated output', () => {
expect(renderResult({ ...base, exitCode: 7, stdout: { text: 'x', truncated: false } }))
.toBe('x\n[exit code: 7]')
})
it('renders signal kills without the timeout marker when not timed out', () => {
expect(renderResult({ ...base, exitCode: null, signal: 'SIGKILL' }))
.toBe('(no output)\n[killed by signal: SIGKILL]')
})
it('reports a timeout that exited 0 (trapped signal) without a kill marker', () => {
expect(renderResult({ ...base, exitCode: 0, signal: null, timedOut: true }))
.toBe('(no output)\n[timed out after 1000ms]')
})
it('orders the timeout marker before a kill marker', () => {
expect(renderResult({ ...base, exitCode: null, signal: 'SIGTERM', timedOut: true }))
.toBe('(no output)\n[timed out after 1000ms]\n[killed by signal: SIGTERM]')
})
it('notes truncation with a fallback when the spill path is missing', () => {
expect(renderResult({ ...base, stdout: { text: 'tail', truncated: true } }))
.toBe('tail\n[output truncated; full output: (unavailable)]')
})
})
describe('status lines', () => {
it('reports kills without a recorded signal (executor raced process exit)', async () => {
const ctx = await setup()
const started = await call(ctx, 'bash', { command: 'sleep 60', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
const task = ctx.bash.get(id)!
await call(ctx, 'bash_kill', { task_id: id })
await task.done
// Simulate the variant where the close event carried no signal.
task.signal = null
const read = await call(ctx, 'bash_output', { task_id: id })
expect(text(read)).toContain('[status: killed]')
})
it('reports completed tasks with a null exit code as exit 0', async () => {
const ctx = await setup()
const started = await call(ctx, 'bash', { command: 'true', description: 'test command', run_in_background: true })
const id = /task (bash-\d+)/.exec(text(started))![1]!
const task = ctx.bash.get(id)!
await task.done
// Defensive: completed tasks always carry an exit code in practice; the
// ?? 0 fallback covers task shapes from other executor implementations.
task.exitCode = null
const read = await call(ctx, 'bash_output', { task_id: id })
expect(text(read)).toContain('[status: completed, exit code: 0]')
})
})
+16
View File
@@ -0,0 +1,16 @@
{
"extends": "../../tsconfig.base.json",
"compilerOptions": {
"rootDir": "src",
"outDir": "lib"
},
"include": ["src"],
"references": [
{ "path": "../../vendor/cosmokit" },
{ "path": "../../vendor/cordis" },
{ "path": "../llm" },
{ "path": "../tools" },
{ "path": "../agent" },
{ "path": "../bash" }
]
}