3.7 KiB
@deepseek-ai/dsh-llm-retry
Function plugin that applies exact-provider retry policy on the agent loop's closed-step recovery seam. It does not wrap ctx.llm.stream(): every adapter call remains one provider attempt, and every retry opens a fresh numbered step.
Each provider adapter owns an optional nested retryPolicy, captured when its route registers on ctx.llm. Omission uses normal mode: two retries for RATE_LIMIT, SERVER, TIMEOUT, and TRANSPORT. A normal policy can change its finite budget, eligible codes, and backoff. Always mode asks downstream recovery first, then retries every model-request failure without an attempt limit; success, cancellation, or plugin disposal stops it.
Both modes use bounded exponential backoff with symmetric jitter. A valid providerRetryAfterMs at or below maxDelayMs replaces local backoff without jitter. An over-cap provider delay makes normal mode delegate, while always mode uses its configured local backoff so it cannot terminate on that instruction.
Before waiting, the plugin appends a non-surface llm/retry event with the provider, mode, failure, and scheduled delay. Normal events include the finite maximum; always events omit it, and UIs render ∞. Cancellation and plugin disposal abort the wait; disposal drains active backoffs, and a callback captured before disposal fails closed.
The separately published ./invariant companion checks that every retry record names the current open turn and latest closed step, matches the failed request's durable provider, has a unique step record and correct provider-policy retry number, and carries a valid mode-specific budget and bounded timer delay. Full jitter may schedule zero milliseconds at its lower boundary.
- name: '@deepseek-ai/dsh-llm-deepseek'
config:
apiKey: !!js process.env.DEEPSEEK_API_KEY
retryPolicy:
mode: always
backoff:
initialDelayMs: 1000
maxDelayMs: 30000
jitterRatio: 0.2
- name: '@deepseek-ai/dsh-llm-retry'
The executor has no policy config. Multi-provider adapters such as dsh-llm-pi-ai place retryPolicy inside each provider profile, avoiding a second provider-name list.
Model Experience
Model-request recovery
What the model sees
No retry event, delay, provider error, or failed partial output is model-visible. The next numbered step reconstructs the same explicit provider/model request from durable surface history unless a downstream recovery policy deliberately changes that surface.
Token effect
Each retry is a new provider request and may repeat input-token billing. Normal mode has a finite budget; always mode can consume unbounded requests until success or cancellation. llm/retry itself contributes no tokens.
KV Cache effect
The reconstructed request preserves the prior prefix and is eligible for provider cache reuse under that provider's rules. The non-surface retry event does not change cache identity.
Known Limitations and Deferred Work
- Agent steps are the only retry boundary — direct
ctx.llm.stream()consumers remain single-attempt because a raw stream cannot separate already-emitted chunks durably. - Always mode retries permanent failures — authentication, quota, invalid-request, protocol, and unrecoverable context errors continue until success, cancellation, or disposal; deployments own provider-specific cost and latency controls.
- Recovery policies compose by waterfall order — always mode accepts a downstream retry before applying its fallback. A later policy that never settles also prevents the fallback from running.
llm/retryrecords scheduling, not completion — later step and turn events establish success, exhaustion, or cancellation.