feat(llm): add per-provider retry policies

This commit is contained in:
Turtle
2026-07-25 10:18:16 +08:00
parent fe84b446b0
commit b58e33268a
60 changed files with 1606 additions and 334 deletions
+4 -3
View File
@@ -10,13 +10,14 @@ An adapter registry plus a single streaming call surface, interceptable via a wa
- `ctx.llm.registerAdapter(providers: string[], adapter: LlmAdapter): () => void` Register one adapter instance for the given provider routes. Registration is all-or-nothing, and is disposed with the calling fiber.
- `ctx.llm.listProviders(): LlmProviderInfo[]` Describe registered provider routes in registration order.
- `ctx.llm.providerRetryPolicy(provider: string): ResolvedRetryPolicy` Return the provider-owned retry policy captured during registration, with normal defaults resolved.
- `ctx.llm.listModels(provider: string): Promise<LlmModelInfo[]>` Discover the models one registered provider currently advertises.
- `ctx.llm.resolveModelContext(provider: string, model: string): Promise<LlmModelContext | undefined>` Resolve authoritative context capacity for one exact route from its owning adapter.
- `ctx.llm.stream(options: GenerateOptions): AsyncIterable<StreamChunk>` Stream one model call as raw chunks (token-level deltas). Consumers assemble the chunks into blocks/messages with `BlockAssembler`.
`LlmService` preserves errors from final adapter selection, synchronous dispatch, iterator construction, and iteration, and binds their provenance to the exact stream handle returned for that model call. `isLlmAdapterFailure(stream, value)` reports only errors from that call's final adapter boundary; `llmFailureOf(stream, value)` returns the adjacent immutable `LlmFailure`. Nested model calls, `llm/stream` middleware, and downstream consumer failures remain unclassified for the outer call. Classification never replaces or mutates the adapter's original coded `Error`.
Provider and model metadata is a discovery surface, not a routing whitelist. `registerAdapter()` still owns provider exclusivity, while an adapter may accept model ids absent from `listModels()`; consumers must not reject a request because its model is unlisted. Returned metadata is detached and invalid or duplicate adapter entries fail with `INVALID_ADAPTER` or `INVALID_CATALOG`.
Provider and model metadata is a discovery surface, not a routing whitelist. `registerAdapter()` still owns provider exclusivity and captures the adapter's retry policy for each route, while an adapter may accept model ids absent from `listModels()`; consumers must not reject a request because its model is unlisted. Returned selector metadata is detached and invalid or duplicate adapter entries fail with `INVALID_ADAPTER` or `INVALID_CATALOG`.
Context capacity is a separate correctness query, not a catalog decoration or global LLM setting. `resolveModelContext()` asks the adapter that owns the exact provider/model route; an adapter can describe an unlisted dynamic model, and `undefined` means only that capacity is unavailable. Invalid returned capacity fails with `INVALID_MODEL_CONTEXT`.
@@ -28,7 +29,7 @@ Context capacity is a separate correctness query, not a catalog decoration or gl
### Extension points
- Subclass `LlmAdapter` and call `ctx.llm.registerAdapter(providers, adapter)` to add one or more provider routes. `GenerateOptions.provider` selects the adapter; `GenerateOptions.model` is adapter-owned and may be resolved dynamically. Override `providerInfo()` and asynchronous `listModels()` to expose selector metadata, and `resolveModelContext()` when exact capacity is known; the defaults use the route id as its name, advertise no models, and return no capacity.
- Subclass `LlmAdapter` and call `ctx.llm.registerAdapter(providers, adapter)` to add one or more provider routes. `GenerateOptions.provider` selects the adapter; `GenerateOptions.model` is adapter-owned and may be resolved dynamically. Override `providerRetryPolicy()` to supply provider-owned recovery configuration, `providerInfo()` and asynchronous `listModels()` to expose selector metadata, and `resolveModelContext()` when exact capacity is known; the defaults use bounded normal retry policy, use the route id as its name, advertise no models, and return no capacity.
- Wrap `llm/stream` via `ctx.on()` waterfall listeners for caching, logging, or routing. A wrapper that retries after emitting a chunk has no durable attempt boundary; shipped agent retry policy therefore uses `agent/request-error` instead.
### Content-block vocabulary (`types.ts`)
@@ -69,7 +70,7 @@ Pass-through; the registry preserves the assembled request prefix, while the sel
## Known Limitations and Deferred Work
- **No default retry/caching/rate-limit policy ships in this service** — `llm/stream` remains a single-attempt call-wrapper seam; the agent loop separately offers proven model-request failures to `agent/request-error`, whose default preserves the original failure. `@deepseek-ai/dsh-llm-retry` is an optional policy plugin loaded by the shared example spine.
- **No retry execution, caching, or rate limiting ships in this service** — provider registration stores retry policy, but `llm/stream` remains a single-attempt call-wrapper seam. The agent loop separately offers proven model-request failures to `agent/request-error`, whose default preserves the original failure; `@deepseek-ai/dsh-llm-retry` is the optional executor loaded by the shared example spine.
- **`GenerateOptions` sampling is `temperature`/`maxTokens`/`stop` only** — no `tool_choice`, `top_p`, or penalty fields; the vocabulary grows when a producer lands ([dropped inert knobs](../../../.agents/notes/implemented/simplification/2026-07-04-drop-inert-request-knobs.md)).
- **Producer-gated variants stay out until produced** — `prefill`, per-tool `strict`, block `cache` hints, and the `agent` message-source variant were pruned as producerless ([Agent Note](../../../.agents/notes/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.md)).
- **`BlockAssembler` handles core block kinds only** — a plugin-added block type whose stream is never closed by `block-end` makes `blocks()` throw.
+5
View File
@@ -38,11 +38,16 @@
"peerDependencies": {
"@deepseek-ai/dsh-brand": "^0.0.1",
"@deepseek-ai/dsh-invariants": "^0.0.1",
"@deepseek-ai/dsh-timeout": "^0.0.1",
"cordis": "^4.0.0-rc.7"
},
"dependencies": {
"schemastery": "^3.18.0"
},
"devDependencies": {
"@deepseek-ai/dsh-brand": "workspace:^",
"@deepseek-ai/dsh-invariants": "workspace:^",
"@deepseek-ai/dsh-timeout": "workspace:^",
"cordis": "^4.0.0-rc.7"
}
}
+43 -4
View File
@@ -16,6 +16,8 @@ import type {
Message,
StreamChunk,
} from './types.ts'
import { resolveRetryPolicy } from './retry-policy.ts'
import type { ResolvedRetryPolicy } from './retry-policy.ts'
import type { ProviderRequestId } from './brand.ts'
import { deepFreeze } from './call-config.ts'
import { HarnessError } from './error.ts'
@@ -27,6 +29,7 @@ export * from './brand.ts'
export * from './never.ts'
export * from './error.ts'
export * from './types.ts'
export * from './retry-policy.ts'
export { BlockAssembler } from './assembler.ts'
export { callConfigEquals, deepFreeze, isAgentLoopRequest, markAgentLoopRequest } from './call-config.ts'
export type { LlmCallConfig } from './call-config.ts'
@@ -119,6 +122,15 @@ export abstract class LlmAdapter {
return { id: provider, name: provider }
}
/**
* Return the provider-owned retry policy captured with this route.
* @param _provider - a route passed to `registerAdapter()` for this instance.
* @returns a resolved policy, or `undefined` to use the normal defaults.
*/
providerRetryPolicy(_provider: string): ResolvedRetryPolicy | undefined {
return undefined
}
/**
* List models this adapter can currently advertise for one owned provider.
* The result is advisory: an adapter may accept unlisted model ids, and
@@ -157,7 +169,11 @@ export abstract class LlmAdapter {
* surface, interceptable via the `llm/stream` waterfall.
*/
export class LlmService extends Service {
private adapters = new Map<string, { adapter: LlmAdapter; provider: LlmProviderInfo }>()
private adapters = new Map<string, {
adapter: LlmAdapter
provider: LlmProviderInfo
retryPolicy: ResolvedRetryPolicy
}>()
constructor(ctx: Context) {
super(ctx, 'llm')
@@ -175,7 +191,11 @@ export class LlmService extends Service {
const dispose = this.ctx.effect(function* (this: LlmService) {
if (providers.length === 0) throw new LlmError('an adapter must register at least one provider', 'INVALID_ADAPTER')
const unique = new Set<string>()
const registrations: { adapter: LlmAdapter; provider: LlmProviderInfo }[] = []
const registrations: {
adapter: LlmAdapter
provider: LlmProviderInfo
retryPolicy: ResolvedRetryPolicy
}[] = []
for (const provider of providers) {
if (provider.length === 0) throw new LlmError('adapter provider names must be non-empty', 'INVALID_ADAPTER')
if (unique.has(provider) || this.adapters.has(provider)) {
@@ -186,7 +206,13 @@ export class LlmService extends Service {
throw new LlmError(`adapter metadata for provider "${provider}" must preserve its id and have a non-empty name`, 'INVALID_ADAPTER')
}
unique.add(provider)
registrations.push({ adapter, provider: { id: info.id, name: info.name } })
const retryPolicy = adapter.providerRetryPolicy(provider)
?? resolveRetryPolicy(undefined, `llm: provider "${provider}" retryPolicy`)
registrations.push({
adapter,
provider: { id: info.id, name: info.name },
retryPolicy,
})
}
for (const registration of registrations) this.adapters.set(registration.provider.id, registration)
yield () => {
@@ -206,6 +232,15 @@ export class LlmService extends Service {
return [...this.adapters.values()].map(({ provider }) => ({ ...provider }))
}
/**
* Resolve the retry policy captured when one provider route was registered.
* @param provider - registered provider route to inspect.
* @returns the provider-owned policy, with normal defaults already resolved.
*/
providerRetryPolicy(provider: string): ResolvedRetryPolicy {
return this.registration(provider).retryPolicy
}
/**
* Discover models advertised by one registered provider. Catalog membership
* is advisory and never changes routing or request validation.
@@ -262,7 +297,11 @@ export class LlmService extends Service {
return { contextWindow: context.contextWindow }
}
private registration(provider: string): { adapter: LlmAdapter; provider: LlmProviderInfo } {
private registration(provider: string): {
adapter: LlmAdapter
provider: LlmProviderInfo
retryPolicy: ResolvedRetryPolicy
} {
const registration = this.adapters.get(provider)
if (!registration) throw new LlmError(`no adapter registered for provider "${provider}"`, 'NO_ADAPTER')
return registration
+184
View File
@@ -0,0 +1,184 @@
/**
* Provider-owned request-retry policy configuration and resolution.
*
* Adapters expose one resolved policy per registered provider route; the
* optional dsh-llm-retry plugin executes it on the agent's failed-step seam.
*
* @module @deepseek-ai/dsh-llm/retry-policy
*/
import z from 'schemastery'
import { MAX_TIMER_DELAY_MS } from '@deepseek-ai/dsh-timeout'
const DEFAULT_MAX_RETRIES = 2
const DEFAULT_INITIAL_DELAY_MS = 500
const DEFAULT_MAX_DELAY_MS = 10_000
const DEFAULT_JITTER_RATIO = 0.1
const DEFAULT_RETRYABLE_CODES = Object.freeze(['RATE_LIMIT', 'SERVER', 'TIMEOUT', 'TRANSPORT'])
/** Bounded exponential backoff with symmetric jitter around each local delay. */
export interface BackoffConfig {
/** Initial local exponential-backoff delay in milliseconds (default 500). */
initialDelayMs?: number
/** Maximum locally scheduled or accepted provider delay in milliseconds (default 10000). */
maxDelayMs?: number
/** Symmetric random multiplier range around one (default 0.1). */
jitterRatio?: number
}
/** Current bounded transient retry behavior for one provider route. */
export interface NormalRetryPolicyConfig {
/** Retry only configured transient failure codes. */
mode: 'normal'
/** Maximum eligible retries after the first request (default 2). */
maxRetries?: number
/** Stable failure codes eligible for this policy. */
retryableCodes?: string[]
/** Local exponential-backoff and jitter configuration. */
backoff?: BackoffConfig
}
/** Unbounded retry behavior for every model-request failure on one provider route. */
export interface AlwaysRetryPolicyConfig {
/** Retry every model-request failure until success, cancellation, or disposal. */
mode: 'always'
/** Local exponential-backoff and jitter configuration. */
backoff?: BackoffConfig
}
/** Provider-owned model-request retry policy configuration. */
export type RetryPolicyConfig = NormalRetryPolicyConfig | AlwaysRetryPolicyConfig
/** Fully resolved backoff shared by both retry modes. */
export interface ResolvedRetryBackoff {
readonly initialDelayMs: number
readonly maxDelayMs: number
readonly jitterRatio: number
}
/** Fully resolved bounded transient retry policy. */
export interface ResolvedNormalRetryPolicy extends ResolvedRetryBackoff {
readonly mode: 'normal'
readonly maxRetries: number
readonly retryableCodes: readonly string[]
}
/** Fully resolved unbounded retry policy. */
export interface ResolvedAlwaysRetryPolicy extends ResolvedRetryBackoff {
readonly mode: 'always'
}
/** Immutable provider policy captured when its adapter route is registered. */
export type ResolvedRetryPolicy = ResolvedNormalRetryPolicy | ResolvedAlwaysRetryPolicy
const backoffSchema: z<BackoffConfig> = z.object({
initialDelayMs: z.number().max(MAX_TIMER_DELAY_MS).default(DEFAULT_INITIAL_DELAY_MS),
maxDelayMs: z.number().max(MAX_TIMER_DELAY_MS).default(DEFAULT_MAX_DELAY_MS),
jitterRatio: z.number().min(0).max(1).default(DEFAULT_JITTER_RATIO),
})
const normalPolicySchema: z<NormalRetryPolicyConfig> = z.object({
mode: z.const('normal').required(),
maxRetries: z.number().step(1).min(0).max(Number.MAX_SAFE_INTEGER).default(DEFAULT_MAX_RETRIES),
retryableCodes: z.array(z.string()).default([...DEFAULT_RETRYABLE_CODES]),
backoff: backoffSchema,
})
const alwaysPolicySchema: z<AlwaysRetryPolicyConfig> = z.object({
mode: z.const('always').required(),
backoff: backoffSchema,
})
/** Cordis schema embedded by each concrete provider configuration. */
export const RetryPolicySchema: z<RetryPolicyConfig> = z.union([
normalPolicySchema,
alwaysPolicySchema,
])
const NORMAL_POLICY_KEYS: ReadonlySet<string> = new Set([
'mode', 'maxRetries', 'retryableCodes', 'backoff',
])
const ALWAYS_POLICY_KEYS: ReadonlySet<string> = new Set(['mode', 'backoff'])
const BACKOFF_KEYS: ReadonlySet<string> = new Set(['initialDelayMs', 'maxDelayMs', 'jitterRatio'])
function validateKeys(value: object, allowed: ReadonlySet<string>, path: string): void {
for (const key of Object.keys(value)) {
if (!allowed.has(key)) throw new Error(`${path}: unknown key "${key}"`)
}
}
function resolveBackoff(config: BackoffConfig | undefined, path: string): ResolvedRetryBackoff {
if (config !== undefined) validateKeys(config, BACKOFF_KEYS, path)
const initialDelayMs = config?.initialDelayMs ?? DEFAULT_INITIAL_DELAY_MS
const maxDelayMs = config?.maxDelayMs ?? DEFAULT_MAX_DELAY_MS
const jitterRatio = config?.jitterRatio ?? DEFAULT_JITTER_RATIO
if (!Number.isFinite(initialDelayMs) || initialDelayMs <= 0 || initialDelayMs > MAX_TIMER_DELAY_MS) {
throw new Error(`${path}.initialDelayMs must be a positive finite number no greater than ${MAX_TIMER_DELAY_MS}`)
}
if (!Number.isFinite(maxDelayMs) || maxDelayMs <= 0 || maxDelayMs > MAX_TIMER_DELAY_MS) {
throw new Error(`${path}.maxDelayMs must be a positive finite number no greater than ${MAX_TIMER_DELAY_MS}`)
}
if (initialDelayMs > maxDelayMs) {
throw new Error(`${path}.initialDelayMs must be less than or equal to maxDelayMs`)
}
if (!Number.isFinite(jitterRatio) || jitterRatio < 0 || jitterRatio > 1) {
throw new Error(`${path}.jitterRatio must be between 0 and 1`)
}
return Object.freeze({ initialDelayMs, maxDelayMs, jitterRatio })
}
/**
* Validate, default, and detach one provider-owned retry policy.
* @param config - optional provider configuration; omission selects normal defaults.
* @param path - diagnostic path naming the provider config that owns the value.
* @returns an immutable policy safe to capture in provider registration state.
*/
export function resolveRetryPolicy(
config: RetryPolicyConfig | undefined,
path: string,
): ResolvedRetryPolicy {
if (config === undefined) {
return Object.freeze({
mode: 'normal',
maxRetries: DEFAULT_MAX_RETRIES,
retryableCodes: DEFAULT_RETRYABLE_CODES,
...resolveBackoff(undefined, `${path}.backoff`),
})
}
switch (config.mode) {
case 'normal': {
validateKeys(config, NORMAL_POLICY_KEYS, path)
const maxRetries = config.maxRetries ?? DEFAULT_MAX_RETRIES
const retryableCodes = config.retryableCodes ?? [...DEFAULT_RETRYABLE_CODES]
if (!Number.isSafeInteger(maxRetries) || maxRetries < 0) {
throw new Error(`${path}.maxRetries must be a non-negative safe integer`)
}
if (retryableCodes.length === 0) {
throw new Error(`${path}.retryableCodes must not be empty`)
}
if (retryableCodes.some(code => code.length === 0)) {
throw new Error(`${path}.retryableCodes must contain only non-empty strings`)
}
if (new Set(retryableCodes).size !== retryableCodes.length) {
throw new Error(`${path}.retryableCodes must not contain duplicates`)
}
return Object.freeze({
mode: 'normal',
maxRetries,
retryableCodes: Object.freeze([...retryableCodes]),
...resolveBackoff(config.backoff, `${path}.backoff`),
})
}
case 'always':
validateKeys(config, ALWAYS_POLICY_KEYS, path)
return Object.freeze({
mode: 'always',
...resolveBackoff(config.backoff, `${path}.backoff`),
})
default:
throw new Error(`${path}.mode must be "normal" or "always"`)
}
}
@@ -0,0 +1,84 @@
import { describe, expect, it } from 'vitest'
import {
resolveRetryPolicy,
RetryPolicySchema,
} from '@deepseek-ai/dsh-llm'
import type { RetryPolicyConfig } from '@deepseek-ai/dsh-llm'
import { MAX_TIMER_DELAY_MS } from '@deepseek-ai/dsh-timeout'
describe('provider retry policy', () => {
it('resolves immutable normal defaults', () => {
const policy = resolveRetryPolicy(undefined, 'provider.retryPolicy')
expect(policy).toEqual({
mode: 'normal',
maxRetries: 2,
retryableCodes: ['RATE_LIMIT', 'SERVER', 'TIMEOUT', 'TRANSPORT'],
initialDelayMs: 500,
maxDelayMs: 10_000,
jitterRatio: 0.1,
})
expect(Object.isFrozen(policy)).toBe(true)
if (policy.mode !== 'normal') throw new Error('expected normal policy')
expect(Object.isFrozen(policy.retryableCodes)).toBe(true)
})
it('resolves and detaches a configured normal policy', () => {
const retryableCodes = ['BUSY']
const config: RetryPolicyConfig = {
mode: 'normal',
maxRetries: 4,
retryableCodes,
backoff: {
initialDelayMs: 25,
maxDelayMs: 100,
jitterRatio: 0,
},
}
const policy = resolveRetryPolicy(config, 'provider.retryPolicy')
retryableCodes.push('LATE')
expect(policy).toEqual({
mode: 'normal',
maxRetries: 4,
retryableCodes: ['BUSY'],
initialDelayMs: 25,
maxDelayMs: 100,
jitterRatio: 0,
})
})
it('resolves always mode with default backoff', () => {
expect(resolveRetryPolicy({ mode: 'always' }, 'provider.retryPolicy')).toEqual({
mode: 'always',
initialDelayMs: 500,
maxDelayMs: 10_000,
jitterRatio: 0.1,
})
expect(RetryPolicySchema).toBeDefined()
})
it.each([
[{ mode: 'normal', maxRetries: -1 }, /maxRetries/],
[{ mode: 'normal', maxRetries: 1.5 }, /maxRetries/],
[{ mode: 'normal', maxRetries: Number.MAX_SAFE_INTEGER + 1 }, /maxRetries/],
[{ mode: 'always', backoff: { initialDelayMs: 0 } }, /initialDelayMs/],
[{ mode: 'normal', backoff: { maxDelayMs: Number.POSITIVE_INFINITY } }, /maxDelayMs/],
[{ mode: 'normal', backoff: { initialDelayMs: MAX_TIMER_DELAY_MS + 1 } }, /initialDelayMs/],
[{ mode: 'always', backoff: { maxDelayMs: MAX_TIMER_DELAY_MS + 1 } }, /maxDelayMs/],
[{ mode: 'normal', backoff: { initialDelayMs: 20, maxDelayMs: 10 } }, /less than or equal/],
[{ mode: 'always', backoff: { jitterRatio: 1.1 } }, /jitterRatio/],
[{ mode: 'normal', retryableCodes: [] }, /must not be empty/],
[{ mode: 'normal', retryableCodes: ['SERVER', 'SERVER'] }, /duplicates/],
[{ mode: 'normal', retryableCodes: [''] }, /non-empty strings/],
[{ mode: 'normal', maxRetires: 1 }, /unknown key "maxRetires"/],
[{ mode: 'always', maxRetries: 1 }, /unknown key "maxRetries"/],
[{ mode: 'always', backoff: { initialDelay: 1 } }, /unknown key "initialDelay"/],
[{ mode: 'sometimes' }, /mode must be "normal" or "always"/],
] as const)('rejects invalid policy %#', (config, message) => {
expect(() => {
resolveRetryPolicy(config as unknown as RetryPolicyConfig, 'provider.retryPolicy')
}).toThrow(message)
})
})
+22
View File
@@ -11,6 +11,7 @@ import LlmService, {
LlmError,
llmFailureOf,
ProviderRequestId,
resolveRetryPolicy,
StreamChunk,
} from '@deepseek-ai/dsh-llm'
import type { LlmModelContext, LlmModelInfo, LlmProviderInfo } from '@deepseek-ai/dsh-llm'
@@ -159,6 +160,27 @@ describe('LlmService', () => {
expect(chunks).toEqual(SCRIPT)
})
it('captures provider-owned retry policy at registration and defaults omission', async () => {
const configured = resolveRetryPolicy({ mode: 'always' }, 'test retryPolicy')
const adapter = new class extends ScriptedAdapter {
override providerRetryPolicy(provider: string) {
return provider === 'configured' ? configured : undefined
}
}(SCRIPT)
const ctx = new Context()
await ctx.plugin(LlmService)
ctx.llm.registerAdapter(['configured', 'defaulted'], adapter)
expect(ctx.llm.providerRetryPolicy('configured')).toBe(configured)
expect(ctx.llm.providerRetryPolicy('defaulted')).toMatchObject({
mode: 'normal',
maxRetries: 2,
})
expect(() => ctx.llm.providerRetryPolicy('missing')).toThrow(
expect.objectContaining({ code: 'NO_ADAPTER' }),
)
})
it('throws NO_ADAPTER for unregistered providers', async () => {
const ctx = new Context()
await ctx.plugin(LlmService)
+3
View File
@@ -19,6 +19,9 @@
},
{
"path": "../../support/invariants"
},
{
"path": "../../util/timeout"
}
]
}