fix(tools): normalize the two class-name joins camelCase's own call misses

camelCase normalized `joined` and then prefixed, so the seam the `Tool`
prefix creates was never covered: `Tool` ends in `l`, a combining-mark
head composes with it, and a name headed by U+0301 was emitted as
`Tool` + U+0301 while CPython compiles `Too` + U+013A. childClassName
has the same shape -- both sides separately NFKC-stable, their join not:
a base ending in a Hangul L jamo or LV syllable composes with a V or T
jamo head. Beyond the declared-name/compiled-symbol mismatch, two
byte-distinct names can fold onto one, and usedClassNames dedupes by raw
bytes, so the collision counter never sees it. Normalize after the
prefix decision and at the join, before the cap. The remaining joins
need nothing: `Args`/`Output` and the digit suffix cannot compose
backwards.

Also record the Unicode-table skew. The predicate reads the engine's
tables (Node 22.23.1: 17.0) and the interpreter reads its own (CPython
3.9.6: 13.0.0), so an interpreter older than the engine takes a bare
name its tokenizer refuses -- U+1C89, U+10570, U+1E290 and U+1E4D0 are
accepted here and rejected there. The other direction only degrades a
legal name to subscript. Closing it needs the CPython floor, which the
backend PR owns; state the asymmetry in the docstring and make the
decision an explicit obligation in the note.
This commit is contained in:
Chinesezjc
2026-08-05 20:39:11 +08:00
parent 2cb0dddb40
commit 8c001d9928
5 changed files with 117 additions and 11 deletions
@@ -42,3 +42,5 @@ Adding a backend language is two table entries — an `SDK_RENDERERS` entry and
The cost is that the Python branch of both tables is unreachable on this base: `CodeRuntime.language` is set by the loaded backend, the only published backend is `dsh-code-runtime-worker` (`'typescript'`), and the registry reads the loaded runtime rather than a config field, so no assembled application can select `renderToolsSdkPy` or `PYTHON_FLAVOR`. The model-visible surface is therefore unchanged by this note's work until a backend reporting `'python'` is published, and this PR's coverage is unit-level — the renderer output plus the dispatch and rejection paths. The keyless snapshot for the Python model interface belongs to the PR that publishes that backend, because only there does a real `cordis.yml` over published plugins produce a Python assembly; a snapshot example that mounted a fixture runtime here would assert against a test double, which [docs/testing.md](../../../../docs/testing.md) rejects as a substitute for the assembled application transcript.
Two runtime contracts the Python SDK text asserts are owed by that same backend PR. First, the instructions tell the model that exactly `tools` and `ToolCallError` are bound and that the declared `TypedDict` classes are not, so the backend must inject those two names — with `ToolCallError.toolName` populated per the seam's `errorClass` contract — and must NOT bind the declared class names into the program's globals; injecting them "helpfully" would make the SDK text false. Second, the language has to be bound to the request: `requireCodeRuntime` resolves `ctx.codeRuntime` separately at assembly and at `run_code` execution, so a reload that swapped the runtime between those two points would hand a program written against one flavor to the other. The split is finer than those two points — `run_code`'s `description` and `parameters` getters each call `resolveFlavor(peekRuntime())`, and `schemaOf` destructures both, so one projection reads the runtime twice; both reads are for `run_code`'s own schema, since the getters are installed on that one definition and every other definition carries plain data properties. A reload between those two reads yields a single schema whose two halves name different languages. Neither is reachable here — one published backend means both reads return the same flavor and no program ever runs against this renderer's output — and the cross-language rejection is not testable until a second language exists.
Third, that PR owns the CPython floor, and with it the Unicode-table skew in `isBareIdentifier`. This renderer decides whether a field or tool name can be emitted bare using the running engine's `\p{XID_Start}`/`\p{XID_Continue}` tables (Node 22.23.1: Unicode 17.0), while the interpreter uses its own (CPython 3.9.6: 13.0.0). An interpreter older than the engine is the failing direction: a character added to `XID_Start` in between is emitted bare and its tokenizer refuses the whole block. The exposure window is exactly the characters added between the two versions, so the PR that names a supported CPython range must decide explicitly between accepting it and tightening the predicate against pinned tables for that floor. Nothing here can decide it: the floor does not exist yet, and a table pinned to a guess would be a deployment-varying constant with no configurability behind it.