From 73339739447904e1c343fefc1373a00271384f5f Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 10:07:12 +1000 Subject: [PATCH 01/13] docs(cache-key): serializer code must be derived, and per-identity (LAB-4351) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The spec's pseudocode already had `SERIALIZER_CODES[serializer_type]`, but it never said where `serializer_type` comes from, and cachekit-py read it from a parameter no caller on its main path passed — so every key ended `:1s` regardless of serializer, and two caches over one function differing only in serializer collided. Spelling out the invariant is what stops the next SDK reimplementing it. - Normative: the code MUST be derived from the configured serializer, never a fixed default. State the failure mode once — shared keyspace, each reader fails the other's serializer-name check, evicts, recomputes, 0% hit rate. - Normative: two identities the wire format records differently MUST NOT share a code, which is why the fallback is derived rather than a constant `x` bucket. `x` is now a prefix plus 2 bytes of blake2b over the identity. - Normative: one identity MUST always produce one code — derived from the same value the wire format records as the serializer name, never from a process-local value such as an object address or a randomised hash. - The reader-side serializer-name check is REQUIRED, not advisory: on a derivation collision it is the only remaining separator. Say plainly that `cache_key` is an AES-256-GCM AAD input, so two serializers sharing a code produce a byte-identical AAD — the cipher is NOT a backstop for a missing name check, which an implementor who knows AAD v0x03 might otherwise assume. - Table carries canonical names only and gains `l` (reference caching), which shipped in cachekit-py but was never documented. Alias spellings and the instance-vs-name distinction move to a clearly-marked Python SDK note, since neither is a cross-SDK concept. - Pseudocode: a `serializer_code()` function with the alias map beside the code table, replacing an unguarded lookup whose trailing comment still said the fallback was `s`. No test vector changes: every vector in `test-vectors/cache-keys.json` uses `serializer_type: "std"` and still expects `:1s`. Vectors for the newly reachable codes need the py fixture re-vendored against a merged revision of this file (it is sha256-pinned), so they follow separately rather than shipping here unverifiable. `cachekit-ts` and `cachekit-rs` do not implement the serializer code at all — they use the shorter `{ns}:{hash}` form — so nothing else in the org needs a matching change today. Refs: cachekit-io/cachekit-py --- spec/cache-key-format.md | 74 +++++++++++++++++++++++++++++++++------- 1 file changed, 62 insertions(+), 12 deletions(-) diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 172bc7d..3dbe478 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -6,7 +6,7 @@ **Deterministic key generation from function identity and arguments.** -*Protocol Version 1.0 · Verified against `cachekit-py` v0.12.0 (`src/cachekit/key_generator.py`)* +*Protocol Version 1.0 · Verified against `cachekit-py` v0.18.0 (`src/cachekit/key_generator.py`)* @@ -42,19 +42,53 @@ ns:{namespace}:func:{module}.{qualname}:args:{blake2b_hash}:{ic_flag}{serializer | `func:{module}.{qualname}:` | Function identifier (module path + qualified name) | `func:myapp.services.get_user:` | | `args:{blake2b_hash}:` | Blake2b-256 hash of normalized, MessagePack-serialized arguments | `args:a3c8d4...f2e1:` | | `{ic_flag}` | Integrity checking: `1` = ByteStorage enabled, `0` = raw MessagePack | `1` | -| `{serializer_code}` | Serializer type (1 char) | `s` | +| `{serializer_code}` | Serializer identity (1 char, or `x` + 4 hex — see below) | `s` | ### Serializer Codes -| Code | Serializer | Cross-language? | -| :---: | :--- | :---: | -| `s` | StandardSerializer (MessagePack) | ✅ Yes | -| `a` | AutoSerializer (Python-specific) | ❌ No | -| `o` | OrjsonSerializer (JSON-based) | ⚠️ Partial | -| `w` | ArrowSerializer (columnar) | ⚠️ Partial | +| Code | Serializer | Canonical name | Cross-language? | +| :---: | :--- | :--- | :---: | +| `s` | StandardSerializer (MessagePack) | `default` | ✅ Yes | +| `a` | AutoSerializer (language-specific types) | `auto` | ❌ No | +| `o` | OrjsonSerializer (JSON-based) | `orjson` | ⚠️ Partial | +| `w` | ArrowSerializer (columnar) | `arrow` | ⚠️ Partial | +| `l` | Reference caching (no serialization) | `local` | ❌ No | +| `x` + 4 hex | Any serializer identity not in this table | — | ❌ No | For cross-SDK interoperability, always use `s` (StandardSerializer). +An identity outside the table gets `x` followed by the first 2 bytes of +`blake2b(identity, digest_size=2)` as lowercase hex — a per-identity code, not a shared +bucket. Codes are therefore 1 character for the table entries and 5 for everything else. + +> [!IMPORTANT] +> **The code MUST be derived from the serializer the cache is configured with, never a +> fixed default.** An SDK that emits one constant code collapses every serializer onto a +> single keyspace: two caches over one function then share a key, each fails the other's +> serializer-name check on read, evicts, and recomputes — a permanent 0% hit rate. +> +> **Two serializer identities that the wire format records differently MUST NOT share a +> code**, which is why the fallback is derived rather than constant. Where a derivation +> collision does occur (≈1 in 2^16), the reader-side serializer-name check in the +> [wire format](wire-format.md) is the only remaining separator, so that check is REQUIRED, +> not advisory. +> +> `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers +> sharing a code therefore produce a **byte-identical AAD**: AAD binding does not separate +> them, and the cipher is not a backstop for a missing name check. +> +> **Conversely, one identity MUST always produce one code.** Derive it from the same value +> the wire format records as the serializer name, never from a process-local value such as +> an object address or a randomised hash, or keys stop being reproducible across processes. + +> [!NOTE] +> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings +> `std`/`standard` for `default` and `pythonic` for `auto`, canonicalizing them before the +> lookup. A serializer passed as an *instance* rather than a name is recorded under its class +> name — built-ins included — and so takes a derived `x`-prefixed code rather than the +> table's. Two instances of the same class are one identity: the SDK documents distinct +> namespaces for per-configuration serializers. + ### Example Keys ``` @@ -203,8 +237,24 @@ Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte- Expand full pseudocode ``` +SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "local": "l"} + +// Alias spellings this SDK accepts -> canonical name. Empty if the SDK accepts only +// canonical names. The SAME map must canonicalize the serializer name the wire format +// records, or a key and its stored envelope can disagree about which serializer wrote it. +SERIALIZER_ALIASES = {"std": "default", "standard": "default", "pythonic": "auto"} + +function serializer_code(serializer_type): + // Resolve any alias spelling this SDK accepts, then look the code up. An identity + // outside the table gets its OWN derived code — never a shared constant, which would + // put every unrecognised serializer on one keyspace. + canonical = SERIALIZER_ALIASES.get(serializer_type, serializer_type) + if canonical in SERIALIZER_CODES: + return SERIALIZER_CODES[canonical] + return "x" + hex(blake2b(canonical.utf8_bytes(), digest_size=2)) + function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, - integrity_checking=true, serializer_type="std"): + integrity_checking=true, serializer_type="default"): // Build key parts parts = [] @@ -222,10 +272,10 @@ function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, parts.append("args:" + hash + ":") - // Metadata suffix + // Metadata suffix. serializer_type is the serializer the cache is CONFIGURED with — + // read it from the decorator/client configuration, never a fixed default. ic_flag = "1" if integrity_checking else "0" - serializer_code = SERIALIZER_CODES[serializer_type] // "s" for standard - parts.append(ic_flag + serializer_code) + parts.append(ic_flag + serializer_code(serializer_type)) key = join(parts) From ebc56e8ed9c9a2aea95a05c5637761b7e2ba1fe6 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 13:47:11 +1000 Subject: [PATCH 02/13] =?UTF-8?q?fix:=20address=20coderabbit=20review=20?= =?UTF-8?q?=E2=80=94=20serializer-code=20provenance,=20collision=20guarant?= =?UTF-8?q?ee,=20read-side=20check,=20hex=20encoding?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Header pins the derivation's verification target to cachekit-py#311 @ 3fd7c35 (open); v0.18.0 emits the constant `s` suffix. Vectors: v0.12.0, unchanged, sha256-pinned in cachekit-py CI at v0.18.0 (Test Vectors + CHANGELOG). - "MUST NOT share a code" is now "by construction"; the 16-bit derived code's collision guarantee is stated as probabilistic and hit-rate-only, scoped to identities the wire format records differently. - Read-side serializer-name comparison is REQUIRED on every read, before decoding, scoped to containers that record a name; an absent name is a mismatch. cachekit-py#311 still exempts "unknown" — recorded as a known divergence in the CHANGELOG. - Derived-code encoding pinned: exactly 4 lowercase zero-padded hex chars, no prefix — the digest's .hex(), not a numeric hex() (blake2b("auto", 2) is 033b, so it matters). Worked examples: cbor → x23d5, :ArrowSerializer → x2263. - Review corrections: `standard` is not an accepted spelling (absent from SERIALIZER_REGISTRY) — dropped from the Python note and pseudocode; the hash input is named `identity` throughout (for an instance it is the `:`- prefixed class name, not the frame tag); MAY → can (RFC 2119). - CHANGELOG entry, per repo convention for every spec change. CodeRabbit-Resolved: spec/cache-key-format.md:9:Align the cache-key vector CodeRabbit-Resolved: spec/cache-key-format.md:62:Make the fallback guarantee CodeRabbit-Resolved: spec/cache-key-format.md:73:Make serializer-name validati CodeRabbit-Resolved: spec/cache-key-format.md:254:Specify the hexadecimal enco --- CHANGELOG.md | 26 ++++++++++++++++ spec/cache-key-format.md | 66 +++++++++++++++++++++++++--------------- 2 files changed, 68 insertions(+), 24 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index e1ce990..6e1626c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,32 @@ All notable changes to the CacheKit Protocol Specification. ## [Unreleased] +### Cache key — serializer code MUST be derived, and per-identity (LAB-4351) + +- [`spec/cache-key-format.md`](spec/cache-key-format.md): the `{serializer_code}` + suffix MUST be derived from the configured serializer, never a constant, and + two identities the wire format records differently MUST NOT be mapped onto + one code by construction. An identity outside the code table gets `x` + the + 2-byte `blake2b` digest of its UTF-8 identity as 4 lowercase hex characters; + the collision guarantee is stated as probabilistic (16 bits). The read-side + comparison of the container's recorded serializer name against the reader's + own is now REQUIRED on every read, before decoding, and a value recording no + name is a mismatch. The table gains `l` (reference caching — shipped, never + documented) and a `Canonical name` column; alias spellings and the instance + identity (`:` + class name) move to a Python SDK note. Documents the + defect corrected by + [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) + (open): through v0.18.0 every auto-mode key ended in `s` regardless of + serializer, so two caches over one function differing only in serializer + shared a key and evicted each other on every read. Known divergence: that + PR's read guard still exempts a frame that records no serializer name. +- Provenance: `test-vectors/cache-keys.json` is unchanged (version 1.0.0, + generated by `cachekit-py` v0.12.0, sha256 `4a0ae13d…86450e`) and is + byte-verified against `CacheKeyGenerator` in cachekit-py CI at v0.18.0 + (`tests/unit/protocol/test_cache_key_vectors.py`, pinned to protocol + `f0672c1c`). Every vector uses `std` → `1s`; vectors for the derived codes + need a released generator that emits them and follow separately. + ### SaaS API — `401` is an authoritative verdict; auth backend faults are `503` (LAB-4093) - [`spec/saas-api.md`](spec/saas-api.md) Error Handling: `401` is emitted only diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 3dbe478..4376726 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -6,7 +6,7 @@ **Deterministic key generation from function identity and arguments.** -*Protocol Version 1.0 · Verified against `cachekit-py` v0.18.0 (`src/cachekit/key_generator.py`)* +*Protocol Version 1.0 · Verified against `cachekit-py` v0.18.0 (`src/cachekit/key_generator.py`); serializer-code derivation verified against [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) @ `3fd7c35` (open); vectors generated by v0.12.0 (see [Test Vectors](#test-vectors))* @@ -57,9 +57,11 @@ ns:{namespace}:func:{module}.{qualname}:args:{blake2b_hash}:{ic_flag}{serializer For cross-SDK interoperability, always use `s` (StandardSerializer). -An identity outside the table gets `x` followed by the first 2 bytes of -`blake2b(identity, digest_size=2)` as lowercase hex — a per-identity code, not a shared -bucket. Codes are therefore 1 character for the table entries and 5 for everything else. +An identity outside the table gets `x` followed by the 2-byte digest +`blake2b(utf8(identity), digest_size=2)` encoded as exactly 4 lowercase hexadecimal +characters — zero-padded, no `0x` prefix, the same encoding as the args hash. Example: +identity `cbor` → `x23d5`. Codes are 1 character for the table entries and 5 for everything +else. > [!IMPORTANT] > **The code MUST be derived from the serializer the cache is configured with, never a @@ -67,27 +69,42 @@ bucket. Codes are therefore 1 character for the table entries and 5 for everythi > single keyspace: two caches over one function then share a key, each fails the other's > serializer-name check on read, evicts, and recomputes — a permanent 0% hit rate. > -> **Two serializer identities that the wire format records differently MUST NOT share a -> code**, which is why the fallback is derived rather than constant. Where a derivation -> collision does occur (≈1 in 2^16), the reader-side serializer-name check in the -> [wire format](wire-format.md) is the only remaining separator, so that check is REQUIRED, -> not advisory. +> **Two serializer identities that the wire format records differently MUST NOT be mapped +> onto one code by construction.** The guarantee is probabilistic, not absolute: the derived +> code carries 16 bits, so two identities it records differently can still collide, at ≈1 in +> 2^16 per pair. Such a collision costs hit rate only — that one pair evicts each other +> exactly as a constant code makes every pair do — and never yields a wrong value, because +> the stored serializer name still differs and the read-side check below rejects it. +> +> **On every read, before decoding the payload, the reader MUST compare the serializer name +> its storage container records (Python: the `s` field of the +> [CK v3 frame](wire-format.md#python-ck-v3-frame)) with the name it would itself record for +> its configured serializer, and MUST reject the entry on mismatch — a miss, never a value. A +> value that records no serializer name is a mismatch.** A colliding entry has the same key as +> the reader's own, so nothing before this comparison can tell them apart. An SDK that offers +> more than one serializer identity MUST record the name in its container, or this check +> cannot exist (`cachekit-ts` and `cachekit-rs` record none and offer one). > > `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers > sharing a code therefore produce a **byte-identical AAD**: AAD binding does not separate > them, and the cipher is not a backstop for a missing name check. > -> **Conversely, one identity MUST always produce one code.** Derive it from the same value -> the wire format records as the serializer name, never from a process-local value such as -> an object address or a randomised hash, or keys stop being reproducible across processes. +> **Conversely, one identity MUST always produce one code.** Derive it from the serializer +> configuration alone — the canonical name the wire format records, or an SDK-defined +> refinement of it that never merges two names (see the Python note below; a refinement +> separates keyspaces, the read-side check still sees only the recorded name) — never from a +> process-local value such as an object address or a randomised hash, or keys stop being +> reproducible across processes. > [!NOTE] -> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings -> `std`/`standard` for `default` and `pythonic` for `auto`, canonicalizing them before the -> lookup. A serializer passed as an *instance* rather than a name is recorded under its class -> name — built-ins included — and so takes a derived `x`-prefixed code rather than the -> table's. Two instances of the same class are one identity: the SDK documents distinct -> namespaces for per-configuration serializers. +> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` for +> `default` and `pythonic` for `auto`, canonicalizing them before the lookup. A serializer +> passed as an *instance* rather than a name is recorded in the frame header under its bare +> class name, built-ins included, and its key identity is `:` + that class name, so it +> takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → `x2263`), never the +> table's. The prefix contains characters no Python identifier can, so a custom class named +> `auto` cannot take AutoSerializer's code. Two instances of the same class are one identity: +> the SDK documents distinct namespaces for per-configuration serializers. ### Example Keys @@ -225,7 +242,7 @@ After key construction, the following characters are replaced: ## Test Vectors -[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; only the args-hash segment must match byte-for-byte. +[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`), so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; only the args-hash segment must match byte-for-byte. Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte-verified against `CacheKeyGenerator` on every default CI run (`tests/unit/protocol/test_cache_key_vectors.py`). A vector failing there is a key-stability break to triage — never silently regenerate: a changed key orphans every existing cache entry and turns the fleet's hits into billed misses. @@ -242,16 +259,17 @@ SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "l // Alias spellings this SDK accepts -> canonical name. Empty if the SDK accepts only // canonical names. The SAME map must canonicalize the serializer name the wire format // records, or a key and its stored envelope can disagree about which serializer wrote it. -SERIALIZER_ALIASES = {"std": "default", "standard": "default", "pythonic": "auto"} +SERIALIZER_ALIASES = {"std": "default", "pythonic": "auto"} function serializer_code(serializer_type): // Resolve any alias spelling this SDK accepts, then look the code up. An identity // outside the table gets its OWN derived code — never a shared constant, which would // put every unrecognised serializer on one keyspace. - canonical = SERIALIZER_ALIASES.get(serializer_type, serializer_type) - if canonical in SERIALIZER_CODES: - return SERIALIZER_CODES[canonical] - return "x" + hex(blake2b(canonical.utf8_bytes(), digest_size=2)) + identity = SERIALIZER_ALIASES.get(serializer_type, serializer_type) + if identity in SERIALIZER_CODES: + return SERIALIZER_CODES[identity] + // digest .hex(): exactly 4 lowercase zero-padded hex chars — never a numeric hex() + return "x" + blake2b(identity.utf8_bytes(), digest_size=2).hex() function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, integrity_checking=true, serializer_type="default"): From 4cfdd8072d6b392f7ed045e7a531f4ee9230954f Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 15:25:33 +1000 Subject: [PATCH 03/13] docs(cache-key): mark the alias map and identity normalisation SDK-supplied MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two review findings on the implementor pseudocode, both cross-SDK correctness. The block hard-coded cachekit-py's `std` and `pythonic` aliases. A non-Python SDK copying it would map those two spellings onto the `s` and `a` table codes even though its own API never accepts them — handing a name a table code where the derived `x` code is what the rule implies. The map is now empty with the Python pair shown as an example, and the comment says not to adopt another SDK's aliases. The Python note documents a serializer passed as an object, whose identity is `:` + its bare class name, but the pseudocode passed `serializer_type` straight into an alias lookup and then called `.utf8_bytes()` on it — undefined for an object, so an implementor following it literally cannot produce the documented `x2263`. A `normalize_identity()` hook now runs before alias and table lookup, defaulting to identity for a names-only SDK, carrying the purity constraint the one-identity-one-code rule already imposes, and naming the prefix as what stops a class called `auto` taking AutoSerializer's code. Matches the shipped derivation: the prefix is applied before the lookup, not inside it. Spec prose and changelog only; the reference verifiers are unaffected and pass. CodeRabbit-Resolved: cache-key-format.md:262:Keep Python aliases o CodeRabbit-Resolved: cache-key-format.md:272:Handle serializer ins --- CHANGELOG.md | 9 +++++++++ spec/cache-key-format.md | 32 ++++++++++++++++++++++++-------- 2 files changed, 33 insertions(+), 8 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6e1626c..52bb8e0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,15 @@ All notable changes to the CacheKit Protocol Specification. serializer, so two caches over one function differing only in serializer shared a key and evicted each other on every read. Known divergence: that PR's read guard still exempts a frame that records no serializer name. +- SDK-implementor pseudocode: the alias map and the identity-normalisation step + are marked SDK-supplied rather than shown as fixed. The block previously + hard-coded cachekit-py's `std`/`pythonic` aliases, which a non-Python SDK + copying it would have used to hand those two names a table code instead of + the derived `x` code its own API implies; and it passed `serializer_type` + straight into a string lookup, leaving the documented instance identity + (`:` + class name) with no place to be produced. A + `normalize_identity()` hook now runs before alias and table lookup, with the + purity constraint the one-identity-one-code rule already requires. - Provenance: `test-vectors/cache-keys.json` is unchanged (version 1.0.0, generated by `cachekit-py` v0.12.0, sha256 `4a0ae13d…86450e`) and is byte-verified against `CacheKeyGenerator` in cachekit-py CI at v0.18.0 diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 4376726..d67a54d 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -256,16 +256,32 @@ Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte- ``` SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "local": "l"} -// Alias spellings this SDK accepts -> canonical name. Empty if the SDK accepts only -// canonical names. The SAME map must canonicalize the serializer name the wire format -// records, or a key and its stored envelope can disagree about which serializer wrote it. -SERIALIZER_ALIASES = {"std": "default", "pythonic": "auto"} +// SDK-SUPPLIED, not fixed by this spec: alias spellings THIS SDK accepts -> canonical +// name. Empty if the SDK accepts only canonical names. The SAME map must canonicalize the +// serializer name the wire format records, or a key and its stored envelope can disagree +// about which serializer wrote it. Do not adopt another SDK's aliases: mapping a spelling +// your API does not accept hands that name a table code instead of the derived `x` code it +// should get. cachekit-py's map is in the Python note above. +SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "auto"} + +// SDK-SUPPLIED, not fixed by this spec: reduce whatever your API accepts as a serializer to +// the canonical STRING identity, before any lookup below. An SDK that accepts only names +// returns the name unchanged. One that also accepts a serializer OBJECT must convert it +// here — the lookups below are string operations and are undefined on an object. This is +// the "SDK-defined refinement" the one-identity-one-code rule permits, so it must be a pure +// function of the serializer's configuration, never of an object address or a randomised +// hash. cachekit-py maps an object to ":" + its bare class name (Python note +// above); the prefix uses characters no identifier can contain, which is what stops a class +// named `auto` from taking AutoSerializer's code. +function normalize_identity(serializer_type): + return serializer_type // names-only SDK; override to handle objects function serializer_code(serializer_type): - // Resolve any alias spelling this SDK accepts, then look the code up. An identity - // outside the table gets its OWN derived code — never a shared constant, which would - // put every unrecognised serializer on one keyspace. - identity = SERIALIZER_ALIASES.get(serializer_type, serializer_type) + // Reduce to a string identity, resolve any alias spelling this SDK accepts, then look + // the code up. An identity outside the table gets its OWN derived code — never a shared + // constant, which would put every unrecognised serializer on one keyspace. + name = normalize_identity(serializer_type) + identity = SERIALIZER_ALIASES.get(name, name) if identity in SERIALIZER_CODES: return SERIALIZER_CODES[identity] // digest .hex(): exactly 4 lowercase zero-padded hex chars — never a numeric hex() From aa4de19caf7d9638ff5b23799b9e1096c3179ba0 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 15:36:39 +1000 Subject: [PATCH 04/13] docs(cache-key): state that the vectors' "std" is read as canonical "default" MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Emptying SERIALIZER_ALIASES in the previous commit left a conformance trap that review caught: all 10 vectors in test-vectors/cache-keys.json record serializer_type "std" and expect the suffix ":1s", but "std" is cachekit-py's alias. With an empty map an SDK following the pseudocode literally treats it as an unaliased identity, derives "x" + blake2b("std"), and fails every vector — while the same pseudocode tells it not to adopt another SDK's aliases. The two instructions contradicted each other with no way out. Both the pseudocode and the Test Vectors section now say the vectors' "std" is to be READ as the canonical "default", not added to the reader's alias map. Verified: 10/10 vectors carry serializer_type "std" and an expected_key ending ":1s". Reference verifiers unaffected and passing. Kody-Resolved: cache-key-format.md:265:Mismatch between SERIAL --- CHANGELOG.md | 7 ++++++- spec/cache-key-format.md | 7 ++++++- 2 files changed, 12 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 52bb8e0..b1adae2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -31,7 +31,12 @@ All notable changes to the CacheKit Protocol Specification. straight into a string lookup, leaving the documented instance identity (`:` + class name) with no place to be produced. A `normalize_identity()` hook now runs before alias and table lookup, with the - purity constraint the one-identity-one-code rule already requires. + purity constraint the one-identity-one-code rule already requires. Because the + shipped vectors record `serializer_type: "std"` — a cachekit-py alias — both + the pseudocode and the Test Vectors section now state that a conforming SDK + reads that as the canonical `default` rather than adding the alias to its own + map; read as an unaliased identity it derives `x` + `blake2b("std")` and fails + all 10 vectors. - Provenance: `test-vectors/cache-keys.json` is unchanged (version 1.0.0, generated by `cachekit-py` v0.12.0, sha256 `4a0ae13d…86450e`) and is byte-verified against `CacheKeyGenerator` in cachekit-py CI at v0.18.0 diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index d67a54d..a281747 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -242,7 +242,7 @@ After key construction, the following characters are replaced: ## Test Vectors -[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`), so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; only the args-hash segment must match byte-for-byte. +[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`) — cachekit-py's alias for the canonical `default`, which an SDK that does not accept that spelling MUST read as `default` rather than add to its own alias map — so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; only the args-hash segment must match byte-for-byte. Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte-verified against `CacheKeyGenerator` on every default CI run (`tests/unit/protocol/test_cache_key_vectors.py`). A vector failing there is a key-stability break to triage — never silently regenerate: a changed key orphans every existing cache entry and turns the fleet's hits into billed misses. @@ -263,6 +263,11 @@ SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "l // your API does not accept hands that name a table code instead of the derived `x` code it // should get. cachekit-py's map is in the Python note above. SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "auto"} +// CONFORMANCE: test-vectors/cache-keys.json records serializer_type "std", which is +// cachekit-py's alias for the canonical "default". An SDK that does not accept "std" MUST +// read those vectors as serializer_type = "default" (-> code "s"); it MUST NOT add the alias +// to its own map to pass them. Read as an unaliased identity, "std" derives +// "x" + blake2b("std") and fails all 10 vectors. // SDK-SUPPLIED, not fixed by this spec: reduce whatever your API accepts as a serializer to // the canonical STRING identity, before any lookup below. An SDK that accepts only names From 3a59b5666e6f36bbc27a6e2725a596f7b98000ba Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 19:03:32 +1000 Subject: [PATCH 05/13] docs(cache-key): fix Test Vectors byte-for-byte matching scope The paragraph's scope statement said only the args-hash segment must match byte-for-byte, contradicting the MUST introduced earlier in the same sentence (an SDK that doesn't accept "std" MUST read it as canonical "default", yielding code "s"). An SDK following the stated scope never compares the `{ic_flag}{serializer_code}` suffix that MUST governs, so the MUST was unenforceable and the vectors couldn't detect a wrong serializer code. Now both the args-hash segment and the trailing `{ic_flag}{serializer_code}` suffix must match byte-for-byte; only the `func:` segment is substituted per SDK. Kody-Resolved: spec/cache-key-format.md:245:Test Vectors byte-for-byte scope --- spec/cache-key-format.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index a281747..107cd6c 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -242,7 +242,7 @@ After key construction, the following characters are replaced: ## Test Vectors -[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`) — cachekit-py's alias for the canonical `default`, which an SDK that does not accept that spelling MUST read as `default` rather than add to its own alias map — so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; only the args-hash segment must match byte-for-byte. +[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`) — cachekit-py's alias for the canonical `default`, which an SDK that does not accept that spelling MUST read as `default` rather than add to its own alias map — so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; the args-hash segment and the trailing `{ic_flag}{serializer_code}` suffix must both match byte-for-byte. Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte-verified against `CacheKeyGenerator` on every default CI run (`tests/unit/protocol/test_cache_key_vectors.py`). A vector failing there is a key-stability break to triage — never silently regenerate: a changed key orphans every existing cache entry and turns the fleet's hits into billed misses. From 1c1c51b1111e499cf42090120aa446d25e803262 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Mon, 21 Sep 2026 19:47:41 +1000 Subject: [PATCH 06/13] docs(cache-key): scope vector byte-match to implemented suffix; require identity injectivity MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Test Vectors: the byte-for-byte match requirement is unconditional only for the args-hash segment. The {ic_flag}{serializer_code} suffix must match too, but only for an SDK that implements serializer codes at all — cachekit-ts and cachekit-rs do not, so an unscoped requirement was unsatisfiable for them and contradicted the vendored fixture's own note field. - normalize_identity(): the purity constraint now also requires injectivity — two configurations that produce different container bytes MUST NOT normalize to the same identity, since the derived code and the frame tag both come from it and a collapsed identity defeats the read-side mismatch check, serving wrong data instead of a miss. An object-to-string refinement MUST use a marker no bare identity can produce, promoted from a cachekit-py-specific example to a requirement on the hook's contract. --- CHANGELOG.md | 12 ++++++++++++ spec/cache-key-format.md | 23 +++++++++++++++++------ 2 files changed, 29 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1adae2..4017947 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -43,6 +43,18 @@ All notable changes to the CacheKit Protocol Specification. (`tests/unit/protocol/test_cache_key_vectors.py`, pinned to protocol `f0672c1c`). Every vector uses `std` → `1s`; vectors for the derived codes need a released generator that emits them and follow separately. +- Test Vectors byte-for-byte scope corrected: the args-hash segment is the + unconditional match requirement; the `{ic_flag}{serializer_code}` suffix is + only required to match for an SDK that implements serializer codes at all + (`cachekit-ts` and `cachekit-rs` do not). An unscoped requirement was + unsatisfiable for those two and contradicted the fixture's own `note` field. +- `normalize_identity()`: the purity constraint gains an explicit injectivity + requirement — two configurations that change the serialized container's + bytes MUST NOT normalize to the same identity, since a collapsed identity + makes the derived code and the frame tag agree and defeats the read-side + mismatch check. An object-to-string refinement MUST use a marker no bare + identity can produce, promoted from a cachekit-py-specific example to a + requirement on the hook's contract. ### SaaS API — `401` is an authoritative verdict; auth backend faults are `503` (LAB-4093) diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 107cd6c..a7da334 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -242,7 +242,7 @@ After key construction, the following characters are replaced: ## Test Vectors -[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`) — cachekit-py's alias for the canonical `default`, which an SDK that does not accept that spelling MUST read as `default` rather than add to its own alias map — so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; the args-hash segment and the trailing `{ic_flag}{serializer_code}` suffix must both match byte-for-byte. +[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and have not changed since; every vector uses `serializer_type: "std"` (→ `1s`) — cachekit-py's alias for the canonical `default`, which an SDK that does not accept that spelling MUST read as `default` rather than add to its own alias map — so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}` — cross-SDK implementations substitute their own module path; the args-hash segment MUST match byte-for-byte. An SDK that implements the `{ic_flag}{serializer_code}` suffix MUST match it too — `cachekit-ts` and `cachekit-rs` implement no serializer codes, so that clause does not apply to them. Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte-verified against `CacheKeyGenerator` on every default CI run (`tests/unit/protocol/test_cache_key_vectors.py`). A vector failing there is a key-stability break to triage — never silently regenerate: a changed key orphans every existing cache entry and turns the fleet's hits into billed misses. @@ -273,11 +273,22 @@ SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "a // the canonical STRING identity, before any lookup below. An SDK that accepts only names // returns the name unchanged. One that also accepts a serializer OBJECT must convert it // here — the lookups below are string operations and are undefined on an object. This is -// the "SDK-defined refinement" the one-identity-one-code rule permits, so it must be a pure -// function of the serializer's configuration, never of an object address or a randomised -// hash. cachekit-py maps an object to ":" + its bare class name (Python note -// above); the prefix uses characters no identifier can contain, which is what stops a class -// named `auto` from taking AutoSerializer's code. +// the "SDK-defined refinement" the one-identity-one-code rule permits. +// +// `normalize_identity()` MUST be a pure function of the serializer's configuration — never +// an object address or a randomised hash — and MUST be injective over any configuration +// dimension that changes the serialized container's bytes: e.g. `arrow+gzip` and +// `arrow+none` MUST NOT both normalize to `"arrow"`. The derived code and the frame tag are +// both computed from this identity, so collapsing two differently-configured serializers +// onto one identity makes them self-report the identical frame tag — the read-side mismatch +// check above can no longer tell them apart, and a mismatched container is served as a hit +// instead of a miss: wrong data, not a recoverable eviction. +// +// A refinement that maps a serializer OBJECT to a string MUST use a marker that no bare +// identity can produce, so a user-named class can never be spelled as a table key. +// cachekit-py maps an object to ":" + its bare class name (Python note above); the +// prefix uses characters no identifier can contain, which is what stops a class named +// `auto` from taking AutoSerializer's code. function normalize_identity(serializer_type): return serializer_type // names-only SDK; override to handle objects From 77c52e1a11ca2f90a3c10b7312b48c44e6ba714d Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sat, 26 Sep 2026 01:53:52 +1000 Subject: [PATCH 07/13] docs(cache-key): separate key identity from recorded serializer name The pseudocode said the derived code and the frame tag are both computed from normalize_identity(), which contradicts the Python note: the frame records the bare class name, the key identity is ":" + that name. State the actual relation instead: the code comes from the identity, the read-side check compares the recorded name, and the identity equals or refines that name. Also state the guarantee boundary: identities that share one recorded name are kept apart by their codes alone. The Python note now covers classes that share a bare __name__ across modules or nesting, not just instances of one class. Such serializers get one code and one recorded name, so the read check cannot separate them; when they write different bytes and share a func: segment, they must be keyed under different ns: namespaces, including across a deploy. --- CHANGELOG.md | 16 ++++++++++------ spec/cache-key-format.md | 23 +++++++++++++++++------ 2 files changed, 27 insertions(+), 12 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9ff99bd..48bfa14 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,7 +16,9 @@ All notable changes to the CacheKit Protocol Specification. own is now REQUIRED on every read, before decoding, and a value recording no name is a mismatch. The table gains `l` (reference caching — shipped, never documented) and a `Canonical name` column; alias spellings and the instance - identity (`:` + class name) are documented in a new Python SDK note. + identity (`:` + class name) are documented in a new Python SDK note, + which requires distinct `ns:` namespaces for serializers whose classes share a + bare name, write different bytes, and share a `func:` segment. Documents the defect corrected by [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) (merged as `ee65250`; ships in cachekit-py 0.20.0): through v0.19.0 every @@ -37,11 +39,13 @@ All notable changes to the CacheKit Protocol Specification. need a released generator that emits them and follow separately. - `normalize_identity()`: the purity constraint gains an explicit injectivity requirement — two configurations that change the serialized container's - bytes MUST NOT normalize to the same identity, since a collapsed identity - makes the derived code and the frame tag agree and defeats the read-side - mismatch check. An object-to-string refinement MUST use a marker no bare - identity can produce, promoted from a cachekit-py-specific example to a - requirement on the hook's contract. + bytes MUST NOT normalize to the same identity, since the identity equals or + refines the recorded name: a collapsed identity gives both configurations one + code and one recorded name, which defeats the read-side mismatch check, and + distinct identities that share one recorded name are separated by their codes + alone. An object-to-string refinement MUST use a marker no bare identity + can produce, promoted from a cachekit-py-specific example to a requirement on + the hook's contract. ### Cache key — 7-segment format is Python SDK convention; server-side requirements diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index cc8c7d5..e24ed1c 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -111,8 +111,14 @@ else. > class name, built-ins included, and its key identity is `:` + that class name, so it > takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → `x2263`), never the > table's. The prefix contains characters no Python identifier can, so a custom class named -> `auto` cannot take AutoSerializer's code. Two instances of the same class are one identity: -> the SDK documents distinct namespaces for per-configuration serializers. +> `auto` cannot take AutoSerializer's code. Beyond that fixed prefix, the identity and the +> recorded name both carry only the bare class name (`__name__`), so any two instances whose +> classes share that name, whatever their module or nesting, get one code and one recorded +> name, and the read-side check cannot separate them. When two such serializers write +> different bytes and share a `func:` segment (one function, or closures from one factory), +> the application MUST key them under different `ns:` namespaces, no namespace counting as +> one; across a deploy, a changed configuration or implementation takes a namespace the old +> one never wrote. ### Example Keys @@ -304,11 +310,16 @@ SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "a // `normalize_identity()` MUST be a pure function of the serializer's configuration — never // an object address or a randomised hash — and MUST be injective over any configuration // dimension that changes the serialized container's bytes: e.g. `arrow+gzip` and -// `arrow+none` MUST NOT both normalize to `"arrow"`. The derived code and the frame tag are -// both computed from this identity, so collapsing two differently-configured serializers -// onto one identity makes them self-report the identical frame tag — the read-side mismatch +// `arrow+none` MUST NOT both normalize to `"arrow"`. The derived code is computed from this +// identity; the read-side check compares the recorded name instead, which the identity +// equals or refines (cachekit-py: `:ArrowSerializer` records `ArrowSerializer`). +// Collapsing two differently-configured serializers onto one identity therefore gives them +// one code — so one key for the same call — AND one recorded name: the read-side mismatch // check above can no longer tell them apart, and a mismatched container is served as a hit -// instead of a miss: wrong data, not a recoverable eviction. +// instead of a miss: wrong data, not a recoverable eviction. Distinct identities that +// share one recorded name are kept apart by their codes alone (for two identities outside +// the table, a 16-bit digest); the hit-rate-only collision guarantee above covers only +// identities recorded differently. // // A refinement that maps a serializer OBJECT to a string MUST use a marker that no bare // identity can produce, so a user-named class can never be spelled as a table key. From 5cf553fe4a9ecbbfcf8a114fb2247bd33fc13315 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sat, 26 Sep 2026 03:01:14 +1000 Subject: [PATCH 08/13] docs(cache-key): the 16-bit code is not a collision-resistant separator (LAB-4351) Identities that share a recorded name are separated only by their codes. A derived code carries 16 bits, so two such identities can collide; key and recorded name then both match and the read-side check serves the other configuration's container as a hit. Require identities that share a recorded name but write different bytes to use different ns: namespaces, consistent with the existing Python note, and align the CHANGELOG entry. --- CHANGELOG.md | 8 +++++--- spec/cache-key-format.md | 10 +++++++--- 2 files changed, 12 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 48bfa14..8e2af6e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -41,9 +41,11 @@ All notable changes to the CacheKit Protocol Specification. requirement — two configurations that change the serialized container's bytes MUST NOT normalize to the same identity, since the identity equals or refines the recorded name: a collapsed identity gives both configurations one - code and one recorded name, which defeats the read-side mismatch check, and - distinct identities that share one recorded name are separated by their codes - alone. An object-to-string refinement MUST use a marker no bare identity + code and one recorded name, which defeats the read-side mismatch check. + Distinct identities that share one recorded name are separated only by their + codes, which are not collision-resistant (16-bit derived digest), so those that + write different bytes MUST be keyed under different `ns:` namespaces, matching + the Python note. An object-to-string refinement MUST use a marker no bare identity can produce, promoted from a cachekit-py-specific example to a requirement on the hook's contract. diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index e24ed1c..0378b67 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -317,9 +317,13 @@ SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "a // one code — so one key for the same call — AND one recorded name: the read-side mismatch // check above can no longer tell them apart, and a mismatched container is served as a hit // instead of a miss: wrong data, not a recoverable eviction. Distinct identities that -// share one recorded name are kept apart by their codes alone (for two identities outside -// the table, a 16-bit digest); the hit-rate-only collision guarantee above covers only -// identities recorded differently. +// share one recorded name are kept apart only by their codes, and a code is NOT a +// collision-resistant separator: two identities outside the table collide at ≈1 in 2^16 +// per pair, and that collision is wrong data too, since key and recorded name then both +// match. The hit-rate-only collision guarantee above covers only identities recorded +// differently. Identities that share a recorded name but write different bytes MUST +// therefore be keyed under different `ns:` namespaces (as the Python note above requires +// for same-name classes). // // A refinement that maps a serializer OBJECT to a string MUST use a marker that no bare // identity can produce, so a user-named class can never be spelled as a table key. From 80b1dc22edf7842f496066cd19df5414fb844907 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sat, 26 Sep 2026 03:25:00 +1000 Subject: [PATCH 09/13] docs(cache-key): one uniqueness rule, required serializer_type, shipped aliases; scope read check to serialized entries (LAB-4351) --- CHANGELOG.md | 43 ++++++++++----------- spec/cache-key-format.md | 82 +++++++++++++++++++++------------------- 2 files changed, 63 insertions(+), 62 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8e2af6e..f04742f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,12 +13,19 @@ All notable changes to the CacheKit Protocol Specification. 2-byte `blake2b` digest of its UTF-8 identity as 4 lowercase hex characters; the collision guarantee is stated as probabilistic (16 bits). The read-side comparison of the container's recorded serializer name against the reader's - own is now REQUIRED on every read, before decoding, and a value recording no - name is a mismatch. The table gains `l` (reference caching — shipped, never - documented) and a `Canonical name` column; alias spellings and the instance - identity (`:` + class name) are documented in a new Python SDK note, - which requires distinct `ns:` namespaces for serializers whose classes share a - bare name, write different bytes, and share a `func:` segment. + own is now REQUIRED on every read of a serialized entry, before decoding, and + a value recording no name is a mismatch; reference caching (`l`) stores no + serialized container and is exempt. One uniqueness rule: an SDK SHOULD make + the identity distinguish configurations that write different bytes, and + wherever it does not, those configurations MUST be keyed under different + `ns:` namespaces (the 16-bit code is not a collision-resistant separator). The + `cache_key` AAD component is identical for serializers sharing a code, so the + cipher is no backstop between serializers that also share a `format` token. + The table gains `l` (reference caching — shipped, never documented) and a + `Canonical name` column; alias spellings (`std`, `standard`, `pythonic`) and + the instance identity (`:` + bare class name, so one class with + different constructor arguments shares one identity) are documented in a new + Python SDK note, which points at that rule. Documents the defect corrected by [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) (merged as `ee65250`; ships in cachekit-py 0.20.0): through v0.19.0 every @@ -30,24 +37,12 @@ All notable changes to the CacheKit Protocol Specification. now defines the table and `serializer_code()`, and adds an alias map and a `normalize_identity()` hook, both marked SDK-supplied; the hook runs before alias and table lookup, with the purity constraint the one-identity-one-code - rule already requires. -- Provenance: `test-vectors/cache-keys.json` is untouched by this change - (generated by `cachekit-py` v0.12.0); cachekit-py CI byte-verifies its - vendored copy against `CacheKeyGenerator` - (`tests/unit/protocol/test_cache_key_vectors.py`, pinned to protocol - `f0672c1c`). Every vector uses `std` → `1s`; vectors for the derived codes - need a released generator that emits them and follow separately. -- `normalize_identity()`: the purity constraint gains an explicit injectivity - requirement — two configurations that change the serialized container's - bytes MUST NOT normalize to the same identity, since the identity equals or - refines the recorded name: a collapsed identity gives both configurations one - code and one recorded name, which defeats the read-side mismatch check. - Distinct identities that share one recorded name are separated only by their - codes, which are not collision-resistant (16-bit derived digest), so those that - write different bytes MUST be keyed under different `ns:` namespaces, matching - the Python note. An object-to-string refinement MUST use a marker no bare identity - can produce, promoted from a cachekit-py-specific example to a requirement on - the hook's contract. + rule already requires, and an object-to-string refinement MUST use a marker + no bare identity can produce. `serializer_type` is a required parameter, and + `serializer_code()` rejects an empty or non-string identity with an error + rather than mapping it to a code. +- Provenance: no vector changed; the derived codes are not yet covered by + vectors. ### Cache key — 7-segment format is Python SDK convention; server-side requirements diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 0378b67..455d315 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -6,7 +6,7 @@ **Deterministic key generation from function identity and arguments.** -*Protocol Version 1.0 · Verified against `cachekit-py` v0.18.0 (`src/cachekit/key_generator.py`); serializer-code derivation verified against [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) @ `ee65250` (merged; ships in v0.20.0); vectors generated by v0.12.0 (see [Test Vectors](#test-vectors))* +*Protocol Version 1.0 · Serializer-code derivation verified against `cachekit-py` @ `ee65250` (the [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) merge)* @@ -84,18 +84,21 @@ else. > exactly as a constant code makes every pair do — and never yields a wrong value, because > the stored serializer name still differs and the read-side check below rejects it. > -> **On every read, before decoding the payload, the reader MUST compare the serializer name -> its storage container records (Python: the `s` field of the +> **On every read of a serialized entry, before decoding the payload, the reader MUST compare +> the serializer name its storage container records (Python: the `s` field of the > [CK v3 frame](wire-format.md#python-ck-v3-frame)) with the name it would itself record for > its configured serializer, and MUST reject the entry on mismatch — a miss, never a value. A > value that records no serializer name is a mismatch.** A colliding entry has the same key as > the reader's own, so nothing before this comparison can tell them apart. An SDK that offers > more than one serializer identity MUST record the name in its container, or this check -> cannot exist (`cachekit-ts` and `cachekit-rs` record none and offer one). +> cannot exist (`cachekit-ts` and `cachekit-rs` record none and offer one). Reference caching +> (`l`) is exempt: it stores the object itself, never a serialized container, so there is no +> recorded name to compare. > > `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers -> sharing a code therefore produce a **byte-identical AAD**: AAD binding does not separate -> them, and the cipher is not a backstop for a missing name check. +> sharing a code therefore share the **`cache_key` AAD component**; when they also share the +> `format` token, AAD binding does not separate them, and the cipher is not a backstop for a +> missing name check. > > **Conversely, one identity MUST always produce one code.** Derive it from the serializer > configuration alone — the canonical name the wire format records, or an SDK-defined @@ -103,22 +106,32 @@ else. > separates keyspaces, the read-side check still sees only the recorded name) — never from a > process-local value such as an object address or a randomised hash, or keys stop being > reproducible across processes. +> +> **An SDK SHOULD make the identity distinguish configurations that write different bytes. +> Wherever its identity does not, configurations that write different bytes and would +> otherwise share a key MUST be keyed under different `ns:` namespaces**, no namespace +> counting as one. Sharing an identity, they share a code, so a key, and a recorded name, so +> the read-side check cannot tell them apart: one is served the other's bytes as a hit — +> wrong data, not an eviction. The `ns:` MUST also covers distinct identities that share one +> recorded name and write different bytes, because the 16-bit code is NOT a +> collision-resistant separator: two such identities collide at ≈1 in 2^16 per pair, and key +> and recorded name then both match. The hit-rate-only collision guarantee above covers only +> identities recorded differently. The Python SDK is the known case where the identity does not distinguish configurations (see the note below). > [!NOTE] -> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` for -> `default` and `pythonic` for `auto`, canonicalizing them before the lookup. A serializer -> passed as an *instance* rather than a name is recorded in the frame header under its bare -> class name, built-ins included, and its key identity is `:` + that class name, so it -> takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → `x2263`), never the -> table's. The prefix contains characters no Python identifier can, so a custom class named -> `auto` cannot take AutoSerializer's code. Beyond that fixed prefix, the identity and the -> recorded name both carry only the bare class name (`__name__`), so any two instances whose -> classes share that name, whatever their module or nesting, get one code and one recorded -> name, and the read-side check cannot separate them. When two such serializers write -> different bytes and share a `func:` segment (one function, or closures from one factory), -> the application MUST key them under different `ns:` namespaces, no namespace counting as -> one; across a deploy, a changed configuration or implementation takes a namespace the old -> one never wrote. +> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` and +> `standard` for `default` and `pythonic` for `auto`, canonicalizing them before the lookup. +> A serializer passed as an *instance* rather than a name is recorded in the frame header +> under its bare class name, built-ins included, and its key identity is `:` + that +> class name, so it takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → +> `x2263`), never the table's. The prefix contains characters no Python identifier can, so a +> custom class named `auto` cannot take AutoSerializer's code. Beyond that fixed prefix, the +> identity and the recorded name both carry only the bare class name (`__name__`), so two +> instances of one class with different constructor arguments, or of any two classes sharing +> that name whatever their module or nesting, get one code and one recorded name. Where they +> write different bytes and share a `func:` segment (one function, or closures from one +> factory), the `ns:` rule above applies — across a deploy too: a changed configuration or +> implementation takes a namespace the old one never wrote. ### Example Keys @@ -297,7 +310,7 @@ SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "l // about which serializer wrote it. Do not adopt another SDK's aliases: mapping a spelling // your API does not accept hands that name a table code instead of the derived `x` code it // should get. cachekit-py's map is in the Python note above. -SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "auto"} +SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "standard": "default", "pythonic": "auto"} // test-vectors/cache-keys.json records serializer_type "std": cachekit-py's alias for the // canonical "default", so every vector's code is "s". @@ -308,22 +321,11 @@ SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "pythonic": "a // the "SDK-defined refinement" the one-identity-one-code rule permits. // // `normalize_identity()` MUST be a pure function of the serializer's configuration — never -// an object address or a randomised hash — and MUST be injective over any configuration -// dimension that changes the serialized container's bytes: e.g. `arrow+gzip` and -// `arrow+none` MUST NOT both normalize to `"arrow"`. The derived code is computed from this -// identity; the read-side check compares the recorded name instead, which the identity -// equals or refines (cachekit-py: `:ArrowSerializer` records `ArrowSerializer`). -// Collapsing two differently-configured serializers onto one identity therefore gives them -// one code — so one key for the same call — AND one recorded name: the read-side mismatch -// check above can no longer tell them apart, and a mismatched container is served as a hit -// instead of a miss: wrong data, not a recoverable eviction. Distinct identities that -// share one recorded name are kept apart only by their codes, and a code is NOT a -// collision-resistant separator: two identities outside the table collide at ≈1 in 2^16 -// per pair, and that collision is wrong data too, since key and recorded name then both -// match. The hit-rate-only collision guarantee above covers only identities recorded -// differently. Identities that share a recorded name but write different bytes MUST -// therefore be keyed under different `ns:` namespaces (as the Python note above requires -// for same-name classes). +// an object address or a randomised hash. The derived code is computed from this identity; +// the read-side check compares the recorded name instead, which the identity equals or +// refines (cachekit-py: `:ArrowSerializer` records `ArrowSerializer`). Whether the +// identity must distinguish configurations that write different bytes is the uniqueness +// rule in Serializer Codes above (SHOULD; where it does not, different `ns:` namespaces MUST). // // A refinement that maps a serializer OBJECT to a string MUST use a marker that no bare // identity can produce, so a user-named class can never be spelled as a table key. @@ -338,6 +340,10 @@ function serializer_code(serializer_type): // the code up. An identity outside the table gets its OWN derived code — never a shared // constant, which would put every unrecognised serializer on one keyspace. name = normalize_identity(serializer_type) + // An empty or non-string identity MUST be rejected with an error, never mapped to a + // code: a fallback code is a shared bucket, and a key computed from it is one nothing wrote. + if name is not a string or name == "": + raise error identity = SERIALIZER_ALIASES.get(name, name) if identity in SERIALIZER_CODES: return SERIALIZER_CODES[identity] @@ -345,7 +351,7 @@ function serializer_code(serializer_type): return "x" + blake2b(identity.utf8_bytes(), digest_size=2).hex() function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, - integrity_checking=true, serializer_type="default"): + integrity_checking=true, *, serializer_type): // Build key parts parts = [] From 54cffcbcabf8c14029e94581df522ad5440fcfa5 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sun, 27 Sep 2026 05:36:33 +1000 Subject: [PATCH 10/13] docs(cache-key): scope the read-side name check, state identity rules in prose, drop stale cross-SDK text - Read-side serializer-name check: the recording and comparison MUSTs now bind an SDK that offers more than one serializer identity, for serialized entries under keys in this format. Single-identity SDKs (cachekit-ts, cachekit-rs) are exempt; Interop Mode entries and in-process live-object caches are outside both rules. "No recorded name is a mismatch" is kept. - Identity derivation (alias resolution, object-identity marker, rejecting an empty identity) moves from pseudocode comments into normative prose; alias resolution now also binds the serializer that writes the bytes. - Python note: cachekit-py accepts std and pythonic; serializer="standard" is rejected by the cache. - Serializer Codes table drops the Cross-language? column; Cross-SDK Key Generation Strategy and the README Quick Start route cross-SDK sharing through Interop Mode only. - wire-format.md: header `s` lists an instance's bare class name; the framing statement excepts reference caching and backend=None. CodeRabbit-Resolved: spec/cache-key-format.md:88:scope missing-name rejection to multi-identity SDKs --- CHANGELOG.md | 30 +++++++++---- README.md | 2 +- spec/cache-key-format.md | 97 +++++++++++++++++++++------------------- spec/wire-format.md | 11 +++-- 4 files changed, 81 insertions(+), 59 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f04742f..dafa9e3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,18 +11,29 @@ All notable changes to the CacheKit Protocol Specification. two identities the wire format records differently MUST NOT be mapped onto one code by construction. An identity outside the code table gets `x` + the 2-byte `blake2b` digest of its UTF-8 identity as 4 lowercase hex characters; - the collision guarantee is stated as probabilistic (16 bits). The read-side - comparison of the container's recorded serializer name against the reader's - own is now REQUIRED on every read of a serialized entry, before decoding, and - a value recording no name is a mismatch; reference caching (`l`) stores no - serialized container and is exempt. One uniqueness rule: an SDK SHOULD make + the collision guarantee is stated as probabilistic (16 bits). An SDK that + offers more than one serializer identity MUST record the serializer name in + the container of each serialized entry under a key in this format, and MUST + compare it against its own on every read of one, before decoding; for such an + SDK, an entry recording no name is a mismatch. An SDK offering a single + identity (`cachekit-ts`, `cachekit-rs`) is exempt from both; Interop Mode + entries and in-process live-object caches are outside both rules. + Identity-derivation rules that lived only in pseudocode comments (alias + resolution, the object-identity marker, rejecting an empty identity) are now + normative prose, and alias resolution now also binds the writer: the code, + the recorded name and the serializer that writes the bytes MUST all resolve + to one canonical name. One uniqueness rule: an SDK SHOULD make the identity distinguish configurations that write different bytes, and wherever it does not, those configurations MUST be keyed under different `ns:` namespaces (the 16-bit code is not a collision-resistant separator). The `cache_key` AAD component is identical for serializers sharing a code, so the cipher is no backstop between serializers that also share a `format` token. The table gains `l` (reference caching — shipped, never documented) and a - `Canonical name` column; alias spellings (`std`, `standard`, `pythonic`) and + `Canonical name` column, and drops its `Cross-language?` column: no + serializer code makes these keys shareable across SDKs, and the Cross-SDK Key + Generation Strategy section (and the README's implementor Quick Start) now + route all cross-SDK sharing through Interop Mode instead of namespace-matched + keys. Alias spellings (`std`, `pythonic`) and the instance identity (`:` + bare class name, so one class with different constructor arguments shares one identity) are documented in a new Python SDK note, which points at that rule. @@ -37,10 +48,13 @@ All notable changes to the CacheKit Protocol Specification. now defines the table and `serializer_code()`, and adds an alias map and a `normalize_identity()` hook, both marked SDK-supplied; the hook runs before alias and table lookup, with the purity constraint the one-identity-one-code - rule already requires, and an object-to-string refinement MUST use a marker - no bare identity can produce. `serializer_type` is a required parameter, and + rule already requires. `serializer_type` is a required parameter, and `serializer_code()` rejects an empty or non-string identity with an error rather than mapping it to a code. +- [`spec/wire-format.md`](spec/wire-format.md): the CK v3 frame's header `s` + table adds a serializer instance's bare class name, and the framing statement + excepts the two in-process modes that store no bytes: reference caching (`l`) + and `backend=None`. - Provenance: no vector changed; the derived codes are not yet covered by vectors. diff --git a/README.md b/README.md index 6a42119..e62bf79 100644 --- a/README.md +++ b/README.md @@ -83,7 +83,7 @@ Building a new SDK? Implement in this order: **1. Key Generation** — [spec/cache-key-format.md](spec/cache-key-format.md) -Generate deterministic cache keys from function identity + arguments. Keys must match across SDKs for cross-language cache sharing. +Generate deterministic cache keys from function identity + arguments. These auto-mode keys are SDK-specific (the 7-segment format is the Python SDK's convention); keys shared across SDKs use [interop mode](spec/interop-mode.md). **2. Wire Format** — [spec/wire-format.md](spec/wire-format.md) diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 455d315..a6bc8b5 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -54,16 +54,17 @@ ns:{namespace}:func:{module}.{qualname}:args:{blake2b_hash}:{ic_flag}{serializer ### Serializer Codes -| Code | Serializer | Canonical name | Cross-language? | -| :---: | :--- | :--- | :---: | -| `s` | StandardSerializer (MessagePack) | `default` | ✅ Yes | -| `a` | AutoSerializer (language-specific types) | `auto` | ❌ No | -| `o` | OrjsonSerializer (JSON-based) | `orjson` | ⚠️ Partial | -| `w` | ArrowSerializer (columnar) | `arrow` | ⚠️ Partial | -| `l` | Reference caching (no serialization) | `local` | ❌ No | -| `x` + 4 hex | Any serializer identity not in this table | — | ❌ No | - -For cross-SDK interoperability, always use `s` (StandardSerializer). +| Code | Serializer | Canonical name | +| :---: | :--- | :--- | +| `s` | StandardSerializer (MessagePack) | `default` | +| `a` | AutoSerializer (language-specific types) | `auto` | +| `o` | OrjsonSerializer (JSON-based) | `orjson` | +| `w` | ArrowSerializer (columnar) | `arrow` | +| `l` | Reference caching (no serialization) | `local` | +| `x` + 4 hex | Any serializer identity not in this table | — | + +No code makes these keys shareable across SDKs; see +[Cross-SDK Key Generation Strategy](#cross-sdk-key-generation-strategy). An identity outside the table gets `x` followed by the 2-byte digest `blake2b(utf8(identity), digest_size=2)` encoded as exactly 4 lowercase hexadecimal @@ -84,16 +85,21 @@ else. > exactly as a constant code makes every pair do — and never yields a wrong value, because > the stored serializer name still differs and the read-side check below rejects it. > -> **On every read of a serialized entry, before decoding the payload, the reader MUST compare -> the serializer name its storage container records (Python: the `s` field of the -> [CK v3 frame](wire-format.md#python-ck-v3-frame)) with the name it would itself record for -> its configured serializer, and MUST reject the entry on mismatch — a miss, never a value. A -> value that records no serializer name is a mismatch.** A colliding entry has the same key as -> the reader's own, so nothing before this comparison can tell them apart. An SDK that offers -> more than one serializer identity MUST record the name in its container, or this check -> cannot exist (`cachekit-ts` and `cachekit-rs` record none and offer one). Reference caching -> (`l`) is exempt: it stores the object itself, never a serialized container, so there is no -> recorded name to compare. +> **An SDK that offers more than one serializer identity MUST record the serializer name in +> the storage container of every serialized entry it stores under a key in this format +> (Python: the `s` field of the [CK v3 frame](wire-format.md#python-ck-v3-frame)). On every +> read of a serialized entry under a key in this format, before decoding the payload, such an +> SDK MUST compare the name the container records with the name it would itself record for +> its configured serializer, and MUST reject the entry on mismatch — a miss, never a value. +> For such an SDK, an entry that records no serializer name is a mismatch.** A colliding +> entry has the same key as the reader's own, so nothing before this comparison can tell them +> apart. An SDK that offers exactly one serializer identity is exempt from both the recording +> and the comparison, because there is no second serializer in its keyspace to confuse with +> the first (`cachekit-ts` and `cachekit-rs` offer one and record none). Neither rule reaches +> a value that is not a serialized entry under a key in this format: +> [Interop Mode](interop-mode.md) entries carry no serializer code and no container, and a +> cache that keeps live objects in process memory stores nothing to compare (in +> `cachekit-py`: reference caching, code `l`, and caches configured with no backend). > > `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers > sharing a code therefore share the **`cache_key` AAD component**; when they also share the @@ -107,6 +113,16 @@ else. > process-local value such as an object address or a randomised hash, or keys stop being > reproducible across processes. > +> **Deriving the identity.** An SDK that accepts alias spellings MUST resolve each accepted +> spelling to exactly one canonical name, and the code, the recorded name, and the serializer +> that writes the bytes MUST all be that canonical name's, or a key, its stored entry and its +> bytes can disagree about which serializer wrote it. An SDK that accepts a serializer object +> MUST reduce it to a string identity carrying a marker that no code-table name, alias +> spelling or accepted serializer name contains, so a user-named class can never take a table +> code (the Python note below gives `cachekit-py`'s marker). An empty or non-string identity +> MUST be rejected with an error, never mapped to a code: a fallback code is a shared bucket, +> and a key computed from it names an entry nothing wrote. +> > **An SDK SHOULD make the identity distinguish configurations that write different bytes. > Wherever its identity does not, configurations that write different bytes and would > otherwise share a key MUST be keyed under different `ns:` namespaces**, no namespace @@ -119,8 +135,9 @@ else. > identities recorded differently. The Python SDK is the known case where the identity does not distinguish configurations (see the note below). > [!NOTE] -> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` and -> `standard` for `default` and `pythonic` for `auto`, canonicalizing them before the lookup. +> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` for +> `default` and `pythonic` for `auto`, canonicalizing them before the lookup. (Its key +> generator's alias map also lists `standard`, which the cache rejects as a serializer name.) > A serializer passed as an *instance* rather than a name is recorded in the frame header > under its bare class name, built-ins included, and its key identity is `:` + that > class name, so it takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → @@ -152,7 +169,7 @@ ns:cache:func:app.views.index:args:0000...0000:0s - If a key exceeds 250 characters: first 50 chars of original key + `:` + first 32 chars of a Blake2b-256 hash of the full key > [!WARNING] -> **Discrepancy with RFC** — The original protocol RFC (Section 3.1.5) specifies a simpler key format: `{namespace}:{hash}`. The actual implementation includes function identity (`func:` prefix) and metadata suffix (`:{ic_flag}{serializer_code}`). **The implementation is authoritative.** For cross-SDK interoperability, SDK implementors must use explicit namespaces (the `func:` segment is language-specific and will differ). +> **Discrepancy with RFC** — The original protocol RFC (Section 3.1.5) specifies a simpler key format: `{namespace}:{hash}`. The actual implementation includes function identity (`func:` prefix) and metadata suffix (`:{ic_flag}{serializer_code}`). **The implementation is authoritative.** For cross-SDK sharing, use [Interop Mode](interop-mode.md) (see [Cross-SDK Key Generation Strategy](#cross-sdk-key-generation-strategy)). --- @@ -179,16 +196,7 @@ generation, invisible to the server. ## Cross-SDK Key Generation Strategy -For multi-language interoperability, all SDKs MUST use **explicit namespaces** rather than auto-generated function signatures. The `func:` segment is inherently language-specific (Python modules vs PHP namespaces vs Go packages), so cross-language cache sharing requires: - -1. All SDKs agree on a namespace string (e.g., `"get_user"`) -2. All SDKs serialize arguments identically (see [Argument Hashing Algorithm](#argument-hashing-algorithm) below) -3. The resulting Blake2b hash in the `args:` segment will be identical - -The `func:` and metadata segments may differ between SDKs — this is acceptable when the key is constructed to match. - -> [!TIP] -> Use [Interop Mode](interop-mode.md) to remove the `func:` segment entirely. Interop mode produces the simplest possible cross-language key: `{namespace}:{operation}:{args_hash}`. +Keys in this format are not shared across SDKs, even under an agreed namespace. The `func:` segment is language-specific, and the values stored under these keys are SDK-internal containers that no other SDK decodes (see [SDK Storage Containers](wire-format.md#sdk-storage-containers-auto-mode)). Cross-language cache sharing uses [Interop Mode](interop-mode.md) exclusively: explicit, language-neutral operation names, keys of the form `{namespace}:{operation}:{args_hash}`, and plain MessagePack values. --- @@ -305,12 +313,12 @@ Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte- SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "local": "l"} // SDK-SUPPLIED, not fixed by this spec: alias spellings THIS SDK accepts -> canonical -// name. Empty if the SDK accepts only canonical names. The SAME map must canonicalize the -// serializer name the wire format records, or a key and its stored envelope can disagree -// about which serializer wrote it. Do not adopt another SDK's aliases: mapping a spelling -// your API does not accept hands that name a table code instead of the derived `x` code it -// should get. cachekit-py's map is in the Python note above. -SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "standard": "default", "pythonic": "auto"} +// name. Empty if the SDK accepts only canonical names. Using this map for the recorded name +// too, with the writer resolving to the same serializer, satisfies "Deriving the identity" +// in Serializer Codes above. Do not adopt another SDK's aliases: mapping a spelling your API +// does not accept hands that name a table code instead of the derived `x` code it should +// get. cachekit-py's accepted aliases are in the Python note above. +SERIALIZER_ALIASES = {} // e.g. cachekit-py accepts: {"std": "default", "pythonic": "auto"} // test-vectors/cache-keys.json records serializer_type "std": cachekit-py's alias for the // canonical "default", so every vector's code is "s". @@ -327,11 +335,9 @@ SERIALIZER_ALIASES = {} // e.g. cachekit-py: {"std": "default", "standard": "d // identity must distinguish configurations that write different bytes is the uniqueness // rule in Serializer Codes above (SHOULD; where it does not, different `ns:` namespaces MUST). // -// A refinement that maps a serializer OBJECT to a string MUST use a marker that no bare -// identity can produce, so a user-named class can never be spelled as a table key. -// cachekit-py maps an object to ":" + its bare class name (Python note above); the -// prefix uses characters no identifier can contain, which is what stops a class named -// `auto` from taking AutoSerializer's code. +// An object-to-string refinement carries a marker no table name, alias or accepted name +// contains ("Deriving the identity" above). cachekit-py maps an object to ":" + its +// bare class name (Python note above). function normalize_identity(serializer_type): return serializer_type // names-only SDK; override to handle objects @@ -340,8 +346,7 @@ function serializer_code(serializer_type): // the code up. An identity outside the table gets its OWN derived code — never a shared // constant, which would put every unrecognised serializer on one keyspace. name = normalize_identity(serializer_type) - // An empty or non-string identity MUST be rejected with an error, never mapped to a - // code: a fallback code is a shared bucket, and a key computed from it is one nothing wrote. + // An empty or non-string identity is an error, never a code ("Deriving the identity"). if name is not a string or name == "": raise error identity = SERIALIZER_ALIASES.get(name, name) diff --git a/spec/wire-format.md b/spec/wire-format.md index 74e5e89..91cfd4d 100644 --- a/spec/wire-format.md +++ b/spec/wire-format.md @@ -502,7 +502,7 @@ Datetime values are encoded as MessagePack maps with sentinel keys: ## SDK Storage Containers (auto mode) Remote backends (Redis, CachekitIO SaaS, Memcached, File) store opaque bytes. (L1 -behavior is SDK-specific: `cachekit-py`'s L1 holds the framed bytes; `cachekit-ts`'s +behavior is SDK-specific: `cachekit-py`'s L1 in front of a backend holds the framed bytes; `cachekit-ts`'s L1 holds live decoded values, not bytes.) What the stored bytes *are* differs per SDK in auto mode: @@ -525,9 +525,11 @@ implementations ([protocol#11](https://github.com/cachekit-io/protocol/issues/11 ### Python: CK v3 frame -Every **auto-mode** value `cachekit-py` stores — all backends, all serializers, -encrypted or not — is framed (interop-mode values are plain MessagePack, never -framed): +Two in-process modes keep live objects and store no bytes at all: `@cache.local` +reference caching (key code `l`) and a cache configured with `backend=None`, whose +keys carry its configured serializer's code. Every other **auto-mode** value +`cachekit-py` stores — all backends, all serializers, encrypted or not — is framed +(interop-mode values are plain MessagePack, never framed): ```text MAGIC b"CK" (0x43 0x4B) | VERSION u8 (0x03) | HDR_LEN u32 big-endian | HEADER | PAYLOAD @@ -543,6 +545,7 @@ MAGIC b"CK" (0x43 0x4B) | VERSION u8 (0x03) | HDR_LEN u32 big-endian | HEADER | | `default`, `auto` | ByteStorage envelope (this document) over MessagePack | | `arrow` | **Arrow envelope**: `[8-byte xxHash3-64 checksum][Arrow IPC file]` (IPC magic `b"ARROW1"` at payload offset 8) | | `orjson` | `[8-byte xxHash3-64 checksum][JSON bytes]` | +| A serializer instance's bare class name (`StandardSerializer`, `ArrowSerializer`, a custom class) | That serializer's own output; a built-in class writes the same payload as its string name above | | any, encrypted | Ciphertext per [encryption.md](encryption.md) | With integrity checking disabled, `default`/`auto` payloads are raw MessagePack (no From 13db17459d61f583c019f5040cf101a06ecd5ece Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sun, 27 Sep 2026 18:19:14 +1000 Subject: [PATCH 11/13] docs(cache-key): scope the collision guarantee to honest writers; AAD never separates a shared code The recorded serializer name sits in the plaintext, unauthenticated CK v3 frame header, so it turns accidental code collisions into misses but is not an integrity control against a writer with backend write access. Say so and point at the frame header caution. The AAD paragraph implied that a differing format token lets AAD binding separate two serializers sharing a code. Readers rebuild the AAD from the stored format, so it never does. Correct the paragraph and its changelog line. No wire or AAD change. --- CHANGELOG.md | 10 +++++++--- spec/cache-key-format.md | 17 +++++++++++------ 2 files changed, 18 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c50136c..febddd4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,7 +11,10 @@ All notable changes to the CacheKit Protocol Specification. two identities the wire format records differently MUST NOT be mapped onto one code by construction. An identity outside the code table gets `x` + the 2-byte `blake2b` digest of its UTF-8 identity as 4 lowercase hex characters; - the collision guarantee is stated as probabilistic (16 bits). An SDK that + the collision guarantee is stated as probabilistic (16 bits). Between honestly + written entries a collision costs hit rate only, never a wrong value; the + recorded serializer name is not an integrity control against a writer with + backend write access. An SDK that offers more than one serializer identity MUST record the serializer name in the container of each serialized entry under a key in this format, and MUST compare it against its own on every read of one, before decoding; for such an @@ -26,8 +29,9 @@ All notable changes to the CacheKit Protocol Specification. the identity distinguish configurations that write different bytes, and wherever it does not, those configurations MUST be keyed under different `ns:` namespaces (the 16-bit code is not a collision-resistant separator). The - `cache_key` AAD component is identical for serializers sharing a code, so the - cipher is no backstop between serializers that also share a `format` token. + `cache_key` AAD component is identical for serializers sharing a code, and a + reader takes `format` from the stored entry, so the cipher is no backstop + between serializers sharing a code. The table gains `l` (reference caching — shipped, never documented) and a `Canonical name` column, and drops its `Cross-language?` column: no serializer code makes these keys shareable across SDKs, and the Cross-SDK Key diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index a6bc8b5..f7083f6 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -81,9 +81,13 @@ else. > **Two serializer identities that the wire format records differently MUST NOT be mapped > onto one code by construction.** The guarantee is probabilistic, not absolute: the derived > code carries 16 bits, so two identities it records differently can still collide, at ≈1 in -> 2^16 per pair. Such a collision costs hit rate only — that one pair evicts each other -> exactly as a constant code makes every pair do — and never yields a wrong value, because -> the stored serializer name still differs and the read-side check below rejects it. +> 2^16 per pair. Between honestly written entries, such a collision costs hit rate only — +> that one pair evicts each other exactly as a constant code makes every pair do — and never +> yields a wrong value, because the stored serializer name still differs and the read-side +> check below rejects it. The recorded name is not an integrity control against a writer with +> backend write access: the CK v3 frame header that carries it is plaintext and +> unauthenticated, even for encrypted entries (see the +> [frame header caution](wire-format.md#python-ck-v3-frame)). > > **An SDK that offers more than one serializer identity MUST record the serializer name in > the storage container of every serialized entry it stores under a key in this format @@ -102,9 +106,10 @@ else. > `cachekit-py`: reference caching, code `l`, and caches configured with no backend). > > `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers -> sharing a code therefore share the **`cache_key` AAD component**; when they also share the -> `format` token, AAD binding does not separate them, and the cipher is not a backstop for a -> missing name check. +> sharing a code therefore share the **`cache_key` AAD component**, and a reader takes the +> `format` component from the entry's stored metadata, not from its own serializer (see +> [Encryption](encryption.md#format-tokens)). AAD binding does not separate them, +> whatever their `format` tokens, and the cipher is not a backstop for a missing name check. > > **Conversely, one identity MUST always produce one code.** Derive it from the serializer > configuration alone — the canonical name the wire format records, or an SDK-defined From 469e30380f9cdbbd67000a3d66fd07619cfcaf9a Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Sun, 27 Sep 2026 18:34:20 +1000 Subject: [PATCH 12/13] docs(cache-key): the serializer name is unauthenticated because it is not an AAD input The previous wording called the whole CK v3 frame header unauthenticated. Header values that feed the AAD are covered by the tag for encrypted entries; only the serializer name is outside it. Say exactly that. --- spec/cache-key-format.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index f7083f6..3eb6c07 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -85,9 +85,9 @@ else. > that one pair evicts each other exactly as a constant code makes every pair do — and never > yields a wrong value, because the stored serializer name still differs and the read-side > check below rejects it. The recorded name is not an integrity control against a writer with -> backend write access: the CK v3 frame header that carries it is plaintext and -> unauthenticated, even for encrypted entries (see the -> [frame header caution](wire-format.md#python-ck-v3-frame)). +> backend write access: in the CK v3 frame it is a plaintext header field that is not an +> [AAD](encryption.md#additional-authenticated-data-aad) input, so nothing authenticates it, +> even for encrypted entries (see the [frame header caution](wire-format.md#python-ck-v3-frame)). > > **An SDK that offers more than one serializer identity MUST record the serializer name in > the storage container of every serialized entry it stores under a key in this format From f9cff445816c249215e225fab76b607145bd8054 Mon Sep 17 00:00:00 2001 From: Ray Walker Date: Tue, 29 Sep 2026 02:57:56 +1000 Subject: [PATCH 13/13] =?UTF-8?q?fix:=20address=20coderabbit=20review=20?= =?UTF-8?q?=E2=80=94=20point=20the=20bare-class-name=20s=20row=20at=20the?= =?UTF-8?q?=20ns:=20rule?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A CK v3 frame records an instance's bare class name in `s`, so same-named classes and differently configured instances of one class record the same value, and the read-side name check cannot tell them apart. The normative rule for that case already lives in cache-key-format.md (Serializer Codes); the wire-format row now points at it instead of reading as if `s` alone identified the serializer. CodeRabbit-Resolved: spec/wire-format.md:539:bare class-name collisions --- spec/wire-format.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/spec/wire-format.md b/spec/wire-format.md index bcd1f0b..735d618 100644 --- a/spec/wire-format.md +++ b/spec/wire-format.md @@ -536,7 +536,7 @@ MAGIC b"CK" (0x43 0x4B) | VERSION u8 (0x03) | HDR_LEN u32 big-endian | HEADER | | `default`, `auto` | ByteStorage envelope (this document) over MessagePack | | `arrow` | **Arrow envelope**: `[8-byte xxHash3-64 checksum][Arrow IPC file]` (IPC magic `b"ARROW1"` at payload offset 8) | | `orjson` | `[8-byte xxHash3-64 checksum][JSON bytes]` | -| A serializer instance's bare class name (`StandardSerializer`, `ArrowSerializer`, a custom class) | That serializer's own output; a built-in class writes the same payload as its string name above | +| A serializer instance's bare class name (`StandardSerializer`, `ArrowSerializer`, a custom class) | That serializer's own output; a built-in class writes the same payload as its string name above. Classes sharing a bare name, and differently configured instances of one class, record the same `s` (see the [`ns:` rule](cache-key-format.md#serializer-codes)) | | any, encrypted | Ciphertext per [encryption.md](encryption.md) | With integrity checking disabled, `default`/`auto` payloads are raw MessagePack (no