diff --git a/README.md b/README.md index 2fd2f83..97d89bd 100644 --- a/README.md +++ b/README.md @@ -84,7 +84,7 @@ Building a new SDK? Implement in this order: **1. Key Generation** — [spec/cache-key-format.md](spec/cache-key-format.md) -Generate deterministic cache keys from function identity + arguments. Keys must match across SDKs for cross-language cache sharing. +Generate deterministic cache keys from function identity + arguments. These auto-mode keys are SDK-specific (the 7-segment format is the Python SDK's convention); keys shared across SDKs use [interop mode](spec/interop-mode.md). **2. Wire Format** — [spec/wire-format.md](spec/wire-format.md) diff --git a/changelog.d/20260924_server-side-requirements.md b/changelog.d/20260924_server-side-requirements.md new file mode 100644 index 0000000..7de7c46 --- /dev/null +++ b/changelog.d/20260924_server-side-requirements.md @@ -0,0 +1,18 @@ +### Cache key — 7-segment format is Python SDK convention; server-side requirements + +- [`spec/cache-key-format.md`](spec/cache-key-format.md): the 7-segment key + structure is marked the Python SDK's internal convention, not a server + contract. A new + [Server-Side Requirements](spec/cache-key-format.md#server-side-requirements) + section lists the only checks the CachekitIO backend enforces and which API + key classes may write each key space — including the open `default` space + that unprefixed keys (TypeScript/Rust `{ns}:{hash}`, interop, bare hashes) + fall in. +- Test Vectors: `test-vectors/cache-keys.json` is Python-SDK-only. The rule + that a cross-SDK implementation substitutes its own module path and matches + the args-hash segment byte-for-byte is withdrawn for these vectors; cross-SDK + conformance uses `test-vectors/interop-mode.json`. The fixture's `note` and + `key_format` fields now say so; no vector changed. +- [`spec/interop-mode.md`](spec/interop-mode.md): the deployed validator + accepts interop-format keys (`{namespace}:{operation}:{args_hash}`, in the + `default` namespace), replacing the warning that it would reject them. diff --git a/changelog.d/20260929_lab-4351.md b/changelog.d/20260929_lab-4351.md new file mode 100644 index 0000000..3c3af8b --- /dev/null +++ b/changelog.d/20260929_lab-4351.md @@ -0,0 +1,57 @@ +### Cache key — serializer code MUST be derived, and per-identity (LAB-4351) + +- [`spec/cache-key-format.md`](spec/cache-key-format.md): the `{serializer_code}` + suffix MUST be derived from the configured serializer, never a constant, and + two identities the wire format records differently MUST NOT be mapped onto + one code by construction. An identity outside the code table gets `x` + the + 2-byte `blake2b` digest of its UTF-8 identity as 4 lowercase hex characters; + the collision guarantee is stated as probabilistic (16 bits). Between honestly + written entries a collision costs hit rate only, never a wrong value; the + recorded serializer name is not an integrity control against a writer with + backend write access. An SDK that + offers more than one serializer identity MUST record the serializer name in + the container of each serialized entry under a key in this format, and MUST + compare it against its own on every read of one, before decoding; for such an + SDK, an entry recording no name is a mismatch. An SDK offering a single + identity (`cachekit-ts`, `cachekit-rs`) is exempt from both; Interop Mode + entries and in-process live-object caches are outside both rules. + Identity-derivation rules that lived only in pseudocode comments (alias + resolution, the object-identity marker, rejecting an empty identity) are now + normative prose, and alias resolution now also binds the writer: the code, + the recorded name and the serializer that writes the bytes MUST all resolve + to one canonical name. One uniqueness rule: an SDK SHOULD make + the identity distinguish configurations that write different bytes, and + wherever it does not, those configurations MUST be keyed under different + `ns:` namespaces (the 16-bit code is not a collision-resistant separator). The + `cache_key` AAD component is identical for serializers sharing a code, and a + reader takes `format` from the stored entry, so the cipher is no backstop + between serializers sharing a code. + The table gains `l` (reference caching — shipped, never documented) and a + `Canonical name` column, and drops its `Cross-language?` column: no + serializer code makes these keys shareable across SDKs, and the Cross-SDK Key + Generation Strategy section (and the README's implementor Quick Start) now + route all cross-SDK sharing through Interop Mode instead of namespace-matched + keys. Alias spellings (`std`, `pythonic`) and + the instance identity (`:` + bare class name, so one class with + different constructor arguments shares one identity) are documented in a new + Python SDK note, which points at that rule. + Documents the defect corrected by + [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) + (merged as `ee65250`; ships in cachekit-py 0.20.0): through v0.19.0 every + auto-mode key ended in `s` regardless of serializer, so two caches over one + function differing only in serializer shared a key and evicted each other on + every read. +- SDK-implementor pseudocode: the block previously defaulted `serializer_type` + to cachekit-py's `std` spelling and indexed a code table it never defined. It + now defines the table and `serializer_code()`, and adds an alias map and a + `normalize_identity()` hook, both marked SDK-supplied; the hook runs before + alias and table lookup, with the purity constraint the one-identity-one-code + rule already requires. `serializer_type` is a required parameter, and + `serializer_code()` rejects an empty or non-string identity with an error + rather than mapping it to a code. +- [`spec/wire-format.md`](spec/wire-format.md): the CK v3 frame's header `s` + table adds a serializer instance's bare class name, and the framing statement + excepts the two in-process modes that store no bytes: reference caching (`l`) + and `backend=None`. +- Provenance: no vector changed; the derived codes are not yet covered by + vectors. diff --git a/spec/cache-key-format.md b/spec/cache-key-format.md index 78f2b4c..3eb6c07 100644 --- a/spec/cache-key-format.md +++ b/spec/cache-key-format.md @@ -6,7 +6,7 @@ **Deterministic key generation from function identity and arguments.** -*Protocol Version 1.0 · Verified against `cachekit-py` v0.12.0 (`src/cachekit/key_generator.py`)* +*Protocol Version 1.0 · Serializer-code derivation verified against `cachekit-py` @ `ee65250` (the [cachekit-io/cachekit-py#311](https://github.com/cachekit-io/cachekit-py/pull/311) merge)* @@ -50,18 +50,110 @@ ns:{namespace}:func:{module}.{qualname}:args:{blake2b_hash}:{ic_flag}{serializer | `func:{module}.{qualname}:` | Function identifier (module path + qualified name) | `func:myapp.services.get_user:` | | `args:{blake2b_hash}:` | Blake2b-256 hash of normalized, MessagePack-serialized arguments | `args:a3c8d4...f2e1:` | | `{ic_flag}` | Integrity checking: `1` = ByteStorage enabled, `0` = raw MessagePack | `1` | -| `{serializer_code}` | Serializer type (1 char) | `s` | +| `{serializer_code}` | Serializer identity (1 char, or `x` + 4 hex — see below) | `s` | ### Serializer Codes -| Code | Serializer | Cross-language? | -| :---: | :--- | :---: | -| `s` | StandardSerializer (MessagePack) | ✅ Yes | -| `a` | AutoSerializer (Python-specific) | ❌ No | -| `o` | OrjsonSerializer (JSON-based) | ⚠️ Partial | -| `w` | ArrowSerializer (columnar) | ⚠️ Partial | +| Code | Serializer | Canonical name | +| :---: | :--- | :--- | +| `s` | StandardSerializer (MessagePack) | `default` | +| `a` | AutoSerializer (language-specific types) | `auto` | +| `o` | OrjsonSerializer (JSON-based) | `orjson` | +| `w` | ArrowSerializer (columnar) | `arrow` | +| `l` | Reference caching (no serialization) | `local` | +| `x` + 4 hex | Any serializer identity not in this table | — | -For cross-SDK interoperability, always use `s` (StandardSerializer). +No code makes these keys shareable across SDKs; see +[Cross-SDK Key Generation Strategy](#cross-sdk-key-generation-strategy). + +An identity outside the table gets `x` followed by the 2-byte digest +`blake2b(utf8(identity), digest_size=2)` encoded as exactly 4 lowercase hexadecimal +characters — zero-padded, no `0x` prefix, the same encoding as the args hash. Example: +identity `cbor` → `x23d5`. Codes are 1 character for the table entries and 5 for everything +else. + +> [!IMPORTANT] +> **The code MUST be derived from the serializer the cache is configured with, never a +> fixed default.** An SDK that emits one constant code collapses every serializer onto a +> single keyspace: two caches over one function then share a key, each fails the other's +> serializer-name check on read, evicts, and recomputes — a permanent 0% hit rate. +> +> **Two serializer identities that the wire format records differently MUST NOT be mapped +> onto one code by construction.** The guarantee is probabilistic, not absolute: the derived +> code carries 16 bits, so two identities it records differently can still collide, at ≈1 in +> 2^16 per pair. Between honestly written entries, such a collision costs hit rate only — +> that one pair evicts each other exactly as a constant code makes every pair do — and never +> yields a wrong value, because the stored serializer name still differs and the read-side +> check below rejects it. The recorded name is not an integrity control against a writer with +> backend write access: in the CK v3 frame it is a plaintext header field that is not an +> [AAD](encryption.md#additional-authenticated-data-aad) input, so nothing authenticates it, +> even for encrypted entries (see the [frame header caution](wire-format.md#python-ck-v3-frame)). +> +> **An SDK that offers more than one serializer identity MUST record the serializer name in +> the storage container of every serialized entry it stores under a key in this format +> (Python: the `s` field of the [CK v3 frame](wire-format.md#python-ck-v3-frame)). On every +> read of a serialized entry under a key in this format, before decoding the payload, such an +> SDK MUST compare the name the container records with the name it would itself record for +> its configured serializer, and MUST reject the entry on mismatch — a miss, never a value. +> For such an SDK, an entry that records no serializer name is a mismatch.** A colliding +> entry has the same key as the reader's own, so nothing before this comparison can tell them +> apart. An SDK that offers exactly one serializer identity is exempt from both the recording +> and the comparison, because there is no second serializer in its keyspace to confuse with +> the first (`cachekit-ts` and `cachekit-rs` offer one and record none). Neither rule reaches +> a value that is not a serialized entry under a key in this format: +> [Interop Mode](interop-mode.md) entries carry no serializer code and no container, and a +> cache that keeps live objects in process memory stores nothing to compare (in +> `cachekit-py`: reference caching, code `l`, and caches configured with no backend). +> +> `cache_key` is an AES-256-GCM AAD input (see [Encryption](encryption.md)). Two serializers +> sharing a code therefore share the **`cache_key` AAD component**, and a reader takes the +> `format` component from the entry's stored metadata, not from its own serializer (see +> [Encryption](encryption.md#format-tokens)). AAD binding does not separate them, +> whatever their `format` tokens, and the cipher is not a backstop for a missing name check. +> +> **Conversely, one identity MUST always produce one code.** Derive it from the serializer +> configuration alone — the canonical name the wire format records, or an SDK-defined +> refinement of it that never merges two names (see the Python note below; a refinement +> separates keyspaces, the read-side check still sees only the recorded name) — never from a +> process-local value such as an object address or a randomised hash, or keys stop being +> reproducible across processes. +> +> **Deriving the identity.** An SDK that accepts alias spellings MUST resolve each accepted +> spelling to exactly one canonical name, and the code, the recorded name, and the serializer +> that writes the bytes MUST all be that canonical name's, or a key, its stored entry and its +> bytes can disagree about which serializer wrote it. An SDK that accepts a serializer object +> MUST reduce it to a string identity carrying a marker that no code-table name, alias +> spelling or accepted serializer name contains, so a user-named class can never take a table +> code (the Python note below gives `cachekit-py`'s marker). An empty or non-string identity +> MUST be rejected with an error, never mapped to a code: a fallback code is a shared bucket, +> and a key computed from it names an entry nothing wrote. +> +> **An SDK SHOULD make the identity distinguish configurations that write different bytes. +> Wherever its identity does not, configurations that write different bytes and would +> otherwise share a key MUST be keyed under different `ns:` namespaces**, no namespace +> counting as one. Sharing an identity, they share a code, so a key, and a recorded name, so +> the read-side check cannot tell them apart: one is served the other's bytes as a hit — +> wrong data, not an eviction. The `ns:` MUST also covers distinct identities that share one +> recorded name and write different bytes, because the 16-bit code is NOT a +> collision-resistant separator: two such identities collide at ≈1 in 2^16 per pair, and key +> and recorded name then both match. The hit-rate-only collision guarantee above covers only +> identities recorded differently. The Python SDK is the known case where the identity does not distinguish configurations (see the note below). + +> [!NOTE] +> **Python SDK specifics.** `cachekit-py` additionally accepts the alias spellings `std` for +> `default` and `pythonic` for `auto`, canonicalizing them before the lookup. (Its key +> generator's alias map also lists `standard`, which the cache rejects as a serializer name.) +> A serializer passed as an *instance* rather than a name is recorded in the frame header +> under its bare class name, built-ins included, and its key identity is `:` + that +> class name, so it takes a derived code (`ArrowSerializer()` → `:ArrowSerializer` → +> `x2263`), never the table's. The prefix contains characters no Python identifier can, so a +> custom class named `auto` cannot take AutoSerializer's code. Beyond that fixed prefix, the +> identity and the recorded name both carry only the bare class name (`__name__`), so two +> instances of one class with different constructor arguments, or of any two classes sharing +> that name whatever their module or nesting, get one code and one recorded name. Where they +> write different bytes and share a `func:` segment (one function, or closures from one +> factory), the `ns:` rule above applies — across a deploy too: a changed configuration or +> implementation takes a namespace the old one never wrote. ### Example Keys @@ -82,7 +174,7 @@ ns:cache:func:app.views.index:args:0000...0000:0s - If a key exceeds 250 characters: first 50 chars of original key + `:` + first 32 chars of a Blake2b-256 hash of the full key > [!WARNING] -> **Discrepancy with RFC** — The original protocol RFC (Section 3.1.5) specifies a simpler key format: `{namespace}:{hash}`. The actual implementation includes function identity (`func:` prefix) and metadata suffix (`:{ic_flag}{serializer_code}`). **The implementation is authoritative.** For cross-SDK interoperability, SDK implementors must use explicit namespaces (the `func:` segment is language-specific and will differ). +> **Discrepancy with RFC** — The original protocol RFC (Section 3.1.5) specifies a simpler key format: `{namespace}:{hash}`. The actual implementation includes function identity (`func:` prefix) and metadata suffix (`:{ic_flag}{serializer_code}`). **The implementation is authoritative.** For cross-SDK sharing, use [Interop Mode](interop-mode.md) (see [Cross-SDK Key Generation Strategy](#cross-sdk-key-generation-strategy)). --- @@ -109,16 +201,7 @@ generation, invisible to the server. ## Cross-SDK Key Generation Strategy -For multi-language interoperability, all SDKs MUST use **explicit namespaces** rather than auto-generated function signatures. The `func:` segment is inherently language-specific (Python modules vs PHP namespaces vs Go packages), so cross-language cache sharing requires: - -1. All SDKs agree on a namespace string (e.g., `"get_user"`) -2. All SDKs serialize arguments identically (see [Argument Hashing Algorithm](#argument-hashing-algorithm) below) -3. The resulting Blake2b hash in the `args:` segment will be identical - -The `func:` and metadata segments may differ between SDKs — this is acceptable when the key is constructed to match. - -> [!TIP] -> Use [Interop Mode](interop-mode.md) to remove the `func:` segment entirely. Interop mode produces the simplest possible cross-language key: `{namespace}:{operation}:{args_hash}`. +Keys in this format are not shared across SDKs, even under an agreed namespace. The `func:` segment is language-specific, and the values stored under these keys are SDK-internal containers that no other SDK decodes (see [SDK Storage Containers](wire-format.md#sdk-storage-containers-auto-mode)). Cross-language cache sharing uses [Interop Mode](interop-mode.md) exclusively: explicit, language-neutral operation names, keys of the form `{namespace}:{operation}:{args_hash}`, and plain MessagePack values. --- @@ -220,7 +303,7 @@ After key construction, the following characters are replaced: ## Test Vectors -[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}`. These vectors are **Python-SDK-only**: the `func:` segment is language-specific, so no other SDK can reproduce these keys or share the cache entries they name. Cross-SDK conformance uses [`test-vectors/interop-mode.json`](../test-vectors/interop-mode.json) (see [Interop Mode](interop-mode.md)). +[`test-vectors/cache-keys.json`](../test-vectors/cache-keys.json) contains 10 auto-mode key vectors (`args` + `kwargs` + metadata → `expected_key`) covering primitives, mixed args/kwargs, `null`, booleans, nested dicts, and the no-namespace form. They were generated by `cachekit-py` v0.12.0 and the `vectors` array has not changed since; every vector uses `serializer_type: "std"` (→ `1s`), cachekit-py's alias for the canonical `default`, so the derived codes above are not yet covered. Keys were generated at top level, so the `func:` segment is `__main__.{qualname}`. These vectors are **Python-SDK-only**: the `func:` segment is language-specific, so no other SDK can reproduce these keys or share the cache entries they name. Cross-SDK conformance uses [`test-vectors/interop-mode.json`](../test-vectors/interop-mode.json) (see [Interop Mode](interop-mode.md)). Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte-verified against `CacheKeyGenerator` on every default CI run (`tests/unit/protocol/test_cache_key_vectors.py`). A vector failing there is a key-stability break to triage — never silently regenerate: a changed key orphans every existing cache entry and turns the fleet's hits into billed misses. @@ -232,8 +315,53 @@ Enforcement: the vectors are vendored (sha256-pinned) into cachekit-py and byte- Expand full pseudocode ``` +SERIALIZER_CODES = {"default": "s", "auto": "a", "orjson": "o", "arrow": "w", "local": "l"} + +// SDK-SUPPLIED, not fixed by this spec: alias spellings THIS SDK accepts -> canonical +// name. Empty if the SDK accepts only canonical names. Using this map for the recorded name +// too, with the writer resolving to the same serializer, satisfies "Deriving the identity" +// in Serializer Codes above. Do not adopt another SDK's aliases: mapping a spelling your API +// does not accept hands that name a table code instead of the derived `x` code it should +// get. cachekit-py's accepted aliases are in the Python note above. +SERIALIZER_ALIASES = {} // e.g. cachekit-py accepts: {"std": "default", "pythonic": "auto"} +// test-vectors/cache-keys.json records serializer_type "std": cachekit-py's alias for the +// canonical "default", so every vector's code is "s". + +// SDK-SUPPLIED, not fixed by this spec: reduce whatever your API accepts as a serializer to +// the canonical STRING identity, before any lookup below. An SDK that accepts only names +// returns the name unchanged. One that also accepts a serializer OBJECT must convert it +// here — the lookups below are string operations and are undefined on an object. This is +// the "SDK-defined refinement" the one-identity-one-code rule permits. +// +// `normalize_identity()` MUST be a pure function of the serializer's configuration — never +// an object address or a randomised hash. The derived code is computed from this identity; +// the read-side check compares the recorded name instead, which the identity equals or +// refines (cachekit-py: `:ArrowSerializer` records `ArrowSerializer`). Whether the +// identity must distinguish configurations that write different bytes is the uniqueness +// rule in Serializer Codes above (SHOULD; where it does not, different `ns:` namespaces MUST). +// +// An object-to-string refinement carries a marker no table name, alias or accepted name +// contains ("Deriving the identity" above). cachekit-py maps an object to ":" + its +// bare class name (Python note above). +function normalize_identity(serializer_type): + return serializer_type // names-only SDK; override to handle objects + +function serializer_code(serializer_type): + // Reduce to a string identity, resolve any alias spelling this SDK accepts, then look + // the code up. An identity outside the table gets its OWN derived code — never a shared + // constant, which would put every unrecognised serializer on one keyspace. + name = normalize_identity(serializer_type) + // An empty or non-string identity is an error, never a code ("Deriving the identity"). + if name is not a string or name == "": + raise error + identity = SERIALIZER_ALIASES.get(name, name) + if identity in SERIALIZER_CODES: + return SERIALIZER_CODES[identity] + // digest .hex(): exactly 4 lowercase zero-padded hex chars — never a numeric hex() + return "x" + blake2b(identity.utf8_bytes(), digest_size=2).hex() + function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, - integrity_checking=true, serializer_type="std"): + integrity_checking=true, *, serializer_type): // Build key parts parts = [] @@ -251,10 +379,10 @@ function generate_cache_key(namespace, func_module, func_qualname, args, kwargs, parts.append("args:" + hash + ":") - // Metadata suffix + // Metadata suffix. serializer_type is the serializer the cache is CONFIGURED with — + // read it from the decorator/client configuration, never a fixed default. ic_flag = "1" if integrity_checking else "0" - serializer_code = SERIALIZER_CODES[serializer_type] // "s" for standard - parts.append(ic_flag + serializer_code) + parts.append(ic_flag + serializer_code(serializer_type)) key = join(parts) diff --git a/spec/wire-format.md b/spec/wire-format.md index 13d3e6d..c8ef4b2 100644 --- a/spec/wire-format.md +++ b/spec/wire-format.md @@ -507,7 +507,7 @@ Datetime values are encoded as MessagePack maps with sentinel keys: ## SDK Storage Containers (auto mode) Remote backends (Redis, CachekitIO SaaS, Memcached, File) store opaque bytes. (L1 -behavior is SDK-specific: `cachekit-py`'s L1 holds the framed bytes; `cachekit-ts`'s +behavior is SDK-specific: `cachekit-py`'s L1 in front of a backend holds the framed bytes; `cachekit-ts`'s L1 holds live decoded values, not bytes.) What the stored bytes *are* differs per SDK in auto mode: @@ -530,9 +530,11 @@ implementations ([protocol#11](https://github.com/cachekit-io/protocol/issues/11 ### Python: CK v3 frame -Every **auto-mode** value `cachekit-py` stores — all backends, all serializers, -encrypted or not — is framed (interop-mode values are plain MessagePack, never -framed): +Two in-process modes keep live objects and store no bytes at all: `@cache.local` +reference caching (key code `l`) and a cache configured with `backend=None`, whose +keys carry its configured serializer's code. Every other **auto-mode** value +`cachekit-py` stores — all backends, all serializers, encrypted or not — is framed +(interop-mode values are plain MessagePack, never framed): ```text MAGIC b"CK" (0x43 0x4B) | VERSION u8 (0x03) | HDR_LEN u32 big-endian | HEADER | PAYLOAD @@ -548,6 +550,7 @@ MAGIC b"CK" (0x43 0x4B) | VERSION u8 (0x03) | HDR_LEN u32 big-endian | HEADER | | `default`, `auto` | ByteStorage envelope (this document) over MessagePack | | `arrow` | **Arrow envelope**: `[8-byte xxHash3-64 checksum][Arrow IPC file]` (IPC magic `b"ARROW1"` at payload offset 8) | | `orjson` | `[8-byte xxHash3-64 checksum][JSON bytes]` | +| A serializer instance's bare class name (`StandardSerializer`, `ArrowSerializer`, a custom class) | That serializer's own output; a built-in class writes the same payload as its string name above. Classes sharing a bare name, and differently configured instances of one class, record the same `s` (see the [`ns:` rule](cache-key-format.md#serializer-codes)) | | any, encrypted | Ciphertext per [encryption.md](encryption.md) | With integrity checking disabled, `default`/`auto` payloads are raw MessagePack (no