Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -568,9 +568,9 @@ def get_data():

See [Changing Serializers](serializers/README.md#changing-serializers-separate-keyspaces)
for the code table, the cold-cache warning, and the data-retention caveat on orphaned entries.
Upgrading from v0.18 or earlier? Keys for non-default serializers change identity without any
Upgrading from v0.19 or earlier? Keys for non-default serializers change identity without any
change on your side — see
[Breaking change in v0.19.0](serializers/README.md#breaking-change-in-v0190-the-key-carries-the-real-serializer).
[Breaking change in v0.20.0](serializers/README.md#breaking-change-in-v0200-the-key-carries-the-real-serializer).

**Best Practice**: Use namespace versioning for zero-downtime migrations:

Expand Down
2 changes: 1 addition & 1 deletion docs/features/l1-invalidation.md
Original file line number Diff line number Diff line change
Expand Up @@ -174,7 +174,7 @@ for uid in [1, 2, 3]:
- User data refresh
- Post cache invalidation

**Effect:** The entry is removed from this process's L1 cache **and**, when an L2 backend is configured, deleted from shared L2. Cache keys are deterministic, so the L2 delete removes the entry no matter which process wrote it. In L1-only mode (`backend=None`) there is no L2 to delete from — the invalidation is purely local. If the L2 delete fails, cachekit logs an ERROR `Failed to delete L2 key` and tracks the key in this process, so a later no-args `invalidate_cache()` from the same process retries it. Tracking is process-local: another process, or this one after a restart, does not retry it.
**Effect:** The entry is removed from this process's L1 cache **and**, when an L2 backend is configured, deleted from shared L2. Cache keys are deterministic, so the L2 delete removes the entry no matter which process wrote it. On a generated key with a non-default serializer it also deletes the key a pre-v0.20.0 release wrote for the same arguments (see [the v0.20.0 key change](../serializers/README.md#breaking-change-in-v0200-the-key-carries-the-real-serializer)). In L1-only mode (`backend=None`) there is no L2 to delete from — the invalidation is purely local. If the L2 delete fails, cachekit logs an ERROR `Failed to delete L2 key` and tracks the key in this process, so a later no-args `invalidate_cache()` from the same process retries it. Tracking is process-local: another process, or this one after a restart, does not retry it.

### Whole-Function Invalidation

Expand Down
55 changes: 39 additions & 16 deletions docs/serializers/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,11 +41,11 @@ For caching Pydantic models, see [Caching Pydantic Models](pydantic.md).

## Migration Guide

### Breaking change in v0.19.0: the key carries the real serializer
### Breaking change in v0.20.0: the key carries the real serializer

Before v0.19.0 the serializer half of the key suffix was a constant: **every key ended `:1s`
whatever serializer was configured** (`:0s` with `integrity_checking=False`). From v0.19.0 the
code reflects the serializer in use — the "Before v0.19.0" column in the table below — so keys
Before v0.20.0 the serializer half of the key suffix was a constant: **every key ended `:1s`
whatever serializer was configured** (`:0s` with `integrity_checking=False`). From v0.20.0 the
code reflects the serializer in use — the "Suffix" column in the table below — so keys
change identity on upgrade, with no change on your side, for:

- `serializer="auto"` / `"pythonic"`, `"orjson"`, `"arrow"` → now `:1a`, `:1o`, `:1w`;
Expand All @@ -60,12 +60,28 @@ pre-upgrade entries are never read again. Two decorators over one function that
in serializer used to evict each other on every call through the `Serializer mismatch` path;
they now coexist.

**Before you deploy:** pre-upgrade entries are unreachable from the upgraded code, so
`invalidate_cache()` cannot delete them. If you cache personal data, or rely on invalidation
reaching entries written before the upgrade, follow the retention warning
[below](#changing-serializers-separate-keyspaces) — **after the last v0.18 replica is
**What invalidation still reaches.** Upgraded code never *reads* a pre-upgrade entry, but
single-key invalidation — `fn.invalidate_cache(*args)` / `await fn.ainvalidate_cache(*args)`
on a decorator with a generated key — deletes both the current key and the pre-v0.20.0
`:{integrity_flag}s` key for the same arguments. An erasure through the SDK therefore removes
the pre-upgrade copy too, including one an old replica wrote during a rolling deploy. The
cost: a default-serializer decorator over the same function, namespace and arguments loses
that entry and recomputes once.

The reverse direction is not covered. An erasure served by a v0.19 replica during the rollout,
or by any replica after a rollback to v0.19, deletes only the `:{integrity_flag}s` key, so a
copy that a v0.20.0 replica wrote under the new key survives. Re-issue any erasure made during
the rollout once the last v0.19 replica is retired, or cover it with the flush below.

**What still needs a backend flush.** The SDK cannot reach a pre-upgrade entry whose
arguments you never invalidate. On a function that takes parameters, no-argument
`invalidate_cache()` / `cache_clear()` does not reach them either: it deletes the keys this process tracked plus, on the tenant-scoped Redis
backend, the keys in the server-side key registry, and releases before v0.20.0 recorded their
keys in neither. So if you cache personal data under `ttl=None`, or otherwise need every pre-upgrade
entry gone rather than aging out, follow the flush procedure in the retention warning
[below](#changing-serializers-separate-keyspaces) — **after the last v0.19 replica is
retired**, not at the start of a rolling deploy, or replicas still on the old release keep
writing `:1s` entries behind your flush. A `namespace=` bump gives an explicit cut-over but
writing `:{integrity_flag}s` entries behind your flush. A `namespace=` bump gives an explicit cut-over but
orphans the old keyspace rather than deleting it; the retention step still applies.

### Changing Serializers: Separate Keyspaces
Expand All @@ -74,7 +90,7 @@ The serializer is part of the cache key. The key's trailing metadata suffix is
`{integrity_flag}{serializer_code}` — `1`/`0` for integrity checking, then one character
for the serializer:

| Configured as | Code | Suffix | Before v0.19.0 |
| Configured as | Code | Suffix | Before v0.20.0 |
| :--- | :---: | :--- | :--- |
| `serializer="std"` / `"default"` / `"standard"` (the default) | `s` | `:1s` | `:1s` |
| `serializer="auto"` / `"pythonic"` | `a` | `:1a` | `:1s` |
Expand Down Expand Up @@ -134,13 +150,20 @@ def get_data():
> On a hot path, roll it out behind your usual warm-up or stampede controls.

> [!WARNING]
> **Orphaned entries are a data-retention question, not just a hit-rate one.** Once the key
> changes, `invalidate_cache()` computes the *new* key and can no longer reach the old copy
> — a deletion for erasure, consent withdrawal or permission revocation will report success
> while the previous entry survives until its TTL expires, or indefinitely if no TTL is set.
> **Orphaned entries are a data-retention question, not just a hit-rate one.** When you
> change a function's serializer, `invalidate_cache()` computes the *new* key and cannot
> reach the old copy — a deletion for erasure, consent withdrawal or permission revocation
> will report success while the previous entry survives until its TTL expires, or
> indefinitely if no TTL is set. The one old key it does reach is the default serializer's
> `:{integrity_flag}s` key, kept for the v0.20.0 upgrade (see
> [above](#breaking-change-in-v0200-the-key-carries-the-real-serializer)). Any move *away*
> from a non-default serializer is not covered — to another one (`"auto"` to `"arrow"`) or
> to the default (`"auto"` to `"default"`, or to `@cache.secure` on its default serializer).
> If you cache personal data, **flush the affected namespace** when you change a serializer
> or upgrade across a release that re-keys it, rather than relying on expiry. The SDK has no
> bulk delete — `cache_clear()` only knows the keys the current process wrote — so flush on
> rather than relying on expiry, and after upgrading to v0.20.0 for any entries that
> single-key invalidation will not reach. The SDK has no
> bulk delete — `cache_clear()` reaches only keys this release tracked, never a pre-upgrade
> one — so flush on
> the backend: on Redis, `SCAN` for the key prefix (`ns:<namespace>:*`) and `UNLINK` the
> matches; the File backend stores one file per hashed key in `cache_dir`, so the only flush
> is the whole directory. Memcached and CachekitIO offer no pattern delete, so old entries
Expand Down
53 changes: 47 additions & 6 deletions src/cachekit/cache_handler.py
Original file line number Diff line number Diff line change
Expand Up @@ -1434,14 +1434,55 @@ def get_cache_key(
>>> key1 == key2 # _bypass_cache doesn't affect key
True
"""
return self._generated_key(
func, args, kwargs, namespace, integrity_checking, self.serialization_handler.serializer_key_name
)

# Pre-0.20.0 releases never passed the serializer to generate_key, so every generated
# key ended in the default code `s` whatever the serializer was (LAB-4351). A
# deployment upgraded from one still holds those entries — and old replicas keep
# writing them during a rolling deploy — at a key the current code never computes.
# Invalidating only the current key would let an erasure return normally while the
# pre-upgrade copy survives to its TTL, or forever at ttl=None, so single-key
# invalidation also deletes this one. Over-deleting costs a `default`-serializer
# decorator on the same function and arguments one recompute; under-deleting leaks
# retained data. Remove only in a major release whose notes declare upgrades from below
# 0.20.0 unsupported: ttl=None entries never age out, so no TTL clock can retire this.
def get_legacy_cache_key(
self,
func: Callable[..., Any],
args: tuple[Any, ...],
kwargs: dict[str, Any],
namespace: str | None,
integrity_checking: bool = True,
) -> str:
"""The key a pre-0.20.0 release wrote for this call; equal to get_cache_key's on the default serializer.

Examples:
>>> from cachekit.key_generator import CacheKeyGenerator
>>> def my_func(x): return x
>>> auto = CacheOperationHandler(CacheSerializationHandler("auto"), CacheKeyGenerator())
>>> auto.get_cache_key(my_func, (1,), {}, None)[-3:], auto.get_legacy_cache_key(my_func, (1,), {}, None)[-3:]
(':1a', ':1s')
>>> std = CacheOperationHandler(CacheSerializationHandler(), CacheKeyGenerator())
>>> std.get_legacy_cache_key(my_func, (1,), {}, None) == std.get_cache_key(my_func, (1,), {}, None)
True
"""
return self._generated_key(func, args, kwargs, namespace, integrity_checking, "default")

def _generated_key(
self,
func: Callable[..., Any],
args: tuple[Any, ...],
kwargs: dict[str, Any],
namespace: str | None,
integrity_checking: bool,
serializer_type: str,
) -> str:
"""generate_key minus reserved kwargs: one filter, so a legacy key hashes the same arguments."""
filtered_kwargs = {k: v for k, v in kwargs.items() if k != "_bypass_cache"}
return self.key_generator.generate_key(
func,
args,
filtered_kwargs,
namespace,
integrity_checking,
serializer_type=self.serialization_handler.serializer_key_name,
func, args, filtered_kwargs, namespace, integrity_checking, serializer_type=serializer_type
)

def _handle_l2_read_error(self, e: SerializationError, cache_key: str) -> None:
Expand Down
42 changes: 30 additions & 12 deletions src/cachekit/decorators/wrapper.py
Original file line number Diff line number Diff line change
Expand Up @@ -1132,8 +1132,16 @@ def _interop_cache_key(call_args: tuple[Any, ...], call_kwargs: dict[str, Any])
flat = bind_flat_args(_interop_sig, call_args, call_kwargs)
return generate_interop_key(namespace, interop, flat)

# The generated key is the only one carrying a serializer code, so it alone has a
# pre-0.20.0 twin. One flag, read by both functions below, so the twin can never be
# computed for a key the write path did not generate.
_generated_key_mode = interop is None and custom_key_func is None and not fast_mode

def _resolve_cache_key(call_args: tuple[Any, ...], call_kwargs: dict[str, Any]) -> str:
"""Single key derivation shared by the read/write and invalidate paths (LAB-4387)."""
# Standard key generation with type-aware handling
if _generated_key_mode:
return operation_handler.get_cache_key(func, call_args, call_kwargs, namespace, integrity_checking)
# Interop mode takes priority (mutually exclusive with key= and fast_mode)
if interop is not None:
return _interop_cache_key(call_args, call_kwargs)
Expand All @@ -1143,13 +1151,18 @@ def _resolve_cache_key(call_args: tuple[Any, ...], call_kwargs: dict[str, Any])
if not isinstance(custom_key, str):
raise TypeError(f"key function must return str, got {type(custom_key).__name__}")
return f"{namespace or 'default'}:{custom_key}"
if fast_mode:
# Minimal key generation - no string formatting overhead (10-50μs savings)
from ..hash_utils import cache_key_hash
# fast_mode: minimal key generation - no string formatting overhead (10-50μs savings)
from ..hash_utils import cache_key_hash

return (namespace or "default") + ":" + func_hash + ":" + cache_key_hash(str(call_args) + str(call_kwargs))
# Standard key generation with type-aware handling
return operation_handler.get_cache_key(func, call_args, call_kwargs, namespace, integrity_checking)
return (namespace or "default") + ":" + func_hash + ":" + cache_key_hash(str(call_args) + str(call_kwargs))

def _resolve_invalidation_keys(call_args: tuple[Any, ...], call_kwargs: dict[str, Any]) -> list[str]:
"""The key _resolve_cache_key derives, plus its pre-0.20.0 twin on the generated-key path."""
cache_key = _resolve_cache_key(call_args, call_kwargs)
if not _generated_key_mode:
return [cache_key]
legacy_key = operation_handler.get_legacy_cache_key(func, call_args, call_kwargs, namespace, integrity_checking)
return [cache_key] if legacy_key == cache_key else [cache_key, legacy_key]

# Track the cache keys this process wrote or read for this function (for no-args
# invalidation). Key normalization (hashing of long keys) makes prefix matching
Expand Down Expand Up @@ -2285,6 +2298,11 @@ def _invalidate_key(cache_key: str) -> None:
elif _l1_cache:
_l1_cache.invalidate(cache_key)

def _invalidate_keys(cache_keys: list[str]) -> None:
"""_invalidate_key per key: each logs its own failure, so one never skips the next."""
for cache_key in cache_keys:
_invalidate_key(cache_key)

def invalidate_cache(*args: Any, **kwargs: Any) -> None:
nonlocal _backend

Expand All @@ -2311,8 +2329,8 @@ def invalidate_cache(*args: Any, **kwargs: Any) -> None:
return

# Single-key invalidation (specific args provided, or zero-param function).
# Same derivation as the write path — one key, not two (LAB-4387).
_invalidate_key(_resolve_cache_key(args, kwargs))
# Same derivation as the write path (LAB-4387), plus the pre-0.20.0 twin (LAB-5288).
_invalidate_keys(_resolve_invalidation_keys(args, kwargs))

async def ainvalidate_cache(*args: Any, **kwargs: Any) -> None:
nonlocal _backend
Expand All @@ -2338,10 +2356,10 @@ async def ainvalidate_cache(*args: Any, **kwargs: Any) -> None:
return

# Single-key invalidation (specific args provided, or zero-param function).
# Same derivation as the write path — one key, not two (LAB-4387). The sync delete runs
# off the loop; never a backend's delete_async, whose pooled client may be bound to an
# earlier event loop (asyncio.run per job).
await asyncio.to_thread(_invalidate_key, _resolve_cache_key(args, kwargs))
# Same derivation as the write path (LAB-4387), plus the pre-0.20.0 twin (LAB-5288). The
# sync deletes run off the loop; never a backend's delete_async, whose pooled client may
# be bound to an earlier event loop (asyncio.run per job).
await asyncio.to_thread(_invalidate_keys, _resolve_invalidation_keys(args, kwargs))

def check_health() -> dict[str, Any]:
"""Check health status of this cached function's infrastructure."""
Expand Down
29 changes: 29 additions & 0 deletions tests/unit/test_error_path_key_redaction.py
Original file line number Diff line number Diff line change
Expand Up @@ -247,6 +247,35 @@ async def cached_func(user: str) -> str:
assert len(backend.received_keys) == 1
_assert_redacted(caplog, backend.received_keys[0])

async def test_legacy_key_failure_logs_redacted_at_error(self, caplog: pytest.LogCaptureFixture) -> None:
"""A non-default serializer also deletes the pre-0.20.0 key (LAB-5288); both failures log redacted at ERROR."""
error = ValueError(f"illegal input: {TENANT_KEY}")

for invoke in ("sync", "async"):
backend = _FailingBackend(error)

@cache(backend=backend, l1_enabled=False, namespace="tenant-42-secret", serializer="auto")
def cached_func(user: str) -> str:
return user

@cache(backend=backend, l1_enabled=False, namespace="tenant-42-secret", serializer="auto")
async def acached_func(user: str) -> str:
return user

caplog.clear()
with caplog.at_level(logging.ERROR):
if invoke == "sync":
cached_func.invalidate_cache("alice")
else:
await acached_func.ainvalidate_cache("alice")

current_key, legacy_key = backend.received_keys
assert legacy_key.endswith(":1s") and not current_key.endswith(":1s")
error_records = [r for r in caplog.records if r.levelno == logging.ERROR]
assert len(error_records) == 2, invoke
_assert_redacted(caplog, current_key)
_assert_redacted(caplog, legacy_key)


class TestKeyCarryingBackendErrorRedaction:
"""A BackendError that carries the raw key must not leak it through ``{e}``.
Expand Down
Loading
Loading