Skip to content

Repository files navigation

cachekit-rs

Caching for Rust — dual-layer L1/L2, zero-knowledge encryption, multi-backend.

Crates.io docs.rs License: MIT MSRV

Features · Quick Start · Encryption · Backends · Architecture


Status: beta — CacheKit is in closed beta ahead of 1.0. APIs are stabilising; minor breaking changes may still occur between 0.x releases.


Overview

cachekit-rs is the Rust SDK for cachekit.io. Pick an intent preset — minimal, production, secure, or io — and get a pre-configured cache in one call, from bare Redis speed to dual-layer with client-side encryption. Pick secure and bytes never leave your process in plaintext.

Component What it does
CacheKit get / set / delete / exists with automatic L1 → L2 layering
SecureCache Transparent AES-256-GCM encryption before storage (zero-knowledge)
Backend Pluggable trait — cachekit.io SaaS, Redis, Memcached, local File, Cloudflare Workers
L1 Cache In-process moka cache with write-through + backfill

Tip

For the Python SDK with decorators, see cachekit. For the low-level compression/encryption primitives, see cachekit-core.


Features

Feature Default Description
cachekitio ✅ HTTP backend for api.cachekit.io via reqwest + rustls
encryption ✅ Zero-knowledge AES-256-GCM via cachekit-core
l1 ✅ In-process L1 cache via moka, with stale-while-revalidate (native)
reliability ✅ Retry with backoff + jitter, circuit breaker, backpressure, distributed fill locks (native only)
redis ❌ Redis backend via fred (native only)
memcached ❌ Memcached backend via rust-memcache (native only)
file ❌ Local filesystem backend, byte-compatible with cachekit-py's File backend (native only)
workers ❌ Cloudflare Workers backend via worker
macros ❌ #[cachekit] proc-macro decorator (mints interop/v1 keys)
tracing ❌ tracing events per cache operation and breaker transition — see Observability
# Defaults: SaaS + encryption + L1
[dependencies]
cachekit-rs = "0.8"

# With Redis backend
[dependencies]
cachekit-rs = { version = "0.8", features = ["redis"] }

# For Cloudflare Workers (no L1, no Redis)
[dependencies]
cachekit-rs = { version = "0.8", default-features = false, features = ["workers", "encryption"] }

Warning

Mutually exclusive features:

  • workers + redis — Workers runtime cannot use fred
  • workers + l1 — moka requires std threads unavailable in wasm32
  • workers + reliability — retry/breaker timers need tokio time, unavailable in wasm32
  • workers + memcached — Workers runtime has no TCP sockets
  • workers + file — Workers runtime has no filesystem

Quick Start

Intent Presets (recommended)

One call that names your use case. Each preset returns a pre-configured builder you can still override before .build():

Preset When to use Backend L1 Encryption Reliability¹ Auto-reconnect² Default TTL
CacheKit::minimal(url) Development, public data, product catalogs — speed first, no extras Redis³ ✅ (no SWR) ❌ ❌ ❌ 300 s
CacheKit::production(url) User sessions, API responses, production services Redis³ ✅ ❌ ✅ ✅ 600 s
CacheKit::secure(url, master_key_hex)⁵ PII, payments, GDPR/HIPAA-sensitive data — zero-knowledge AES-256-GCM Redis³ ✅ ✅ ✅ ✅ 600 s
CacheKit::io(api_key)⁴ Serverless, edge compute, managed caching without running Redis cachekit.io ✅ ❌ ✅ n/a (HTTP) 3 600 s

¹ Retry with backoff + jitter, circuit breaker, backpressure — the reliability stack. Requires the default-on reliability feature. ² See the resilience contract below. ³ Requires the redis feature flag; secure also needs the default-on encryption feature. ⁴ Or CacheKit::io_from_env() to read CACHEKIT_API_KEY. ⁵ Or CacheKit::secure_from_env(url) to read CACHEKIT_MASTER_KEY, plus the decrypt-only rotation keys in CACHEKIT_PREVIOUS_MASTER_KEYS (see Key Rotation). Both take the key as a hex string and decode it the same way every CacheKit SDK does. Use exactly 32 bytes (64 hex chars, openssl rand -hex 32) — the only length every SDK accepts.

use cachekit::prelude::*;

#[tokio::main]
async fn main() -> Result<(), CachekitError> {
    // Needs: cachekit-rs = { version = "0.8", features = ["redis"] }
    let cache = CacheKit::production("redis://localhost:6379").await?
        .namespace("api")
        .build()?;

    cache.set("greeting", &"Hello, world!").await?;
    let val: Option<String> = cache.get("greeting").await?;
    println!("{val:?}");

    Ok(())
}

Resilience contract — connection failures, at construction and mid-run:

  • production / secure auto-reconnect: a dropped connection is re-established with exponential backoff (100 ms → 30 s cap), retrying indefinitely.
  • minimal is fail-fast: a dropped connection is not re-established — every subsequent operation that reaches Redis errors until you rebuild the client. Reads served from a warm L1 entry still return without contacting Redis.
  • Initial connections fail fast for every Redis preset: a bad URL or unreachable Redis errors immediately at construction, never enters a retry loop. io opens no connection at construction: an empty API key fails at construction, while an invalid key or unreachable endpoint surfaces at the first request.
  • secure / secure_from_env validate the master key (and secure_from_env any previous keys) before any Redis connection is attempted — a missing, non-hex or short key, or a malformed CACHEKIT_PREVIOUS_MASTER_KEYS, is a deterministic local error, never masked by (or paying for) network I/O, and never a fallback to plaintext.
  • Auto-reconnect is connection-level repair, distinct from the per-operation reliability stack (retry, circuit breaker, backpressure) that production / secure / io also enable. minimal has neither — every failure is yours to handle.

From Environment Variables

use cachekit::prelude::*;

#[tokio::main]
async fn main() -> Result<(), CachekitError> {
    let cache = CacheKit::from_env()?.build()?;

    cache.set("greeting", &"Hello, world!").await?;
    let val: String = cache.get("greeting").await?.unwrap();
    println!("{val}");

    Ok(())
}

Builder API

use std::sync::Arc;
use std::time::Duration;
use cachekit::prelude::*;
use cachekit::backend::cachekitio::CachekitIO;

let backend = CachekitIO::builder()
    .api_key("ck_live_...")
    .build()?;

let cache = CacheKit::builder()
    .backend(Arc::new(backend))
    .default_ttl(Duration::from_secs(600))
    .namespace("myapp")
    .l1_capacity(5000)
    .build()?;

Important

Never hardcode API keys or master keys. Use environment variables or a secrets manager.


Zero-Knowledge Encryption

Call .secure_cache() to get an encrypted cache handle. All values are encrypted client-side with AES-256-GCM before hitting any backend. The backend only ever sees ciphertext.

// Env: CACHEKIT_MASTER_KEY=<64 hex chars>
let cache = CacheKit::secure_from_env("redis://localhost:6379").await?.build()?;
let secure = cache.secure_cache()?;

// Encrypt → store (backend sees only ciphertext)
secure.set("user:42:ssn", &"123-45-6789").await?;

// Retrieve → decrypt (transparent to caller)
let ssn: String = secure.get("user:42:ssn").await?.unwrap();
┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│  Your Code   │────>│  SecureCache  │────>│   Backend    │
│              │     │  AES-256-GCM  │     │  (cachekit.io│
│  plaintext   │     │  encrypt /    │     │   or Redis)  │
│              │<────│  decrypt      │<────│              │
└──────────────┘     └──────────────┘     └──────────────┘
                      L1 stores ciphertext
                      (zero-knowledge preserved)
Security Properties
Property Implementation
Encryption AES-256-GCM (AEAD) via cachekit-core (ring on native, aes-gcm on wasm32)
Key Derivation HKDF-SHA256 — per-tenant cryptographic isolation
AAD Binding Cache key bound to ciphertext (prevents substitution attacks)
Memory Safety zeroize on drop for all key material
L1 Guarantee L1 stores ciphertext, never plaintext
Cache-key path encoding (CWE-22) Keys are percent-encoded into the CachekitIO request path; the empty key, and a key encoding to a reserved segment (., .., health, ttl, lock), are rejected rather than sent

Cache-key path encoding (CWE-22): CachekitIO keys are percent-encoded (urlencoding::encode) so a key can only ever address /v1/cache/{key}. The empty key, and a key whose encoded form is one of the five reserved path segments — ., .., health, ttl, lock — are rejected with a permanent error rather than sent (protocol spec/saas-api.md § Cache-Key Path Encoding, rule 2). The empty key encodes to an empty segment, so /v1/cache/{key} becomes /v1/cache/ and /v1/cache/{key}/ttl becomes /v1/cache//ttl, neither of which addresses a stored entry. ./.. are dot segments that reqwest's WHATWG URL parser (rust-url) strips before the request leaves the process (/v1/cache/.. → /v1/); health/ttl/lock are live route tokens (/v1/cache/health is the health endpoint, a trailing ttl/lock selects a sub-resource). Encoding can't neutralise either — WHATWG collapses %2E%2E too — so the SDK refuses rather than emit a request whose path was rewritten. This matches the cachekit-ts twin and is stricter than cachekit-py's older %2E rewrite; none of these is ever a canonical CacheKit key (those are non-empty and contain :), so nothing legitimate is affected. Every other key encodes to the same bytes as cachekit-py; cachekit-ts may leave ! * ' ( ) raw, and every conformant form decodes once to the same key (spec rule 4).

AAD v0x03 wire format:

[version(0x03)][len(4)][tenant_id][len(4)][cache_key][len(4)][format][len(4)][compressed]

Each field is length-prefixed with a 4-byte big-endian u32 to prevent boundary-confusion attacks. Cross-SDK compatible — ciphertext produced by the Python SDK decrypts with the Rust SDK and vice versa.

Is AES hardware-accelerated on this host? cache.secure_cache()?.hardware_acceleration_enabled() (also on EncryptionLayer) forwards cachekit-core's detection. Informational only: ring/aes-gcm pick their implementation independently, so use it to explain secure-cache latency, not to change behaviour. The per-architecture semantics are core's — as of cachekit-core 0.6.0 a runtime AES-NI probe on x86/x86_64, true on every aarch64 build (it tests NEON, which all aarch64 targets enable, not the Crypto Extension — a Raspberry Pi 4, a Cortex-A72 without the Crypto Extension, reports true while running software AES), and false on wasm32.

Key Rotation

Rotate the master key without invalidating existing entries: promote the new key to current and keep the old one as a decrypt-only previous key during a grace window (max 3, per the protocol keyring spec). Writes always use the current key; reads attempt the current key first, then each previous key in order. Old entries age out via TTL or re-encrypt on the next write — no bulk re-encryption.

// Env: CACHEKIT_MASTER_KEY=<k2-hex> CACHEKIT_PREVIOUS_MASTER_KEYS=<k1-hex>
let cache = CacheKit::from_env()?.build()?;

// The Redis `secure` preset reads the same two variables:
let cache = CacheKit::secure_from_env("redis://localhost:6379").await?.build()?;

// Or explicitly on the client builder, with exactly 32 raw bytes per key
// (decoded, never the ASCII of a hex string; any other length is an error):
let cache = CacheKit::builder()
    .backend(backend)
    .encryption_from_bytes_with_previous(&k2_bytes, &[&k1_bytes], "tenant")?
    .build()?;

Rotation is forward-only: a retired key is never re-promoted (re-promoting would resume a used AES-GCM nonce budget), and a config listing the current key among the previous keys is rejected at load. For the three-phase zero-miss rollout and compromise response, see the key rotation runbook.

Knowing when to drop the old key. Every read served by a previous key is counted against that key's position; cache.secure_cache()?.previous_key_hits() returns the counts (hits[i] for previous_keys[i], current-key reads not counted, no key material). The signal confirms a grace window has drained; it does not shorten one. Follow the protocol's scheduled-rotation runbook: audit for non-expiring entries, add the incoming key as decrypt-only fleet-wide, then promote it. The clock starts only when the promotion deploy has completed on every instance — a lagging instance still writes under the retiring key and reads it silently as its current key. From then, wait at least the longest TTL in use (including any explicit set_with_ttl values), aggregating counts across every instance (they are per process and reset on restart). Once the retiring key's count has stayed flat over that whole window, every live entry has aged out or been re-encrypted on write, and the key can be dropped from CACHEKIT_PREVIOUS_MASTER_KEYS without a hard cut-over.


Cross-SDK Interop Mode

Interop mode (interop/v1) lets the Python, TypeScript, and Rust SDKs share cache entries: keys are {namespace}:{operation}:{args_hash} with an explicit operation name (no language-specific function path), and values are plain MessagePack — no envelope, readable by any MessagePack library.

use cachekit::interop::{interop_key, InteropValue};

// Every SDK computes this exact key for get_user(42)
let key = interop_key("users", "get_user", &[InteropValue::from(42i64)])?;

cache.set_with_ttl(&key, &user, ttl).await?;          // plain MessagePack — already interop
let user: Option<User> = cache.interop_get(&key).await?; // strict read: exactly one document

ns and nsapi are reserved as namespaces — the CachekitIO server parses a key starting ns: or nsapi: as namespace-prefixed — so interop_key rejects them with InvalidKey and #[cachekit(namespace = ...)] with a compile error; operations are unaffected. Neither segment may contain .. (a..b is rejected the same two ways; a lone ., as in app.v1, is fine), because the server rejects .. anywhere in a key.

Argument hashing is byte-identical across SDKs (canonical MessagePack + Blake2b-256), verified against the shared protocol test vectors (interop-mode.json, vendored) in this repo's test suite. interop_get (also on SecureCache) rejects trailing bytes and Python-internal CK frames instead of silently misreading them. Every decode of backend-supplied bytes (get and interop_get alike) first passes a header-only structural walk that rejects, before anything is decoded, a document nested deeper than serializer::MAX_DECODE_DEPTH (100 levels, matching the TypeScript SDK; every collection header counts, empty ones included) or declaring more elements or bytes than the input can back, verified against the protocol's shared decode-bounds.json vectors (vendored) — a forged nested-header entry is a bounded Serialization error, not a memory blow-up or a stack overflow. Encryption works unchanged — interop keys are identical across SDKs, so the AAD verifies cross-SDK.

Important

Use interop keys on a client without .namespace() — a client prefix would rewrite the storage key to {prefix}:{interop_key}, which no other SDK computes. interop_get fails closed with a config error rather than silently missing; interop keys already carry their own namespace segment.

Interop mode in production: Skyline — the canonical example project — runs this SDK on wasm32 (Cloudflare Workers) deriving interop keys and verifying payload integrity for the namespace the Python and TypeScript SDKs share (public aggregate reads stay on the TypeScript edge — see the example page).


Backends

cachekit.io SaaS (default)

HTTP backend targeting api.cachekit.io with session tracking, L1 metrics headers, SSRF-safe URL validation, distributed locking, and TTL inspection.

use cachekit::backend::cachekitio::CachekitIO;

let backend = CachekitIO::builder()
    .api_key("ck_live_...")
    .api_url("https://api.cachekit.io")  // optional, this is the default
    .build()?;

Redis

Native Redis via fred with cluster support, TTL inspection, and distributed locking (SET NX PX acquire, atomic Lua compare-and-delete release, <key>:lock namespace shared with cachekit-py). Requires the redis feature flag.

cachekit-rs = { version = "0.8", features = ["redis"] }
use cachekit::backend::redis::RedisBackend;

let backend = RedisBackend::builder()
    .url("redis://localhost:6379")
    .build()?;
backend.connect().await?;  // explicit connect required

Memcached

Memcached via rust-memcache (single server, connection-pooled, per-socket timeouts — a hung server errors one operation instead of wedging the backend). Keys are validated against protocol metacharacters (whitespace/control bytes) before anything reaches the wire, keeping the key space identical to cachekit-py's.

TTL capability, precisely: memcached's protocol cannot read a key's remaining TTL, so this backend does not implement TtlInspectable — matching cachekit-py, where Memcached is likewise not TTL-inspectable. Both SDKs do ship a bare refresh_ttl (wrapping the memcached touch command) callable directly on the backend, outside the capability trait — so TTL-refresh works, but TTL-driven features that need to read TTLs never engage on memcached in any SDK.

TTLs above memcached's 30-day ceiling are clamped (larger values would be misread as absolute timestamps); values above the item-size limit (default 1 MiB) fail loudly client-side, and a server-side "object too large" classifies as permanent (never retried). Requires the memcached feature flag.

cachekit-rs = { version = "0.8", features = ["memcached"] }
use cachekit::backend::memcached::MemcachedBackend;

let backend = MemcachedBackend::builder()
    .url("tcp://localhost:11211")
    .connect()          // eager: verifies the server is reachable
    .await?;

File (local filesystem)

Local disk cache, byte-compatible with cachekit-py's File backend — a py and an rs process pointed at the same directory read each other's entries (Blake2b-128 hashed filenames, shared 14-byte header, atomic write-then-rename, lazy expiry). Implements TtlInspectable (TTL read off the on-disk header, in-place refresh). Concurrency matches py: same-process operations serialize on a backend-wide lock (py's RLock); on unix, reads and in-place TTL rewrites take advisory flock while writes stay lock-free via atomic rename; and expired-entry unlinks are inode-validated so a stale read decision doesn't delete a concurrent writer's fresh entry. On unix the cache directory must be owned by you and not group/other-writable. Not yet ported from py: LRU eviction and size caps — the directory grows until entries expire or you clear it. Requires the file feature flag and a tokio runtime (I/O runs via spawn_blocking).

cachekit-rs = { version = "0.8", features = ["file"] }
use cachekit::backend::file::FileBackend;

let backend = FileBackend::builder()
    .cache_dir("/var/cache/myapp")  // default: <system temp dir>/cachekit
    .build()?;

Cloudflare Workers

wasm32-unknown-unknown backend using worker::Fetch, with distributed locking and TTL inspection against the SaaS lock/TTL endpoints. Requires the workers feature with default features disabled.

cachekit-rs = { version = "0.8", default-features = false, features = ["workers", "encryption"] }
Custom Backend

Implement the Backend trait to plug in any storage:

use async_trait::async_trait;
use cachekit::backend::{Backend, HealthStatus};
use cachekit::error::BackendError;
use std::time::Duration;

struct MyBackend;

#[async_trait]
impl Backend for MyBackend {
    async fn get(&self, key: &str) -> Result<Option<Vec<u8>>, BackendError> { todo!() }
    async fn set(&self, key: &str, value: Vec<u8>, ttl: Option<Duration>) -> Result<(), BackendError> { todo!() }
    async fn delete(&self, key: &str) -> Result<bool, BackendError> { todo!() }
    async fn exists(&self, key: &str) -> Result<bool, BackendError> { todo!() }
    async fn health(&self) -> Result<HealthStatus, BackendError> { todo!() }
}

Optional extension traits: TtlInspectable (TTL queries), LockableBackend (distributed locking).


Dual-Layer Caching

When the l1 feature is enabled (default), CacheKit maintains an in-process moka cache in front of the backend:

┌─────────────────────────────────────────────────────────┐
│                     CacheKit Client                     │
├─────────────────────────────────────────────────────────┤
│                                                         │
│  GET path:                                              │
│  L1 fresh hit (~50ns) ──► return immediately            │
│  L1 stale hit ──► return + background refresh (SWR)     │
│  L1 miss ──► L2 backend ──► backfill L1 (30s cap)      │
│                                                         │
│  SET path:                                              │
│  write to L2 backend ──► write-through to L1            │
│                                                         │
│  DELETE path:                                           │
│  invalidate L1 first ──► delete from L2 backend         │
│                                                         │
├─────────────┬───────────────────────────────────────────┤
│  L1 (moka)  │  L2 (cachekit.io / Redis / Workers)      │
│  ~50ns      │  ~2–50ms                                  │
└─────────────┴───────────────────────────────────────────┘
Behavior Detail
Write-through set() writes to L2 first, then L1
Backfill on miss L2 hits populate L1 with a capped 30s TTL
Invalidate-first delete() evicts L1 before touching L2
Encrypted L1 SecureCache stores ciphertext in L1 (never plaintext)
Evict on decrypt failure A SecureCache read that fails decryption drops the key's L1 copy and returns the error; the backend entry is kept, so the next read reaches the backend: a hit once the entry is replaced with ciphertext the client can decrypt, a miss once it expires or is deleted
Default capacity 1,000 entries (configurable via .l1_capacity())
Live counters cache.stats() reports L1 hits / L2 hits / misses; cache.l1_entry_count() the current occupancy — see Observability
Stale-while-revalidate On by default (native; minimal turns it off): #[cachekit] serves an L1 hit past swr_threshold_ratio × entry TTL (default 0.5, ±10% jitter) immediately and refreshes it in the background — see below

Stale-while-revalidate (SWR)

With SWR (default when l1 is on, native targets), an L1 entry has two phases before it disappears: fresh until swr_threshold_ratio of its TTL has elapsed, then stale until hard expiry. A #[cachekit]-wrapped call that hits a stale entry returns it immediately — no caller ever blocks on a merely-stale value — while exactly one background task re-executes the function. If the same-key mutation token is still current, the task rewrites both cache layers and renews L1 hard expiry with the full write-path TTL; if a newer set() or delete() landed through the same client (or one of its clones) while the origin ran, that explicit mutation wins and the older refresh result is discarded before it can touch L2. Entry expiry and capacity eviction do not invalidate the token, so a valid slow refresh can still repopulate both layers. Refresh dedup rides the same single-flight as the cold-miss path (in-process, plus distributed fill locks on lock-capable backends), so N concurrent stale readers cost one origin execution — misses are billable; stampedes are not acceptable. A hard-expired entry always takes the normal blocking miss path: SWR never serves past hard expiry.

let cache = CacheKit::builder()
    .backend(backend)
    .swr_threshold_ratio(0.25) // stale after 25% of entry TTL (default 0.5)
    // .swr_enabled(false)     // restore strict expire-or-serve behaviour
    .build()?;

Semantics mirror cachekit-py (swr_threshold_ratio = elapsed-lifetime fraction; enabled by default) and cachekit-ts (getWithSwr). Worth knowing:

  • The freshness window derives from each entry's own TTL. A backfilled entry (L2 hit → L1, 30 s cap) goes stale at ~ratio × 30 s — the cap still bounds staleness of L2-derived data, but SWR replaces its expiry cliff with a background refresh that restores the full write-path TTL. The configured ratio is never silently clamped; the window follows the entry.
  • Refresh completion is version-guarded and same-key ordered. Each L1 stale read receives a mutation token. A concurrent explicit write or delete through that client or a clone invalidates the token, so an older origin result cannot clobber the new value or resurrect the deletion in either layer. The guard is intentionally process-local; cross-instance invalidation is outside SWR's serving-policy scope.
  • Jitter is fixed per entry. The ±10% threshold jitter is drawn when an entry is inserted, not on every hit; hot L1 reads do not perform entropy work and the entry's freshness boundary stays stable for its lifetime.
  • The background refresh needs a tokio runtime (Handle::try_current). On other executors the stale value is still served and the refresh is skipped — behaviourally SWR-off, never a panic.
  • Refresh failures are absorbed: the stale value keeps serving, a later stale read retries, and once the entry hard-expires the blocking path surfaces errors normally.
  • Native only: on wasm32 (workers excludes l1) and under unsync there is no SWR; the builder knobs don't exist there, so misuse is a compile error rather than a silent no-op. A sync function under #[cachekit] is likewise a clear compile-time error.

Reliability

With the reliability feature (default, native only), the production, secure, and io presets wrap every backend operation in a reliability stack; minimal stays bare for maximum throughput:

Layer What it does Defaults
Retry Truncated exponential backoff + jitter on transient/timeout errors (BackendErrorKind::is_retryable); permanent and auth errors propagate immediately 3 attempts, 100 ms base, 5 s cap, jitter ×[0.5, 1.5)
Circuit breaker closed → open after N retryable failures in a rolling window; fails fast (BackendErrorKind::CircuitOpen) while open; half-open probes recovery threshold 5, window 60 s, open 5 s, 3 probes, close after 3 successes
Backpressure Bounds concurrent backend data ops with a semaphore + bounded waiting queue; over-limit calls are shed with BackendErrorKind::Backpressure before reaching the backend — a slow backend can't exhaust the caller's connection pool or memory 100 concurrent, 1 000 queued, 100 ms wait (Python SDK parity)
Graceful degradation On outage-class backend failure (transient, timeout, open breaker, backpressure shed), #[cachekit]-wrapped functions run uncached (fail-open); permanent/auth errors propagate — a wrong API key fails loudly. #[cachekit(secure)] paths fail closed on everything — encrypted workloads never silently degrade built into the macro
Single-flight Concurrent misses of one key collapse to a single execution: per-key in-process lock, plus a distributed fill lock across processes on lock-capable backends (cachekit.io, Redis) in-process always on; cross-process 5 s lock, 100 ms polls
Stale-while-revalidate Stale-but-unexpired L1 hits are served immediately while one single-flight-deduplicated background task re-executes the function (details) on by default with l1 (native); threshold 0.5 × entry TTL ±10% jitter

Retry sits inside the breaker (one exhausted retry sequence = one breaker failure) and backpressure sits outside both — one permit per logical operation, held across the whole retry sequence, so retry amplification is bounded and shed calls never skew breaker state. Degradation, single-flight, and SWR sit in the #[cachekit] macro around the read path — the same composition as the TypeScript SDK's ReliabilityExecutor and the Python decorator.

use std::time::Duration;
use cachekit::{CacheKit, ReliabilityConfig, RetryConfig};

// Presets enable it — override or disable per client:
let cache = CacheKit::production("redis://localhost:6379").await?
    .reliability(ReliabilityConfig {
        retry: Some(RetryConfig { max_attempts: 5, ..RetryConfig::default() }),
        ..ReliabilityConfig::default()
    })
    .build()?;

// Opt a preset out: a disabled config applies no wrapping.
let bare = CacheKit::production("redis://localhost:6379").await?
    .reliability(ReliabilityConfig::disabled())
    .build()?;

Requires a tokio runtime for backoff timers (the redis and cachekitio backends already do).


Observability

Every client counts its reads, with no configuration:

let stats = cache.stats();                 // cachekit::L1Stats — live, shared by all clones
println!(
    "L1 {} / L2 {} / miss {} — L1 hit rate {:.1}%",
    stats.l1_hits, stats.l2_hits, stats.misses, stats.l1_hit_rate() * 100.0,
);
let occupancy = cache.l1_entry_count();    // Option<u64>: None when L1 is off
let breaker = cache.circuit_state();       // Option<CircuitState>: Closed / Open / HalfOpen
Surface What you get
CacheKit::stats() L1Stats { l1_hits, l2_hits, misses, l1_enabled } for every value read (get, interop_get, SWR and SecureCache variants). exists and reads that fail with a backend error are not counted.
CacheKit::l1_entry_count() Exact L1 occupancy (runs moka's pending housekeeping first — poll it, don't put it on a hot path).
CacheKit::circuit_state() Live breaker state (reliability feature); None when the client has no breaker.
SaaS telemetry headers The cachekit.io backends send X-CacheKit-L1-Hits / L2-Hits / Misses / L1-Hit-Rate from the same counters, wired automatically by CacheKitBuilder::build(). A .metrics_provider(..) set on the backend builder still takes precedence. One backend instance reports one client — the first built over it; once that client is gone the headers fall back to disabled.
tracing feature One debug event per completed operation on the cachekit target, and breaker transitions on cachekit::reliability (warn on open, info for half-open / closed).

With the tracing feature, point your subscriber at the crate:

RUST_LOG=cachekit=debug cargo run
DEBUG cachekit: op=get key_hash=bcb35ae6f64fa65b2770ab3af631b1ce outcome=miss
DEBUG cachekit: op=set key_hash=bcb35ae6f64fa65b2770ab3af631b1ce ttl_secs=3600
DEBUG cachekit: op=get key_hash=bcb35ae6f64fa65b2770ab3af631b1ce outcome=l1_hit
 WARN cachekit::reliability: circuit breaker opened breaker=1 seq=1 from=Closed to=Open

Fields: op (get | set | delete), outcome (l1_hit | l1_stale | l2_hit | miss), ttl_secs, existed. Breaker events carry breaker, seq, from, to, and (breaker, seq) is the ordering key: breaker is a process-unique id assigned when the breaker is built (stable for its lifetime, not a key or secret), seq counts that breaker's transitions and is assigned under the breaker lock. Events are emitted after the lock is released (so a subscriber may call circuit_state() safely), which means two transitions can arrive out of order under contention, and several clients in one process each restart seq at 1 — group by breaker, order by seq, never by arrival. For fleet-wide correlation combine the pair with the host/process fields your subscriber adds. Events carry key_hash — Blake2b-128 of the namespaced storage key (cachekit::metrics::key_hash) — never the key itself: keys routinely embed user identifiers (CWE-532). The digest is a correlator, not a redaction: it is unkeyed and deterministic, so it matches the File backend's on-disk filename (a log line names the cache file it touched), and for the same reason a low-entropy key like user:42 can be recovered from it by enumeration. Treat cachekit=debug output with the care you give the keys themselves.

Prometheus exposition and OpenTelemetry spans are deliberately not built in: Rust services bring their own registry and bridge tracing themselves.


Environment Variables

Variable Required Description
CACHEKIT_API_KEY ✅ API key for cachekit.io (from_env() and CacheKit::io_from_env())
CACHEKIT_API_URL ❌ Override API endpoint (default: https://api.cachekit.io)
CACHEKIT_MASTER_KEY ❌ Hex-encoded master key for encryption (CacheKit::secure_from_env() and from_env()); use exactly 32 bytes (64 hex chars) — shorter is rejected
CACHEKIT_PREVIOUS_MASTER_KEYS ❌ Comma-separated hex-encoded decrypt-only previous master keys for key rotation (CacheKit::secure_from_env() and from_env(); max 3; a blank value is treated as unset)
CACHEKIT_DEFAULT_TTL ❌ Default TTL in seconds (min 1, default: 300)

Caution

CACHEKIT_API_URL must use HTTPS and must not point to a private IP address. Both constraints are enforced at configuration time.


Architecture

cachekit-rs/
├── crates/
│   ├── cachekit/              # Main SDK crate
│   │   └── src/
│   │       ├── lib.rs         # Public API + prelude
│   │       ├── client.rs      # CacheKit, SecureCache, CacheKitBuilder
│   │       ├── config.rs      # CachekitConfig + from_env()
│   │       ├── encryption.rs  # AES-256-GCM + AAD v0x03
│   │       ├── error.rs       # CachekitError, BackendError
│   │       ├── interop.rs     # interop/v1 cross-SDK keys + strict reads
│   │       ├── metrics.rs     # Live counters, SaaS telemetry headers, tracing events
│   │       ├── session.rs     # SDK session tracking
│   │       ├── url_validator.rs # SSRF-safe URL validation
│   │       ├── serializer/    # MessagePack serialization
│   │       ├── l1/            # moka-based L1 cache (feature = "l1")
│   │       └── backend/
│   │           ├── mod.rs     # Backend + TtlInspectable + LockableBackend traits
│   │           ├── cachekitio.rs      # cachekit.io HTTP backend
│   │           ├── cachekitio_lock.rs # Distributed locking
│   │           ├── cachekitio_ttl.rs  # TTL inspection
│   │           ├── saas_wire.rs       # SaaS lock/TTL JSON wire bodies
│   │           ├── redis.rs           # Redis backend (feature = "redis")
│   │           └── workers.rs         # Workers backend (feature = "workers")
│   │
│   └── cachekit-macros/       # Proc-macro crate
│       └── src/lib.rs         # #[cachekit] decorator
│
├── Cargo.toml                 # Workspace root
└── Makefile                   # Development commands

Development

make quick-check   # fmt + clippy + test (run before every commit)
make security      # cargo deny + cargo audit (the CI supply-chain gate)
make test          # cargo test --features $(NATIVE_FEATURES) (CI's list; see Makefile)
make build         # cargo build --release
make build-wasm    # wasm32-unknown-unknown (workers feature)

make security runs the same two enforcement commands as the supply-chain job in .github/workflows/security.yml, with cargo audit in its strictest CI form (--deny yanked) — so a local pass means a pass on every CI event, with two asymmetries: the weekly run additionally proves the yank check actually executed (see the guard in security.yml), so with crates.io unreachable a local run warns and passes where the weekly run goes red; and the job's final step, the gate tamper check below, is PR-context-only and has no local equivalent. It needs cargo-deny and cargo-audit installed, and it reaches the network to refresh the RustSec advisory database — which is why it is not folded into quick-check.

Both tools are required, because they answer different questions. "Fails" below means it turns the check red — anything else is reported but not enforced:

cargo deny --locked --all-features check cargo audit
Reads feature-resolved dependency graph Cargo.lock verbatim
Licence allowlist, banned crates, registry/source policy fails not checked
Vulnerabilities in crates no enabled feature activates not seen (pruned) fails
Yanked crates in Cargo.lock warns (feature-resolved graph only, so lockfile-only crates are missed) warns on PR and push runs; fails only the weekly scheduled run (--deny yanked)
Unsound / unmaintained advisories on transitive deps not seen — deny.toml narrows unmaintained to workspace; unsound already defaults to that scope reports only, does not fail — deliberate (see deny.toml)

--all-features is load-bearing: the default feature set excludes the memcached, redis, file and macros backends, so a banned crate reintroduced behind an optional feature passes a bare cargo deny check.

deny.toml is the policy — notably a hard ban on openssl-sys, native-tls and toxiproxy_rust, because this SDK is rustls-only. Run make deny before adding or bumping a dependency.

Gate tamper-evidence

The supply-chain check reads both its policy (deny.toml) and its own definition (security.yml) from the PR head, so a PR could weaken the gate it is being graded by — delete a [bans] entry, or drop --all-features while keeping the job name green. Two properties make that visible:

  • Deletion fails closed. supply-chain is a required status check: if nothing reports it, the PR cannot merge. The check is pinned to the GitHub Actions app, so a status posted from outside Actions cannot satisfy it.
  • Modification trips a wire. The job's final step fails the required check when a PR changes deny.toml or security.yml relative to its base, or changes any other workflow file that mentions supply-chain, unless the PR body contains the exact, case-sensitive string [gate-change-approved] (add it after human sign-off, then push a commit — the marker is read from the push-time event, so a body edit alone does not re-trigger).

Minimum Supported Rust Version

Rust 1.85 or later (Edition 2021).

License

MIT — see LICENSE for details.


About

Rust caching SDK for CacheKit (beta) — dual-layer L1/L2, zero-knowledge encryption, multi-backend

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages