Module: internal/telemetry · Milestone: later · Effort: ~1.5w (feature #12)
Let a small OSS team learn how devstack behaves in the wild — which commands run, where they fail, on which OS/arch and Docker/compose versions — without ever betraying user trust. The whole point is to answer questions like "WSL2 trust install fails 30% of the time" that issue trackers can't, while emitting nothing that could identify a user, a repo, or a secret. Telemetry is strictly opt-in, default OFF, fully documented, and best-effort: a telemetry failure must never affect a command's exit code, output, or latency.
Builds on: the spec 14 update-notifier substrate — the only other thing in
devstackthat talks to the network on a normal command. It reuses that spec's TTY/CI/env suppression discipline, the cache-not-ledger storage split, and the hard time-budget pattern. Deferred to "later" because it touches no pillar, carries real reputational risk if rushed, and is only worth building once the v1 command surface (spec 07) is stable enough that event names mean something.
- Default is OFF. No event is collected, queued, or sent until the user has affirmatively enabled telemetry. There is no silent-on, no "anonymous by default", no soft-opt-out.
- First-run prompt: on the first interactive (TTY) invocation of a mutating command, devstack prints exactly what would be collected (a one-screen summary, plus
devstack telemetry showto see a real sample) and asks a single yes/no. The default answer is No. A declined or skipped prompt records consentfalseand is never asked again that major version. - Consent lives in config, not the ledger. Persist
telemetry.enabled: true|false(+telemetry.consentAt,telemetry.consentVersion) underXDG_CONFIG_HOME/devstack/telemetry.yaml, not inworkspace.yaml(consent is per-user/per-machine, not per-repo, and must never be committed) and not instate.db(it's user policy, not Docker-context ledger state). A missing file means "never decided" → treat as OFF. - Revocable at any time, single command:
devstack telemetry disableflips it OFF and is honored on the very next invocation;devstack telemetry enableopts back in;devstack telemetry statusprints the current state, endpoint, and the stable install id. - One-flag / one-env disable:
--no-telemetry(global flag) andDEVSTACK_NO_TELEMETRY=1force OFF for that run regardless of config; the env opt-out, like the spec 14 notifier env opt-out, skips collection and the network call entirely (nothing queued, nothing sent). There is intentionally no--telemetryflag that turns it on for one run — enabling is always an explicit, persisted consent action.
Telemetry is forced OFF — and the first-run prompt is never shown — when any of these hold, even if consent is true:
| Condition | Detect via | Rationale |
|---|---|---|
| Non-interactive stdout | golang.org/x/term.IsTerminal false |
can't obtain real consent; scripts shouldn't phone home |
| CI environment | CI env set (+ common GITHUB_ACTIONS/GITLAB_CI/BUILDKITE…) |
CI runs would swamp and skew data, and aren't a human |
DO_NOT_TRACK set to a truthy value |
env DO_NOT_TRACK=1 |
honor the cross-tool consoledonottrack.com convention |
--json / --quiet machine modes |
global flags (spec 07) | machine-driven, non-interactive by contract |
--no-telemetry / DEVSTACK_NO_TELEMETRY |
flag/env | explicit per-run kill switch |
The prompt is also suppressed when stdin isn't a TTY (no way to read an answer). CI auto-off means the env opt-out is belt-and-suspenders, not the primary guard.
The collected fields are an exhaustive allowlist; anything not on this list is structurally impossible to send (the event type has no field for it). This is the inverse of scrubbing-after-the-fact, which always leaks.
| Field | Example | Notes |
|---|---|---|
command |
up, down, doctor, shared gc |
the cobra command path only — never args/flags values |
subcommand_flags |
["--build","--profile"] |
flag names that were set, never their values |
outcome |
ok / error |
|
error_category |
docker_daemon_unreachable, compose_too_old, port_in_use, git_auth_prompt, state_locked |
a closed enum mapped from the wrapped-error signatures (ARCHITECTURE §7.6) — never the raw error string |
duration_ms |
1840 |
bucketed (e.g. nearest 100ms) to blunt fingerprinting |
os / arch |
linux/arm64, darwin/amd64 |
runtime.GOOS/GOARCH |
is_wsl2 |
true |
from the existing spec 08 WSL detection — the single most useful dimension |
tool_version |
1.3.0 |
ldflags version (spec 14) |
docker_version / compose_version |
27.5.1 / 2.31.0 |
already probed by doctor/preflight (spec 03) |
install_id |
random UUIDv4 | generated once, stored beside consent; rotatable via telemetry reset; not derived from anything machine-identifying |
Never collected, by construction: repo names, project names, paths (workspace/CWD/binary), file contents, template names or contents, secret names or values, env-var names or values, ${ref} keys, git remotes/branches, hostnames, usernames, IP addresses (the endpoint may log the source IP — see threat model), MAC addresses, or any free-form string. There is no metadata: map[string]string escape hatch.
devstack telemetry show runs the real collection path for a synthetic invocation (or against the last command if --last) and prints the exact event(s) that would be sent, as pretty JSON, plus the destination endpoint and whether sending is currently enabled. This is the trust primitive: a skeptical user can verify, byte-for-byte, that nothing sensitive leaves the machine. It performs no network I/O. --format json emits the raw OTLP body for the truly paranoid.
- Pure-Go OTLP/HTTP exporter (
go.opentelemetry.io/otel+otlptracehttp/otlploghttpover protobuf-or-JSON) so the staticCGO_ENABLED=0binary is preserved. Events are modeled as OTel log records (or span events) on adevstack.cli.commandscope — structured, queryable, and standard. No gRPC exporter (heavier transitive deps); HTTP keeps the binary lean. - Self-hostable endpoint. The default endpoint is a project-run OTLP collector; a self-hosting OSS team overrides it with
telemetry.endpoint(config) orOTEL_EXPORTER_OTLP_ENDPOINT(honor the OTel standard env) and points it at their own collector/Grafana/Honeycomb-compatible sink. No vendor lock-in; the wire format is the open standard. - Best-effort + hard budget. Fire the export in a goroutine with a short budget (≤500ms) and a small in-process queue; if the command finishes first, the export is abandoned (no flush-on-exit blocking). Network error, DNS failure, timeout, non-2xx → swallowed, logged only under
--debug. A telemetry failure never changes the command's exit code. - Sampling. Default head sampling at a configurable rate (
telemetry.sampleRate, default e.g.0.25) so a busy user'sup/statusloop doesn't flood the collector;erroroutcomes are always sent (sample rate 1.0) because failures are the high-value signal. Sampling is deterministic per(install_id, day)so a user is consistently in-or-out within a window (cleaner aggregates, less fingerprinting than per-event coin flips). - No daemon, no on-disk spool. Consistent with the stateless-CLI model (ARCHITECTURE §1): events are best-effort fire-and-forget within the process lifetime; dropped events are acceptable (telemetry is statistical, not an audit log). A durable spool + batch flush would want the v2 daemon (Q-DAEMON) — out of scope, see open questions.
- PII hides in error strings and paths. A raw
dial tcp 10.0.0.5:5432: connect: connection refusedoropen /home/alice/work/acme-secrets/.env: no such fileleaks an IP, a username, and a private repo name. Mitigation: we never transmiterr.Error()or any path — only the closederror_categoryenum. The mapping from wrapped error → category happens ininternal/telemetry, and the category set is small and reviewed; an unmatched error becomeserror_other(no text). - Allowlist, not denylist. The event struct has only the fields above. There is no generic string map and no reflection-based "dump everything" path, so a future field can't accidentally start carrying PII — adding a field is a deliberate, reviewed code + spec change.
- Source IP is unavoidable at the network layer. The OTLP endpoint sees the client IP like any HTTP server. We document this honestly (it is not in the payload, and the project collector is configured to drop/not-log source IPs); a user who considers their IP sensitive should leave telemetry OFF (the default).
- Install id is not identity.
install_idis a random UUID with no machine derivation, rotatable withtelemetry reset(regenerates the id) and erased bytelemetry disable. It correlates events from one install for funnel analysis, nothing more. - Consent is unambiguous and persisted. Default-No prompt, explicit enable, single-command revoke, honored next run. We never re-prompt within a major after a decision, and never treat "ignored the prompt" as consent.
- The notifier is a separate call with its own opt-out. The spec 14 update check is not telemetry: different endpoint (GitHub Releases), different purpose, different (default-ON, opt-out) policy, different env switches (
DEVSTACK_NO_UPDATE_NOTIFIER). Disabling one never silently disables the other;doctorreports both network behaviors distinctly so users aren't surprised by two outbound calls.
- OTLP exporter must stay pure-Go.
go.opentelemetry.io/otel+ theotlptracehttp/otlploghttpexporters are pure-Go and cross-compile to all four targets withCGO_ENABLED=0; do not pull the OTLP gRPC exporter or any collector-runtime packages (heavier deps, slower builds). Pin the otel modules — they version in lockstep and move fast (dependency-risk register). runtime.GOOS == "linux"does not distinguish WSL2 — reuse the existing spec 08/ARCHITECTURE §8 WSL detection foris_wsl2; that dimension is the highest-value field, so don't approximate it.DO_NOT_TRACKis a real, widely-honored convention (consoledonottrack.com); honoring it costs one env check and signals good faith — treatDO_NOT_TRACK∈ {1,true} (case-insensitive) as auto-off.- Never block the command. Like the notifier's 800ms budget, the exporter runs detached with a hard timeout; a hung collector or captive-portal DNS must not add latency to
up. Do not callTracerProvider.Shutdown()(it flushes synchronously) on the command's hot exit path. - Consent must not be inferable from the network. The first-run prompt and all collection are TTY-gated and CI-gated, so an automated environment never both prompts and sends — avoiding the classic "CI silently enabled telemetry" incident.
- Don't conflate sample-out with opt-out. A sampled-out event is silently dropped (normal); an opted-out user sends nothing and the code path short-circuits before constructing any event — these are different branches and both need tests.
- A telemetry config file must never be created by
import/generation. It is written only by an explicittelemetry enable/disable(or the first-run prompt answer); the generation pipeline (spec 02) never emits it, keeping consent out of committed artifacts.
- With no consent file present, no command emits any network call to the telemetry endpoint, and
telemetry statusreports OFF. - The first-run prompt appears only on an interactive TTY mutating command, defaults to No, and is never shown again after any answer; declining persists
enabled: false. -
--no-telemetry,DEVSTACK_NO_TELEMETRY=1,DO_NOT_TRACK=1,CI, non-TTY stdout, and--json/--quieteach independently force OFF with no network call, even when consent istrue. -
telemetry showprints the exact event JSON that would be sent and the endpoint, performing zero network I/O. - A CI test asserts the event payload contains no path, repo/project name, secret name/value, env name/value,
${ref}key, or raw error string — only allowlisted fields (mirrors the spec 04 no-secret-in-output test). - A wrapped daemon-unreachable / compose-too-old error maps to its
error_categoryenum and the rawerr.Error()never appears in the payload. - A telemetry endpoint that hangs, times out, or returns 500 leaves the command's exit code, stdout, and latency unchanged (best-effort proven under a fault-injected exporter).
-
telemetry disableis honored on the very next invocation;telemetry resetrotatesinstall_id. -
erroroutcomes are sent regardless ofsampleRate; a0.0sample rate still sends errors and never sendsokoutcomes. - The built binary remains
CGO_ENABLED=0static with the OTLP/HTTP exporter linked in (no CGO, all four targets).
Consumes internal/version (the tool_version stamp — spec 14), internal/docker (the already-probed docker/compose versions — spec 03), internal/xdg (consent-file path + WSL2 detection), and internal/config (the telemetry.* keys + the global --no-telemetry/error-render contract — spec 07). Consumed by internal/cli (the telemetry command group, the first-run prompt hook on mutating commands, and the per-command emit at exit) and internal/doctor (spec 13 reports telemetry + notifier network behavior). It deliberately does not touch internal/lock or internal/state — telemetry never mutates shared ledger state, so it stays off the concurrency spine (spec 08). Thinner v2 vs full: a thin first cut (consent + auto-off + allowlisted event + best-effort OTLP export + telemetry show/enable/disable/status) lands in ~1w; the full version (deterministic sampling, error_category enum coverage across every subsystem's wrapped errors, telemetry reset, and the self-host collector config + docs) is ~1.5w.
- Q-DAEMON — a durable event spool with batched flush (so events survive a quick command and aren't dropped) would want the v2 background agent; v1/v2 telemetry stays fire-and-forget within the process. Revisit alongside the dashboard.
- Q-TELEMETRY (new) — where does the default OTLP endpoint live, who operates it, and what is its data-retention + source-IP-dropping policy? This must be answered (and published) before telemetry ships, because the default endpoint is the project's standing privacy promise. If no project-operated collector exists at ship time, default the endpoint to empty (telemetry can be enabled but goes nowhere until a self-hoster sets
telemetry.endpoint).