Skip to content
FlopsstuffPublic

About

Claude Code OpenTelemetry

Resources

Stars

11 stars

Watchers

0 watching

Forks

Repository files navigation

Flopsstuff logo

cotel — Claude Code Telemetry

One Docker container. OTLP ingest on :4318, interactive analytics dashboard on :8080. No cloud dependencies, no sign-up.

cotel Overview

What you get

  • Overview dashboard — one range switcher in the header (All / Year / Month / Week / Day, default 30 days) that every figure on the page obeys: the KPI cards for sessions, users, total cost and token counts, and the Span activity / Users / Spans & cost / Tools / Models / Sessions blocks below them. The Span activity grid is a GitHub-style block of cells counting spans, and the range picks how much time a cell is — a day over a year, four hours over a month, an hour over a week, ten minutes over a day. The Spans & cost block charts spans and spend together on one field — spans in blue against the left axis, cost in amber against the right — so a spend spike lands under the activity that caused it. The Users block ranks your top 5 principals by spend in the selected range. Arriving with ?user_id= scopes the whole page to one user, with a chip in the header to clear it
  • Sessions — live table of every Claude Code session with user, model, duration, cost, and status (OK / ERROR); search by user and click any user to filter the table to their sessions
  • History — time-series and daily-activity heatmaps for sessions and token spend over time. The day, week and month series keep charting after retention has rolled raw spans into daily totals; the hourly series and both heatmaps need a per-span timestamp, so they cover raw days only and say from when
  • Costs — cumulative spend chart + breakdown table by model
  • Tools — call counts, average duration, and error rate per tool (Bash, Read, Edit, …), scoped to a range switcher (All / Year / Month / Week / Day, default 30 days), plus a per-command breakdown of Bash calls
  • Models — token and cost breakdown across all Claude model variants
  • Users — one sortable, searchable list with per-user cost and sessions scoped to a range switcher (All / Year / Month / Week / Day, default 30 days); click any row for that user's page — token, rotate, delete, and links into their activity and sessions
  • Setup — step-by-step onboarding guide with copy-paste settings.json snippets pre-filled with your ingest URL and token
  • Export / Import — download all data as a versioned ZIP/CSV archive; restore it on a fresh instance
  • Cloudflare Tunnel — publish cotel over HTTPS with a single env var; bearer-token auth for OTLP, Zero Trust for the dashboard

Screens

Every screenshot here is one instance seeded with synthetic telemetry — see Demo data.

Sessions — every session with its user, model, tokens, cost and OK / ERROR status
Sessions
Costs — daily spend and cost by model across the window you pick
Costs
Users — cost and session count per principal for the selected range
Users
Tools — call count, average duration and error rate per tool
Tools

Quick start

docker run -d \
  --name cotel \
  -p 4318:4318 \
  -p 8080:8080 \
  -v cotel-data:/data \
  ghcr.io/flopsstuff/cotel:latest

Available tags: :latest and :0.3 (current release), :0.x.y (patch), :main (tip of main branch).

Open http://localhost:8080 → Setup for the guided onboarding.

Demo data

An empty cotel shows empty charts, which makes it hard to judge. scripts/seed-demo.py fills a throwaway instance with a synthetic team — seven users, three models, 90 days of sessions and tool calls — over the real OTLP endpoint, so what you see is what ingest actually produces:

python3 scripts/seed-demo.py --dash-url http://localhost:8080 \
                             --ingest-url http://localhost:4318

It only creates users and ingests spans; it never deletes. Point it at an instance you are willing to throw away, not at one holding real telemetry. The screenshots in this README come from exactly this seed — the recipe is in docs/operations/screenshots.md.

Point Claude Code at cotel

Add to your ~/.claude/settings.json:

{
  "env": {
    "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
    "CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
    "OTEL_TRACES_EXPORTER": "otlp",
    "OTEL_EXPORTER_OTLP_PROTOCOL": "http/json",
    "OTEL_EXPORTER_OTLP_TRACES_ENDPOINT": "http://localhost:4318/v1/traces"
  }
}

Restart Claude Code. Telemetry starts flowing immediately.

Users and authentication

cotel ships with token-based authentication for OTLP ingest. Tokens are tied to named users.

Open the dashboard → Users (people icon in the sidebar) → Add user. Give the user a name (e.g. the machine or agent name sending telemetry) and copy the token from the banner. The token is retrievable any time from that user's page (click the row); rotating or deleting a user revokes its token immediately.

By default, cotel accepts spans without an Authorization header (allow-anonymous mode). To enforce strict auth: open Setup → Settings tab → disable Allow anonymous OTLP. Requests without a valid token then receive 401 Unauthorized.

See docs/operations/users-and-auth.md for the full guide: multi-user setup, token rotation, and security considerations.

Publishing with Cloudflare

Use Cloudflare Tunnel to expose cotel over HTTPS without opening inbound ports. The dashboard is protected by Cloudflare Zero Trust; OTLP ingest is protected by a bearer token you create in the cotel Users page.

Prerequisites

  • A Cloudflare account (free tier is sufficient)
  • A domain managed by Cloudflare DNS
  • Cloudflare Zero Trust enabled on your account (free for up to 50 users)

1. Create a tunnel

  1. Go to Cloudflare Zero Trust → Networks → Tunnels → Add a tunnel.
  2. Choose Cloudflared, name the tunnel (e.g. cotel), and click Save tunnel.
  3. Copy the tunnel token shown on the next screen — you will use it as CLOUDFLARE_TUNNEL_TOKEN.

2. Configure public hostnames

In the tunnel's Public Hostnames tab, add two entries:

Subdomain Domain Service
cotel yourdomain.com http://localhost:8080
cotel-ingest yourdomain.com http://localhost:4318

Replace the subdomains and domain with your own values.

3. Protect the dashboard with Zero Trust

In Zero Trust → Access → Applications, create a Self-hosted application for your dashboard hostname (e.g. https://cotel.yourdomain.com). Add an access policy — for example "Allow email ending in @yourcompany.com".

Leave the OTLP ingest hostname (cotel-ingest.yourdomain.com) unprotected in Zero Trust. Authentication for ingest is handled by the cotel bearer token instead.

4. Start cotel with the tunnel token

Pass CLOUDFLARE_TUNNEL_TOKEN when running the container:

docker run -d \
  --name cotel \
  -v cotel-data:/data \
  -e CLOUDFLARE_TUNNEL_TOKEN=<your-tunnel-token> \
  ghcr.io/flopsstuff/cotel:latest

No -p flags are needed — Cloudflare Tunnel uses outbound connections only.

5. Create an agent token

Open your dashboard URL (e.g. https://cotel.yourdomain.com), go to Users in the sidebar, and click Add user. Give it a name (e.g. the agent's hostname or machine name) and copy the token.

6. Configure Claude Code

Add to your ~/.claude/settings.json:

{
  "env": {
    "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
    "CLAUDE_CODE_ENHANCED_TELEMETRY_BETA": "1",
    "OTEL_TRACES_EXPORTER": "otlp",
    "OTEL_EXPORTER_OTLP_PROTOCOL": "http/json",
    "OTEL_EXPORTER_OTLP_TRACES_ENDPOINT": "https://cotel-ingest.yourdomain.com/v1/traces",
    "OTEL_EXPORTER_OTLP_HEADERS": "Authorization=Bearer cotel_<your-token>"
  }
}

Replace cotel-ingest.yourdomain.com with your actual OTLP hostname, and cotel_<your-token> with the token you copied above. Restart Claude Code to apply the changes.

Token enforcement: cotel accepts anonymous spans until the first user with a token is created. Once any token exists, every OTLP request must carry a valid Authorization: Bearer cotel_... header — unless you explicitly re-enable Allow anonymous OTLP in Setup → Settings.

Ports

Port Purpose
4318 OTLP/HTTP trace ingest (POST /v1/traces)
8080 Analytics dashboard

HTTP-only: cotel speaks OTLP/HTTP (application/x-protobuf and application/json). There is no gRPC listener on port 4317. If spans are silently dropped, ensure OTEL_EXPORTER_OTLP_PROTOCOL=http/json (or http/protobuf) is set — the OTel SDK default is grpc, which will fail against cotel.

Startup and deploys

On boot, cotel opens the DuckDB file (WAL replay + schema migration) before it can serve queries. On a large production database this can take minutes. To avoid dropping telemetry during that window (and on every deploy — a merge to main restarts the container), both ports bind and start accepting connections immediately, before the database is opened:

  • While storage is initialising, the ingest (:4318) and dashboard (:8080) ports answer every request with 503 Service Unavailable and a Retry-After: 5 header. OTLP/HTTP exporters retry on 503, so spans are held by the client and delivered once cotel is ready — nothing is lost. (The old behaviour bound the ports only after the open finished, so clients got a connection reset and their spans were dropped.)
  • Once the database is ready the gates open and both ports serve live traffic.

Schema migrations are version-guarded: storage.Open skips the whole of schema.sql only when the database records both the current version (schema_version) and the exact file hash (schema_sql_sha256 in settings), so migrations run once (on the deploy that introduces them) instead of on every boot. The hash is a safety net — any edit to schema.sql re-applies it (once, idempotently) even if the version row was forgotten, so a new migration is never silently skipped. This is correctness hygiene, not the reason a start is slow — measured against a copy of production, the whole schema apply is ~165 ms. The multi-minute open is DuckDB replaying the WAL, which happens before the version can even be read; keeping that fast is a separate concern (see ADR-0010).

Startup is logged so the open duration is visible (no more guessing from a bare Up container):

listening on ingest :4318 and dashboard :8080 (storage initialising, serving 503 until ready)
opening db /data/cotel.duckdb
db ready: schema/migrations applied in 2m58s
startup checkpoint complete in 8ms
ready: serving live traffic on ingest :4318 and dashboard :8080

After schema/migrations apply, cotel runs the same DuckDB CHECKPOINT used on shutdown, before the gates open and before the retention worker starts. ALTER and CREATE INDEX have already run, so this fold serializes the indexes. A hard failure exits 3 and writes <db path>.checkpoint-failed without ever serving live traffic, so the deploy health gate fails the rollout instead of discovering a damaged file days later. A checkpoint that merely hits its 8s deadline stays benign: the process continues and serves traffic, the same distinction shutdown makes between a failed fold (exit 3) and a timeout (exit 0). On a healthy 35 000-span file the extra start checkpoint is <1 ms when schema is skipped, and 8 ms immediately after a schema re-apply that rebuilds the four ART indexes.

Stopping — checkpoint on shutdown

Since the slow part of a start is replaying the WAL, the fix is to not leave a WAL to replay: cotel checkpoints on the way out. A file whose WAL has been folded into the main database opens in milliseconds. On SIGTERM/SIGINT (what docker stop and a deploy send) it stops accepting ingest, runs a DuckDB CHECKPOINT to fold the WAL into the main file, and exits:

received terminated: stopping listeners and checkpointing before shutdown
checkpoint complete in 30ms; exiting

Because the running process has already applied the log in memory, this checkpoint is cheap and completes well inside the container's stop grace period (Docker's default 10s). The next start then finds an empty WAL and is ready in a second or two instead of minutes. The container entrypoint forwards the stop signal to cotel and waits for it, so the checkpoint runs before the container is torn down.

Two caveats:

  • The first deploy that ships this behaviour still pays the full replay — the process being killed is the old binary without a shutdown handler. The win starts from the restart after that. (A merge to main is itself a deploy.)
  • A hard kill is safe but slow to recover. docker kill, an OOM, or a SIGKILL after the grace period expires skips the checkpoint. Nothing is corrupted and committed spans are not lost — DuckDB replays the WAL on the next open — but that open pays the replay again. COTEL_WAL_AUTOCHECKPOINT (below) bounds how large the WAL, and therefore that worst-case replay, can get: it folds the log into the main file once it passes the threshold, so a hard kill finds little left to replay. The exception is a damaged database — see below.

When the checkpoint fails

A checkpoint that runs out of time is benign; one that fails outright is not. It fails because the database itself is damaged, and the WAL it leaves behind can then abort the next open inside libduckdb (a C++ abort() Go cannot catch), so the database never opens again until someone repairs it by hand. Startup and shutdown share this reporting:

When Outcome Log Exit code Marker
Shutdown succeeded checkpoint complete in …; exiting 0 removed if present
Shutdown 8s deadline checkpoint on shutdown timed out …, WAL left for replay on next start 0 —
Shutdown failed checkpoint on shutdown FAILED …, the WAL left behind may not be replayable 3 <db path>.checkpoint-failed
Startup succeeded startup checkpoint complete in … continues removed if present
Startup 8s deadline startup checkpoint timed out …, continuing continues —
Startup failed startup checkpoint FAILED … 3 <db path>.checkpoint-failed
docker inspect --format '{{.State.ExitCode}}' cotel    # 3 = a checkpoint did not fold the WAL

The marker file sits next to the database on the /data volume and outlives the container's logs; cotel logs a warning about it on the next start, just before it opens the database, and clears it after the next clean fold (startup or shutdown). If you find it, or the container restart-loops with exit code 134, follow docs/operations/duckdb-recovery.md.

Health check

The image ships a Docker HEALTHCHECK that probes the dashboard /healthz (GET, returns 200 only once the database is open). During the initial open the check returns non-zero, so the container reports health: starting rather than a misleading healthy state while ingest is still dark; it flips to healthy when the database is ready. The probe is the cotel binary itself (cotel -healthcheck), so no extra tooling is needed in the runtime image:

docker inspect --format '{{.State.Health.Status}}' cotel   # starting → healthy

What /healthz returns

{"ok": true, "spans": 128402, "last_ingest_at": "2026-10-04T03:12:44.118Z", "newest_span_age_seconds": 97}
Field Meaning
ok false when the database could not be queried at all
spans Rows in spans
last_ingest_at When the newest span was accepted (RFC 3339), or null if nothing was ever ingested
newest_span_age_seconds Seconds since last_ingest_at, or null if nothing was ever ingested

The age is measured from ingested_at, not start_time: importing an archive replays the original ingest time, so a restore cannot make a dark instance look alive, and a span arriving late with an old start_time still counts as fresh traffic. An empty database reports null rather than 0, because "nothing ever arrived" is not "something arrived just now".

A growing age is the signal that the process is up but spans are not reaching it — a state that otherwise looks identical to a healthy idle instance. It does not change the status code: staleness answers 200 like any other readable database, because the right threshold ("nothing for 2 hours" vs "a quiet weekend") belongs to whoever polls, not to the container. Only a failed query is a failure: that answers 503 with {"ok": false}, so the probe cannot report a dead database as healthy. The same two fields are on GET /api/v1/health — see docs/operations/api-reference.md.

The retention and snapshot workers' health is deliberately not here. This endpoint is what cotel --healthcheck, the Docker HEALTHCHECK and the deploy's wait-for-healthy.sh read, so a failing backup reported through ok would mark a working container unhealthy and fail the next deploy. Those claims live on GET /api/v1/health, which the health probe reads separately — see Production health probes.

The deploy waits for healthy

docker compose up -d returns as soon as the container has started, which is not the same as working: a deploy whose storage.Open dies reports success identically to one that serves traffic. So the Deploy workflow does not stop at up -d — it runs scripts/wait-for-healthy.sh, which blocks until the container reports healthy and fails the deploy on any of:

Condition Gate result
healthy within the timeout pass
Container exited, or is in restarting (crash loop) fail, immediately
Container restarted during the wait (crash loop) fail, immediately
Health probe reports unhealthy fail, immediately
Still starting when the timeout expires fail
Service defines no HEALTHCHECK fail

On failure it dumps docker compose ps, the last health-probe output, the crash cause, and docker compose logs --tail=200, so the reason lands in the workflow run log instead of needing shell access to the runner. The healthcheck itself is defined in the Dockerfile and compose inherits it from the image; the "no HEALTHCHECK" row means the gate cannot be quietly defeated by dropping it.

The crash cause is reported separately from the tail because a tail cannot carry it. A Go crash names its fault on the first line — fatal error: …, panic: …, [signal SIGABRT…] — and then prints a goroutine dump thousands of lines long, so --tail=200 reliably starts mid-stack, past the only line that says why the process died. The gate therefore searches the whole log for that header and prints the first one in context, plus how many headers the log holds in total — more than one means the container died repeatedly. The excerpt reaches backwards as well as forwards, because a crash inside cgo is reported twice: the C++ runtime says why it aborted before Go's handler prints SIGABRT, and that earlier line is the real cause. Tune with LOG_TAIL, CRASH_CONTEXT, CRASH_CONTEXT_BEFORE.

Only restarts observed during the wait count against a deploy. A restart count of its own does not: up -d leaves an already-current container in place, and a container that crashed once and recovered carries that count for the rest of its life — including while it legitimately replays a WAL, which is exactly when the wait is longest and the gate most needs to hold.

Run it by hand against a local stack the same way:

scripts/wait-for-healthy.sh cotel 120     # service, timeout in seconds

Timeout. The default is 120 s, and a normal deploy is far inside it: a graceful stop checkpoints the WAL, so the next open takes milliseconds, and measured against a 109 MB copy of production the container reports healthy in ~6 s — that figure is the probe cadence, not the database. A start that follows a hard kill replays the WAL instead and can take minutes; for that deploy use Run workflow and raise the health_timeout input rather than widening the default, which would blunt the gate for every other deploy.

Data

Data lives in the named volume (cotel-data) at /data/cotel.duckdb. You can query it directly:

docker run --rm -v cotel-data:/data ubuntu \
  duckdb /data/cotel.duckdb "SELECT model, COUNT(*) FROM spans GROUP BY model"

Snapshots

cotel exports its whole database to Parquet on a timer and keeps the newest COTEL_SNAPSHOT_KEEP exports in a second volume (/snapshots), so there is a recovery point that does not depend on the live DuckDB file being readable. The export runs inside cotel, on the same connection as everything else, and costs 0.15-0.3 s per run on a 152 MB database; the format is portable Parquet plus a plain-text schema.sql, which any DuckDB build can read (ADR-0023).

# what is on disk
docker run --rm -v cotel-snapshots:/snapshots debian:bookworm-slim ls -1 /snapshots

# restore one into an empty volume, verified against the snapshot's manifest
docker run --rm --entrypoint /usr/local/bin/cotel \
  -v cotel-snapshots:/snapshots -v cotel-data-restore:/data \
  ghcr.io/flopsstuff/cotel:latest --db-import /snapshots/2026-10-06T12-00-00Z

The worker's last outcome is reported on GET /api/v1/health under a snapshot object, and a failed export degrades the top-level status. The scheduled health probe reads that object on every tick and pages when the backup stops producing restore points, so nobody has to open the endpoint by hand. Full procedure, including promoting a restored volume and what the probe calls red: Database Snapshots and Restore.

Retention defaults

Tier Period Storage
Raw spans 30 days spans table
Daily aggregates 90 days daily_usage table

Override with environment variables:

Variable Default Description
COTEL_RETENTION_RAW_DAYS 30 Delete raw spans older than this many days (rounded down to a whole day — see below)
COTEL_RETENTION_AGGREGATE_DAYS 90 Delete daily aggregate rows older than this many days
COTEL_RETENTION_INTERVAL 6h How often the retention worker runs (Go duration string, e.g. 1h, 30m)

Roll-up consumes whole days only

The worker ticks several times a day, but it only ever rolls up and purges complete calendar days: the raw-span cutoff is snapped back to midnight before use. Rolling up a day in slices would risk a slice's raw spans being purged before the rest of the day is aggregated.

Those are UTC calendar days, and the cutoff is UTC midnight, whatever timezone the server itself runs in. Aggregates are bucketed by UTC day everywhere, so daily figures do not shift with the host's zone.

The practical effect: a raw span survives up to one day longer than COTEL_RETENTION_RAW_DAYS before it is aggregated away.

Late-arriving spans accumulate

A day's aggregate is built by accumulation: each roll-up cycle adds its sum to the existing daily_usage row (ON CONFLICT DO UPDATE) rather than replacing it. So a span dated to a day that was already rolled up and purged — a backfill, or an import (POST /api/v1/import) of telemetry older than COTEL_RETENTION_RAW_DAYS — is added to that day's total instead of overwriting it with itself alone. The accumulate and the raw-span purge run in a single transaction, so a crash between them cannot double-count.

Unattributed usage — the unknown sentinel

daily_usage is keyed by (day, session_id, model, tool_name). A raw span that carries no value for one of these (a missing/empty model, session_id, or tool_name) is rolled up under the sentinel string unknown rather than being dropped. This keeps the usage countable — e.g. how much spend has no model attached:

SELECT SUM(total_cost_usd) FROM daily_usage WHERE model = 'unknown';

Raw spans are left untouched (they keep their original empty/NULL value); the sentinel only exists in the daily rollup. See ADR-0009.

Token totals survive roll-up (including cache tokens)

daily_usage records total_input_tokens, total_output_tokens, total_cache_read_tokens, and total_cache_write_tokens. In real Claude Code traffic cache tokens are the overwhelming majority of volume (often ~99%), so without the cache columns any analysis of data older than the raw-span window (COTEL_RETENTION_RAW_DAYS, default 30 days) would undercount tokens by two orders of magnitude. Cost is unaffected either way — total_cost_usd is SUM(spans.cost_usd) and each span's cost already accounts for all four token kinds.

The cache columns were added later (schema v8). Rows rolled up before that migration keep NULL in the two cache columns — honestly "unknown", not 0, and not recoverable (the raw spans were already purged). Rows rolled up after the migration carry real sums.

Retention worker health

The retention worker's last outcome is reported on GET /api/v1/health under a retention object (status: ok | error | unknown, plus last_run_at and last_error). If the last roll-up failed, the top-level health status becomes degraded and the failure is logged at ERROR level — the worker keeps running and self-heals on the next tick.

Development

Building needs Go 1.24.0 or newer and a C toolchain — the DuckDB driver is CGo (ADR-0018). docker compose build brings its own; a local go build does not.

# Build locally
docker compose build

# Run locally
docker compose up

# Run tests
CGO_ENABLED=1 go test ./...

The runtime image is ≈ 294 MB. Most of that is the statically linked DuckDB engine (it was 213 MB on the DuckDB 1.1.3 driver, which loaded the ICU extension from the network at startup instead of bundling it).

Environment variables

Variable Default Description
COTEL_DB_PATH /data/cotel.duckdb DuckDB file path
COTEL_INGEST_ADDR :4318 Ingest listener address
COTEL_DASH_ADDR :8080 Dashboard listener address
COTEL_RETENTION_RAW_DAYS 30 Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer)
COTEL_RETENTION_AGGREGATE_DAYS 90 Daily aggregate retention in days
COTEL_RETENTION_INTERVAL 6h Retention worker tick interval (Go duration)
COTEL_SNAPSHOT_DIR (unset - snapshots off) Directory the snapshot worker exports the whole database into, one dated subdirectory per snapshot. Empty disables snapshots; docker-compose.yml sets /snapshots, backed by its own volume. See Database Snapshots and Restore.
COTEL_SNAPSHOT_INTERVAL 6h How often a snapshot is taken (Go duration). A snapshot is only taken when the newest complete one is older than this, so a restart cannot churn through the retained window.
COTEL_SNAPSHOT_KEEP 56 How many complete snapshots to keep; older ones and incomplete ones are pruned after each successful export. At the default interval, 56 is 14 days of reach for about 235 MB. The last snapshot standing is never pruned.
COTEL_SNAPSHOT_VOLUME cotel-snapshots Read by docker-compose.yml, not by the binary: the Docker volume mounted at /snapshots. Note that docker volume prune on a stopped deploy deletes it - the volume counts as in use only while the container exists.
COTEL_WAL_AUTOCHECKPOINT 4MB DuckDB checkpoint_threshold: the write-ahead log is folded into the main file once it grows past this size. Lower values bound how much WAL an ungraceful kill leaves to replay on the next open; higher values checkpoint less often during ingest. DuckDB's own default is 16MB.
CLOUDFLARE_TUNNEL_TOKEN (unset) When set, starts cloudflared tunnel run before cotel; enables public HTTPS access via Cloudflare Tunnel
TUNNEL_EDGE_IP_VERSION 4 in token mode Read by cloudflared, not by cotel: the address family used to reach the Cloudflare edge (4, 6 or auto). cloudflared's own default became auto in 2026.4.0, which tries whichever family the resolver answers with first and falls back only after a connection has failed; token mode pins 4 unless you set this. Not set in local-config mode, where config.yml owns the setting. See token mode
COTEL_PUBLIC_INGEST_URL (unset) Absolute http/https URL of the public OTLP ingest endpoint (e.g. https://cotel-ingest.yourdomain.com). When set, the Setup page substitutes this URL into the copy-paste Claude Code snippets.
SNAPSHOT_CHECK auto Read by scripts/probe-healthz.sh, not by the binary: whether a green /healthz is followed by a snapshot check against /api/v1/health. auto pages only on affirmative failure (status: error, or a snapshot older than the window); require also pages when the instance makes no snapshot claim at all, which is what you want wherever snapshots are expected; off never asks. See Production health probes.
SNAPSHOT_STALE_AFTER_SECONDS 43200 Read by scripts/probe-healthz.sh, not by the binary: how old snapshot.last_run_at may get before the probe goes red. Two COTEL_SNAPSHOT_INTERVAL periods at the shipped 6h; the probe reads the server's answer, not its configuration, so move this when you move the interval.
API_HEALTH_URL (derived from the /healthz URL) Read by scripts/probe-healthz.sh, not by the binary: overrides the /api/v1/health address the probe derives from the /healthz URL it was given.
COTEL_DATA_VOLUME cotel-data-repaired-20261004 Read by docker-compose.yml, not by the binary: the Docker volume mounted at /data. Point it at a restored copy to bring an instance up without writing to the volume being restored from — Docker has no volume rename, so the only other way to serve repaired data under the expected name is to overwrite the damaged original, which is also the forensic evidence. The default is a restored copy, not the cotel_cotel-data name compose derives by itself: that volume holds a database the binary can no longer open and is kept untouched as evidence. A fresh install that has neither volume gets the default created empty, which is correct.

Architecture

Claude Code  →  OTLP/HTTP (port 4318)  →  ingest handler
                                           ↓
                                        DuckDB (named volume)
                                           ↓
dashboard (port 8080)  ←─────────────────┘

Single Go binary, single DuckDB file, single named volume. No sidecars.

About

Claude Code OpenTelemetry

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages