Metrics
shed samples every running container and the host itself every 10 seconds, keeps a week of history in SQLite, and serves it as 180 evenly spaced buckets per chart.
Container sampling
A collector ticks every 10 seconds. It lists the running containers that carry the shed.service label and asks Docker for a one-shot stats reading of each. Those readings are cumulative counters, so a number only exists once there are two of them: shed keeps the previous reading per container and turns the difference into a rate. A container's first reading yields nothing, and neither does one where a counter went backwards, as after a restart.
cpu = Δcpu_ns / Δsystem_ns × online_cpus × 100 # percent of ONE core (fallback: Δcpu_ns / Δwall_ns × 100) memory = usage − inactive_file # a level, not a rate: page cache that can be dropped is not counted net = Δbytes / Δseconds # rx and tx, summed over all interfaces disk = Δbytes / Δseconds # block reads and writes
CPU is a percentage of one core, so a busy two-thread app can show 200%. A service's limit in the same units is cpuLimit × 100. During a zero-downtime deploy two containers of a service run side by side, and their rates are summed into one row per service per tick. Samples older than 7 days are deleted hourly.
Host metrics
The host is read on the same tick from the kernel, not from Docker, so it includes everything on the box, not just shed's containers.
| Metric | Source | Computation |
|---|---|---|
| cpu | /proc/stat | Total is the first eight fields of the cpu line (user through steal). Busy is total minus idle and iowait, clamped to the elapsed total. Percent is busy / total × CPUs × 100. |
| memory | /proc/meminfo | MemTotal − MemAvailable. |
| net, disk I/O | /sys/class/net, /proc/diskstats | Bytes per second summed over physical devices only: those with a device link in sysfs. Docker bridges, veth pairs, loop and device-mapper devices are left out so traffic is not counted twice. Sectors are 512 bytes. |
| disk used | statfs | Used space of the filesystem that holds data.dir. |
cpus, memoryTotal, and diskTotal are read live when you query, not stored per sample.
Ranges and buckets
A query asks for a range: 1h, 6h, 24h, or 7d. The answer always has 180 buckets, each the average of the samples that fall in it, or null if there are none. The response gives start and step, and bucket i covers [start + i × step, start + (i + 1) × step).
| Range | Bucket width | Samples per bucket |
|---|---|---|
1h | 20s | 2 |
6h | 120s | 12 |
24h | 480s | 48 |
7d | 3360s | 336 |
Buckets are aligned to multiples of their width since the Unix epoch, so a bucket's boundaries are the same on every query and charts do not jitter as time passes. The window normally ends with the bucket that contains now. But a bucket that has just started may not have a sample yet, and a chart that ends in a hole looks like an outage. In that case the whole window shifts back one step, to end at the last complete bucket. To make that possible the server fetches 181 buckets, one before the window.
| Field | Value |
|---|---|
| now | 2026-01-01 00:12:17 UTC |
| step | 20s |
| start | 2025-12-31 23:12:20 UTC |
| end | 2026-01-01 00:12:20 UTC |
| start / step | 88361137 (a whole number: epoch-aligned) |
The bucket containing now already has samples, so the window ends with it.
Charts refetch about once per bucket, between 10 seconds and a minute. The same bucketing applies to service metrics and host metrics. See CPU and memory limits for the limit lines drawn on service charts.