USE Method for Resource Saturation

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-05-04 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

When a system is slow and you don't know why, ask three questions of every resource it depends on: how busy is it, what is queued waiting for it, and is it returning errors. Brendan Gregg's USE method is the diagnostic primitive for performance problems.

The three signals per resource

For every relevant hardware and logical resource:

  • Utilization — % of time the resource was busy over an interval. CPU at 80%, disk at 60%, network link at 40%.
  • Saturation — work waiting because the resource cannot keep up. Run-queue length, disk request queue, TCP retransmits, GC pause time. Even when utilization is sub-100%, saturation can be non-zero (variance + queueing).
  • Errors — error events from the resource. ECC corrections, dropped packets, failed disk reads, failed memory allocations.

The order matters. Utilization tells you whether the resource is the candidate. Saturation tells you whether the candidate is actually the bottleneck. Errors tell you whether the resource is misbehaving even when nominal.

The resource list (Linux server, illustrative)

ResourceUtilizationSaturationErrors
CPU%CPU from top, mpstatrun-queue length (uptime, vmstat)thermal throttle events
Memoryfree -m, MemAvailableswap-in rate, oom-kill eventsECC errors (dmesg, mcelog)
Disk I/Oiostat -x, %utilawait, aqu-szI/O errors in dmesg
Networkinterface bps vs link captc -s qdisc, drops, retransmitsRX/TX errors (ip -s link)
File descriptorscat /proc/sys/fs/file-nrEMFILE errors in app logsper-process limits hit
Connection poolactive vs maxwait time for free conntimeout errors in app logs

This is the working set for a checklist. Walk it top-to-bottom when you don't know where to look.

Why the method works

The signal that a resource is the bottleneck is not utilization alone — it is utilization plus saturation. A CPU at 95% utilization with run-queue length 0.3 is busy but not bottlenecked. A CPU at 70% utilization with run-queue length 12 is the bottleneck — variance + scheduling delay are queueing work.

This is why "the box has CPU headroom" reasoning fails. Average utilization understates queueing. Saturation does not.

Where USE pairs with RED

USE is resource-driven (server-side). RED (Rate, Errors, Duration) is request-driven (service-side).

  • Use RED to detect that a service is slow.
  • Use USE to find which resource on which host is causing it.

A symptom-based alert (see Symptoms over Causes for Alerting) fires on RED. The investigation that follows uses USE.

What the method misses

  • Logical contention. Lock contention inside the application is not visible at the resource level. Tools: perf lock, eBPF profilers, application metrics.
  • Coherent failure. When multiple resources are saturated together (CPU + memory + disk), the cause is usually upstream — a thundering-herd retry storm, a runaway batch job. USE finds the symptoms; the cause is in the workload.
  • The cloud-resource layer. Network ENI bandwidth, EBS IOPS limits, instance-level CPU credits are saturation events that look like the OS is fine. Always check provider-side metrics in parallel.

See also

Symptoms over Causes for Alerting · CPU Cache Hierarchy and Speculative Execution · 02-scalability · 03-observability