Skip to content
Machine Behavior

Local LLM

serviceupdated 2026-10-09edit the YAML

The self-hosted model server on the home server, metered by a pass-through proxy (the ollama-exporter) for time to first token, decode speed, token counts and resident models. The only service with defined SLOs and burn-rate alerts.

Owner
Stefan Coetzee
Tier
Tier 3 Tooling a reader does not meet directly. Waits for the next working session.
Lifecycle
Production
System
Research tooling
Repository
uncovertechtalent/agent-observability
Dashboard
Local LLM (Grafana)
Current state
on the status page

Docs

Runbooks

No runbook yet.

Service level objectives

SLOTargetGood event
Availability99%A request without an upstream error
Latency95%A streamed chat request with a first token within 4 s

Defined and measured in Prometheus; see Alerts and SLOs.

Dependencies

Depends on

Outside the catalog: Ollama.

Used by

Service catalog · built from services/local-llm.yml by build_inside.py · services.json