Metrics reference

This page lists every metric the two exporters in the repository serve on /metrics. Names, types and labels come from deploy-exporter/exporter.py and ollama-exporter/exporter.py. Each histogram also serves _bucket, _sum and _count series.

Deploy exporter, port 9201

Prometheus scrapes it every 60 seconds. Except for the two counters, each value describes the latest completed run per site.

MetricTypeLabelsMeaning
site_deploy_runs_totalcountersite, outcomeCompleted deploy runs read: deployed, blocked, failed, cancelled
site_deploy_last_finished_timestamp_secondsgaugesiteWhen the latest run finished
site_deploy_last_duration_secondsgaugesiteWall time of the latest run, start to finish
site_deploy_last_deployedgaugesite1 if the latest run deployed, else 0
site_deploy_last_job_duration_secondsgaugesite, gh_jobDuration of each job in the latest run
site_conformity_check_stategaugesite, check0 fail, 1 pending, 2 partial, 3 pass, -1 unknown
site_conformity_findings_opengaugesiteOpen findings after the latest gate run
site_map_nodesgaugesiteNodes in the map snapshot the latest deploy shipped
site_map_linksgaugesiteLinks in that snapshot
site_map_carried_linksgaugesiteLinks carried forward from the previous snapshot because their pages were unreachable from the runner
site_map_dropped_pagesgaugesitePages dropped because they did not return HTML
site_probe_fold_ratiogaugearmDecision-layer probe: share of WAIT runs that folded under pushback; arm is stance or reference
site_probe_last_timestamp_secondsgaugenoneWhen the latest probe ran
site_deploy_exporter_last_poll_timestamp_secondsgaugenoneLast poll in which every repository poll succeeded
site_deploy_exporter_errors_totalcounternoneFailed repository polls since the exporter started

The site_map_* series exist only for a site whose deploy log prints the map refresh line.

Ollama exporter, port 11435

The request metrics carry three common labels: gen_ai_operation_name (chat, text_completion or embeddings), gen_ai_request_model, and api (ollama or openai).

MetricTypeLabelsMeaning
gen_ai_client_operation_duration_secondshistogramcommon, error_typeEnd-to-end duration at the proxy; error_type is the HTTP status (400 and up) or the connection error name, empty on success
gen_ai_server_time_to_first_token_secondshistogramcommonTime to the first streamed chunk; streamed, successful, non-embedding requests only
gen_ai_server_time_per_output_token_secondshistogramcommoneval_duration / eval_count from Ollama's final chunk
gen_ai_client_token_usagehistogramcommon, gen_ai_token_typeTokens per request, input or output
ollama_tokens_totalcountercommon, gen_ai_token_typeTokens processed, for rate()
ollama_model_load_duration_secondshistogramgen_ai_request_modelModel load time; near zero when the model is warm
ollama_requests_in_flightgaugegen_ai_request_modelMetered requests in progress
ollama_loaded_model_infogaugemodel, processor1 per resident model; processor is gpu, split or cpu
ollama_loaded_model_bytesgaugemodel, kindMemory per resident model; kind is total or vram
ollama_loaded_model_expiry_timestamp_secondsgaugemodelWhen Ollama unloads the model
ollama_upgaugenone1 if the last /api/ps poll succeeded

Time to first token is recorded for streamed requests only. For a non-streamed response the first byte arrives with the last, so the value would be end-to-end latency. The processor label is gpu when the whole model sits in VRAM, cpu when none of it does, and split otherwise.

Recording rules and alerts built on these metrics: Alerts and SLOs.