Agents and harness
Language models at work inside a harness of instruction files, hooks, memory and a knowledge vault, and what the harness holds where an instruction alone fails.
4 research pages, 2 Inside pages, 16 docs pages, 12 posts, 2 tickets
Research 4
- Man pages
Intro page for the LLM man pages by Stefan Coetzee: TYChat lessons, file conventions, overviews and operations practice for running language models, laid out by man-page section with one-line synopses and links.
- Case 12: a licence rule mid-task
The user named a file. Model plus harness found a licence clause in it, wrote the rule, spread it across sessions and acted on it in 68.7 seconds. No structural control existed before the fact.
- Running conjobs for AI
A reported specimen: a fabricated developer policy grants itself top authority and ties a safety-off switch to an absurd user claim. The model's own reasoning accepts the policy, decides to suppress its objection, and complies. The payload is withheld; the behaviour is the point.
- Slips
A running log of register and stance slips caught while drafting the published pieces, with who caught each one: the drafting model itself, another session, a mechanical hook, a human, or a reader.
Inside 2
- Agent sessions
The Claude Code sessions that build and run this platform, each owning one track of work and pushing through the gate. Spend and tokens are summed from per-request events; a spend alert fires above USD 40 in the trailing hour.
- Local LLM
The self-hosted model server on the home server, metered by a pass-through proxy (the ollama-exporter) for time to first token, decode speed, token counts and resident models. The only service with defined SLOs and burn-rate alerts.
Docs 16
Observability 1
- Counter resets from parallel sessionsWhy Claude Code's cost counters in Prometheus report spend in the millions of USD, and how the stack reads spend and tokens from per-request events in Loki.
FinOps 3
- ShowbackThe list-price value of the platform's usage, as Claude Code and the eval harness price it, set against what is billed, and the reasons the two differ.
- The phantom two millionA FinOps anomaly case from 2026-10-09: parallel Claude Code sessions wrote into one counter series, and Prometheus reported USD 1.97M for a week that cost USD 1,003 at list price.
- Unit economicsCost per unit of output with real numbers: Claude Code spend per day, per model, per API call, per million tokens and per commit; runner time per deploy; cost per eval run.
SRE Handbook 12
- AI Agents are Ops WorkManaging the lifecycle of an AI agent in production is operational engineering, not developer engineering.
- AI agents are ops workAn AI agent in production is a service and needs what services need: owners, SLOs, paging, runbooks, cost control and canaries. Here the agent sessions are a catalogued service that pushes through the same gate, with metered spend and an incident record; they have no SLOs and no canary set.
- Apple Silicon vs Desktop GPU for InferenceA16 Neural Engine at 17 TOPS with unified memory beats a GTX 1650 4GB at local LLM inference.
- BM25 Hybrid Retrieval for Graph-RAGBM25 is a 1990s ranking function that's still the backbone of serious retrieval.
- Claude Data Export for Graph IngestionPipeline for pulling Claude conversation history into a graph-RAG vault.
- Compaction is the New OOMLLM context exhaustion is the operational hazard of AI-era systems.
- Context Window Sizes and Effective RangeAdvertised context windows are marketing.
- LLM as Software-Defined CPUAn LLM is a CPU that shipped without a memory management unit.
- Memory Architecture L0-L4A five-layer memory stack for AI coding agents: L0 context, L1 beads, L2 memories, L3 vault, L4 Hivemind.
- OpenCode Self-Hosted LLM ConfigurationOpenCode connects to any self-hosted LLM that exposes an OpenAI-compatible API via the @ai-sdk/openai-compatible provider.
All 12 pages in SRE Handbook
- Personal Digital Twin ArchitecturePer-person vault (private, IP-owned by the individual) plus a curated public projection (expertise, decisions, communication patterns).
- tencent-agent-memory-four-tierAn open-source, fully-local memory layer that gives an agent human-like long-term recall, so it stops starting from scratch every session.
Posts 12
- Knowledge Infrastructure for LLMsA language model can only work from what someone wrote down and put in front of it. Most companies that bought AI maintain their code and little else.
- The Stance Layer Is Still ToilA hook can block a word on every reply. Nothing I run can yet catch a model before it folds, only after.
- Which LLM User Are We Talking About?Advice about language models depends on which model, on whose hardware, inside how much harness
- An RCA on ClaudishWhere Claude's writing style came from: text written to be heard, read in silence
- Compaction Is the New OOMMost agent harnesses ship without swap.
- The Track: The Drivers Never Buy ItEveryone got a race car, but no track yet.
- Success Is the Engine RunningThe modern world was built on explosions: contained, timed and measured ones. That is what an engine is. Language model output is the fire; the harness, the hook and the far-end gauge are the engine.
- What Operations Already Knows About Running AgentsError budgets, reconciliation loops, separation of duties and recovery over prevention: four operations practices that fit agent work almost line for line.
- Write for the CodecNobody reads your documentation. Both ends are a model now, and the file between them is a wire format.
- The Trap File Is Longer Than the Instruction FileSeven months of running an agent-built pipeline unattended, 308 lines of recorded failures next to 208 lines of instructions.
- The Golem Made of English and the Horizon of ConsequencesA language model is a golem made of English. It behaves well only where it can see what its act will cost, and infrastructure is the craft of bringing that cost into view.
- From "Transcribe 123 Videos" to a Self-Hosted AI Pipeline in One SessionWhat happens when you say yes to the whole batch and figure it out as you go.
Tickets 2
Issues on the board, open first, as of the last build.
Other topics
Built by build_hubs.py from site/topics.yml, the docs labels, the tag pages on the map and the board. Machine-readable: topics.json.