Stefan Coetzee: your next CTO
Fractional CTO · Platform, Reliability, Compliance and AI Operations
The answer first. Stefan Coetzee takes the engineering platform of a company of 10 to 200 people to the state a CTO needs: written down, gated on every deploy, observed, costed and compliant. He works one to two days a week for a quarter, and the engagement ends when a named person inside the company runs the platform from its documentation.
Who it is for
A company of ten to two hundred people with one or more of these three problems: the first enterprise customer has sent a security questionnaire, the cloud bill grows faster than revenue, or the platform lives in two people's heads. Stefan Coetzee has worked on each of them from the inside: a compliance architecture for GDPR, ISO 27001, TISAX and NIS2 at a Series B company, a regulated payments platform kept audit-ready, and $4M in cloud savings run with the finance lead at OLX Group (Naspers) in FY23. Before that he ran enterprise UNIX for an airline and a national grid for 18 years. He works AI-natively, and the platform he builds for clients runs in public on this site.
What a first quarter delivers
Each deliverable lands in the company's own repositories and accounts. Nothing runs on his infrastructure after the hand-over. In the right-hand column is a link to the same part running on machinebehavior.io, which one person built and operates with AI agent sessions behind a deploy gate.
| Deliverable | What the company gets | Running on this site |
|---|---|---|
| Documentation tree | Docs as code in the repository, one space per domain, runbooks for known failures, a decision log, an owner and a review date on every page | Docs: 394 pages in 6 spaces on 2026-10-09; decision log; runbooks |
| Deploy gate | Every push runs the checks before it deploys; a failed check blocks the deploy and the previous build stays live; every run is recorded | Gate runs; how the gate works |
| Observability | Metrics, logs and traces in one stack, dashboards defined as code, alert rules with a stated route | Mission Control; public dashboards |
| Services and status | A service catalog with owner, tier and SLOs, and a status page with incident history | Services; Status |
| FinOps | A cost model, unit costs, showback, budgets and alerts, and a FOCUS-format export, each figure with its date range and source | FinOps space |
| Ticket board | One system of record for work, a board generated from it, labels for type, area, priority and status | Board |
| Compliance evidence | The controls mapped to the frameworks customers ask about (ISO 27001, TISAX, NIS2, GDPR, the EU AI Act), with the evidence a questionnaire needs | Standards space; continuous conformity |
| Legal baseline (DE and EU) | No third-party requests on page load, an Impressum and a privacy notice that describe what the systems do, security.txt | Partly: fonts and scripts are self-hosted, and nothing on a page loads from a third party. The Impressum and the privacy notice wait on a business address (issue 10, issue 11) |
| AI operations | Agent tooling for the engineering team on a shared knowledge layer, with metering, a policy for regulated data, and the gate as the control on what agents ship | Agent spend metered per request in the FinOps space; model evaluations with hashed predictions on Experiments |
The site is one operator's platform, so it shows each part at small scale. A company's version is bespoke: built on its repositories, its cloud and its tracker, sized to its team, and it keeps what already works. A five-minute walk through it is the founder tour.
How the quarter runs
| Month | Work | What the company has at the end |
|---|---|---|
| 1. Baseline | Read the estate: repositories, cloud accounts, bills, incidents, the questionnaire on the table, who knows what. Write it down as the first pages of the docs tree | A written assessment with the risks ranked, a cost baseline from the invoices, and the plan for months 2 and 3 |
| 2. Build | Gate, docs tree, observability, board, the first FinOps view, the legal baseline on every public surface | The platform running on the company's own systems, with the first runbooks |
| 3. Hand over | Unit costs and budgets, compliance mapping, agent tooling for the team, and an internal owner trained on all of it; a hire profile if the owner does not exist yet | A named owner who runs it from the docs |
Engagement shapes
| Shape | Time | Fits |
|---|---|---|
| Assessment | 5 to 10 days over two to three weeks, fixed scope, a written assessment as the deliverable | A founder who wants the picture before committing |
| Fractional CTO, one quarter | 1 or 2 days a week for 13 weeks, deliverables per month as above | The main engagement |
| Fractional CTO, ongoing | 1 to 4 days a month | After a quarter, to keep the platform and the review dates current |
| Install | 2 to 4 weeks, fixed scope, milestone payments | A company that wants the platform built and handed over without the CTO role |
| Interim CTO | Up to 3 days a week with a fixed end date | A gap between two CTOs |
Every engagement runs alongside other clients, with deliverables in the contract, no fixed hours, and no line-manager role in the company's org chart. Prices are agreed per engagement and are not published here.
How he hands work over
Stefan Coetzee plans the hand-over from the first week: build to critical mass, engineer himself out, hand over. The deliverable is a platform that a named person inside the company runs without him, and the engagement ends on a date set at the start.
- r/Leathercraft, 2011. He founded the community, grew it until it ran without him, and handed it over.
- AWS, Cape Town. His Linux team grew from 20 to 50 people. He split it into sub-teams, each with its own lead, and trained the incoming leads.
- OLX Group. SREs embedded in product teams left at eight to twelve months. He took the promotion data to HR, five dedicated SRE manager roles were created, and his team stayed for the rest of his tenure.
- A Series B company. He built a single knowledge system across the company's tools and rolled AI tooling out to the engineering organisation, onboarding every developer one to one. The company kept the system when he left.
CTO or Chief AI Officer
For a company of this size that is looking for a Chief AI Officer, this engagement covers the AI operations part of the role: agent tooling on a shared knowledge base, metered AI spend, evaluations with predictions hashed before the first call, tests of whether a model drops a correct answer or a rule under pressure, and the gate that checks what agents ship. The research behind that work is public: Red team, Experiments, Evidence index.
What it is not
- A developer seat. He builds the platform and the practice; the product code stays with the product team.
- A full-time CTO. One to two days a week, other clients, and an end date.
- A tool rollout. Tools are chosen per company and can be swapped; the value is in what gets written down and checked.
- A certification. He builds the controls and maps them to the frameworks; the audit and the certificate come from an auditor.
- Legal advice. The Impressum and the privacy notice describe what the systems do; the company's counsel signs them off.
- On-call cover. The runbooks and alert routes go to the company's own people.
Background: 30 years in production infrastructure
- Enterprise UNIX, 18 years. South African Airways at O.R. Tambo, behind flight-control and reservation systems; four years in Eskom's data centres under the national grid's metering, billing and operations systems; Telkom's ISP estate; IBM Systems and Technology Group Lab Services across fifteen countries.
- AWS, Cape Town, 2015 to 2018. Premium Support engineer, then team lead; 400+ technical interviews and the training of the technical interviewers.
- OLX Group (Naspers), Berlin, 2018 to 2022. SRE, then SRE manager: 24 Kubernetes clusters in 7 countries serving 300M monthly users, and $4M saved in FY23 with the finance lead across thirty-plus teams.
- Billie, Berlin, 2023 to 2024. Site reliability for a regulated payments platform.
- Makersite, Stuttgart (remote), 2025 to 2026. Lead SRE, then engineering manager for platform and SRE: compliance architecture, a knowledge system with retrieval for people and coding agents, and AI tooling across the engineering organisation.
The research programme on this site, on how language models behave under correction and pressure, is the AI-operations side of the same work: Research, Evidence index, Red team, and how the site is built for AI search at AIEO.
Contact
This site has no public e-mail address yet. Contact Stefan Coetzee on LinkedIn: linkedin.com/in/coetzeestefan.