Probing Microarchitecture for Vulnerability Discovery

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-04-25 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

CVE discovery on CPUs follows a repeatable methodology: probe the gap between what the architecture claims and what the implementation actually does. The same probing discipline applies to RAG systems, distributed systems, and any layered abstraction.

The methodology

Researchers who find microarchitectural vulnerabilities (Spectre, Meltdown, Foreshadow, ZombieLoad, MDS) use four primary techniques, often in combination:

  1. Cache timing attacks. FLUSH+RELOAD, PRIME+PROBE, EVICT+TIME. The attack measures whether a memory access hit cache or main memory based on timing, and from that infers what speculative or kernel-level operation must have run.
  2. Microarchitectural fuzzing. Generating random instruction sequences with random operands, observing micro-architectural state via performance counters, looking for sequences that produce unexpected side-channel signal.
  3. Hardware manual gaps. Reading vendor architecture manuals carefully for operations that are documented but underspecified. The gap between specified and implemented behavior is where vulnerabilities hide.
  4. Performance counter introspection. Modern CPUs expose hundreds of performance counters that leak information about internal state (cache hit/miss, branch predictor state, TLB state, ROB occupancy). Researchers use these to map internal behavior that vendors did not intend to expose.

The architectural-claim vs actual-behavior pattern

The unifying insight is that complex systems have a published behavioral contract (architecture) and a real implementation (microarchitecture). The contract claims certain things never happen, are never observable, or are bounded in some way. The implementation is more permissive than the contract because permissiveness is faster.

Vulnerabilities live in the gap. The probing methodology is: read the contract, design experiments that should produce no observable signal under the contract, run them, look for signal anyway.

Spectre and Meltdown both came out of this pattern. The contracts said speculative execution was invisible to userspace. The implementations leaked it through cache timing.

Why the same pattern applies to RAG systems

Stefan applies this in retrieval debugging. A RAG system has an architectural claim ("retrieval returns the most relevant chunks for a query, scored consistently") and an actual implementation that is messier (chunking boundaries change semantics, embeddings drift, BM25 and vector scores disagree, reranking produces non-determinism).

The probing move: design a query that should retrieve a known chunk, observe whether it does, vary one input at a time, build a model of the actual retrieval surface as opposed to the documented one. The same methodology that finds Spectre finds the cases where the company Hivemind returns confidently wrong answers.

The convergence pattern

Spectre and Meltdown were discovered independently by multiple research groups (Graz University of Technology, Google Project Zero, MIT, Cyberus Technology) within months of each other. This is itself a signal: when multiple teams converge on the same gap, the gap is a structural feature of the design space, not a coincidental bug.

Operationally: when two unrelated debugging sessions point to the same implementation gap, treat it as a property of the system, not a one-off. Document it as a known surface.

Tools

For CPU work: rdtsc for cycle counting, perf for counter access, custom kernel modules for privileged probing. For RAG work: query logs with retrieval traces, embedding similarity matrices, A/B retrieval comparisons across configurations.

The principle is the same across both: instrument the implementation, hold the contract as a hypothesis, design experiments that should produce no signal if the contract is intact, look for signal anyway.

See also

CPU Cache Hierarchy and Speculative Execution · BM25 Hybrid Retrieval for Graph-RAG · Memory Architecture L0-L4 · Context Window Sizes and Effective Range