Horizontal vs Vertical Scaling

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-05-04 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

The choice is not "which is better" — it is "where does the bottleneck actually live, and which dimension can absorb it." Vertical scaling is bounded by hardware. Horizontal scaling is bounded by your willingness to engineer for it.

The two axes

Vertical (scale up): bigger box. More CPU, more RAM, faster disk on the same instance. Cheap at small sizes, exponentially expensive at the top end. Caps out at the largest instance type your cloud sells.

Horizontal (scale out): more boxes. Linear cost in capacity, no hard ceiling, but every shared resource between the boxes (database, session store, cache) becomes a coordination problem.

The trap: teams scale vertically until the largest instance is exhausted, then try to engineer for horizontal scaling under deadline pressure. Sharding a hot database while it's on fire is the classic version.

When vertical is correct

  • Single-writer databases where consistency cost dominates network cost
  • Workloads with sub-millisecond inter-component latency requirements (in-memory analytics)
  • Stateful systems where the state is too large or too coupled to partition cleanly
  • Early-stage products where engineering hours cost more than instance hours

When horizontal is correct

  • Stateless request-handlers (web, API tiers)
  • Anything fronted by a load balancer that can do consistent-hash routing
  • Read-heavy workloads where read replicas absorb most traffic
  • Workloads whose peak is 5x+ their median (autoscaling earns its complexity)
  • Anything you cannot afford to take down for a vertical resize

What actually limits horizontal scaling

It is not "the architecture." It is one of these, almost always:

BottleneckFailure mode
Shared databaseConnection pool exhaustion, write contention
Shared cacheHot-key thundering herd
Sticky sessionsUneven load, broken failover
Centralized config serviceSingle point of failure for N replicas
Cross-replica chatterCoordination cost grows non-linearly

Each one is a project to solve. Horizontal scalability is engineering work paid up-front to buy capacity later. Vertical scaling is engineering work deferred until the next instance class is impossible.

The hybrid that wins

Real systems are almost always N horizontally-scaled stateless tiers in front of a small number of vertically-scaled stateful tiers. The web tier scales out; the primary database scales up; the read replicas scale out; the analytics warehouse scales up. Knowing which tier is which is the design judgment.

See also

README · Apple Silicon vs Desktop GPU for Inference · 01-reliability · 07-performance