Orvathis

Orvathis

News and analysis from the world of production systems.

Engineering

The Hidden Cost of Chatty Microservices

September 11, 2026

Splitting a monolith into services trades local complexity for distributed complexity, and nowhere is that clearer than in fan-out. A request that used to make three function calls now makes three network calls, each with its own timeout, retry policy, and failure mode.

The arithmetic is unforgiving. If each hop has a 99.9% success rate and your request touches ten services, end-to-end success drops below 99%. Users do not experience your architecture diagram - they experience the product of every dependency's reliability.

Continue reading →

Operations

What Good Observability Actually Looks Like

September 1, 2026

Monitoring consoles sprawl uncontrollably while offering little insight during live incidents. True observability operates under inverted priorities: an on-call engineer gets paged, and telemetry systems must identify the root diff within sixty seconds.…

Engineering

The Operator's Guide to Load Testing

May 10, 2026

Benchmark simulations repeatedly fail to anticipate live incidents because synthetic request topologies overlook messy reality. Evenly distributed traffic aimed at single endpoints provides isolated micro-benchmarks. Real degradation occurs when synchronized retry floods slam backends after a moment…

Infrastructure

Why Edge Caching Still Matters in 2026

May 24, 2026

Every few years someone declares the edge cache obsolete: bandwidth is cheap, compute is fast, so why bother? Yet p99 latency keeps telling a different story. The round trip from a user in Sao Paulo to an origin in Frankfurt costs around 200 ms on a good day, and no amount of application optimisatio…

Operations

Zero-Downtime Deployments Without the Drama

September 8, 2026

Popular engineering lore depicts seamless releases through elaborate blue-green switches and instantaneous traffic flips. Production reality centers on disciplined fundamentals: synthetic probes verifying realistic workloads, graceful connection draining on retiring nodes, and database schema mutati…

More reading

About us

Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.

More about the project →