Aug. 2, 2026 · 7 min read
Squirrels have shut down Nasdaq trading twice. Beavers, cows, and sharks have taken down real networks too, but misconfiguration, sabotage, and copper theft cause far more damage than any animal ever has.
Jul. 26, 2026 · 5 min read
Luck decides more than most people are willing to admit. Hard work and grit determine how ready you are for it.
Jul. 19, 2026 · 7 min read
Hyrum’s Law says that with enough users, every observable behavior becomes a depended-upon feature. Your implementation details become your API contract, regardless of documentation or intent.
Jul. 9, 2026 · 6 min read
Traditional accident models assume every failure traces back to a broken part. STAMP treats safety as a control problem instead, which explains disasters where every component worked as designed.
Jul. 3, 2026 · 6 min read
In 1986, a network slowdown of nearly a thousandfold forced Van Jacobson to invent TCP congestion control. Reno, CUBIC, and BBR are the forty years of guesses that followed, and the guessing continues today.
Jun. 26, 2026 · 7 min read
A Bloom filter checks whether something belongs to a set without storing the set. Firefox’s CRLite version compresses revocation data for nearly a billion certificates into 4 megabytes, guaranteeing only what’s absent.
Jun. 21, 2026 · 6 min read
Your p99 latency might not be a percentile. Averaging percentiles across replicas produces a meaningless number, and t-digest is the sketch algorithm that fixes it at scale.
Jun. 14, 2026 · 7 min read
Working notes on how LLMs train and generate text. Probably mostly right, certainly incomplete.
Jun. 7, 2026 · 5 min read
Goodhart’s Law says every metric used to track progress eventually gets used to fake progress instead.
Jun. 1, 2026 · 5 min read
Distributed causation means complex failures have many causes spread across time, systems, and organizational boundaries. Fixing just one leaves the rest in place.
May. 21, 2026 · 5 min read
Alignment distributes decisions to whoever has the context. Authority centralizes them with whoever has the title, which is almost never the same person.
May. 11, 2026 · 5 min read
LRU caching evicts the segments live viewers need when they rewind. Netflix’s patent US 12,621,504 B2 fixes this with segment age and distributed ownership, without bloating every edge server.
May. 7, 2026 · 4 min read
Primacy bias, recency bias, sycophancy, and anchoring are predictable distortions in how LLMs weight information. Understanding them changes how you prompt, evaluate, and trust model outputs.
May. 2, 2026 · 4 min read
Why efficiency improvements in technology often increase total consumption rather than reduce it, and how leaders can anticipate and manage the inevitable rebound.
Apr. 30, 2026 · 5 min read
How John Boyd’s OODA Loop (Observe, Orient, Decide, Act) applies to technology leadership, and practical strategies for eliminating the bottlenecks that slow your competitive cycles.
Jan. 8, 2023 · 28 min read
Southwest Airlines has grown while desperately trying to be scrappy, creative and humorous. The challenge is that scrappy doesn’t scale well.
Sep. 24, 2020 · 4 min read
Learn how you can move faster and focus on the things that matter by using incident analysis as your secret weapon. Operating at speed and at scale tests the capabilities of even the most experienced engineering teams. In this software world, it is inevitable that things will break. When they do, what do you do? Pick up the pieces and carry on? What if that’s not enough? Learning from incidents has taught us that broken things can lead to powerful opportunities, but only when we’re looking at them through the right lens.
Aug. 17, 2020 · 16 min read
Think about your team for a moment. How well is it functioning? Are you currently on a high-performing team, having found your groove and flow state as a group? Are you on a team that isn’t quite in that magical state of being yet, but it feels like you’re on your way? Are you feeling some friction—frustration, confusion—with your team? Or have you just joined a new team, so none of these apply yet?
Jul. 14, 2020 · 1 min read
Jun. 25, 2020 · 4 min read
Chaos Engineering builds confidence in distributed systems by deliberately introducing failures before they occur on their own. The discipline shifts teams from reactive incident response to proactive resilience building.