technology••5 min read

Why Software Engineers Are Looking at Building Failures to Save Distributed Systems

Software architects are increasingly adopting civil engineering principles to address systemic fragility. By analyzing 'progressive collapse,' experts are learning how to isolate components and prevent small errors from triggering catastrophic outages.

Why Software Engineers Are Looking at Building Failures to Save Distributed Systems

The Engineering Parallel

In civil engineering, a progressive collapse is a devastating event where a localized failure—like a structural column giving way—triggers a chain reaction, leading to the disproportionate collapse of an entire building. Think of the 1968 Ronan Point tower disaster. Today, software architects are realizing that their distributed systems are increasingly prone to the exact same phenomenon.

Sam Newman explores the intersection of structural engineering and system resilience at QCon.
Sam Newman explores the intersection of structural engineering and system resilience at QCon.

Why Modern Systems are Fragile

As development velocity increases, so does the complexity of our systems. Modern DevOps and agile methodologies have supercharged the pace at which we deploy code, but this speed often comes at the cost of oversight. Traditional quality assurance practices struggle to keep up with the emergent behaviors of microservices and complex cloud architectures.

  • Increased development velocity introduces more frequent failure modes.
  • System complexity often obscures the potential impact of small component failures.
  • High interdependency between services creates pathways for cascading outages.

Strategies for Resilience

To prevent a 'progressive collapse' in digital infrastructure, engineers must look beyond basic uptime monitoring. The focus is shifting toward compartmentalization and failure isolation. Key tactics include implementing strict timeouts for external service calls and rigorous retry policies to prevent a single slow service from clogging the entire pipeline. By reducing tight interconnections between services, teams can ensure that if one component fails, the rest of the building remains standing.

The more we change our applications and systems, the more likely we are to introduce failure modes that lead to emergent, system-wide behavior.

— Resilience Engineering Expert

Key Takeaways

  • Progressive collapse describes how a small, localized failure can trigger a massive system-wide shutdown.
  • Modern development velocity often outpaces traditional testing, making systems more susceptible to complex failures.
  • Strengthening components and isolating services is essential for maintaining stability in distributed environments.
  • Strict timeout and retry policies serve as critical circuit breakers in high-traffic cloud architectures.
  • Learning from physical structural engineering provides a blueprint for building more durable software.

FAQ

What is a progressive collapse in a software context?

It is a scenario where a single service failure creates a chain reaction, causing an entire distributed system to collapse disproportionately.

Why does increased development velocity impact reliability?

Faster deployment cycles introduce more frequent changes and complex dependencies, which makes it harder for traditional testing to catch emergent failure modes.

How can developers prevent cascading failures?

Engineers can use strategies like implementing strict timeouts, properly configuring retries, and reducing unnecessary interdependencies between microservices.

Are there lessons from physical building failures relevant to software?

Yes, civil engineering concepts like load path redundancy and blast resistance are being applied to create 'resilience engineering' strategies for software.

Sources