Brad Doleman
← Writing

Redundancy Is Not Resilience

Why duplicate components do not necessarily protect the business from failure

Executive Summary

Infrastructure diagrams often communicate resilience through duplication. Two circuits, two routers, two firewalls, two power supplies, two data centers, or two cloud regions create an immediate visual impression that the design can survive failure. Sometimes that impression is accurate. Sometimes the duplicate components are exposed to the same underlying dependency and therefore fail together.

That distinction is the difference between redundancy and resilience. Redundancy provides additional components or paths. Resilience is the ability of a service to continue operating, or to recover within an acceptable period, when something actually fails. A design can contain substantial redundancy and still lack resilience if the redundant elements share power, physical pathways, carriers, upstream providers, control planes, configuration errors, management systems, authentication services, DNS, human procedures, or other common dependencies.

This is why infrastructure leaders should evaluate failure domains rather than count devices. The important question is not how many components appear on the diagram, but which credible failures the architecture can tolerate without unacceptable business impact.

True resilience is therefore a property of the complete service, not a feature of an individual device. It depends on architecture, operations, testing, recovery procedures, observability, and a clear understanding of what the business actually needs to survive.

Redundancy Is Easy to Draw

Redundancy is attractive because it is visible. A second firewall can be placed beside the first. A second circuit can be drawn beside the primary connection. Dual power supplies can be shown feeding a device. Two data centers can appear on opposite sides of an architecture diagram. The design immediately looks stronger.

The problem is that diagrams tend to emphasize components while hiding dependencies. Two circuits may enter the same building through the same conduit. Two carriers may ultimately rely on the same local access provider. Two firewalls may depend on the same switch, authentication service, management platform, or configuration. Two data centers may rely on the same identity provider or cloud control plane. A pair of redundant devices may share a software defect that causes both to fail under the same condition.

None of these examples makes redundancy useless. Redundant components remain an important tool for improving availability. The mistake is treating the presence of duplication as proof that the service is resilient.

A resilient design begins by asking what can fail together.

Start With the Failure Domain

A failure domain is the set of components or services that can be affected by the same failure. Thinking in failure domains changes the architecture conversation because it shifts attention away from individual devices and toward shared dependencies.

Consider a branch location with two internet circuits. If both services enter through the same underground pathway and construction damages that pathway, the organization may lose both connections at once. The circuits are redundant from a logical perspective, but the physical failure domain remains common.

A resilient design begins by asking what can fail together.

The same principle applies inside the facility. Two edge devices connected to the same upstream switch may still depend on that switch. Dual power supplies connected to the same power source may protect against a failed power supply while doing nothing for a building power failure. Two application paths that rely on the same DNS or identity service may both become unusable when that shared service fails.

Resilience improves when the architecture identifies these shared dependencies deliberately and decides which ones the business can accept.

Not Every Common Dependency Must Be Eliminated

A discussion of failure domains can quickly become unrealistic if the goal is to remove every possible shared dependency. Nearly every service ultimately depends on something common, and eliminating each dependency would be technically difficult, operationally complicated, and financially wasteful.

The appropriate objective is risk-based design. A small office may tolerate a longer outage and therefore justify a simpler connectivity model. A distribution facility, hospital, manufacturing site, trading environment, or other critical operation may require a much stronger design because connectivity loss has a greater business consequence.

The architecture should therefore begin with business impact rather than an abstract desire for maximum redundancy. How long can the service be unavailable? What functions must continue during a failure? How much degraded operation is acceptable? Which failure scenarios are credible enough to justify investment?

Resilience is strongest when the cost of protection is proportional to the consequence of failure.

Carrier Diversity Is More Complicated Than Two Logos

WAN resiliency illustrates the problem particularly well. Ordering circuits from two different carriers may appear to provide diversity, but the carrier names on the invoice do not necessarily describe the entire physical path.

Providers may use another company’s local facilities, share building entrances, lease common fiber, traverse the same conduit, or converge on the same upstream infrastructure. The enterprise may believe it purchased carrier diversity while still retaining a physical dependency that neither the network diagram nor the invoice makes obvious.

This does not mean perfect physical diversity is always available or economically justified. In some buildings and geographic areas, it may be impossible. The important point is that the limitation should be understood. If a location cannot obtain sufficiently diverse wired services, the architecture may choose another access technology, such as cellular or satellite connectivity, to create a different failure characteristic.

The objective is not simply to buy two circuits. It is to understand what kinds of failure those two circuits are expected to survive.

Device High Availability Solves a Specific Problem

High-availability pairs are another common source of false confidence. Two firewalls, routers, controllers, or load balancers can protect against hardware failure and, depending on the design, provide continuity during maintenance. That is valuable, but the pair addresses only the failures it was designed to handle.

Both devices may run the same software version and therefore share the same defect. Both may receive the same erroneous configuration. Both may depend on the same management system, routing policy, upstream network, or authentication service. A synchronization problem can even allow a fault to propagate from one member to the other.

The presence of an active and standby device should therefore prompt a second question: what remains common between them? The answer may be entirely acceptable, but it needs to be understood.

High availability is a mechanism within a resilience strategy. It is not the strategy itself.

Configuration Can Be a Shared Failure Domain

Infrastructure teams often think of common-mode failure in physical terms, yet configuration is one of the most powerful shared dependencies in a modern network. Centralized management, automation, and templating allow organizations to operate large environments consistently, but the same capabilities can distribute a mistake rapidly.

A routing change, firewall policy, software upgrade, authentication configuration, or automation error can affect redundant systems simultaneously. In that situation, the organization may have duplicate hardware and diverse physical paths while still experiencing a service outage because the failure occurred in the logical control of the environment.

This does not argue against centralized management or automation. Those capabilities are essential to operating at scale. It means resilience planning should include staged deployment, peer review, validation, rollback procedures, and controls that limit the blast radius of change.

Redundancy protects poorly against a mistake that is intentionally delivered to every redundant component.

People and Process Can Defeat Technical Redundancy

Resilience is also affected by how people operate the environment. A technically redundant service may depend on one engineer who understands the failover procedure, one vendor escalation path, one undocumented recovery step, or one set of credentials that becomes unavailable during an incident.

This is especially important during high-pressure events. A design may be capable of failing over automatically, but operations teams still need to recognize what happened, verify that the backup path is healthy, understand any degraded conditions, and restore normal operation safely. If the architecture is poorly documented or rarely exercised, the existence of redundant technology may provide less protection than expected.

Operational resilience therefore includes documentation, cross-training, escalation procedures, access management, vendor contacts, monitoring, and recovery practice. These elements rarely receive the visual prominence of hardware on an architecture diagram, but they can determine whether the organization actually recovers.

A resilient service must survive the loss of knowledge as well as the loss of equipment.

Failover That Has Never Been Tested Is an Assumption

Many organizations invest in redundant infrastructure but hesitate to test it because intentionally introducing failure feels risky. The concern is understandable. Testing can disrupt production if the design does not behave as expected.

That is also the reason testing matters. A backup path that has never carried production traffic may contain routing problems, stale configuration, insufficient capacity, expired credentials, missing security policy, or dependencies that were never identified. The first real outage is a poor time to discover those conditions.

Resilience testing does not have to mean reckless disruption. Tests can be staged, scoped, scheduled, observed, and designed with rollback criteria. Some components can be validated individually before larger service-level exercises are attempted. The appropriate method depends on business criticality and the architecture involved.

The principle is straightforward: confidence in failover should come from evidence, not from the fact that the diagram contains a second path.

Capacity Matters After Failover

A backup path may be technically available and still fail to preserve the business if it cannot carry the required workload. This is common when secondary connectivity is intentionally smaller or less expensive than the primary service.

There is nothing inherently wrong with designing degraded-mode capacity. Full duplication can be unnecessary and expensive. The business may be able to operate acceptably with reduced bandwidth, fewer applications, or prioritized traffic during an outage.

Confidence in failover should come from evidence, not from the fact that the diagram contains a second path.

The important point is that the degraded state should be designed intentionally. Which applications receive priority? What traffic can be limited? How many users can continue working? Will voice, payment processing, remote access, cloud applications, or operational systems remain usable? How long can the business operate in that condition?

Resilience does not always mean maintaining one hundred percent of normal capability. It means preserving the level of service the business has determined it needs during failure.

Monitoring Must See the Backup Path

Observability is another part of resilience that can be overlooked. Primary services receive constant use, which naturally provides evidence that they are functioning. Backup components may remain idle for long periods.

A secondary circuit can fail silently while the primary circuit continues operating. A standby device can develop a problem that does not become visible until failover is attempted. A backup power source can lose capacity. A replication process can stop. Without monitoring, the organization may believe it has redundancy that no longer exists.

This creates an unusual operational requirement: resilience mechanisms need to be monitored even when they are not actively carrying the normal workload. Health checks, synthetic tests, status telemetry, periodic validation, and alerting can provide evidence that the backup capability remains available.

A redundant component that has failed unnoticed is not redundancy when the primary component finally needs it.

Geographic Redundancy Can Still Share Logical Dependencies

Geographic separation is often used to reduce physical risk. Applications may be distributed between data centers, availability zones, cloud regions, or other facilities so that a local event does not interrupt the entire service.

Geography solves only part of the problem. Two locations can still depend on the same DNS platform, identity service, certificate infrastructure, management system, database, security service, network control plane, or external provider. A logical dependency can therefore create a failure domain that spans physically separate facilities.

The same issue can occur with operational procedures. If both environments are managed through the same process and a configuration error is deployed everywhere, physical distance provides little protection.

Geographic diversity is valuable, but resilience analysis must follow the service end to end. The question is not merely whether infrastructure exists in two places. It is whether a credible failure in one dependency can make both places unusable.

Complexity Can Reduce Resilience

Adding redundancy also adds components, state, configuration, monitoring, licensing, support relationships, and failure modes. Beyond a certain point, an architecture intended to improve availability can become difficult enough to operate that complexity itself creates risk.

This is an important counterweight to the instinct to duplicate everything. A simpler design that is well understood, well tested, and appropriately monitored may be more resilient than an elaborate design whose interactions are poorly understood.

Every additional resilience mechanism should therefore have a clear purpose. What failure does it address? How likely or consequential is that failure? How will the mechanism be tested? Who will support it? What new dependencies does it introduce?

Resilience is not measured by the number of boxes, circuits, vendors, regions, or recovery products an organization can afford. It is measured by whether the service can withstand the failures that matter.

Recovery Is Part of Resilience

Some failures cannot be absorbed through instantaneous failover. A destructive configuration error, cyber incident, widespread provider outage, data corruption event, or physical disaster may require recovery rather than redundancy.

This distinction matters because organizations sometimes invest heavily in high availability while giving less attention to restoration. High availability is designed to keep a service running through certain failures. Recovery is designed to restore the service when continuity mechanisms are insufficient.

A mature resilience strategy considers both. It identifies which failures should be handled automatically, which may result in degraded operation, and which require restoration from backups, rebuilding infrastructure, changing providers, or executing a disaster-recovery procedure.

The business does not ultimately care whether continuity came from failover or recovery. It cares how much capability was lost and how long it took to restore an acceptable service.

A Practical Resilience Review

A useful resilience review begins with the business service rather than an individual device. Trace the service from the user to the application and identify the dependencies that could interrupt it. Then determine which failures are protected, which are merely duplicated, and which remain accepted risks.

Questions to ask, and what each one tells you
QuestionWhat It Reveals
What business service are we trying to protect?Resilience objective
How long can the service be unavailable?Recovery requirement
What degraded level of service is acceptable?Minimum operating capacity
Which components are redundant?Visible duplication
Which dependencies are shared by those components?Common failure domains
Do redundant circuits have meaningful carrier and path diversity?WAN resilience
Do redundant devices share software, configuration, management, or upstream dependencies?Logical common-mode risk
Can the backup path carry the required workload?Failover capacity
Are standby components monitored while idle?Backup health
When was failover last tested successfully?Operational evidence
Can the service be recovered if failover does not work?Recovery capability
Are recovery procedures documented and accessible during an incident?Operational readiness
How many people can execute the recovery process?Knowledge resilience
Which external providers or services remain single dependencies?Third-party exposure
What additional complexity does the resilience design introduce?Operational burden
Which remaining risks have leadership consciously accepted?Risk ownership

The final question is the one that matters most: Which realistic failure can still interrupt the service despite everything we have duplicated?

That question often reveals more about resilience than counting redundant components ever will.

Resilience Must Be Proportional to Business Need

Not every system deserves the strongest possible resilience architecture. Designing every service for near-zero interruption would be prohibitively expensive and operationally complex. Some applications can tolerate hours of downtime. Some branch locations can operate temporarily on reduced connectivity. Other services may have consequences that justify substantial investment in diversity, failover, recovery, and testing.

Infrastructure leaders should therefore connect resilience spending to business impact. Recovery objectives, critical business processes, regulatory requirements, safety concerns, revenue exposure, customer impact, and operational dependencies all help determine how much protection is appropriate.

This approach also improves executive conversations. Instead of asking leadership to fund ’more redundancy,’ the infrastructure team can explain which business failure the investment is intended to prevent, what exposure remains without it, and what level of service the proposed design is expected to preserve.

That is a much stronger basis for investment than the assumption that more duplicate technology is always better.

Key Takeaways

  1. Redundancy and resilience are related, but they are not the same. Redundancy adds components or paths; resilience describes the service’s ability to continue or recover when failure occurs.
  2. Failure domains matter more than device counts. Redundant components can still fail together when they share physical, logical, operational, or human dependencies.
  3. Carrier diversity requires more than two provider names. Physical pathways and underlying access dependencies should be understood where the business requirement justifies the effort.
  4. High-availability pairs address specific failures, not every failure. Shared software, configuration, management, and upstream dependencies can defeat duplicate hardware.
  5. Testing converts resilience from an assumption into evidence. Backup paths, failover behavior, degraded-mode capacity, and recovery procedures should be validated at a level appropriate to business risk.
  6. Operational processes are part of resilience. Documentation, monitoring, credentials, cross-training, escalation, and recovery practice can determine whether technical redundancy actually protects the business.
  7. More complexity does not automatically create more resilience. Every additional mechanism should address a defined failure and justify the operational burden it introduces.
  8. Resilience should be proportional to business need. The objective is not maximum redundancy; it is sufficient continuity and recovery for the consequences the organization is trying to manage.

Closing Thought

Infrastructure diagrams make redundancy easy to see. Resilience is harder because much of it lives outside the boxes and lines. It exists in physical pathways, software behavior, configuration controls, carrier dependencies, monitoring, recovery procedures, documentation, and the people who have to make the system work when normal conditions disappear.

That is why a second device or circuit should never end the architecture conversation. It should begin another one: What failure does this protect us from, and what can still fail with it?

Sometimes the answer will reveal a dependency that should be removed. Sometimes the dependency cannot be removed economically and should simply be documented and accepted. Sometimes the business does not require enough protection to justify further investment. Those are all legitimate outcomes when they are understood.

The objective of resilient architecture is not to create a network in which nothing can fail. That network does not exist. The objective is to understand how the service can fail, decide which failures matter enough to address, and build enough technical and operational capability to keep the business functioning when those failures occur.

Redundancy gives us alternatives. Resilience determines whether those alternatives will actually be there when the business needs them.

The examples in this article are generalized. It is grounded in professional experience with distributed enterprise networks, high-availability design, WAN connectivity, carrier services, infrastructure operations, failover, monitoring, and recovery planning, and is written to communicate architectural lessons without identifying a specific employer, client, carrier arrangement, proprietary network, or individual project.

Brad Doleman leads enterprise infrastructure and IT operations, and wrote A Field Guide for Engineers Moving Into IT Leadership. The professional record More writing