All insights
operating-modelresilienceriskconcentrationoperations

The blast radius: the operating risk your cost model can't see

Every operation gets mapped by cost and by the org chart. Neither of those maps tells you the one thing that decides your worst day: how far a single failure spreads. And the efficiency moves everyone rewards — consolidating onto one provider, one system, one source of truth, and stripping out the slack — quietly maximise that spread. This is a way to see the blast radius your cost model hides, to price the trade-off honestly, and to shrink it at the few nodes that matter without giving up the efficiency everywhere else.

Yasir Aheer15 September 20269 min read

A single supplier has a bad morning, and half your business stops. Not because anyone was careless — because, over years of sensible decisions, more and more of what you do came to rest on that one thing. One cloud region. One vendor. One "single source of truth." Each move made the operation cheaper and simpler. Together they quietly wired the whole thing to fail as a unit.

We have watched this play out at national scale. Analysing the US payment system, the European Central Bank reports modelling in which a cyber incident at just one of the five largest participants would, on average, take down more than a third of the entire network [1]. One node. A third of the system.

38%

Share of the US payment network — measured as a share of US banking-system assets — that a cyber incident at a single one of the five largest participants would affect, on average, in Federal Reserve Bank of New York pre-mortem modelling reported by the ECB.

Source: European Central Bank, Macroprudential Bulletin, Feb 2025 [1]

That number is not an indictment of the institution that fails. It is a property of the shape of the system around it. And most organisations have never drawn that shape.

Two maps, and the one you're missing

Every operation gets mapped at least twice. There's the org chart — who owns what — and there's the cost map — what each part spends. Leaders spend their lives optimising against those two pictures: consolidate the tooling, standardise on one platform, cut the duplicated effort, trim the spare capacity nobody could point at a use for.

Neither map shows you the variable that decides your worst day. That variable is the blast radius: if this node fails, how much of the operation goes with it? You can run an operation for years, hit every efficiency target, and never once have looked at it through this lens — because nothing in the cost report or the reporting lines makes it visible.

Why efficiency quietly maximises the damage

Here's the uncomfortable part: the moves that widen the blast radius aren't mistakes. They're the textbook. Consolidate onto one provider to gain leverage and lower unit cost. Route everything through a single system of record so the data is consistent. Strip out redundancy and spare capacity because idle resources look like waste. Each is defensible on its own. Each also takes a failure that used to be contained and gives it somewhere further to travel.

This isn't a software quirk; it's a general law of operations, and it's visible in physical supply chains too. The OECD finds that roughly 30% of exported products already sit under high concentration in a few trading partners, that import concentration is rising, and — crucially — that this concentration "is also a potential source of vulnerability to shocks" [4]. Their guidance is not "stop concentrating." It is to stop optimising for a single objective "without assessing possible trade-offs" [4]. That is the whole game: you cannot manage a trade-off you refuse to look at.

So look at it. Here is the same small operation you'd find anywhere — a few customer-facing functions sitting on a few shared dependencies. Detonate one shared node and watch how much goes with it.

The Blast Radius

Detonate One Shared Node. See What Goes With It.

Same operation, three shared dependencies — pick one to fail

When the shared identity / sso fails, 3 of 4 functions stop, taking down about 75% of the operation.

Share of the operation down

75%

3 of 4 functions stop

Why it's so wide: One sign-on service gates most functions. Convenient — and a wide blast radius. The more you consolidate onto it, the cheaper it runs — and the more of the operation it can take down at once.

Notice what the map teaches that the cost report never could: the node with the widest blast radius is usually the one you were proudest of consolidating. The efficiency and the exposure are the same decision, seen from two sides.

Price the trade-off, don't just feel it

"Resilience versus efficiency" is easy to nod along to and easy to ignore, because it stays abstract. It stops being abstract the moment you put numbers on it. Dial up the two ordinary efficiency levers — how much you route through one shared node, and how much slack you strip out — and watch the saving and the blast radius move together.

What does the efficiency actually buy?

Routing 70% of functions through one shared node and removing 60% of slack yields an efficiency index of 65 and a blast radius of 59% — about 5 of 8 functions. This is past the line where the saving is no longer worth the exposure.

Efficiency gained

65 idx

lower cost to run

Blast radius

59%

down on a single failure

Functions exposed

5 / 8

stop at the same moment

Over-consolidated. One failure now takes down 59% of the operation — 5 of 8 functions at once. Past this line the marginal saving is not worth the exposure. Put a firebreak or a second source on the shared node before you bank the efficiency.

Illustrative model. The absorb-able line is a design choice; regulators size it for whole markets, operators for a single business.

The model is built to tell you when to stop, not just when to worry. Keep the concentration modest and it blesses the design: efficient and survivable, leave it alone. Push past the line — the share of the operation you could actually absorb losing at once — and the next increment of saving is no longer worth what it costs you in exposure. That line is a design choice. Regulators draw it for whole markets; you draw it for one business. Either way, the point is to draw it on purpose rather than discover it during an outage.

Not every node deserves a firebreak

The wrong lesson here is "add redundancy everywhere." That just rebuilds the cost you cut, and most nodes don't warrant it. The blast radius only matters where two things are true at once: a lot depends on the node, and you can't swap it out quickly. That intersection — wide reach, slow substitution — is precisely what financial regulators sat down to define.

The EU's Digital Operational Resilience Act exists because of "the potential systemic risk entailed by increased outsourcing practices and by the ICT third-party concentration." Tellingly, it does not impose hard caps — the drafters judged "it is not considered appropriate to set out rules on strict caps and limits" — and instead directs supervisors to find the specific concentrations "likely to put a strain" on the system [2]. That is the right posture for an operator too: don't ban concentration, locate the concentration that would hurt, and treat only that.

So triage. For any node, two questions settle it.

Does this node deserve a firebreak?

Answer two questions about one node

1. How much of the operation depends on it?
2. How fast could you switch it out?
Answer both to see where this node lands.

Most nodes come back "accept it" — narrow reach, easy to replace, leave them efficient. A handful come back "firebreak it." Those few are where your entire resilience budget should go.

Shrinking the blast radius without losing the efficiency

Resilience is not the opposite of efficiency, and it is not a mood. It is a defined engineering property. NIST's standard names it precisely: the ability to "anticipate, withstand, recover from, and adapt to adverse conditions, stresses, attacks, or compromises" [3]. Those four verbs are not vibes — they're a checklist of design moves, and each one shrinks a blast radius in a different way without forcing you to abandon efficiency everywhere else.

Four moves that shrink the blast radius

Resilience isn't the opposite of efficiency — it's designed back in, node by node

Draw the dependency graph and mark every node by how much of the operation it can take down. This is a different picture from the org chart or the cost centre view — and it is the one that predicts your worst day.

Why it works

Supervisors now build exactly this: a "cyber map identifying operational and technology connections" to detect concentration before it bites. What is prudent for a regulator is prudent for an operator.

In practice

One diagram, nodes sized by blast radius. The shared node quietly sitting under half your functions is your top priority, whatever it costs to run.

Read them together and the pattern is clear: you keep the efficient, consolidated design almost everywhere, and you spend deliberately at the two or three nodes where a failure would be both wide and hard to undo. Firebreaks contain it. A warm second source shortens the recovery. A pre-agreed degraded mode means "less" instead of "nothing." None of this requires un-consolidating the whole operation. It requires knowing which nodes are load-bearing — which you only learn from the map you weren't drawing.

Draw the map before someone draws it for you

The financial sector got here first, and not by choice — DORA "came into application in January 2025", and supervisors now build exactly the picture I'm describing: a "cyber map identifying operational and technology connections" to spot concentration before it detonates [1]. Every other sector is on the same road, whether pushed by regulation, insurers, or a large customer who has done the arithmetic on your dependencies even if you haven't.

You don't need a regulator to make this worth doing. The blast radius is already there in your operation, decided by a hundred reasonable efficiency calls, waiting for the one bad morning that reveals it. The only question is whether you've seen it. Draw the operation by how far a failure spreads — not by what each part costs — and the few nodes that could take you down stop being a surprise and start being a choice.

Efficiency asks a good question: what's the cheapest way to run this? It just isn't the only question. The other one — when this breaks, how much breaks with it? — is the one your cost model can't answer, and the one your worst day will.

Sources

  1. European Central Bank. Cyber resilience stress testing from a macroprudential perspective (Macroprudential Bulletin). February 2025.On concentration as a systemic vulnerability: 'Concentration arises when there is a reliance on a small number of providers of a given service, meaning that an incident at one provider could have a disproportionate impact on the system.' Reporting Federal Reserve Bank of New York pre-mortem modelling, the bulletin states that a cyber incident in the wholesale payments network 'at one of the five largest participants in the US payment system would affect, on average, 38% of the network (as a share of US banking system assets).' It notes 'a cyber map identifying operational and technology connections may make it easier to detect concentration risks', records that a February 2023 ransomware incident left the trading-services group ION 'unable to process its transactions until the issue had been resolved', and confirms DORA 'came into application in January 2025.'View source
  2. European Union (Official Journal). Regulation (EU) 2022/2554 — Digital Operational Resilience Act (DORA). December 2022.Recital 31 establishes the Oversight Framework 'Taking into account the potential systemic risk entailed by increased outsourcing practices and by the ICT third-party concentration'. On why the remedy is management rather than prohibition: 'it is not considered appropriate to set out rules on strict caps and limits to ICT third-party exposures.' Instead a Lead Overseer must 'discover specific instances where a high degree of concentration of critical ICT third-party service providers in the Union is likely to put a strain on the Union financial' system.View source
  3. National Institute of Standards and Technology (NIST). SP 800-160 Vol. 2 Rev. 1 — Developing Cyber-Resilient Systems. 2021.Defines the property as 'the ability to anticipate, withstand, recover from, and adapt to adverse conditions, stresses, attacks, or compromises that use or are enabled by cyber resources.' From a risk-management perspective it is 'intended to help reduce the mission, business, organizational, enterprise, or sector risk of depending on cyber resources.' The four verbs — anticipate, withstand, recover, adapt — structure the discipline.View source
  4. OECD. OECD Supply Chain Resilience Review. 2025.'Approximately 30% of exported products are subject to high levels of concentration in few trading partners', and 'import concentration is on the rise, as countries are increasingly sourcing products from fewer suppliers than is globally possible.' Such concentration 'is also a potential source of vulnerability to shocks.' The report's policy guidance is to 'balance sustainability, efficiency and resilience: Focus on the performance of global supply chains as an ecosystem, and not target a single objective (such as assurance of supply) without assessing possible trade-offs with economic outcomes.'View source
Y
Yasir Aheer· Founder, OpsTeam

Yasir Aheer is the founder of OpsTeam. He writes about engineered operations, operating-model design, and the business and organisational implications of running People + Engineered Platforms + Production AI as one integrated system.

Ready to put this thinking into practice?

We design and operate integrated operating models for organisations ready to compound efficiency. Let's discuss yours.

Schedule Discovery Call