When “All-in-One” Becomes a Single Point of Failure

I was reminded of an old client story during a recent podcast recording. It’s one of those stories that starts with a bit of confidence. Then gets uncomfortable. Then ends with a lesson you carry into every design conversation afterwards. At the time, we had a medical clinic running six virtual machines from a single host. On paper, it was all fine. We had migrated them to a cloud backup solution. There was a proper disaster recovery plan. Backups were happening. Recovery options were documented. The boxes were ticked, the diagrams looked sensible, and everyone could sleep soundly. Or so we thought.

Nick Clift

7/19/20263 min read

I was reminded of an old client story during a recent podcast recording.

It’s one of those stories that starts with a bit of confidence. Then gets uncomfortable. Then ends with a lesson you carry into every design conversation afterwards.

At the time, we had a medical clinic running six virtual machines from a single host.

On paper, it was all fine.

We had migrated them to a cloud backup solution. There was a proper disaster recovery plan. Backups were happening. Recovery options were documented. The boxes were ticked, the diagrams looked sensible, and everyone could sleep soundly.

Or so we thought.

Then the RAID array on the host lost two drives.

Out of four.

Not the sort of alert you want to see while enjoying a quiet cup of coffee.

The recovery plan kicked in. We restored the environment onto the DR appliance, and technically, everything worked as designed. The virtual machines came back. Data was intact. The client could continue operating.

So far, so good.

But then we met the problem that nobody had properly challenged during the design phase.

Could that one DR appliance run the client’s production workload and handle recovery activity at the same time?

The answer was a very clear no.

It simply did not have the horsepower.

That’s the thing with a neat, all-in-one solution. It can look beautifully simple right up until the moment it has to do more than one important job. Then simplicity can turn into a bottleneck rather quickly.

A bit like asking one poor person at an MSP event to run registration, make coffee, fix the Wi-Fi and somehow enjoy the evening. It’s ambitious.

We worked closely with the vendor and moved the workload into cloud-based disaster recovery. That took the immediate pressure off the local appliance and kept the clinic running while we rebuilt the host environment.

Then came the slow, careful part.

As the repaired host came back online, we migrated the VMs back one at a time. No drama. No shortcuts. No attempt to rush the process just because everyone wanted the incident behind them.

The client didn’t lose data.

They experienced no meaningful downtime.

And, importantly, they never realised quite how close things had come to being much more disruptive.

But it still took nearly six weeks to fully restore their environment back onto the repaired host.

Six weeks.

That is a long time to be carrying the operational weight of an incident, even when the client is protected from most of the noise.

The lesson wasn’t really about whether we had backups. We did.

It wasn’t even about whether our disaster recovery plan worked. It did, broadly speaking.

The real lesson was about capacity, bandwidth, and asking better questions before the failure happens.

A single appliance can be a single point of failure in the obvious sense: if it breaks, you have a problem.

But it can also be a single point of constraint.

If one device is expected to store backups, run production workloads during an outage, manage recovery tasks, and perhaps support ongoing replication, it may be doing far too much. Not because the technology is bad, but because the design has asked it to be a hero.

And heroes, as it turns out, need a day off occasionally.

After that incident, we changed our standard.

For client environments above 2TB, we split the workload across multiple appliances. The overall cost was broadly the same. There was a little more to manage. There were a few more moving parts.

But the resilience improved massively.

That is often the trade-off in technology: simpler is not always simpler when things go wrong.

A single-box solution might make the initial deployment easier. It may reduce the number of devices in a rack. It may make a proposal look cleaner.

But if that one box becomes responsible for everything, it deserves a very honest design conversation.

I’m still figuring this out after more than 30 years around the MSP world, but experience has made one thing clearer: resilience is rarely about buying the cleverest product. It is about designing for the awkward day.

The day when two drives fail.

The day when recovery needs to happen while production is still running.

The day when a client needs you to be calm, not surprised.

If you manage critical client infrastructure, it may be worth taking a look this week at the devices doing “everything.”

Not to overcomplicate things for the sake of it.

Just to ask whether one appliance is being asked to carry more than it reasonably can.

Sometimes the best improvement is not a shiny new platform.

Sometimes it is simply giving an important job a little more room to breathe.

Connect

Clarity + Focus = Results

© 2026. All rights reserved.