Multi-Cloud Is Usually Overkill: Cost, Complexity, and the SLA Math
New! Listen to Concept to Cloud - Real stories from the trenches of software engineering
Multi-Cloud Is Usually Overkill
Cloud Architecture

Multi-Cloud Is Usually Overkill

TB
Tom Barber
July 6, 2026
0 min read

Multi-cloud promises resilience but usually delivers doubled cost and complexity for an outage that may never come. A look at the DNS trap, AWS's 300+ SLAs, the October 2025 US-EAST-1 outage, and why multi-region discipline beats a second provider.

In today’s post, I would like to discuss the whole premise behind multi-cloud and why, in most cases, it is overkill for the situation you find yourselves in. For those of you who don’t know, multi-cloud is the practice of spreading your cloud resources across multiple cloud environments to better handle high availability and failures within a single cloud provider. Now the whole idea with the cloud, of course, is that you offload your hosting burden to another service to allow you to concentrate on shipping products or whatever you’re trying to host inside of your cloud, but in doing so, you also greatly increase the complexity and the likelihood of mistakes being made for an outage that may never materially impact your organisation.

Of course, that’s not to say that multi-cloud should never be done. If, for example, the longevity of your organisation was at risk from any serious outage, then of course, multi-cloud might indeed be the answer.

The important stuff you have to bear in mind, for example, is DNS. If you host the DNS with a cloud provider and that provider goes down and you can’t change your DNS, then you’ve not increased your reliability in the slightest. If you offload your DNS to a third-party provider, you’re adding yet more services and a greater support burden on your staff to maintain them.

Say that you had a set-up that involved Azure, AWS, and Cloudflare for DNS. That gives you three different services, yet you still have a single point of failure in the DNS inside Cloudflare, which, you know, would always be the case because that’s DNS for you. At the same time, you then have:

  • firewalls
  • networking cross-boundary issues

and that’s before you even start thinking about database replication or the data locality issues that come with having two copies of everything. That then, of course, takes you on to the third point, which is the cost, because if you’ve got hot replicas of every data point in both systems (which I would assume, from a failure perspective, you would do), this is doubling your cost for an eventuality that may never come to pass. I have worked with a number of organisations over the years that thought that multi-cloud was the direction they wanted to go. Once you sit down and really discuss the reasoning behind why people want to go in a multi-cloud direction, the case often is for a better-architected solution rather than flipping straight to multi-cloud. That, of course, is not to say you don’t put some things in other cloud services in disaster recovery scenarios. For example, you may want to store your DR backups in a different provider, not just a different region. That would be fair enough and understandable. It doesn’t have to be a full-on hot copy of your existing cloud services to provide an always-on environment.

To try and add some colour to this otherwise quite dry blog post, let’s look at AWS’s published SLAs as an example. AWS publishes over 300 separate service SLAs rather than one platform-wide guarantee, as you can imagine.

The most common components that you would use, for example, are EC2 and EBS, under the hood, a lot of stuff runs on EC2 and EBS. With two or more AZs deployed concurrently in the same region, you have 99.99% monthly uptime. If you use a single EC2 instance, that drops to 99.5 under the instance-level SLA. If an S3 standard is 99.9% uptime, which is the uptime commitment, compared to the durability commitment AWS cites at 99.99999999999%, due to replication. RDS for databases across multiple AZs is 99.95%, and DynamoDB is 99.99% for standard tables and 99.999% for global tables.

At 99.99%, this equates to 4.38 minutes of permitted downtime per month. It’s worth stressing that these are credit-triggering commitments, not guarantees, and service credits are not automatically applied. Customers must, of course, raise a claim. AWS, at the end of the day, had at least one outage every month of 2025 except the last two. The shortest was 14 minutes, and the longest was 15 hours. If you know AWS, like we know AWS, the region with the highest number of outages in 2025 is US East 1, often attributed to its being the oldest, busiest region and to hosting many control plane services.

Now, for anyone who remembers, in October 2025, DynamoDB DNS race conditions caused thousands of websites and other services to be taken down in US East 1, with a reported blast radius ranging from 75 AWS services alone to more than 140 AWS services, depending on the count. There were over 6.5 million Downdetector reports spanning more than 1,000 services globally. For us to report, they identified this as the fourth major outage for US East 1 in five years, calling the concentration risk a “dangerously powerful yet routinely overlooked systemic vulnerability”. It wasn’t so isolated: the region suffered two outages in October 2025, followed by a May 2026 thermal event that knocked Coinbase offline for roughly seven hours.

Now, a single region at 99.99 already means 4.4 minutes of expected downtime per month. Two independent regions in active-active theoretically push you to 99.998%, adding a second cloud pushes a theoretical number higher still, but the marginal availability gain over multi-region is vanishingly small, while the complexity cost is enormous, with two IAM models, two networking stacks, as I mentioned before, the lowest common denominator services, and doubled operational services. I could go on.

More importantly, the independence assumption that justifies the math is false. In October 2025, many multi-region applications kept their primary databases in US East 1, only making them single points of failure. Applications in healthy regions couldn’t start because they relied on the secrets manager or the parameter store in US East 1. Global services, including IAM authentication, CloudFront, Route 53, and DynamoDB global tables, depend on US East 1 endpoints even for resources deployed in other regions. The fixes there were all multi-region discipline, not a second provider.

Of course, you really have to bear in mind that multi-cloud does not necessarily protect against AWS outages, because your other SaaS dependencies may not be multi-cloud and thus won’t be as resilient. Even a perfectly mirrored two-cloud app fails if it depends on a shared third-party service that runs on the cloud that is down, and that is not uncommon.

So I’ll wrap this blog post up, but obviously, when it comes to defaulting to multi-cloud, ensure that you understand the reasoning behind it. Ensure that you understand the complexity that goes with it because in switching to multi-cloud there is a distinct possibility that you’ll have more outages, not less, if you don’t execute the multi-cloud roll out as required.

TB
Written by Tom Barber

Ex-NASA engineer and cloud architect with over a decade of experience building scalable systems for startups and enterprises.

Work with Tom →

Related Articles

Tips

Should I be using Kubernetes?

Examine whether Kubernetes is the right choice for your infrastructure needs. Learn when Kubernetes excels and when simpler alternatives like Docker Compose, AWS ECS, or Google Cloud Run make more sense.

Read More →
Strategy

Cloud Stack Decisions for Early-Stage Startups: What Actually Matters

Most early-stage cloud stack debates are arguments about the wrong question. The tool doesn't matter nearly as much as the cost of switching, who fixes it at 3am, and whether your second engineer can read it.

Read More →
Strategy

The Essential Guide to Cloud Migration: Navigating Your Digital Transformation

A comprehensive cloud migration roadmap for 2025, covering foundational assessment through legacy system decommissioning, migration strategies, planning, execution, and long-term success factors.

Read More →

Ready to Build Your Product?

Let's discuss how we can help you bring your vision to life with expert cloud solutions

Get Started