SLA Impact on Cost: Uptime → Spend

Published by Bles Software, a custom software and AI company based in Yehud-Monoson, Israel, building web apps, AI agents and API integrations for clients in Israel, the US, the UK and the EU.

Every nine of uptime, and the spend that buys it

An availability target is a purchase order. Each extra nine removes most of the downtime you are allowed and roughly doubles some line on the infrastructure bill, and the providers publish both halves. This page puts four things next to each other: the downtime each nine allows, the commitment the provider actually publishes, the service credit you get when it misses, and the price of the engineering that buys the nine, from a Multi-AZ database to an on-call rota. A target can then be chosen on arithmetic rather than ambition.

What each nine actually allows

There is no authoritative published table for this, so it is stated as what it is: arithmetic, on a 30 day month of 43,200 minutes and a 365 day year of 525,600 minutes [calc].

Four and a half minutes a month is the number worth staring at. At that allowance you cannot do a careless deploy, you cannot wait for a human to wake up, and you cannot restart a database by hand. That is why the cost of this target is organisational before it is technical.

What the providers actually promise, and what a breach pays you

Three things in that list matter more than the percentages.

A service credit is a discount, not a remedy. Google caps aggregate monthly credits at the amount due for the covered service [3], Cloudflare caps a year of Enterprise credits at six months of fees [6], and AWS RDS will not issue a credit worth less than a dollar [2]. If an hour of your downtime costs more than a month of your hosting bill, and for most products it does, the SLA is not insurance. It is an accountability statement with a small refund attached.

Stripe is the honest outlier and worth knowing about before you design around it. It publishes no SLA, its services agreement states it does not warrant uninterrupted or error-free use, and the 99.999 per cent figure you may have seen is a historical uptime claim on its marketing pages rather than a commitment [8]. Any availability number your product promises on top of a payment provider is your promise, not theirs.

Read the geography too. Google's own SLA commits 99.99 per cent for multi-zone Compute Engine but 99.95 per cent in Mexico and Stockholm [3]. The target you can buy depends on where you run.

The price of each nine, as the vendors publish it

The jump from one availability class to the next has a published price, and it is the cleanest number on this page.

So the database half of the first real nine costs exactly 100 per cent more [9][10], and both vendors say so on their own price lists. That is a far more useful planning number than any estimate: moving from a single instance to a highly available pair doubles that line, and the SLA you get in exchange moves from 99.5 per cent to 99.95 per cent on RDS [2].

Cross-region is the next step up and it is priced differently, as a per-gigabyte flow rather than a doubled instance [12]. A replica in another region costs the instance again plus continuous transfer, which is why the next nine up is a different order of spend rather than another doubling.

The support contract, which is part of the SLA whether you like it or not

An availability target implies someone answering the phone, and the cloud vendors price that as a percentage of spend.

AWS Business Support+ is the greater of $29 a month per account or 9 per cent of monthly AWS charges up to $10,000, then 7 per cent from $10,000 to $80,000, 5 per cent from $80,000 to $250,000 and 3 per cent above that [13]. Enterprise Support is the greater of $5,000 a month or 10 per cent up to $150,000, then 7, 5 and 3 per cent through the higher bands [13]. Google Cloud Customer Care Enhanced has a $100 a month minimum or 10 per cent of the first $10,000, then 7, 5 and 3 per cent, and Premium has a $15,000 a month minimum or 10 per cent up to $150,000, then 7, 5 and 3 per cent [14].

At enterprise scale that is a tenth of the infrastructure bill before any engineering, and it is a hard requirement for a credible four-nine commitment, because the response times are what you are buying.

The humans, and the tooling that makes the number true

The on-call rota is the line people forget, and it is not the seat price. A rotation needs enough engineers that nobody is on call every other week, which for 24 hour coverage is realistically six or more people. Priced against public wage data, where the United States Bureau of Labor Statistics puts software developers at a mean $71.20 per hour [21], the compensation and attrition cost of that rota dwarfs the $21 a month seat [15].

Whether the nine is worth it

The only published figures on the cost of downtime are self-reported surveys, so here is one with its method stated. ITIC's 2024 Hourly Cost of Downtime survey, polled from November 2023 to mid-March 2024, reports that 97 per cent of large enterprises of more than 1,000 employees say one hour of downtime costs over $100,000, and 41 per cent report $1 million to over $5 million an hour, from a sample that was 27 per cent small and midsize firms, 28 per cent mid-size enterprises and 45 per cent large enterprises [22].

That is what respondents believe, not audited loss, and we quote it as that. The honest way to use it is to do the same calculation on your own numbers: revenue per hour in your peak window, plus the support cost of the backlog, plus whatever a contractual penalty costs you. Compare that against the doubling in the price list above [9][10]. For a business where an hour's outage costs less than a month of doubled database spend, three nines is the right answer and the money should go into deploy discipline and observability instead.

How to choose a target

  1. Work out what an hour costs you, in your own numbers, not from a survey [22].
  2. Read what your providers actually commit, including the geography and the credit cap [1][3][5][6].
  3. Check what the providers underneath you promise. If one of them publishes no SLA at all [8], your number cannot exceed theirs honestly.
  4. Price the next nine as the doubling it is [9][10], plus the support percentage [13][14], plus the rota.
  5. Only then write the number into a contract, and write the measurement method next to it. An SLA without a stated measurement window and a stated exclusion list is an argument waiting to happen, which is exactly what the vendors' own documents avoid by stating both [2][3].

Sources

More Costs and Timelines from Bles Software