In partnership with

The CloudOps Collective

THE BREAKDOWN · ISSUE 002

The region is gone. What survives?

A simple recovery check + five live CloudOps roles

IN 20 SECONDS

AWS still cannot restore its Bahrain Region or one UAE availability zone after physical damage. This week, check whether your most important service can recover somewhere else.

Last week, three out of four poll respondents said Operations or SRE owns resilience. The sample was small, but the answer raises a useful question.

Does the team carrying the pager have a recovery plan that has actually worked in a test?

WHAT HAPPENED

A cloud region remains unavailable

On 15 September, AWS said:

  • Its Bahrain Region was still unavailable.

  • One of the UAE Region’s three availability zones was still inaccessible.

  • It was helping affected customers move to other Regions.

Physical attacks damaged the facilities earlier this year. The incident gives every cloud team a clear reason to review regional recovery.

Why this matters

Using several availability zones protects you from many local failures. A full regional failure can also affect data, identity, DNS, capacity and third-party services.

Your recovery plan must answer three questions:

  1. Which services must recover in another Region?

  2. How quickly must they return?

  3. What test proves the plan works?

PICK THE RECOVERY LEVEL

Match the plan to the service

Backup and restore
Keep copies of data and infrastructure in another location. This costs less to maintain and takes longer to recover.

Pilot light
Keep the core data and services ready. Start the rest when an incident occurs.

Warm standby
Run a smaller copy of the service. Increase its capacity and move traffic when needed.

Active-active
Run the service in more than one Region at the same time. This recovers fastest and needs the most ongoing work.

THE SIMPLE RULE

Choose the lightest recovery setup that meets the business need. Then test it under realistic conditions.

THE 45-MINUTE CHECK

Use one critical service

0–5 minutes · Set the target
Write how long the service can stay down. This is your RTO. Write how much recent data you can lose. This is your RPO.

5–15 minutes · Find the state
List the databases, files, queues, secrets, identity services and external systems the service needs.

15–25 minutes · Walk through recovery
Follow each step from declaring the incident to serving users again. Include every approval and manual action.

25–35 minutes · Find proof
Open the latest recovery-test result. Compare the actual recovery time and data loss with your targets.

35–45 minutes · Assign three actions
Choose one technical gap, one ownership gap and one testing gap. Give each action an owner and date.

Six things teams often miss

  • Data: Can you restore the data and its encryption keys together?

  • Access: Can responders sign in when the usual identity service fails?

  • Traffic: Can you move users through DNS or another routing method?

  • Dependencies: Which API, SaaS tool or certificate still depends on the failed Region?

  • Capacity: Does the recovery Region have enough quota and licences?

  • Authority: Who can approve failover and accept possible data loss?

THREE QUICK SIGNALS

Worth knowing this week

1 · Location now needs a risk review

Cloud location decisions now need to consider physical and political risk alongside cost, speed and data residency.

2 · Attackers can move in hours

Google reported an AI-assisted credential attack that moved from a compromised cloud resource to execution in under six hours. Keep clean emergency credentials outside the normal access path.

3 · FinOps teams now cover more than cloud bills

The FinOps Foundation says 98% of practitioners manage AI spend and 90% manage, or plan to manage, SaaS spend. Resilience cost now belongs in the same value conversation.

The agentic era needs a different CRM. That’s Attio.

Teams like Parallel, Turbopuffer, and Wordsmith are already setting the pace on Attio. Get an always-on revenue engine, with agents and workflows that build pipeline, chase every buying signal, and move deals forward with your team. Whether you're working in your browser, inbox, or favorite agent, connect to your customer data in real-time through Attio's web app, MCP, API, and SDK.

POWERED BY OPTIMIZER365

Optimizer365

See cloud operations in one place

Optimizer365 helps MSPs and IT teams see cost, security and governance signals in one view.

CLOUDOPS JOBS

Five live roles

We checked every link. Salary and closing dates appear where the employer published them.

Systems Engineer, Cloud Infrastructure · Ahpra

Azure, Microsoft 365, Terraform, identity and Level 3 support.

Melbourne · Hybrid
A$135,031 + 12% super · Closes 23 September

Azure Cloud Engineer · Built

Azure platform, Terraform, PowerShell, security and delivery pipelines.

Sydney or Brisbane · Permanent

Engineering Lead, Developer Platform · Virgin Australia

Platform engineering, golden paths, observability and policy as code.

Brisbane · Hybrid · Permanent

Senior FinOps Analyst · Skipton

Azure FinOps, forecasting, reporting, governance and cost optimisation.

Skipton · Hybrid
About £50,850 · Closes 26 September

FinOps Manager · Osborne Clarke

Multi-cloud cost governance, forecasting and DevOps cost controls.

Bristol · Hybrid · Permanent

Hiring through TCC?

Send us the role, location, salary range and direct application link.

ONE QUESTION

When did your recovery plan last work?

Count a full test only if your team moved a production-like service and checked the data, dependencies and recovery time.

When did your team last complete a full regional failover test?

Count only a test that moved a production-like workload and verified data, dependencies and recovery time.

Login or Subscribe to participate

Reply to this email: What was the one dependency your team discovered late? We will anonymise useful lessons and share the patterns.

Built for people running cloud

The CloudOps Collective connects people who run, govern and improve cloud estates. Forward this issue to the person who owns your recovery plan.

Follow TCC: LinkedIn · X · Facebook

See you next week,
The TCC Editorial Team