THE BREAKDOWN · ISSUE 002
The region is gone. What survives?
A simple recovery check + five live CloudOps roles
IN 20 SECONDS
AWS still cannot restore its Bahrain Region or one UAE availability zone after physical damage. This week, check whether your most important service can recover somewhere else.
Last week, three out of four poll respondents said Operations or SRE owns resilience. The sample was small, but the answer raises a useful question.
Does the team carrying the pager have a recovery plan that has actually worked in a test?
WHAT HAPPENED
A cloud region remains unavailable
On 15 September, AWS said:
Its Bahrain Region was still unavailable.
One of the UAE Region’s three availability zones was still inaccessible.
It was helping affected customers move to other Regions.
Physical attacks damaged the facilities earlier this year. The incident gives every cloud team a clear reason to review regional recovery.
Why this matters
Using several availability zones protects you from many local failures. A full regional failure can also affect data, identity, DNS, capacity and third-party services.
Your recovery plan must answer three questions:
Which services must recover in another Region?
How quickly must they return?
What test proves the plan works?
PICK THE RECOVERY LEVEL
Match the plan to the service
Backup and restore
Keep copies of data and infrastructure in another location. This costs less to maintain and takes longer to recover.
Pilot light
Keep the core data and services ready. Start the rest when an incident occurs.
Warm standby
Run a smaller copy of the service. Increase its capacity and move traffic when needed.
Active-active
Run the service in more than one Region at the same time. This recovers fastest and needs the most ongoing work.
THE SIMPLE RULE
Choose the lightest recovery setup that meets the business need. Then test it under realistic conditions.
THE 45-MINUTE CHECK
Use one critical service
0–5 minutes · Set the target
Write how long the service can stay down. This is your RTO. Write how much recent data you can lose. This is your RPO.
5–15 minutes · Find the state
List the databases, files, queues, secrets, identity services and external systems the service needs.
15–25 minutes · Walk through recovery
Follow each step from declaring the incident to serving users again. Include every approval and manual action.
25–35 minutes · Find proof
Open the latest recovery-test result. Compare the actual recovery time and data loss with your targets.
35–45 minutes · Assign three actions
Choose one technical gap, one ownership gap and one testing gap. Give each action an owner and date.
Six things teams often miss
Data: Can you restore the data and its encryption keys together?
Access: Can responders sign in when the usual identity service fails?
Traffic: Can you move users through DNS or another routing method?
Dependencies: Which API, SaaS tool or certificate still depends on the failed Region?
Capacity: Does the recovery Region have enough quota and licences?
Authority: Who can approve failover and accept possible data loss?
THREE QUICK SIGNALS
Worth knowing this week
1 · Location now needs a risk review
Cloud location decisions now need to consider physical and political risk alongside cost, speed and data residency.
2 · Attackers can move in hours
Google reported an AI-assisted credential attack that moved from a compromised cloud resource to execution in under six hours. Keep clean emergency credentials outside the normal access path.
3 · FinOps teams now cover more than cloud bills
The FinOps Foundation says 98% of practitioners manage AI spend and 90% manage, or plan to manage, SaaS spend. Resilience cost now belongs in the same value conversation.
The agentic era needs a different CRM. That’s Attio.
Teams like Parallel, Turbopuffer, and Wordsmith are already setting the pace on Attio. Get an always-on revenue engine, with agents and workflows that build pipeline, chase every buying signal, and move deals forward with your team. Whether you're working in your browser, inbox, or favorite agent, connect to your customer data in real-time through Attio's web app, MCP, API, and SDK.
POWERED BY OPTIMIZER365
See cloud operations in one place
Optimizer365 helps MSPs and IT teams see cost, security and governance signals in one view.
CLOUDOPS JOBS
Five live roles
We checked every link. Salary and closing dates appear where the employer published them.
Systems Engineer, Cloud Infrastructure · Ahpra
Azure, Microsoft 365, Terraform, identity and Level 3 support.
Melbourne · Hybrid
A$135,031 + 12% super · Closes 23 September
Azure Cloud Engineer · Built
Azure platform, Terraform, PowerShell, security and delivery pipelines.
Sydney or Brisbane · Permanent
Engineering Lead, Developer Platform · Virgin Australia
Platform engineering, golden paths, observability and policy as code.
Brisbane · Hybrid · Permanent
Senior FinOps Analyst · Skipton
Azure FinOps, forecasting, reporting, governance and cost optimisation.
Skipton · Hybrid
About £50,850 · Closes 26 September
FinOps Manager · Osborne Clarke
Multi-cloud cost governance, forecasting and DevOps cost controls.
Bristol · Hybrid · Permanent
Hiring through TCC?
Send us the role, location, salary range and direct application link.
ONE QUESTION
When did your recovery plan last work?
Count a full test only if your team moved a production-like service and checked the data, dependencies and recovery time.
When did your team last complete a full regional failover test?
Reply to this email: What was the one dependency your team discovered late? We will anonymise useful lessons and share the patterns.
Built for people running cloud
The CloudOps Collective connects people who run, govern and improve cloud estates. Forward this issue to the person who owns your recovery plan.
See you next week,
The TCC Editorial Team




