Multi-Cloud
AI
Amazon Bedrock
Resiliency

Beyond the Outage Notification: Building Cloud Status Hub with AI Insight

October 7, 20267 min read

An outage notification tells you something is broken. What you actually need in that moment is to know what it means for you — and what to do next.

In my professional life, I own and manage a multi-cloud strategy. That means my day-to-day doesn't live inside a single provider's console. Workloads, teams, and dependencies are spread across AWS, Azure, Google Cloud, and Oracle Cloud, and each of those providers publishes its health in its own way, on its own status page, in its own format, with its own vocabulary.

When something goes wrong, the first question is always the same: is it us, or is it them? Answering that quickly used to mean opening four browser tabs — AWS, Azure, Google Cloud, and Oracle Cloud — scanning four different status pages, and mentally stitching together whether a regional event at one provider lined up with what our own monitoring was showing.

That frustration is what led me to build Cloud Status Hub.

One View Across Every Provider

At its core, Cloud Status Hub is a single, unified dashboard showing the real-time operational status of AWS, Azure, GCP, and OCI. It refreshes every 60 seconds and gives you:

  • Current health for all four providers side by side, with a region × service breakdown
  • Active and recently resolved incidents, with links back to each provider's official status page
  • A 90-day incident history, so you can look back at what happened and how it played out

For anyone responsible for a multi-cloud footprint, being able to provide a single view of status across every provider is tremendously valuable. It turns a scattered investigation into a glance. It gives operations teams, architects, and leadership the same picture at the same time. And it removes a surprising amount of noise from the first few minutes of an incident — which are often the minutes that matter most.

But a single view on its own isn't new. Plenty of sites aggregate outages and availability. What I wanted was something more.

Where AI Changes the Game

Cloud Status Hub is one of the first projects I've shared with the community where AI is integrated directly into the application itself — not just used to help build it. Every incident that comes through the platform is analyzed by Anthropic's Claude running on Amazon Bedrock, using a custom prompt I've developed specifically to turn a provider's status update into something actionable.

This is what I see as the major differentiator from other outage and availability sites. They tell you that something happened. Cloud Status Hub attempts to tell you what it means and what to do about it.

For each incident, the AI Insight panel produces two separate briefs, written for two very different audiences:

The Technical Brief

Written for the engineers and architects who have to respond. It covers:

  • What we know — a clear explanation of what the issue actually is, beyond the provider's often-terse status text
  • Next actions — prioritized by urgency (immediate, high, medium, monitor), so you know what to do first
  • Services to check — guidance on how to review your existing workloads for exposure to this issue
  • Resiliency questions — recovery considerations and resiliency patterns worth evaluating so the same kind of event has less impact next time (if you want to go deeper, the AWS Well-Architected Reliability Pillar is a great place to start)

The Executive Brief

Written for the leaders who need to make decisions without wading through technical detail:

  • The bottom line — is this an act-now situation, a decide-if-confirmed situation, or simply awareness?
  • What's happening and how serious it is, in plain language
  • Likely customer impact
  • Decisions to consider if the impact is confirmed

Both briefs are shown right on the dashboard and can be downloaded as PDFs — easy to drop into an incident channel or forward to a stakeholder who just needs the summary.

I spent a lot of time on the prompt itself. Rather than letting the model improvise from a thin vendor paragraph, the prompt includes a curated resiliency reference for each category of service, so the guidance draws on reviewed patterns rather than invented specifics. It's also deliberately careful with disaster recovery advice — framing failover recommendations conditionally ("if a failover path exists, consider…") rather than assuming every organization's architecture is the same. And as an incident evolves — new updates, scope changes, resolution — the briefs are regenerated so they keep pace with what the provider is actually reporting.

It's like having a full-time engineer reviewing every outage and incident for you, and handing you a complete summary of next steps to deal with the issue and its potential impact — at every level of the organization.

Why Real-Time Insight Matters

During an active incident, information is the most valuable thing you have.

Every minute spent translating a status update, figuring out which of your workloads could be affected, or writing the first summary for leadership is a minute not spent on recovery.

A simple outage notification starts that clock. High-value context, delivered in real time, shortens it.

A typical outage alert

Service X is degraded in us-east-1.

Now the clock starts — and the investigation is all yours.

Cloud Status Hub + AI Insight
  • ✓ What the issue actually is
  • ✓ Which of your workloads to check
  • ✓ Prioritized recovery steps
  • ✓ Resiliency patterns for next time
  • ✓ An executive summary, ready to share

That's the gap I wanted Cloud Status Hub to close: going from notification to understanding to action, as quickly as possible.

Subscribe to Email Alerts

The site also includes email alerts, and they're open to anyone. Click Get alerts on the dashboard and subscribe to one provider, a few, or all of them — whatever matches your environment. You'll get an email when a provider you follow reports a new incident and again when it resolves, with a link that takes you straight to that incident and its AI Insight briefs on the dashboard.

Sign-up uses a double opt-in confirmation, and every email includes a link to change your providers or unsubscribe at any time.

What's Next

This is just the beginning. I have a list of future updates I'm looking forward to sharing for Cloud Status Hub — including more ways to tailor alerts to what matters to you — and I'm also working on some other AI Insight–based solutions that apply this same idea of turning raw signals into actionable guidance.

If you manage workloads across more than one cloud, or you just want a faster read on provider health when things go sideways, I hope you'll give it a try. I'd love to hear what you think and what you'd like to see next.

Cloud Status Hub is free and live at cloudstatus.synepho.com. The source is available on GitHub. AI Insight briefs are generated guidance to support your own assessment — always validate against your provider's official status page and your own monitoring.

About the Author

John Xanthopoulos is a cloud architect and web application developer. He writes about technology, systems design, and multi-cloud operations at synepho.com.