⚠️ SYSTEM CRITICAL: PROD CLUSTER OUTAGE DETECTED ⚠️

SRE On-Call Arcade

It is 3:00 AM on Sunday. You are the On-Call DevOps Engineer. An alert siren blares! The production K3s Kubernetes cluster serving 50,000 active users just threw a 502 Bad Gateway and CPU utilization spiked to 100%!

SLA Budget
60.00s
Cluster Status
🔴 CRITICAL
Traffic
50,000 req/min
Error Rate
84%
[3:00 AM] INITIATING INCIDENT RESPONSE...
Goal: Diagnose the root cause of the 502 errors and CPU spike.
[3:01 AM] DIAGNOSIS COMPLETE.
Findings: Pod `payment-service-v2-7f8b9c` is crash-looping and hogging memory.
Goal: Contain the runaway pod and stabilize the cluster.
[3:02 AM] CONTAINMENT SUCCESSFUL.
Status: Rollback succeeded, but Redis cache is clogged with stale dead sessions.
Goal: Clean up the cache without dropping active user carts.

🎉 INCIDENT RESOLVED!

Claim Your Award

CERTIFICATE OF EXCELLENCE
This certifies that
Name
has successfully resolved a critical production incident under extreme pressure, maintaining system SLA and demonstrating exceptional Cloud Emergency Operations skills.
Lakshay Walia
Chief Cloud Architect

💥 SYSTEM CRASHED

SLA Violated. Users are furious. The company stock dropped 15%.