AWS Outage This Morning: What Happened & Our Response (Oct 20, 2025)

AWS experienced a broad incident this morning that triggered elevated errors and latency across popular services. Most Reliable Penguin clients were unaffected; where needed, we executed DR runbooks and failed over to alternate regions.

Table of Contents

Final status: RESOLVED. AWS reports all services returned to normal operations by 6:01 p.m. ET (3:01 p.m. PDT) on Oct 20. A formal post‑event summary is forthcoming from AWS.

Summary: Early this morning, a widespread Amazon Web Services (AWS) incident in US‑EAST‑1 (N. Virginia) caused elevated error rates and latency that rippled across multiple services. Most Reliable Penguin clients were not impacted; in several cases we executed disaster recovery (DR) runbooks and region failovers to maintain availability and performance.

High-level timeline (Eastern Time)

  • 3:11 a.m. ET — AWS begins investigating increased error rates/latencies in US‑EAST‑1.
  • 5:01 a.m. ETTrigger identified: DNS resolution issues for regional DynamoDB endpoints.
  • 5:24 a.m. ET — DynamoDB DNS issue resolved; services begin recovering.
  • 6:35 a.m. ET — DNS fully mitigated; most operations normal; EC2 new launches still erroring; some services (CloudTrail/Lambda) clearing backlogs.
  • 9:42 a.m. ET — AWS rate‑limits new EC2 launches to aid recovery across AZs.
  • 11:43 a.m. ET — Root cause narrowed to an internal subsystem monitoring Network Load Balancer (NLB) health; ongoing throttling of select operations.
  • 12:38 p.m. ETNLB health checks recovered; broad connectivity/API health improves.
  • 4:03 p.m. ETLambda invocation errors fully recovered; SQS→Lambda polling scaled back to pre‑event rates.
  • 5:48 p.m. ETEC2 launch throttles restored to pre‑event levels; dependent services clearing backlogs.
  • 6:01 p.m. ETAll AWS services returned to normal operations.
  • 6:53 p.m. ET[RESOLVED] Final note; some analytics/reporting backlogs (Config/Redshift/Connect) continue to process for a few hours.

What we saw

Impact to Reliable Penguin customers

Most Reliable Penguin clients remained up and serving traffic. In several environments we followed runbooks to shift traffic and/or promote standby resources in alternate AWS regions. Customer-facing availability and performance stayed within acceptable thresholds.

Key takeaways

  1. Multi‑region readiness pays off. Workloads with warm standbys or active‑active patterns avoided prolonged downtime.
  2. Upstream dependencies matter. Even if your application is healthy, SaaS/APIs you call may be degraded when their regions are impacted.
  3. Expect backlog drain. After status pages go green, queues and background jobs can take time to normalize.

Recommendations

  • Audit failover paths quarterly: DNS, health checks, data replication, feature flags, and message queues.
  • Validate RTO/RPO against a full regional impairment scenario.
  • Map external dependencies (payments, auth, messaging, analytics) and confirm each provider’s DR stance.
  • Runbook drills: rehearse traffic shifting (Route 53/Global Accelerator), database promotion, and queue rebalancing.

Service‑specific notes from this incident

Amazon EC2

New instance launches in US‑EAST‑1 experienced elevated error rates for several hours due to an EC2 internal subsystem impairment and subsequent NLB health‑check issues. AWS rate‑limited new EC2 launches during recovery and restored throttles to pre‑event levels by 5:48 p.m. ET. Launches not targeting a specific AZ generally succeeded sooner as capacity returned in some zones. Dependent services (RDS, ECS, Glue, Redshift) saw knock‑on effects while launch success rates were degraded.

  • Post‑incident: review Auto Scaling activity history for failed launches/throttling; confirm instance health across AZs; validate scaling policies.
  • Runbook: keep cross‑AZ launch templates; avoid AZ‑pinned launches during partial‑AZ impairments; ensure warm capacity in a secondary region.

AWS Lambda

Polling delays for SQS Event Source Mappings and async invocation throttling occurred during recovery due to impaired NLB health checks and dependent subsystems. Lambda invocation errors fully recovered by ~4:03 p.m. ET, with SQS→Lambda polling scaled back to pre‑event rates and queued events processed thereafter.

  • Post‑incident: verify DLQ depth, replays, and idempotency metrics; check event source mapping iterator age and function error rates.
  • Runbook: allocate reserved concurrency for critical functions; monitor async dwell time and queue depth.

Event services

EventBridge and CloudTrail accumulated backlogs early in recovery; AWS indicated new events were delivered normally while historical backlogs processed (from ~8:48 a.m. ET).

Root cause (final per AWS Health)

AWS states the event started with DNS resolution issues for regional DynamoDB endpoints (trigger identified at 3:26 a.m. ET). After DNS recovery (5:24 a.m. ET), a subsequent impairment in an EC2 internal subsystem responsible for launching instances caused additional issues. As EC2 launch impairments progressed, NLB health checks became impaired, leading to network connectivity issues across services such as Lambda, DynamoDB, and CloudWatch. AWS recovered NLB health checks by 12:38 p.m. ET, then gradually removed temporary throttles (EC2 launches, SQS→Lambda ESM processing, async Lambda invocations) until full service recovery at 6:01 p.m. ET. A small number of services (e.g., Config, Redshift, Connect) continued processing backlogs for a few hours after.

Live resources

Published: Oct 20, 2025, 6:16 a.m. ET (America/New_York)
Last updated: Oct 20, 2025, 6:58 p.m. ET


Need a DR check‑up? Reliable Penguin can pressure‑test your plan and run a table‑top exercise tailored to your stack.

Have a project or a problem?

Talk with a senior engineer for practical recommendations—no obligation.

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts

Categories

Get a free consultation from Reliable Penguin

Submit the form—or for immediate service call 866-649-7984.