Amazon Web Services confirms cause of global service outage and details subsequent response

Amazon Web Services confirms cause of global service outage and details subsequent response

AWS has released a statement revealing details of its response to the major service disruption that struck its Northern Virginia data centres, affecting key cloud services including DynamoDB, EC2 and Lambda.

Amazon Web Services (AWS) has confirmed the cause of a widespread service disruption that affected its Northern Virginia data centres, one of its largest and most critical regions.

The statement also outlines in detail the AWS response that brought services back on-line.

The incident, which began late on October 19 and lasted through the afternoon of October 20, caused elevated error rates, failed instance launches and connection issues across key AWS products including DynamoDB, EC2, Lambda and Network Load Balancer (NLB).

While the outage was contained to the US-EAST-1 region, the ripple effects reached customers globally who rely on AWS for storage, computing and networking.

According to AWS, the disruption unfolded in three overlapping phases over roughly 15 hours, with different services experiencing distinct impacts and recovery timelines.

A chain reaction beginning with DynamoDB

AWS says the first signs of trouble emerged at 11:48pm  PDT on October 19, when Amazon DynamoDB – the company’s fully managed NoSQL database service – began returning elevated API error rates. For nearly three hours, customers and internal AWS systems that depended on DynamoDB were unable to connect to the database.

AWS engineers traced the issue to a latent defect in DynamoDB’s automated DNS management system, which governs how endpoints resolve traffic across massive fleets of load balancers. The defect triggered a rare ‘race condition’ between two independent components known as DNS Enactors. Each component is designed to apply DNS configuration plans to Amazon Route 53, AWS’s internal DNS service, across multiple availability zones.

Under normal operation, these Enactors prevent conflicting updates, but in this instance, one Enactor’s delayed processing overlapped with another’s rapid plan cleanup. The timing led the older plan to overwrite the newer one just before it was deleted, inadvertently erasing all IP addresses associated with the regional DynamoDB endpoint (dynamodb.us-east-1.amazonaws.com). This left the system in an inconsistent state that could not automatically self-correct.

AWS says, as a result, all requests to DynamoDB in the region began failing DNS lookups. Customer applications, as well as dependent AWS services such as Lambda, Redshift and EC2’s orchestration systems, were unable to reach the database. Global table replicas in other regions continued to operate but experienced significant replication lag.

By 12:38am PDT on October 20, AWS engineers had identified the faulty DNS state as the root cause. Manual interventions were implemented to restore DNS configurations, and by 2:25am, endpoint records were repaired. As cached DNS entries expired, service connectivity gradually returned, and by 2:40am DynamoDB operations were fully restored.

EC2 instance launches hampered by cascading failures

AWS acknowledges that even after DynamoDB recovered, the fallout continued. Amazon EC2 – the backbone of AWS’s virtual server offerings – suffered degraded performance and new instance launch failures between 11:48pm October 19 and 1:50pm October 20.

While running instances remained healthy, AWS says attempts to launch new EC2 instances failed with ‘insufficient capacity’ or ‘request limit exceeded’ errors. The problem traced back to EC2’s DropletWorkflow Manager (DWFM) subsystem, which manages the physical servers, or ‘droplets’, that host EC2 instances. DWFM depends on DynamoDB to track server leases.

When the database became unreachable, DWFM’s lease checks began failing, causing existing leases to expire and marking many droplets as unavailable for new workloads.

AWS says after DynamoDB was restored, DWFM attempted to re-establish leases across thousands of servers, but the sudden volume of requests led to system congestion. Engineers described the system as entering a ‘congestive collapse’, unable to clear its queue of lease re-establishment tasks fast enough.

To resolve the bottleneck, AWS says its engineers throttled incoming requests and selectively restarted DWFM hosts to clear backlogs. By 5:28am the system had re-established leases but network configuration propagation remained delayed due to a second subsystem – Network Manager – struggling to process a backlog of network state updates. Newly launched instances were operational but could not communicate properly within virtual networks. Full network propagation recovered by 10:36am and EC2 services were restored completely by 1:50 pm.

Load balancers compound the impact

The third wave of disruption involved the Network Load Balancer (NLB) service, which routes incoming traffic across EC2 instances. Between 5:30am and 2:09pm. PDT on October 20 some NLBs in the Northern Virginia region began returning increased connection errors.

AWS says the issue stemmed from the same network propagation delays that had affected EC2. As NLB’s health-check subsystem attempted to validate newly launched instances whose network configurations were still incomplete, it misinterpreted them as unhealthy and temporarily removed them from service.

This, says AWS, led to a ‘vicious cycle’ with healthy nodes incorrectly marked unhealthy, removed from DNS rotation and then reinstated moments later when the next health check succeeded. The rapid oscillation overloaded the health-checking system itself, which began to degrade, causing DNS failovers between availability zones. Some multi-AZ load balancers lost capacity, leading to connection timeouts for customer applications.

AWS engineers intervened at 9:36am by disabling automatic health-check failovers, allowing all available capacity to return online. The NLB service was fully restored after EC2 stabilised later that afternoon.

Collateral effects across the AWS ecosystem

Because DynamoDB, EC2 and NLB are core building blocks of the AWS ecosystem, dozens of dependent services were affected.

The statement says AWS Lambda, which relies on DynamoDB and EC2 infrastructure for scaling, experienced API errors and invocation delays from 11:51pm October 19 through 2:15pm October 20.

Function creation and updates failed early in the event, followed by processing delays for event sources like Amazon SQS and Kinesis. AWS throttled background Lambda workloads to preserve synchronous invocations until sufficient capacity was restored by midday.

Container services including Amazon ECS, EKS and Fargate suffered cluster scaling failures and delayed container launches during the same period, recovering by 2:20pm.

AWS says Amazon Connect, the company’s contact-centre platform, faced widespread errors handling calls, chats and tasks. Inbound callers encountered busy signals and dropped connections, while agents experienced login and routing failures. Service availability was not fully restored until 1:20pm October 20 – with data-sync delays continuing into October 28.

The statement says AWS Security Token Service (STS), responsible for temporary access credentials, was impaired twice – initially by DynamoDB connectivity failures and later by NLB health-check issues. Other services such as IAM, Redshift Airflow and Outposts also experienced intermittent disruptions tied to the same underlying dependencies.

Root causes and next steps

AWS says it has already disabled the automation systems that triggered the DNS error and will implement additional safeguards before re-enabling them. Specifically, the company plans to fix the race condition in DynamoDB’s DNS management system and add new protections to prevent stale DNS plans from overwriting newer ones.

For EC2, AWS will expand its internal scale testing to cover DWFM recovery workflows and improve throttling mechanisms to handle future load spikes. NLB will receive new ‘velocity controls’ to limit the capacity a single load balancer can remove during failover events.

In the statement, AWS apologised for the widespread impact saying: “AWS knows how critical our services are to our customers, their applications and their businesses,” the company said. “We will do everything we can to learn from this event and use it to improve our availability even further.”

While AWS maintains one of the industry’s highest uptime records, it acknowledges the October 19–20 incident underscores the complex interdependencies of its cloud infrastructure. The US-EAST-1 region – the oldest and largest AWS region – supports countless enterprise workloads and often acts as a control plane for global operations.

In the statement, AWS acknowledges that even isolated issues, as this outage showed, can ripple across the cloud ecosystem.

Browse our latest issue

Intelligent CIO North America

View Magazine Archive