Published: October 23, 2025
35
213
1.2k

The detailed AWS incident report is out, and it’s worth a read - DNS records managed by 2 systems; a race condition led to regional record getting unset - EC2 lease establishment was borked as it depends on DynamoDB - fluctuating NLB health checks leading to EC2 DNS entry purges

Image in tweet by Cindy Sridharan
Image in tweet by Cindy Sridharan
Image in tweet by Cindy Sridharan
Image in tweet by Cindy Sridharan

@copyconstruct Where is it? Searched for phrases and so on, no luck

@copyconstruct @oieduardorabelo Very interesting, the summary only mentions the *public* endpoints. Am I to understand that none of those other affected AWS services use VPC endpoints? Also, why was the first Enactor “ unusually delayed“? This didn’t happen before in this severity, apparently.

@copyconstruct Always so much effort put into technical root cause and zero reflection on oeganizational root causes

@copyconstruct “DNS plan.. automation”? Was it AI? Did “it was DNS” just pass the torch to “it was AI?”

@copyconstruct They're really committing to the "it's always DNS" bit

@copyconstruct Do I skim correctly that dns depends on dynamodb and dynamodb depends on dns?

@copyconstruct For anyone needing help understanding like I did: https://x.com/i/grok/share/bpC...

@copyconstruct @grok Why hadnt this AWS problem occurred before? Whats new that deleted the DNS address?

@copyconstruct The way they wrote this is so confusing that it is no small wonder the system crashed.

@copyconstruct Ooof, the health check flapping 😭

@copyconstruct @grok Explain this to me using a post office analogy like Im a fifth grader

@copyconstruct fluctuating nlb health checks purging dns is a wild example of cleanup logic having too much power. yikes.

Image in tweet by Cindy Sridharan

@copyconstruct AWS and any big cloud environment is the definition of a fragile system. Why people flock to these systems and pay a premium for the privilege to guarantee their business is fragile vs. anti-fragile is beyond me. At a certain point, complexity becomes the single point of failure

@copyconstruct @cyb3rops To Err is human. If you want to really fuck it up, give it to Amazon.

@copyconstruct I was kinda surprised the lack of CAS on per-endpoint plan version or rejecting stale writes via 2PC or single-writer lease per endpoint like patterns. Definitely a painful one and kudos to AWS for being so transparent and detailed :hugops:

@copyconstruct It really was DNS

@copyconstruct if you're ever having a bad day, just remember that at least you're (probably) not "in a state of congestive collapse"

@copyconstruct A cloud-Chernoble

@copyconstruct @fatih @grok what do you know about the AWS incident, explain it to a junior engineer

@copyconstruct tl;dr: wouldn't happen if they sticked to bind.

@copyconstruct 2systems but both Amazon and One should be someone else

@copyconstruct It’s only due to poor @aws management. These problems have existed for years and they made no investment to avoid this failure.

@copyconstruct I need a video with diagrams. This is too much text to read with a cold 🤧

@copyconstruct So inept h1b and offshore resources. With a poorly designed infrastructure that couldn't troubleshoot in a reasonable amount of time. Hire experienced Americans and this crap wouldn't happen.

@copyconstruct Seems like a classic case of interdependencies gone wrong. Definitely highlights the fragility of complex systems. Curious about the organizational factors too-what have they learned to prevent similar issues?

@copyconstruct Thanks for sharing this. I'd been wondering what caused the DNS record to get wiped in the first place. The details of that race condition make for painful reading.

@copyconstruct Looks like a classic issue of bootstrapping a system and adding a dependency that sits a level above it. Happens all the time when knowledge is lost over time.

Share this thread

Read on Twitter

View original thread

Navigate thread

1/32