The detailed AWS incident report is out, and it’s worth a read - DNS records managed by 2 systems; a race condition led to regional record getting unset - EC2 lease establishment was borked as it depends on DynamoDB - fluctuating NLB health checks leading to EC2 DNS entry purges
@copyconstruct Where is it? Searched for phrases and so on, no luck
@copyconstruct @oieduardorabelo Very interesting, the summary only mentions the *public* endpoints. Am I to understand that none of those other affected AWS services use VPC endpoints? Also, why was the first Enactor “ unusually delayed“? This didn’t happen before in this severity, apparently.
@copyconstruct Always so much effort put into technical root cause and zero reflection on oeganizational root causes
@copyconstruct “DNS plan.. automation”? Was it AI? Did “it was DNS” just pass the torch to “it was AI?”
@copyconstruct They're really committing to the "it's always DNS" bit
@copyconstruct Do I skim correctly that dns depends on dynamodb and dynamodb depends on dns?
@copyconstruct For anyone needing help understanding like I did: https://x.com/i/grok/share/bpC...
@copyconstruct @grok Why hadnt this AWS problem occurred before? Whats new that deleted the DNS address?
@copyconstruct The way they wrote this is so confusing that it is no small wonder the system crashed.
@copyconstruct Ooof, the health check flapping 😭
@copyconstruct @grok Explain this to me using a post office analogy like Im a fifth grader
@copyconstruct fluctuating nlb health checks purging dns is a wild example of cleanup logic having too much power. yikes.
@copyconstruct AWS and any big cloud environment is the definition of a fragile system. Why people flock to these systems and pay a premium for the privilege to guarantee their business is fragile vs. anti-fragile is beyond me. At a certain point, complexity becomes the single point of failure
@copyconstruct @cyb3rops To Err is human. If you want to really fuck it up, give it to Amazon.
@copyconstruct I was kinda surprised the lack of CAS on per-endpoint plan version or rejecting stale writes via 2PC or single-writer lease per endpoint like patterns. Definitely a painful one and kudos to AWS for being so transparent and detailed :hugops:
@copyconstruct It really was DNS
@copyconstruct if you're ever having a bad day, just remember that at least you're (probably) not "in a state of congestive collapse"
@copyconstruct A cloud-Chernoble
@copyconstruct @fatih @grok what do you know about the AWS incident, explain it to a junior engineer
@copyconstruct tl;dr: wouldn't happen if they sticked to bind.
@copyconstruct @Securityblog lol split brain
@copyconstruct 2systems but both Amazon and One should be someone else
@copyconstruct It’s only due to poor @aws management. These problems have existed for years and they made no investment to avoid this failure.
@copyconstruct I need a video with diagrams. This is too much text to read with a cold 🤧
@copyconstruct So inept h1b and offshore resources. With a poorly designed infrastructure that couldn't troubleshoot in a reasonable amount of time. Hire experienced Americans and this crap wouldn't happen.
@copyconstruct Seems like a classic case of interdependencies gone wrong. Definitely highlights the fragility of complex systems. Curious about the organizational factors too-what have they learned to prevent similar issues?
@copyconstruct Thanks for sharing this. I'd been wondering what caused the DNS record to get wiped in the first place. The details of that race condition make for painful reading.
@copyconstruct Looks like a classic issue of bootstrapping a system and adding a dependency that sits a level above it. Happens all the time when knowledge is lost over time.





