View Transcript
Episode Description
Welcome to episode 369 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and (eventually) Matt are in the studio this week to bring you all the latest news in AI and Cloud, including a new local zone in Vegas, a 20th birthday, and some OAuth news thanks to Cloudflare. There’s a lot to cover, so let’s get into it!
Titles we almost went with this week
- What Happens In Local Zones Stays Low-Latency
- When Git Push Comes to Scaling Shove
- Twenty Policies Walk Into a Role
- AWS Bets Big on Latency in Vegas Local Zone
- AWS Hits the Jackpot with New Local Zone
- Two Decades of Instances, Zero Midlife Crisis
- EC2 Turns 20, Still Refuses to Retire
- Happy Birthday EC2, Now With 1,200 Candles
- Lambda Finally Lets IAM Policies Multitask Like Adults
- Cloudflare’s OAuth Diet: Trimming the Permission Fat
- Hugging Face Squeezes Out a 13 Billion Dollar Valuation
- Bedrock Slashes GPT-5.6 Sol Prices, Wallets Rejoice
- GitHub’s Capacity Crisis Sparks Retry Storm Reckoning
A big thanks to this week’s sponsors:
We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info.
Follow Up01:45 The August 17 outage, and the work ahead
- Update on GitHub’s August outages: root cause analysis published for the August 17 incident, which lasted nearly 8 hours and followed an earlier August 6 Actions failure.
- Root cause identified as a capacity failure, not a code or configuration change: a critical infrastructure component in the Central US data center failed to scale at a new traffic peak, triggering authentication failures and cascading disruption across services including Copilot, which was prolonged by a client-side retry loop.
- Since April, GitHub has added over 3 million CPU cores and 120 petabytes of storage, and accelerated Azure migration; Azure now handles approximately 58 percent of platform load and half of Git operations, up from 12 percent in May.
- Monthly commit volume has roughly doubled since April, from 1.4 billion to 2.9 billion, underscoring the scaling pressure behind both incidents and explaining, though not excusing, per GitHub, the repeated failures.
- Concrete remediation steps include consistent retry limits and budgets across service-to-service calls to prevent retry storms, a review of lower-priority CPU and memory alerts, and continued work isolating critical systems to reduce shared dependencies and blast radius.
03:07 Justin – “It felt a little ‘woe is me, capacity is a problem,’ but it feels like more of the same lip service from them… maybe we need to rethink some core fundamentals of how Git works. Git was designed for humans… around human speed and human scale. ”
General News14:03