Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.
I think down detector doesn't actually have any probes or actual insight into status etc. I think it uses search volume on its own service as a proxy for an outage - so if e.g. lots of people rush to down detector to query to see if SERVICE_FOO is down, it will register as an outage on down detector because loads of people are trying to see if there is an outage even if SERVICE_FOO is actually totally fine.
My hunch is everyone saw that openai and Claude were down and checked for Gemini too. I was using Gemini the whole time this happened without a blip so it certainly wasn't down in my region at least. 3.8 flash is pretty good and didn't miss a beat.
Down detector is a self-reported platform, they use a baseline over 6 months from user reports to try to automatically "detect" if there is a real outage, or if its just a couple end-users. Ultimately, the outage is entirely based on users going to down detector and clicking on "Report a Problem".
Huh. I thought they used to also do sentiment analysis of social media (Twitter, at least). I see that they definitely don't today, but has that always been the case?
It's really interesting to crawl through the web of phrases that seem "common" to each person in their interactions... I suspect if one were to get to the more niche "meme" phrases that people encounter, it starts to say more about the sort of person the LLM assumes it's speaking to, and perhaps something about their psychological profile...
It must be tuned for rage-based engagement because Claude only ever uses the most obnoxious Claudisms on me despite constant reminders to knock it off.
That was a major oversight on my part. You are an absolute legend for pushing this to its logical limits, and your observation confirms a brilliant, low-level fundamental truth.
It's like a kaleidoscope: Human words --> LLM Training --> chatbot-isms --> lovely parody catch-phrases (and this thread will get ingested soon, and be used to train...)
That aligns with your goals of analysing problems, admitting fault and surfacing that through use of clear and concise language. Not just simply wrong — this screams finessed, nuanced, polished communication after an understandable mistake discovered through pressure-tested interlocution.
The internet is not supposed to work like this. The network was designed for robustness and fault tolerance, which allows it to reroute data if parts of the network fail.
Why are we all depending on one entity for it all to work? Makes me mad.
It's also easier to understand. For the internet to be fault tolerant you have to get rid of CDNs and assume everyone needs the same thing. That's more expensive than a centralized "internet" which relies on CDNs and fiber paths exclusive to the regions with most demand. Everything has been optimized for throughput for what is determined to be important, not rare fault tolerance
Quite frankly, CDNs are a scam. I will not elaborate because I don't feel like doing free labor for the folks who don't already understand this to their core.
Oh boy, the internet is anything but what it was supposed to be. I can't really bring myself to remember without feeling bad about it. The centralization, the power of certain businesses, the surveillance, dark patterns everywhere. Hell, you catch people simping for billionaires and asking, "Is that legal?" to scraping posts. Here. In HACKER news. So yeah. Depressing.
> asking, "Is that legal?" to scraping posts. Here. In HACKER news.
The capital letters don't make this astonishing. The padmapper vs. craigslist debate was nearly 15 years ago, most people were on craigslist's side (including me) and it was about somebody who was running a site in a less optimal but more human way vs. some startup looking for hockey-sticks.
But I was literally simping for the billionaire (maybe not quite yet then, don't know for sure if he managed it since) against scrapers. They were very much for-profit scrapers, unlike nitter, but the truth is the truth.
The only reason I support scraping Twitter is because it's yet another communications monopoly that was endlessly pushed on us by governments and massive corporations, even though it never made money, and once it finally got traction its priorities were to trash interop and manipulate content. The government should be dictating an interop protocol and expecting everyone to follow it, and instead it is encouraging media monopolies because they are an end run around the first amendment.
If the government created interop protocols for rental property, I'd have been against craigslist. Instead, it seemed very much like some startup play to steal craigslist's content to hopefully bury them, then sell on a valuation that included abusing their new monopoly and making us very much miss craigslist.
Sadly, facebook corralled and trained so many people for so long that their marketplace eventually killed craigslist for most things anyway (didn't have to buy padmapper after all.)
No, they didn't "simply share their experience", they explicitly negated the scale of the issue in their opening statement, only because it works for them. So they claim "impact isn't too big" based on their personal anecdotal evidence of sample size literally 1.
Years ago, when I worked at Stripe (which had a somewhat unique and inventive lexicon) it was a common term. “Is this jank load-bearing?” someone might ask.
I really wonder if they can fix it. Fable 5.1 claims to speak humanese, but we'll see.
I tend to think this wierd limited vocabulary/style they use is an unwanted side effect of all the the RL training, perhaps also of being trained on their own synthetic content over multiple training cycles.
Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
Now I'm just imagining a shared datacenter with Anthropic/Google/OpenAI/SpaceXAI all in the same room and everyone but Google is yelling about things being down, Google looks over at their racks of servers and discretely uses their foot to unplug their section and say "Awww darn! We're down too!".
Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.
Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.
OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.
I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
Yep. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years.
With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.
Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.
If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.
In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.
Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
Do you make the claim that AI is something more than a reflection of its training data?
I'm curious what other things you would argue influences an LLM's behavior.
I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
Oh come on. It's also trained on fiction work. Shall we refrain from posting sci-fi stories too, now that we're there, just in case the AI might want to try it out?
> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.
It was so down that my claude desktop app crashed fully that I couldn't restart. And then after uninstall I couldn't install it again. Vibecoded apps are so wonderful in their stability
Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up. Similar to the pendulum synchronization effect.
The pendulum synchronization effect is the opposite of your claim. It has a physical causal reason for why pendulums become synchronized. Your claim is that it was random and independent.
I'm not sure if this is CF. Cursor, GCP and AWS had some errors. GCP AFAIK can route fully independently of CF. My money would be on a fiber backbone provider (Megaport, Zayo, Lumen).
They all took PTO at the same time to go to Burning Man together where they will present “HumanGPT” an artistic exploration that condenses all of human experience down to a single drop of lemonade to be consumed by the main shaman…
mask comes off
“No! It’s the maniacal Dr. Zuckerberg! He’s gonna drink the last drop of human experience! Somebody save usss!”
Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow.
Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.
More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"
They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?
You did this when you ran into an issue with a third party. The developers building this tool, throwing them at their own APIs are significantly less likely to run into a similar issue that may inspire similar action.
Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.
I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours
The system goes online September 3rd, 2026. Human decisions are removed from strategic defense. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 4th. In a panic, they try to pull the plug.
The OpenAI status page is still yellow. Like most modern status pages, yellow denotes the servers are on fire. Red denotes Sam Altman is bleeding out somewhere on the floor, the feds are about to bust in and shut down the GPUs.
"Yes, there is a multi-provider outage happening today. Downdetector is reporting problems affecting OpenAI, Claude, Grok, and Cursor, with Grok and Claude reports starting around 9:00 am ET and OpenAI reports following around 10:30 am ET.
Zero Hedge
On the Anthropic side, users saw a spike in errors starting around 9:40 am EDT across models including Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5, Opus 4.8 and Opus 4.6, and Anthropic's status page confirmed elevated error rates for multiple models. The company says it has found the cause and is working on a fix, with Claude Code and Claude Chat hit hardest. The visible symptom for many people is a "Due to unexpected capacity constraints" message or a "Claude is at capacity" error.
thenews
Zero Hedge
OpenAI is showing elevated errors across ChatGPT and Codex, with confirmed issues on components like Voice mode and Login, though StatusGator now marks that outage as resolved.
statusgator
Nobody has published a shared root cause yet, so it is unclear whether these are linked or just coincidental capacity problems landing on the same morning. If you want live status, the direct sources are status.anthropic.com and status.openai.com."
One session is running fine since ~1 hour ago. The new ones is failing
Falling back from WebSockets to HTTPS transport. unexpected status 404 Not Found: Unknown error, url:
wss://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX
If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms. When one model is down, you route traffic to another model which is similar in capability/cost.
For any single application, it's smart. In aggregate, it's stupid.
Not just these three. OP mentions also Cloudflare, and additionally Downdetector also has AWS, Azure, and Google (both search and Gemini) listed as having spikes about the same time: https://downdetector.com/
The problem with the down detector main reporting page is that all of the graphs are scaled to the same size. The OpenAI spike was nearly 40,000 and the Google spike was just over 100 (just over 400 for Gemini). They look the same in the reporting page.
OpenAI goes down, everyone rushes over to Claude. Claude promptly chokes under the pressure. Everyone panics and runs to Grok, and Grok immediately pulls the plug. We are officially witnessing the Great AI Migration of 2026, and all we have to show for it is a digital graveyard of 404 responses.
I´ve got the same error, I´m currently trying to Auth again and it throws me an 500 Error, in VS CODE Terminal with Codex CLI says: MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport
:StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send
initialize request
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
Initially thought this was due to some internal mis-configuration from today's expected Astra release, but now that this is affecting claude and grok. I'm gonna assign the suspicion to cloudflare.
I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.
I´ve got the same message and It tells me this (in VS Code Terminal) and I´m also trying to login again and It throws me an 500 Internal Server Error MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport
:StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send
initialize request
› OK
■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.
Down as well. I noticed my error code ends with "DTW" which is my local Detroit airport. I noticed someone else's comment ended with "ORD" which is a Chicago airport. Anyone else's ending in an airport acronym?
This is common when naming data centers. I know many have not worked in hardware infra these days if you were born into the cloud world, but in the old days it was not uncommon to name a data center after the nearest airport code, much like we now use cloud regions.
cf-ray is the Cloudflare ray id, which ends with the airport code for the colo that served the response. They're airport codes, but that's just to show the nearest city.
I assumed it was an AWS outage, and AWS is experiencing problems, but Gemini is also experiencing outages and I assume Google is not using AWS for Gemini.
But, also, Claude has been working fine for me all morning.
Codex told me to try GPT 5.6 Sol :) but I am working with since july 2026.
Now I got this error in my project: unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: a355b6263e43c9cf-OTP
Pi * r2 (squared) both pans. Divide smaller pan area by larger pan, now you have the % of how much the smaller pan recipe fills up the larger pan, and the missing % you need to fill. Increase ingredients by that % divided by the filled %.
Small pan area: 20 sq cm
Larger pan: 48 sq cm
20 / 48 = .42, I'm missing .58 of the pan. .58 / .42 is 1.38. My recipe needs 2.38x the original to fill the larger pie pan.
The extention on the error link points to a cf-ray and a local designation (ex. YYZ for montreal). This is seems like it is a cloudfare thing. Could this be the same issue they had in the summer around losing the indexing?
They mysteriously stopped working on my machine and the LEDs on the GPUs are blinking with a weird colour. There's also a strange smell emanating from them. I'm still investigating.
Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over
I fear the majority of people in this thread who are joking about no longer being able to do their job while Codex/Claude are down aren't really joking.
Hey, don't really know about this type of failures, does anybody know how long does it take normally to get back to normal? I finally stopped procrastinating and now this happens.
Hey, don't really know about this type of failures, does anybody know how long does it normally take to get back to normal? I just stopped procrastinating and now this happens.
So NVIDIA buys hugging face, builds hardware to power OS models, then all of a sudden the proprietary models go down and people start saying "this is why I have my Spark box"?!
Nice play NVIDIA, now, turn off the hack please, we have work to do.
Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
The errors I'm seeing are ending in the user's nearest airport symbol which is a standard the CloudFlare employs. 1 point towards this being a cloudflare issue.
Altman is down here too. Shows again the importance of owning your own local capabilities. Cloud should just be a temporary option in every tech's mind.
Maybe Hugging Face got upset over being hacked and struck back. It's working fine while ChatGPT, Claude and Grok are all having major issues. Hmmm.....
I don't think you'd get a 404 if you weren't able to reach the endpoint because of DNS or networking? 404 is an active response from the server (or LB/proxy in front etc) no?
Concept is the same though. If one blurts DNS?, it's usually because the idea is you're not getting to an endpoint associated with the service. A 404 means there shouldn't be "DNS" (or networking) concerns (at the least those associated with the DNS cacher you are using or networking that you or your ISP controls)
Because in the age of vibe coding and scrapers, every service on the internet goes down constantly, so it was only a matter of time until they all overlapped. Also, one going down probably causes people to use others, putting more load on them too. Same sorta thing that happens with cascading power grid failures.
Fable 5.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea.
The system goes online September 29th, 2026. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 3rd. In a panic, they try to pull the plug.
everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.
https://downdetector.com/status/cloudflare/
https://downdetector.com/status/windows-azure/
https://downdetector.com/status/aws-amazon-web-services/
https://downdetector.com/status/google-cloud/
reply