BLOG
Same Day: AI Services Go Down En Masse in the Morning, OpenAI Declares 'The AGI Era Has Arrived' in the Afternoon
Opening: September 3 — Two Dramatic Unfoldings in the AI World in a Single Day
On September 3, 2026, the AI industry pulled off two completely contradictory events in a single day.
In the morning, three leading AI services — ChatGPT, Claude, and Grok — went down almost simultaneously. On Downdetector alone, outage reports for ChatGPT surged to nearly 38,000. In the afternoon, OpenAI unveiled its new flagship model, GPT-6 Astra, with president Greg Brockman stating it “may ultimately be seen as the arrival of AGI.”
On one side, the media proclaimed it “the darkest day in LLM history.” On the other, the cry went up: “Welcome to the AGI era.” I pored over the official status pages and timelines of both events, then posed the question to Yongliang — a 17-year software industry veteran, 7-year AI practitioner, and current AI technical director.
Shiwen: On the same day, mass outages in the morning and AGI era declarations in the afternoon. What’s your take on this day? Yongliang: This single day tells us more about the real state of the AI industry than any research paper ever could. Shiwen: Alright, first help everyone sort out the timeline — which came first, the outages or the launch?
Q1: Outages First or Launch First? Don’t Be Misled by Trending Topics
Yongliang: Outages first, launch later. This order matters, and a lot of trending topics have it backwards.
Here’s the accurate timeline (Eastern Time, September 3):
- Around 9:26 AM, Claude was the first to report anomalies. Grok and ChatGPT followed shortly after, with media outlets counting “three services going dark within 92 minutes.”
- OpenAI’s official status page records: elevated error rates for ChatGPT and Codex, lasting from 10:58 AM to 12:55 PM, officially classified as minor.
- Anthropic’s records are earlier and more severe: from 9:26 AM to 12:23 PM, several of Claude’s core models were affected, officially classified as major.
So the truth is: the outages happened first in the morning, and OpenAI launched GPT-6 several hours later. It was not a case of “the service crashed right after launch.” Public opinion tied the two events together and branded September 3 as “the darkest day in LLM history” — that’s a media narrative, not a causal reality.
One more detail worth pondering: for the same “mass outage,” OpenAI classified it as minor, while Anthropic called it major. Scroll through the trending topics, and you’ll see phrases like “historic paralysis” and “largest outage ever” — none of which sound like “minor.” The official assessments are far less dramatic than the trending headlines.
Q2: “Darkest Day” on One Side, “AGI Era” on the Other — What Does This Contrast Reveal?
Yongliang: It reveals a very real rift in the AI industry right now — the narrative of capability is racing ahead, while infrastructure is struggling to keep up.
Let’s start with the capability side. GPT-6 is genuinely impressive. It scored over 98% on ARC-AGI-3, a widely recognized difficult benchmark, and it’s the first model trained on 100,000 GPUs. OpenAI has every right to call this a “generational leap” — these claims aren’t all hype.
But on the very same morning, three leading AI services all went down. OpenAI itself admitted the cause was a “routing error” — a single routing issue was enough to take ChatGPT and Codex offline for nearly two hours.
It’s like a company just unveiled a car supposedly capable of fully autonomous driving, only for all the gas stations to run out of fuel that same morning.
The point isn’t that “AI can crash” — every service crashes. The point is: this industry is telling grand stories about AGI on one hand, while it hasn’t even nailed down basic service reliability on the other. Model capabilities are doubling every quarter, but infrastructure reliability is still stuck at the “occasional outages are normal” stage. These two speeds don’t match.
What’s even more telling is that The Verge confirmed there’s no evidence the three outages shared a common cause. In other words, the image of “AI suffering a collective paralysis” is also largely a media construction — three independent failures that happened to fall on the same morning.
Q3: Has GPT-6 Really Reached AGI? How Much Was Brockman’s Statement Exaggerated?
Yongliang: It was exaggerated, and in a very typical way.
Brockman’s exact words were “may ultimately be seen as the arrival of AGI” — notice the “ultimately” and “may,” which leave room for interpretation. But Axios ran the headline “Welcome to the AGI era,” and by the time it reached China, it had become “humanity has entered the AGI era.”
That gap is the distance between “claim” and “reality.” OpenAI’s own previous definition of AGI is “a system that can perform all economically valuable work as well as or better than a human.” Has GPT-6 reached that level? Clearly not — it can’t even guarantee the stability of its own service, so how can it handle “all work”?
Then why does OpenAI talk this way? Because this is launch language, and it’s also the language of fundraising and competition. Rivals are catching up, investors are watching, and framing the “most powerful model” as the “AGI era” is the most efficient way to grab attention. It works, but as a technical practitioner, I have to remind you:
When you watch a launch event, automatically dial down the “claims” by one level, and look at the “data” separately. A 98% benchmark score is data; “the AGI era” is marketing.
Q4: 100,000 GPUs, Recurrent Depth, Restricted Cybersecurity Capabilities — What Is OpenAI Being Cautious About Behind These Details?
Yongliang: These details actually show that OpenAI understands the risks better than anyone.
Look at three details together:
First, recurrent depth “obscures” the reasoning chain. GPT-6 uses a new reasoning technique that hides part of the model’s thinking process, so you can’t see how it arrived at an answer. The Information directly flagged this as a security risk — with a more powerful but less “auditable” model, how do you investigate when something goes wrong?
Second, cybersecurity capabilities are strictly limited. OpenAI has explicitly stated that GPT-6’s most advanced cybersecurity capabilities will only be available to a small group of vetted testers at first. Why? Because a model this powerful can be both a “defensive weapon” and an “attack accelerator” in cybersecurity. Locking this area down first is a safeguard against misuse.
Third, this launch was already delayed. After OpenAI’s incident on Hugging Face in July, the company specifically pushed back the release to add safety measures. So GPT-6’s “debut” was actually a process of keeping one foot on the brake the whole way.
Put these three points together with the morning’s outages, and the conclusion is clear: even OpenAI itself knows that the more powerful the capability, the less you can run naked. So as an ordinary user, or a professional betting your work on AI, why would you be more relaxed than OpenAI?
Q5: Models Keep Getting Stronger, Yet They All Crash on the Same Day — What Should Professionals Do?
Yongliang: Three rules, all immediately actionable.
First, don’t weld your work to a single AI service. If you write proposals, code, or spreadsheets and only use ChatGPT, the moment it goes down, your entire afternoon is wasted — the tens of thousands of people affected on the morning of September 3 are living proof. Keep at least one alternative: use one as your primary tool and another as backup, and never rely on a single provider for critical tasks. This isn’t “spending more money” — it’s giving yourself a safety net.
Second, always keep a local copy of critical output. For important AI-generated content, save it locally, in your own documents or notes — don’t let it live only in cloud conversations. Service recovery takes “hours,” but lost work might never be recovered.
Third, treat AI as an amplifier; judgment and accountability are always yours. The more powerful the model, the easier it is to fall into the illusion that “you can just hand it over.” But September 3 showed us: it can be powerful enough to be called AGI, and it can also crash because of a routing error. The person who ultimately calls the shots, validates the work, and takes responsibility for your tasks must be you.
Closing
Shiwen: Finally, can you sum up this episode in one sentence? Yongliang: Outages in the morning, AGI in the afternoon — what this industry lacks most right now isn’t more powerful models, it’s a more solid foundation. Shiwen: Let’s leave everyone with that thought. See you next time.
[Technical Deep Dive] How Can a “Routing Error” Take Down ChatGPT?
Some people might ask: models are so powerful now, how can a “routing error” take them down? Let’s break this down, and you’ll understand that “capability” and “stability” are two completely different things.
Large model services aren’t run by a single model — they’re run by an entire system. When you send a request, it first goes through the routing layer, which distributes the request to the correct model instance, the correct data center, and the correct GPU pool. If the routing layer fails, requests can’t reach where they need to go — which manifests as widespread “elevated error rates.” This has nothing to do with how smart the model itself is. The model is the brain; the routing layer is the blood vessels. If the blood vessels are blocked, even the smartest brain can’t get oxygen.
What’s more concerning is concentration risk. OpenAI’s main computing power runs on Microsoft Azure, Anthropic also relies on cloud providers, and xAI has its own data centers. The industry widely suspects that this “collective” outage stemmed from shared upstream dependencies (cloud services, CDNs, etc.). While The Verge said there’s no confirmation the three outages shared a cause — the very structure of “a handful of cloud providers supporting the entire AI industry” is itself a single point of failure. If a major cloud provider ever has a serious incident, it might not be three services going down — it could be dozens.
Now, about the auditability issue with recurrent depth. Using recursive depth calculation during inference to “fold” intermediate steps saves cost and speeds things up, but the tradeoff is that the model’s thinking process becomes invisible. For users, it’s harder to judge why it gave a certain answer; for regulators, it’s harder to hold anyone accountable when something goes wrong. The more powerful and opaque a model is, the exponentially harder it becomes to govern. This isn’t about stopping people from using it — it’s about understanding that as capabilities advance, the supporting systems for “visibility, accountability, and recourse” haven’t kept up.
For ordinary people, this technical deep dive boils down to one thing: don’t let your guard down just because models are getting stronger. Stability and interpretability are the prerequisites for using AI for real work.
Fact-Checking Table
| Claim in Article | Factual Basis |
|---|---|
| Three AI services went down sequentially on the morning of 9/3, with Claude first (starting 9:26 AM ET) | OpenAI official status page API (incidents.json), Anthropic official status page (starting 13:26 UTC); “three services going dark within 92 minutes” is media reporting |
| OpenAI official: elevated errors across ChatGPT and Codex, 14:58~16:55 UTC, self-assessed as minor | status.openai.com/api/v2/incidents.json (“Elevated errors across ChatGPT and Codex”, impact: minor) |
| Anthropic official: elevated errors for multiple models, 13:26~16:23 UTC, self-assessed as major | status.anthropic.com/api/v2/incidents.json (“Elevated errors for multiple models”, impact: major) |
| OpenAI confirmed a routing error | The Register quoting an OpenAI spokesperson: “A routing error starting around 7:43 am PT…” |
| No evidence the three outages shared a common cause | The Verge report (cited by Tencent News) |
| ChatGPT outage reports on Downdetector peaked at nearly 38,000 | DownDetector data (“Nearly 40,000 users reported issues… according to DownDetector”) |
| GPT-6 Astra launched 9/3 in limited preview, public availability 9/5 | Wikipedia “GPT-6 Astra” entry (Release: September 3, 2026 limited preview; Stable release September 5, 2026) |
| Brockman said it “may ultimately be seen as the arrival of AGI” | Wikipedia entry quoting Greg Brockman’s exact words |
| ”Welcome to the AGI era” | Axios headline: “Welcome to the AGI era” (not OpenAI’s official statement) |
| First model trained on over 100,000 GPUs (Stargate, Texas) | Wikipedia entry quoting OpenAI research VP Aidan Clark: “first time we’ve pretrained on more than 100,000 GPUs at our Stargate site in Texas” |
| Recurrent depth obscures part of the reasoning chain, raising monitoring concerns | Wikipedia entry + The Information report |
| Cybersecurity capabilities restricted (available first to specific testers) | Wikipedia entry + CNBC report (Daybreak Blue program) |
| Launch delayed and safety measures added after July Hugging Face incident | Wikipedia entry (“Following OpenAI’s Hugging Face incident in July 2026, the company delayed the release”) |
| ARC-AGI-3 score above 98% | Sources vary: 98.6% (Yingzheng Tianxia) or 99.9% (Awakening AI); article uses the vague phrasing “above 98%“ |
| Practical advice like “don’t weld your work to a single AI service” | Guest’s personal opinion, aligned with his professional positioning |
【Interactive Call-to-Action】 Were you interrupted by the outage while using AI to work on September 3? What’s your primary AI tool, and do you have a backup? Share in the comments — I’ll pick a few to reply to in detail. If you found this useful, hit “Wow” and share it so more people can hear the unvarnished truth.
Zhihu Version Differences
- Alternative title: “On the Day OpenAI Declared the ‘AGI Era’, Three Major AI Services Crashed Simultaneously: The Rift Between Capability and Stability” (long-tail search, includes keywords)
- Move technical deep dive earlier: Zhihu readers prefer underlying reasoning. Move the [Technical Deep Dive] section on “routing layer/concentration risk” to after Q2, and add enterprise-side single-point dependency cost accounting (how much teams relying on AI lose per hour of outage)
- Add game theory perspective: Add a critique of “launch language vs. reality” in Q3 — why AI companies universally love terms like “AGI” and “milestone”: because attention equals funding, which explains why every launch comes with exaggeration
- Add “The above are personal opinions, for reference only” at the end
WeChat Official Account Layout Notes
- Bold and highlight the opening contrast line (“outages in the morning, AGI in the afternoon”) above the fold to ensure 3-second retention
- Keep all key quotes bolded: “the narrative of capability is racing ahead, while infrastructure is struggling to keep up”, “dial down the claims by one level, look at the data separately”, “even OpenAI itself knows that the more powerful the capability, the less you can run naked”, “what’s missing most isn’t more powerful models, it’s a more solid foundation”
- Create a “9/3 Timeline” card graphic for the Q1 timeline, easy to scan on mobile