Between August 26 and September 3, 2026, the vendors that AI automations depend on had a bad nine days. Across the five model and voice providers we track, seven incidents in that window carry the vendor’s own severity label of major, spread across three of them — Anthropic, OpenAI and Telnyx. Another 17 at lower severity bring in Deepgram and ElevenLabs, for 24 incidents across five vendors in all. Every one is published by the vendor itself — Anthropic, OpenAI, Telnyx, Deepgram and ElevenLabs — and we re-fetched each status page rather than working from coverage.
This is not a post about AI being unreliable, and nobody should finish it frightened of the technology. Outages are normal. Every platform your business runs on has them. The useful question is narrower: what does your automation do when its model provider returns errors for hours? For most mid-market automations we see, the answer is nothing. They fail quietly, and nobody finds out until someone notices the work did not happen.
Outages are normal. An unrouted dependency is the defect.
The same material, narrated, is on YouTube — the incident table, the GitHub quote and the dashboard mismatch, in the order they appear below.
The nine days, in the vendors’ own words
Severity is the vendor’s label, not ours. The seven marked major, with UTC start times and reported durations:
- 2026-08-28 17:22, 2h59m — Anthropic, “Elevated errors on Claude Code and Claude Cowork”
- 2026-08-31 15:04, 5h23m — OpenAI, “ChatGPT Work seeing elevated errors and latency”
- 2026-09-01 06:04, 1h02m — Telnyx, “Cloud Storage Partially Unavailable in All Regions”
- 2026-09-02 21:17, 0h26m — Anthropic, “Elevated errors for Claude Sonnet 5”
- 2026-09-03 00:04, 0h06m — OpenAI, “ChatGPT Work Mode High Error Rates”
- 2026-09-03 12:37, 0h18m — Anthropic, “Elevated errors for Claude Sonnet 5”
- 2026-09-03 13:26, 2h57m — Anthropic, “Elevated errors for multiple models”
We will not rank these vendors against each other, and you should not either. All five declared incidents in the same window, and three of them declared major ones. That is the finding.
The anchor incident recovered one model at a time
The last entry is the one worth reading closely, because it is the one coverage flattens into “Anthropic went down.” The affected components — claude.ai, the Claude API, Claude Code and Claude Cowork — moved together. The model layer underneath them did not.
At 13:26 UTC: “We are investigating elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.” At 13:50: “An exhaustive list of affected models: Mythos/Fable 5.1, Mythos/Fable 5, Opus 5, Opus 4.8, Opus 4.6.” At 15:25 the list had shrunk: “The only affected models right now are Opus 4.8 and Opus 5. The rest of the models have recovered to baseline error rate.”
The incident degraded and recovered unevenly, model by model. Sonnet 5 was not on the affected list at all — it had already had its own separate 18-minute outage that ended less than an hour before this one opened. And Mythos/Fable 5.1, which had shipped days earlier, was in a major incident within days of general availability.
Read those together and you have the lesson of the nine days. Model choice was not a mitigation. Model routing was. No model picked in advance would have carried you through. There were only windows in which some models worked and others did not, and what separated riding it out from stopping work was whether anything could move a request.
What we will not tell you is why it happened. Anthropic said it had identified the cause partway through and never published what it was. No post-incident review exists, and every update says only “elevated errors.”
GitHub published our argument for us
In the same nine days, GitHub declared three Copilot incidents it attributed not to itself but to the layer beneath it. The first, on August 27, carried GitHub’s CRITICAL severity:
“We are experiencing degraded availability for the Kimi K3 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting ‘Auto’ to continue using Copilot.”
That last sentence is the argument of this post, written by GitHub rather than by us. Choosing another model, or selecting Auto, is fallback routing — and ‘Auto’ is that capability shipped as a product feature.
The other two say it less quotably. On August 31, from 08:37 to 09:41 UTC, GitHub reported “Elevated rate of errors for OpenAI models provided by Copilot,” naming gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5 and the gpt-5.6 family. At 09:48: “One of our model providers has confirmed an incident on their end.” On September 3, Grok 4.6 and 4.5 degraded, again “due to an issue with an upstream model provider.”
The morning the status page was green
Now try to match that August 31 GitHub window to an OpenAI incident. When we checked OpenAI’s public status history, on the evening of September 3, 2026 — 01:33 UTC on September 4 — it returned 25 incidents spanning July 31 to September 3. Exactly two are dated August 31: one at 15:04:53Z, major, “ChatGPT Work seeing elevated errors and latency,” and one at 22:27:56Z, minor, “Elevated latency in the Responses API.” Nothing covers the 08:37–09:41 hours GitHub reported.
The window reaching back to July 31 is what makes that worth saying. August 31 sits well inside the returned history, not at its edge, which rules out the obvious rebuttal that an entry aged out of a rolling window. There was none to age out.
Now the caveat, and it is load-bearing. GitHub may route GPT models through Azure OpenAI rather than OpenAI directly, in which case the incident belonged to a different operator and OpenAI omitted nothing. From outside we cannot tell which reading is right, and we are not saying OpenAI concealed or failed to acknowledge anything. The claim we can support is narrow: no corresponding entry was visible when we checked.
What survives either reading is what matters to you. A customer refreshing status.openai.com that morning saw no incident posted while their GPT calls were failing. The status page was not the instrument that would have told them.
The failure that costs money is the one nobody sees
None of the incidents above is the expensive one. This is.
Zapier opened an incident on September 1 at 12:03 UTC titled “Freshservice - Webhook delivery disruption.” At September 3, 21:39 UTC it was still marked investigating, the latest update reading: “We’re still investigating the disruption with Freshdesk.” The title says Freshservice, the update says Freshdesk; we quote both as published.
Consider what a webhook that does not fire looks like from inside a business. No error message. No page turns red. Nothing lands in an inbox. The downstream steps never run, so there is no failed run to look at — only an absence, and absences raise no alarms. That incident had been open for more than two days, and was still open when we looked.
Compare the other Zapier item in the window, “Intermittent instant trigger failures” from August 26, which resolved with a line every operator wants to see: “We have successfully replayed the affected triggers.” That is recovery when the system knows what it missed. Silence is what it looks like when nothing counted.
The dashboard that disagrees with its own dependencies
One more, found only by looking. Vapi, a voice-AI platform, publishes a provider-health panel for the services it depends on. Read in a browser on the evening of September 3, 2026, over its “90 days ago to Today” range, it reports OpenAI 99.993%, Anthropic 100%, Google Gemini 100%, Deepgram 100%, ElevenLabs 100%, Cartesia 99.988%, Daily.co 100%, Gladia 100% and Soniox 100%.
Inside that window sits a stretch we can check directly. Anthropic’s own status API reaches back 45 days, July 21 to September 3, entirely within the 90 days the panel covers. Across those 45 days Anthropic declares 17 incidents at major or critical severity, three of them critical. The panel reports 100%. We cannot see the earlier half of the panel’s window and claim nothing about it.
Vapi is not lying, and this is not an accusation. Those tiles almost certainly measure Vapi’s own probes of the endpoints it calls, a different measurement from a provider’s declared incident state — a probe can succeed while a specific model or workload is erroring. The mismatch is the finding: the dependency dashboard you would think to look at does not reflect your dependency’s declared outages. The question is not whether to trust Vapi. It is what your own dashboards count.
Update — September 2026: the next day
This post carries a September 3 date and went live at 04:48 UTC on September 4. In the eleven hours after it did, three more incidents opened across two of the same providers. None of them is in the count above, which stays bounded to August 26 through September 3.
OpenAI opened an incident at 07:00:26 UTC affecting users in the APAC region across “ChatGPT, Work, image generation, file upload, Voice, and Codex Cloud.” It applied a mitigation at 09:47:58 and closed the incident at 10:46:54 UTC with “All impacted services have now fully recovered.” Three hours forty-six minutes, across chat, voice and a coding agent at once.
ElevenLabs then opened two separate speech-to-text incidents two hours and fourteen minutes apart, in two different data residencies. The first, at 11:41:29 UTC, was “a latency increase on the EU Residency server for Speech-to-Text models, affecting a small subset of requests”; it resolved at 12:12:17. The second, at 13:55:26 UTC, was the same symptom on “the Global Residency server for Speech-to-Text and Dubbing models.” At 14:34:50 ElevenLabs wrote “The issue has been fixed, the error rate should be zero. We are currently monitoring the situation.” Read again at 15:50 UTC on September 4, that incident was still open at monitoring rather than resolved.
That pair is the part worth keeping. Data residency is chosen for a compliance reason — where the audio is processed, and under whose law. It buys nothing in availability terms. Both residencies degraded the same morning, independently, with the same symptom on the same class of model.
And the counterweight, which belongs here for the same reason the rest does: Anthropic opened nothing at all in those fourteen hours. Its status API still returned the September 3 multi-model incident as the newest entry. Telnyx and Deepgram were clean as well, newest incidents on both dated September 2. The provider that had the worst forty-eight hours in the record above had an uneventful next day. This is not a league table, and one day is not a sample.
One more, outside the model layer entirely. Cloudflare opened “Cache Purging Errors” at 14:51:45 UTC — purge calls through the dashboard or API returning “unable to purge. internal error” — and it was still at monitoring when we read it at 15:50. A purge API that quietly refuses is exactly the shape of failure this post is about: nothing errors in your application, and a stale configuration simply stays served.
Update — September 5, 2026: both open incidents closed, and the “fix” was not the end
Two of the incidents above were still open when we published, and we said so. Both have now closed, and the way they closed is more useful than the fact that they did.
The ElevenLabs Global Residency speech-to-text incident, opened 13:55:26 UTC, resolved at 21:01:16 UTC. That is seven hours and six minutes. But look at where the time went. At 14:16:58 ElevenLabs had “identified the root cause” and was rolling out a fix. At 14:34:50 — thirty-nine minutes in — it wrote “The issue has been fixed, the error rate should be zero. We are currently monitoring the situation.” The incident then stayed open for another six hours and twenty-six minutes, closing with “After completing the monitoring stage, we have confirmed that STT latency has returned to baseline stability.”
Cloudflare’s cache purging incident ran the same shape on a shorter clock: opened 14:51:45, “a fix has been implemented and we are monitoring the results” at 15:31:44, resolved at 16:58:46. Forty minutes to a fix, then another hour and twenty-seven minutes before the incident closed.
Here is why that matters more than the durations. If you are watching a status page to decide whether it is safe to run the batch, ship the release, or stop telling customers there is a problem, the message you are waiting for is the one that arrives early. “The issue has been fixed, the error rate should be zero” reads like an all-clear. In both of these it was the beginning of the longest phase of the incident, not the end of it. The vendors were not being evasive — monitoring is exactly the right thing to do before declaring a system stable, and both said plainly that was what they were doing. The problem is on the reading end.
There is one more thing worth naming, because it undercuts a metric people quote. Both of these were labelled minor for their whole lifespan. A seven-hour speech-to-text degradation carries the same severity tag as a thirty-one minute one. If your vendor review counts major incidents, or your contract defines credits against a severity label the vendor assigns, that is the number you are relying on — and it does not describe what your customers experienced on a phone line that afternoon.
The operational version of this is short. Do not treat “fix implemented” as resolved; treat resolved as resolved, and if you cannot wait that long, verify against your own telemetry rather than the vendor’s narrative. And do not let a severity label stand in for impact — measure the impact where it lands, which is in your application, not on their dashboard.
This update, narrated — separate from the longer video above, which covers the original nine-day window:
The overlap we did not report, and what it does to our own advice
Re-reading September 3 against the vendors’ own APIs on September 5 turned up something the account above missed. It lands on our conclusion rather than on a vendor.
OpenAI opened “Elevated errors across ChatGPT and Codex” at 14:58:23 UTC on September 3, wrote “we have applied the mitigation and are monitoring the recovery” at 15:50:41, and resolved it at 16:55:49. Anthropic’s multi-model incident — the anchor incident taken apart above — was open from 13:26:04 until 16:23:12, with impact ending, in its own words, “as of 9:16 PT / 16:16 UTC.”
Lay those over each other. For one hour and twenty-five minutes both vendors were carrying open, declared incidents at the same time. Measured the stricter way — OpenAI’s incident against Anthropic’s stated end of impact rather than its resolution — the overlap is still one hour and eighteen minutes.
That window is the awkward one, because of what we concluded above: model choice was not a mitigation, model routing was. That holds across the nine days. It does not hold here. For most of two hours on September 3, both of the destinations a routing rule would pick between were degraded at once. Routing remains the right design and we are not withdrawing it. But a fallback that only fails over between these two providers has a window where it fails over into the same weather, and the options that survived that window were the other two on the list below: queue it, or hand it to a person.
The severity labels belong side by side as well. Anthropic carried its incident as major. OpenAI carried the overlapping one as minor. Same afternoon, same class of failure — elevated error rates on model requests — two different tags, set by two vendors against two internal scales that were never meant to agree. The section above argues a severity label does not describe your impact. Read across vendors, it does not describe the same thing twice.
They do not even publish the same shape of fact. Anthropic stated when impact ended. OpenAI’s record carries an opened time, a mitigation time and a resolved time, and no statement of when impact ended at all. Two durations from two vendors are not the same measurement, and a vendor review that lines them up in one column is comparing things that were never defined against each other.
And the gap between “fix applied” and “resolved” — the thing the previous section warns about — is not one number. Inside this single window, Anthropic ran about sixteen minutes from its monitoring update to resolution. OpenAI ran an hour and five. The ElevenLabs incident in the same forty-eight hours ran six hours and twenty-six minutes. That sharpens the earlier point rather than repeating it: the early message is not a lie, and the tail behind it is not a quantity you can plan against. In one week it ranged from a quarter of an hour to most of a working day.
One correction to our own bookkeeping, which the same evidence forces. The count above — 24 incidents across five vendors, seven of them major, August 26 to September 3 — was collected on September 3, and this OpenAI incident opened at 14:58 UTC that day. It is inside the window and it is not inside the collection. We are leaving the number where it is rather than quietly restating it, because the honest reading is the one this entire post is about: a status API is a rolling window, and a count drawn from it is a reading taken at a moment, not a total. Treat 24 as a floor. Any count of incidents, including ours, is bounded by when somebody looked.
What to do about it on Monday
None of this calls for a project. It calls for an inventory and a decision per line.
- List the automations that call a model provider. Not the vendors — the automations. It is the same inventory that decides what an integration costs you after it ships, and most teams have never written it down. Most operators can name their tools and cannot name which workflows stop when one degrades.
- Decide, per automation, what the right behavior on error is. Somebody has to own that call — an agent needs a manager, not just an API key. Four honest options, not interchangeable: retry, route to a second provider, queue the work for later, or fail loudly to a human. A low-stakes summarizer can queue. A customer-facing voice agent cannot.
- Assume degradation is partial. Some models failed while others were fine, and the list changed while the incident ran. A design that assumes a provider is up or down does not survive that.
- Make sure something fails loudly. Do this first. The Zapier case is the reason: the failures that cost real money produce no error anyone sees.
- Do not treat a status page as monitoring. It is a vendor’s account of its own service, on its own schedule, sometimes for a layer you do not call. Your error rates are the instrument.
Nine days of status pages cannot tell you how reliable AI is; the sample is too short and the question is the wrong one. What they show is what normal looks like: providers degrade in parts, recover unevenly, and do not always explain themselves. That is the operating environment, not an aberration in it. The automations that keep working through it are built as though that were true.
Related reading: stop asking which AI model is best — the question this post turns into a routing decision, and what an API integration costs you after it ships.
Related service: AI integration — building the automations, and the routes they fall back to when a provider degrades.