Today there was an issue with one of our GPU providers that caused it run at about 50% capacity. Requests above capacity return 503s, which we route to another, fallback provider. This fallback provider was not able to rise to the occasion and also returned 503s.
In the very near future we will be onboarding additional providers to avoid these kinds of issues.
Posted Aug 06, 2026 - 19:51 PDT
Monitoring
A burst of requests increased the failure rate. Both primary and secondary model providers failed to fulfill the requests at the time; we are root-causing the issue.
We are continuing to monitor.
Posted Aug 06, 2026 - 14:08 PDT
Update
We are actively investigating elevated error rate on LiveKit Inference (google/gemma-4-31b-it)
Posted Aug 06, 2026 - 12:51 PDT
Identified
Between approximately 18:45 and 19:15 UTC today, a subset of LiveKit Inference requests using the google/gemma-4-31b-it model returned errors. Affected agent sessions would have seen an LLM request error on those requests. Other models and other LiveKit services were not affected.
No action is needed on your part. If you continue to see errors, please reach out to support. We apologize for the disruption.