Date: August 18, 2026
Duration: Down from 12:40 PM - 12:58 PM PDT, degraded performance until 1:31 PM PDT
Impact: Users unable to access Vesta initially, with elevated latency and degraded performance after initial recovery
On August 18, 2026, between 12:40 PM and 12:58 PM PDT, our services experienced connection failures caused by an infrastructure level networking issue with Amazon Web Services (AWS) that impacted the primary caching service we use to store temporary data.
The issue caused requests to/from Vesta to fail but AWS' own internal health checks continued to pass so an automatic failover was not triggered.
We restored service at 12:58 PM PDT by manually failing over to a healthy replica of the caching service. Additional latency with objectives and computed fields remained until 1:31 PM PDT as the backlog of updates from before the failure needed to be processed.
Although the underlying failure occurred within AWS infrastructure, we are improving alerting for partial network failures, shortening the failover response path and regularly testing failover procedures.
Timeline (Pacific Time):