Production outage 8/18

Incident Report for Vesta

Postmortem

Production outage 8/18

Date: August 18, 2026
Duration: Down from 12:40 PM - 12:58 PM PDT, degraded performance until 1:31 PM PDT
Impact: Users unable to access Vesta initially, with elevated latency and degraded performance after initial recovery

Summary

On August 18, 2026, between 12:40 PM and 12:58 PM PDT, our services experienced connection failures caused by an infrastructure level networking issue with Amazon Web Services (AWS) that impacted the primary caching service we use to store temporary data.

The issue caused requests to/from Vesta to fail but AWS' own internal health checks continued to pass so an automatic failover was not triggered.

We restored service at 12:58 PM PDT by manually failing over to a healthy replica of the caching service. Additional latency with objectives and computed fields remained until 1:31 PM PDT as the backlog of updates from before the failure needed to be processed.

Resolution and Next Steps

Although the underlying failure occurred within AWS infrastructure, we are improving alerting for partial network failures, shortening the failover response path and regularly testing failover procedures.

Timeline (Pacific Time):

  • 12:40 PM: Primary caching service stopped report key metrics on our dashboards
  • 12:43 PM: First connection failures appeared
  • 12:58 PM: Initiated failover to backup caching service
  • 1:02 PM: Access to Vesta is restored and most of platform recovers but objective/computation latency remains high due to backlog
  • 1:31 PM: All latency recovered
Posted Aug 19, 2026 - 10:46 PDT

Resolved

The system has fully recovered.
Posted Aug 18, 2026 - 14:10 PDT

Identified

We identified network connectivity issues affecting key infrastructure on our end. We failed over to a secondary option, and service has recovered, though some elevated latency may continue. We're still investigating the root cause of the original issue.
Posted Aug 18, 2026 - 13:16 PDT

Investigating

We have been alerted to an outage affecting our production environments. Our team is actively investigating the issue and working to identify the cause. We will provide additional updates as they become available.
Posted Aug 18, 2026 - 12:53 PDT
This incident affected: Production Environment (Core Platform).