Resolved -
The database is fully recovered, and all systems are operational. We apologize for any inconvenience this might have caused. We are currently implementing new measures to ensure this won't happen again.
Jul 22, 04:38 CEST
Monitoring -
All systems are operational. We expect the database to fully recover in 4 hours (1 am UTC).
We are monitoring the platform's performance. We apologize for any inconvenience this might have caused.
Jul 21, 23:28 CEST
Update -
Actors are now fully operational, and the system is processing all the runs in time.
Jul 21, 22:52 CEST
Update -
We are continuing to work on a fix for this issue.
Jul 21, 22:34 CEST
Identified -
We have identified the issue and are working towards recovery. The system is operational, but the capacity is capped, resulting in Actor runs being buffered in the "ready" state and a large portion of them timing out.
Jul 21, 22:33 CEST
Investigating -
Our database cluster performance has further degraded, and the whole platform is facing a major outage.
We will keep informing you as the situation progresses.
Jul 21, 19:47 CEST
Monitoring -
We are now monitoring the situation. Worker capacity improved and is now able to handle starting Actor runs.
Jul 21, 19:11 CEST
Update -
We are continuing to work on a fix for this issue.
Jul 21, 19:09 CEST
Update -
Slow database operations continue to impact our systems. API performance has improved, but Actor runs are currently timing out due to reduced worker capacity.
Jul 21, 18:43 CEST
Identified -
We are currently investigating timing out operations on our API caused by slow database operations due to cluster changes.
Jul 21, 16:49 CEST