On Oct. 20, between 06:52 and 09:22 UTC, all US Taktile customers were impacted by a major AWS outage in the US East (N. Virginia) region. Taktile APIs returned 5xx error codes. Our current understanding is that no other regions were affected.
We will conduct post-mortem analysis to understand how we could respond to this incident better.
Resolved
On Oct. 20, between 06:52 and 09:22 UTC, all US Taktile customers were impacted by a major AWS outage in the US East (N. Virginia) region. Taktile APIs returned 5xx error codes. Our current understanding is that no other regions were affected.
We will conduct post-mortem analysis to understand how we could respond to this incident better.
Monitoring
AWS is reporting recovery. We are observing the first signs of recovery on our end as well.
Monitoring
Incident still ongoing. AWS has identified the root cause, but services have not recovered yet.
AWS update:
[02:01 AM PDT] We have identified a potential root cause for error rates for the DynamoDB APIs in the US-EAST-1 Region. Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1. We are working on multiple parallel paths to accelerate recovery. This issue also affects other AWS Services in the US-EAST-1 Region. Global services or features that rely on US-EAST-1 endpoints such as IAM updates and DynamoDB Global tables may also be experiencing issues. During this time, customers may be unable to create or update Support Cases. We recommend customers continue to retry any failed requests. We will continue to provide updates as we have more information to share, or by 2:45 AM.
Monitoring
Since ~6:15am UTC, all US Taktile customers are impacted by an ongoing AWS regional outage. Taktile APIs return 5xx error codes.
Blast radius: all APIs in the us-east-1 AWS region. Other regions are not affected at this time. We lack clear visibility into the impact because the AWS console itself is having major issues.
Next steps:
We are tracking the AWS incident and service recovery, and will keep updating our status page as we have more to share.
Monitoring
Socure seems to be down because of the AWS incident too. It is likely that customers using Socure Connections might experience issues.
Identified
AWS has confirmed the incident, multiple other services are impacted as well. We are monitoring the incident and waiting for recovery/more information from AWS's end.
Investigating
There appears to be an ongoing AWS incident. Our on-call team is investigating.