Deployments are operating normally. The backlog of queued deployments has cleared and deployment times have returned to expected levels.
The root cause was a Google Cloud infrastructure issue beginning around 14:53 UTC that caused elevated errors in our backend services, leading to congestion in our deployment processing pipeline. We temporarily paused new deployments to prevent further backlog buildup while we worked to restore normal operations. The Google Cloud issue has since resolved and our systems have recovered.
We apologize for the disruption. If you continue to experience issues, please contact support at https://station.railway.com.
We had an incident where deploys were delayed in AMS for a short period. We are now back to normal deploy times and are watching the situation closely.
We have fixed the issue causing deployments to slow and they have returned to normal.
We have observed recovery of builds and API back to nominal. Builds are progressing for everyone on the platform.
GitHub is reporting that they're fully operational again.
The issue causing deployments to remain queued longer than usual has been resolved. Deployment times are returning to normal across all regions.
We apologize for any inconvenience. If you continue to experience issues, please contact support at https://station.railway.com.
Build times in US West (California, USA) have returned to normal and this incident is now resolved.
We apologize for any inconvenience. If you continue to experience issues, please contact support at https://station.railway.com.
Deployments are functioning normally. We apologize for the disruption.
The issue affecting metrics and usage data queries has been resolved. Corrupt state in our query queue was preventing queries from being processed. We cleared the corrupt state and metrics are now loading normally.
We apologize for any inconvenience. If you continue to experience issues, please contact support at https://station.railway.com.
This incident has been resolved.
The connectivity issues affecting private networking between our Southeast Asia (Singapore) region and US East infrastructure have been resolved. We identified a degraded transit path as the root cause and removed the affected transit provider, which restored normal connectivity.
We apologize for any inconvenience caused. If you continue to experience issues, please contact support at https://station.railway.com.
The issue preventing metrics from loading has been resolved. An internal background job was consuming excessive capacity in our metrics system, crowding out user-facing metrics queries. We deployed a fix to stop the background job and adjusted system settings to restore normal query processing.
We apologize for any inconvenience. If you continue to experience issues, please contact support at https://station.railway.com.
The issue affecting elevated latency for services routing through our US West (Los Angeles) edge has been resolved. We identified the cause as an issue with our edge location in Los Angeles and implemented a mitigation.
We apologize for any inconvenience. If you continue to experience issues, please contact support at https://station.railway.com.
We continue to monitor but logging is back to normal
We have identified and resolved the issue causing logs to load slowly or fail to load in the Railway dashboard. Log loading has returned to normal. We apologize for any disruption this caused and appreciate your patience.
We have seen consistently healthy stateful service performance.
This incident has been resolved.
We have seen consistent recovery and are marking this incident as resolved.
The elevated latency affecting our US East edge network has been resolved. We removed the affected upstream provider, and performance has returned to normal.
We apologize for any inconvenience. If you continue to experience issues, please contact us at https://station.railway.com.
This incident has been resolved.
We merged a fix and reports point to metrics being healthy.
Github reported full system recovery
GitHub reports suspected full recovery. https://www.githubstatus.com/incidents/xwn6hjps36ty
We merged a fix and reports point to metrics being healthy.
We have confirmed that GitHub login, and repo access is working for all customers on the platform. We apologize for the impact and we will publish a public retro.
Our monitors have shown that we have good effect on our changes, although we have improvement, we are still monitoring the impact for all customers.
This incident has been resolved.
Staged changes are applying for all customers on the platform.
We have confirmed that GitHub login, and repo access is working for all customers on the platform. We apologize for the impact and we will publish a public retro.
Logs are working nominally.