Monitor and debug Ray workloads with fully persisted Cluster and Actor dashboards on Anyscale
Enable the new Cluster and Actor Dashboards to persist Ray metrics and debug post‑mortem after cluster shutdown.
Activate the Cluster and Actor Dashboards in your Anyscale account and start persisting cluster events for post‑mortem analysis.
Summary
Anyscale has released the Cluster and Actor Dashboards, completing the rollout of a fully persisted Ray Dashboard that now covers infrastructure, task, and actor views alongside the existing Train, Data, and Task dashboards. The new dashboards allow developers to monitor training progress, data pipeline bottlenecks, distributed task execution, and actor lifecycles, and to retain all metrics after a cluster shuts down for post‑mortem analysis. The previous Ray dashboard suffered from ephemerality—data vanished when a cluster terminated—and limited retention, keeping only 10 minutes of dead‑node data and the most recent 100,000 killed actors, which proved insufficient for large‑scale AI workloads. The Cluster and Actor Dashboards overcome these constraints by streaming and persisting events off‑cluster into Anyscale‑managed storage and query engines, enabling long‑term analysis of thousands of nodes and millions of actors. Key improvements include full persistence of all cluster events, scalability to handle large deployments, a faster, more intuitive UX with enhanced filtering and visualizations, and a unified navigation experience that lets users jump seamlessly between workload and system dashboards. Powered by the Ray Event Export Framework, the dashboards provide a single, interactive view of the entire Ray workload lifecycle, from high‑level overview to node‑level details. This new observability stack allows teams to debug failures, analyze performance, and compare workloads even after clusters have terminated, without maintaining their own infrastructure. The release is available today and promises to reduce debugging time and cost for any Anyscale user running Ray workloads.
Key changes
- Full persistence of cluster events after shutdown for post‑mortem analysis
- Scalability to thousands of nodes and millions of actors
- Enhanced UX with faster filtering, search, and new visualizations
- Unified navigation between workload (Train, Data) and system (Cluster, Task, Actor) dashboards
- Powered by Ray Event Export Framework, streaming events off‑cluster to Anyscale storage
- Supports long‑term analysis without maintaining own infrastructure
- Improves debugging of failures and performance bottlenecks