Anyscale
19 articles in the TruthFoundry index. Each links out to the original.
- Self-hosted LLM serving cuts coding agent costs 90% with Ray and vLLM · Anyscale demonstrates self-hosted LLM serving reduces coding agent costs by 90% versus cloud APIs.
- Anyscale on Azure Launches to Enable Full AI Loop Ownership · Anyscale on Azure is now generally available, allowing enterprises to own their full AI loop on Microsoft's cloud.
- Ray Data LLM Delivers 2x Throughput Over vLLM Synchronous Engine · Ray Data LLM library achieves double throughput for production LLM batch inference workloads.
- Ray Summit 2026 Highlights Production Reinforcement Learning Infrastructure · Ray Summit 2026 showcased production RL and physical AI deployments running on Ray infrastructure.
- Anyscale Launches KubeRay Connect to Integrate with Kubernetes · Anyscale introduces KubeRay Connect to enable observability and orchestration on existing KubeRay stacks without migration.
- Ray Data introduces Shuffle V2 for faster, fault-tolerant joins · Ray Data launches Shuffle V2 to improve stability and scalability for large-scale data operations.
- KubeRay v1.6 and v1.7 Releases Add Beta History Server and Security Features · KubeRay v1.6 and v1.7 releases introduce a beta History Server, NetworkPolicy, mTLS, and RBAC authentication for Ray on Kubernetes.
- Anyscale launches GPU Health Observability to link hardware faults to Ray jobs · Anyscale announces private preview of GPU Health Observability to correlate hardware signals with Ray workloads.
- SkyRL Achieves On-Policy FP8 Sync to Preserve RL Policy Consistency · SkyRL introduces on-policy FP8 weight synchronization to maintain policy consistency across training and rollout in reinforcement learning.
- Anyscale ships GPU-native operators in Ray Data 2.58 with cuDF and RapidsMPF · Ray Data 2.58 adds cuDF batch format and RapidsMPF GPU shuffle, delivering up to 4x faster pipelines and 3.1x TCO improvement.
- Ray History Server Beta Enables Post-Mortem Debugging for Kubernetes Clusters · Ray promotes History Server to beta, allowing post-mortem observability for ephemeral Kubernetes clusters.
- Ray Scales AI Workloads to 10,000 Node Clusters · Ray Core optimizations enable AI training clusters up to 10,000 nodes with significant performance gains.
- Anyscale Outlines Learning Loops Maturity Curve for Building AI Intelligence Moats · Anyscale blog argues companies must build learning loops—data curation, custom training, inference—to own differentiated AI intelligence.
- Ray Serve LLM Introduces Token-Load-Aware Routing to Optimize Serving · Ray Serve LLM introduces a router that balances KV cache reuse with token load to improve latency.
- Ray 2.58 introduces native gVisor-based sandboxing in Google partnership · Ray 2.58 adds native sandboxing via gVisor, letting users run isolated OCI environments scheduled and scaled through Ray.
- CISA Lists Ray CVE-2025-62593 as Exploited; Users Urged to Upgrade · CISA added a Ray vulnerability to its KEV catalog, urging users to upgrade to version 2.52.0 and enable token authentication.
- Ray Serve Async Inference Benchmarked Against Amazon SageMaker · Ray Serve's async inference pipeline handles 67k video requests with zero failures under heavy load.
- Ray Direct Transport RDMA APIs cut RL weight-syncing overhead by up to 7.5x · Ray's RDT feature streamlines RDMA weight syncing in RL, claiming up to 7.5x speedup over naive implementations.
- NVIDIA Introduces Ray NVLink Domain-Aware Placement for GB300 Systems · NVIDIA launches Ray placement groups that automatically schedule workloads within single NVLink Domains for GB300 NVL72 racks.