vLLM
23 articles in the TruthFoundry index. Each links out to the original.
- vLLM Optimizes DeepSeek-V4.1-Flash for 5x Agentic Throughput · vLLM engineers optimized DeepSeek-V4.1-Flash to achieve 5x throughput on agentic serving benchmarks.
- vLLM Disaggregated Serving Separates Prefill and Decode for Lower Latency · vLLM's disaggregated serving splits prefill and decode stages to reduce latency and improve throughput.
- vllm-metal Brings Concurrent LLM Serving to Apple Silicon · vllm-metal launches v0.28.0, enabling concurrent LLM serving on Apple Silicon with vLLM's scheduler and MLX execution.
- vLLM Achieves 5000 TPS/GPU on Qwen3.8-2.4T with GB300 Cluster · Researchers demonstrate vLLM serving Qwen3.8-2.4T on GB300 NVL72 clusters, achieving 5000 tokens per second per GPU.
- vLLM Adds NVIDIA GPU Video Decoding to Scale Multi-GPU Captioning · vLLM integrates PyNvVideoCodec to shift video decoding from CPU to NVIDIA GPUs, removing bottlenecks in multi-GPU inference.
- Novita AI releases Chord, faster INT4 MoE operator for vLLM · Novita AI open-sources Chord, a high-performance MoE CUDA operator accelerating Kimi K2.x inference on vLLM.
- vLLM Team Trains DSpark Speculator for Kimi K3 Frontier Model · vLLM researchers extended their Speculators library to train a DSpark draft model for the 2.8T-parameter Kimi K3.
- vime and RL-Kernel Achieve Bitwise Train–Rollout Consistency on AMD ROCm · Researchers achieved zero numerical mismatch between training and rollout engines on AMD Instinct MI300X GPUs using vime and RL-Kernel.
- vLLM Optimizations Boost Kimi K3 Serving Throughput by 2.8x · New vLLM optimizations reduce Kimi K3 latency by 56% and increase throughput by up to 2.8 times.
- vLLM launches tiered KV cache offloading for improved LLM serving · vLLM implements tiered KV cache offloading to reduce recompute and boost serving capacity.
- MiniMax M3 Inference Optimized on AMD Instinct MI355X · MiniMax M3 achieves 342.4 output tokens/s/GPU on AMD Instinct MI355X via sparse attention and speculative decoding.
- vLLM Optimizes Agentic Serving with New KV Cache and Parallelism Features · vLLM introduces hybrid KV cache management and disaggregated serving to handle growing agentic AI workloads.
- vLLM Introduces Hybrid HiSparse for GLM 5.3 Inference · vLLM launches Hybrid HiSparse to enable full 1M context length for GLM 5.3 on memory-constrained hardware.
- Tenstorrent Adds Mesh-Aware vLLM Plugin for LLM Serving · Tenstorrent introduces a vLLM plugin enabling LLM serving on its mesh hardware via standard plugin mechanisms.
- MiniMax H3 real-time video serving via vLLM-Omni and FastH3 four-step generation · vLLM-Omni and FastVideo's FastH3 reduce MiniMax H3 video generation latency to below playback duration on eight NVIDIA B300 GPUs.
- vLLM Implements Speculative Decoding for AMD GPU Optimization · vLLM explores speculative decoding methods to increase output-token throughput on AMD Instinct GPUs.
- vLLM Implements Sharded Weight Transfer with Ray Direct Transport · vLLM launches a native sharded weight transfer engine using Ray Direct Transport (RDT) and NIXL for efficient RL model syncing.
- IsoExec Uses Unified Execution Contract to Eliminate Train-Inference Mismatch in SkyRL · IsoExec unifies training and inference execution via an execution contract and aligned kernels to eliminate bitwise mismatch in SkyRL.
- VeRL-Omni v0.2.0 accelerates diffusion RL and stabilizes omni training · VeRL-Omni v0.2.0 introduces faster diffusion RL via vLLM-Omni and stable omni training with reusable adapters.
- vLLM-Omni Introduces Distributed Layerwise Offload for 200B+ Video Models · vLLM-Omni launches Distributed Layerwise Offload to run large video models across GPUs with minimal memory overhead.
- vLLM Adds Adaptive Verification to Optimize Speculative Decoding · vLLM integrates DSpark's adaptive verification to dynamically adjust speculative decoding based on confidence and system load.
- Inferact Announces Day-0 vLLM Support for Qwen3.8-2.4T-A95B Open-Weight Model · Inferact releases Qwen3.8-2.4T-A95B with immediate vLLM support and optimized quantized weights.
- NVIDIA Announces Day-0 Support for Nemotron 3.5 Lightning on vLLM · NVIDIA enables immediate deployment of Nemotron 3.5 Lightning on vLLM with speculative decoding and MoE optimizations.