vLLM
3 articles in the TruthFoundry index. Each links out to the original.
- MiniMax H3 Serving Optimized with vLLM-Omni and FastH3 for Real-Time Generation · MiniMax H3 video generation achieves real-time serving via vLLM-Omni system optimization and FastH3 denoising reduction.
- vLLM Tests Speculative Decoding Performance on AMD GPUs · vLLM experiments show speculative decoding throughput varies by draft method and model family on AMD hardware.
- SkyRL Introduces IsoExec to Eliminate RL Training-Inference Mismatch · SkyRL launches IsoExec, a unified execution framework ensuring bitwise consistency between RL training and inference engines.