There was an error while loading. Please reload this page.
A high-throughput and memory-efficient inference and serving engine for LLMs
Python 91.5k 22.1k
A framework for efficient model inference with omni-modality models
Python 6.8k 1.7k
Common recipes to run vLLM
JavaScript 1k 414
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python 3.8k 660
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python 825 218
A programmable Mixture-of-Models router for heterogeneous LLM inference
Go 5.7k 935
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
Community maintained hardware plugin for vLLM on Ascend
Cost-efficient and pluggable Infrastructure components for GenAI inference
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
The vLLM XPU kernels for Intel GPU
Stateful API logic for agentic applications using vLLM