There was an error while loading. Please reload this page.
A high-throughput and memory-efficient inference and serving engine for LLMs
Python 90.7k 21.5k
A framework for efficient model inference with omni-modality models
Python 6.5k 1.6k
Common recipes to run vLLM
JavaScript 1k 399
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python 3.7k 644
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python 787 206
A programmable Mixture-of-Models router for heterogeneous LLM inference
Go 5.5k 876
vLLM plugin for attention-ffn disaggregation support
Cost-efficient and pluggable Infrastructure components for GenAI inference
Community maintained hardware plugin for vLLM on Intel Gaudi
TPU inference for vLLM, with unified JAX and PyTorch support.
An LLM post-training framework with vLLM for RL Scaling
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
Community maintained hardware plugin for vLLM on Ascend
Loading…