
An open, programmable decision layer for models and compute.
Documentation | Playground | Blog | Publications | Hugging Face | Slack
About
Intelligence beyond any one model.
Give your agent harness one API for many models. vLLM Semantic Router selects or combines models for each call, guided by your policy.
Your harness keeps the agent loop, tools, and task state. The Router chooses among configured backends across local, private, and cloud compute.
| Dimension | Fragmented today | With vLLM SR |
|---|---|---|
| Models | Different models excel at different tasks. | Select or combine models. |
| Compute | Hardware varies in speed and capacity. | Choose among configured backends. |
| Location | Edge, private, and cloud. | Keep calls within approved locations. |
| Preference | Priorities change by task. | Set quality, latency, and cost priorities. |
Getting Started
Install
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable
For pip, uv, or agent-driven installation, see the Installation Guide.
Connect your agent harness
Point your harness at the Router's inference endpoint. Use a public model ID such as vllm-sr/auto.
Follow Connect an agent harness for setup and compatibility.
Online playground
Try the online playground at .
Credentials:
- Username:
love@vllm-sr.ai - Password:
vllm-sr-read
Latest News
- [2026/10/06] Decision 2.0 reached #1 on Hugging Face Trending Collections.
- [2026/10/06] Vela 2.0: Towards Open Foundation Routing Models
- [2026/09/24] vLLM Semantic Router v0.4 Hermes: Many Models, One Improving System
- [2026/09/22] Introducing Decision 1.0: Open Decision Foundation Models
- [2026/09/18] Introducing Vela 1.0
- [2026/08/24] Find Your Focus: How to Join and Work Together
- [2026/08/05] LettuceDetect v2 in Semantic Router: Generative Hallucination Detection as a vLLM Endpoint
Earlier announcements
- [2026/07/21] Beyond a Single Model: Building Mixture-of-Models Systems with vLLM Semantic Router
- [2026/07/09] Adding Cursor-Style Auto Model Selection to OpenCode with vLLM Semantic Router
- [2026/06/29] Micro-Agent: Beat Frontier Models with Collaboration inside Model API
- [2026/06/16] Beyond One Model: Fusion in vLLM Semantic Router
- [2026/06/05] vLLM Semantic Router v0.3 Themis: From Signals to Stateful Production Routing
- [2026/03/24] Vision Paper Released: The Workload-Router-Pool Architecture for LLM Inference Optimization
- [2026/03/10] v0.2 Released: vLLM Semantic Router v0.2 Athena Release
- [2026/02/27] White Paper Released: Signal Driven Decision Routing for Mixture-of-Modality Models
- [2026/01/05] Iris v0.1 Released: vLLM Semantic Router v0.1 Iris: The First Major Release
- [2025/12/16] Collaboration: AMD × vLLM Semantic Router: Building the System Intelligence Together
- [2025/12/15] New Blog: Token-Level Truth: Real-Time Hallucination Detection for Production LLMs
- [2025/11/19] New Blog: Signal-Decision Driven Architecture: Reshaping Semantic Routing at Scale
- [2025/11/03] Paper Published: Category-Aware Semantic Caching for Heterogeneous LLM Workloads
- [2025/10/27] New Blog: Scaling Semantic Routing with Extensible LoRA
- [2025/10/12] Paper Accepted: When to Reason: Semantic Router for vLLM
- [2025/10/08] Collaboration: vLLM Semantic Router with vLLM Production Stack Team.
- [2025/09/01] Released the project: vLLM Semantic Router: Next Phase in LLM inference.
More announcements are available on the Blog and Publications pages.
Community
For questions, feedback, or to contribute, please join the #semantic-router channel in vLLM Slack.
Track contributors, workgroups, and weekly activity at community.vllm-sr.ai.
Community Meetings
We host two monthly community meetings across APAC and the Americas:
- APAC-friendly meeting — second Wednesday of the month: 9:00-10:00 AM Singapore time (UTC+8; the same local time in Beijing)
- Americas-friendly meeting — fourth Wednesday of the month: 8:00-9:00 PM Eastern Time (
America/New_York) / 5:00-6:00 PM Pacific Time
Contributing
If you want to contribute, start with CONTRIBUTING.md.
For repository-native development workflow and validation commands, use AGENTS.md as the entrypoint and tools/agent/docs/README.md as the canonical index.
Citation
If you find Semantic Router helpful in your research or projects, please consider citing it:
@misc{semanticrouter2025,
title={vLLM Semantic Router},
author={vLLM Semantic Router Team},
year={2025},
howpublished={\url{https://github.com/vllm-project/semantic-router}},
}
Ecosystem & partnerships
An open ecosystem spanning research, infrastructure, and enterprise adoption.
