← Open Source
vllm-project

semantic-router

An open, programmable decision layer for models and compute.

InfrastructureGatewayGo
Open on GitHub
Momentum
+4stars in 24 hours+0.1%
6.07k
Stars
1.02k
Forks
+64
This week
100
Contributors
Created 2025-08-26 · Updated 2026-10-10 · #2013 today
Top developers
README

vLLM Semantic Router

An open, programmable decision layer for models and compute.

Documentation | Playground | Blog | Publications | Hugging Face | Slack

vllm-project%2Fsemantic-router | Trendshift Decision 2.0 — #1 on Hugging Face Trending Collections, October 6, 2026

Main GitHub Release Go Ask DeepWiki


About

Intelligence beyond any one model.

Give your agent harness one API for many models. vLLM Semantic Router selects or combines models for each call, guided by your policy.

Your harness keeps the agent loop, tools, and task state. The Router chooses among configured backends across local, private, and cloud compute.

Dimension Fragmented today With vLLM SR
Models Different models excel at different tasks. Select or combine models.
Compute Hardware varies in speed and capacity. Choose among configured backends.
Location Edge, private, and cloud. Keep calls within approved locations.
Preference Priorities change by task. Set quality, latency, and cost priorities.

Explore how it works →

Getting Started

Install

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable

For pip, uv, or agent-driven installation, see the Installation Guide.

Connect your agent harness

Point your harness at the Router's inference endpoint. Use a public model ID such as vllm-sr/auto.

Follow Connect an agent harness for setup and compatibility.

Online playground

Try the online playground at .

Credentials:

  • Username: love@vllm-sr.ai
  • Password: vllm-sr-read

Latest News

Earlier announcements

More announcements are available on the Blog and Publications pages.

Community

For questions, feedback, or to contribute, please join the #semantic-router channel in vLLM Slack. Track contributors, workgroups, and weekly activity at community.vllm-sr.ai.

Community Meetings

We host two monthly community meetings across APAC and the Americas:

  • APAC-friendly meeting — second Wednesday of the month: 9:00-10:00 AM Singapore time (UTC+8; the same local time in Beijing)
  • Americas-friendly meeting — fourth Wednesday of the month: 8:00-9:00 PM Eastern Time (America/New_York) / 5:00-6:00 PM Pacific Time

Contributing

If you want to contribute, start with CONTRIBUTING.md.

For repository-native development workflow and validation commands, use AGENTS.md as the entrypoint and tools/agent/docs/README.md as the canonical index.

Citation

If you find Semantic Router helpful in your research or projects, please consider citing it:

@misc{semanticrouter2025,
  title={vLLM Semantic Router},
  author={vLLM Semantic Router Team},
  year={2025},
  howpublished={\url{https://github.com/vllm-project/semantic-router}},
}

Ecosystem & partnerships

An open ecosystem spanning research, infrastructure, and enterprise adoption.

vLLM Semantic Router's growing ecosystem: AMD, Hugging Face, Microsoft, Intel, NVIDIA, Red Hat, IBM, Liquid, DaoCloud, Delta, MBZUAI, McGill, KR Labs, University of Chicago, UC Berkeley, UMass Boston, University of Illinois Chicago, National Taiwan University, New York University, UBS, AI21, Bayer, Dell, and Nutanix.