Load-balanced model pools
Spread one model across several GPU servers by live queue length, and keep each conversation on the server that holds its cache.
What it is
Put several servers running the same model behind one alias. Janus reads each vLLM or llama.cpp server's live load and sends each request where it fits.
Why you want it
Self-hosted GPU fleets waste capacity when traffic is spread blindly. Pools keep queues short, keep conversations on the server that holds their prompt cache, and route around a failing node.
How it works
- Policies: failover, weighted round robin, least loaded and context aware
- Live load scraped from each server's /metrics every 2 seconds
- Session affinity from X-Janus-Session, X-Session-Id, prompt_cache_key or the conversation's opening
- A server that fails before answering is skipped within the same request; three failures remove it on every replica with backoff
- Responses carry X-Janus-Pool-Member and X-Janus-Pool-Reason; a live 'Servers right now' panel shows queue and cache use
POST /v1/chat/completions
X-Janus-Session: review-4471
{ "model": "code-assist", … }
HTTP/1.1 200 OK
X-Janus-Pool-Member: qwen3-coder-30b @ gpu-node-2
X-Janus-Pool-Reason: affinityIn the product
Live queue, context and prompt-cache use for each GPU server in a pool, refreshed every few seconds.
Actual Janus interface · Demo identities and synthetic usage
Pool members, the balancing policy and how strongly conversations stick to the server that holds their cache.
Actual Janus interface · Demo identities and synthetic usageSee it on your own network.
The Community edition is free for up to 25 people. The 30-day Business trial unlocks every Business feature.