Models, providers and routing

Load-balanced model pools

Spread one model across several GPU servers by live queue length, and keep each conversation on the server that holds its cache.

What it is

Put several servers running the same model behind one alias. Janus reads each vLLM or llama.cpp server's live load and sends each request where it fits.

Why you want it

Self-hosted GPU fleets waste capacity when traffic is spread blindly. Pools keep queues short, keep conversations on the server that holds their prompt cache, and route around a failing node.

Two requests for code-assist reach Janus. Janus reads each GPU server's live queue every two seconds. Turn 6 of conversation s-42 stays on gpu-node-1, which holds that conversation's prompt cache. A new conversation goes to gpu-node-2, which has the shortest queue. gpu-node-3 failed before answering, so it is skipped within the same request. REQUESTS FOR code-assist Turn 6conversation s-42 New chatno history Janus least loaded + affinity gpu-node-1 3 waiting has s-42's cache → turn 6 stays gpu-node-2 0 waiting shortest queue → new chat gpu-node-3 failed before answering skipped withinthe same request Live queue and cache use are read from each server every 2 seconds.

How it works

  • Policies: failover, weighted round robin, least loaded and context aware
  • Live load scraped from each server's /metrics every 2 seconds
  • Session affinity from X-Janus-Session, X-Session-Id, prompt_cache_key or the conversation's opening
  • A server that fails before answering is skipped within the same request; three failures remove it on every replica with backoff
  • Responses carry X-Janus-Pool-Member and X-Janus-Pool-Reason; a live 'Servers right now' panel shows queue and cache use
Every pooled response explains its routing
POST /v1/chat/completions
X-Janus-Session: review-4471
{ "model": "code-assist", … }

HTTP/1.1 200 OK
X-Janus-Pool-Member: qwen3-coder-30b @ gpu-node-2
X-Janus-Pool-Reason: affinity

In the product

Janus Servers right now screen with demo data
Servers right now

Live queue, context and prompt-cache use for each GPU server in a pool, refreshed every few seconds.

Actual Janus interface · Demo identities and synthetic usage
Janus Pool settings screen with demo data
Pool settings

Pool members, the balancing policy and how strongly conversations stick to the server that holds their cache.

Actual Janus interface · Demo identities and synthetic usage

See it on your own network.

The Community edition is free for up to 25 people. The 30-day Business trial unlocks every Business feature.