Models, providers and routing
Hosted and self-hosted providers
Cloud APIs and your own GPU servers behind one endpoint: OpenAI, Anthropic, Bedrock, Vertex, Gemini, vLLM, Ollama and more.
What it is
Native adapters for OpenAI-compatible APIs, Anthropic, AWS Bedrock, Google Vertex AI, Google Gemini, Ollama, vLLM, llama.cpp and Hugging Face TEI. Any other OpenAI-compatible server, such as LM Studio, works through the generic adapter.
Why you want it
Most organizations use a mix of cloud APIs and their own GPU servers. One gateway in front of all of them gives one set of controls and one bill view.
How it works
- Registered adapter types: openai_compatible, anthropic, bedrock, vertex, gemini, ollama, vlm, llama_cpp, tei (plus codex and copilot used by personal subscriptions)
- Upstream credentials encrypted at rest with AES-256-GCM
- Bedrock and Vertex use your own cloud account and credentials
- Google Gemini uses an AI Studio key, meters thinking tokens as output and records implicit cache hits
See it on your own network.
The Community edition is free for up to 25 people. The 30-day Business trial unlocks every Business feature.