Metering, cost and limits
Per-person rate limits
Requests-per-minute limits that hold across every replica protect a shared upstream from one heavy user.
What it is
Requests-per-minute rules per person, service token or endpoint. The limit is a cluster-wide ceiling, divided across live replicas.
Why you want it
Rate limits protect a shared upstream from one heavy user without a separate API gateway.
How it works
- Token buckets in memory, divided by the live replica count from the instance heartbeat
- Rules can target one subject or every user individually
- Enforcement is switched on with the per_user_rate_limits_enabled runtime flag
See it on your own network.
The Community edition is free for up to 25 people. The 30-day Business trial unlocks every Business feature.