Metering, cost and limits

Per-person rate limits

Requests-per-minute limits that hold across every replica protect a shared upstream from one heavy user.

What it is

Requests-per-minute rules per person, service token or endpoint. The limit is a cluster-wide ceiling, divided across live replicas.

Why you want it

Rate limits protect a shared upstream from one heavy user without a separate API gateway.

How it works

  • Token buckets in memory, divided by the live replica count from the instance heartbeat
  • Rules can target one subject or every user individually
  • Enforcement is switched on with the per_user_rate_limits_enabled runtime flag

See it on your own network.

The Community edition is free for up to 25 people. The 30-day Business trial unlocks every Business feature.