Skip to main content
An in-flight request is a request Requesty has accepted and not finished yet. Requesty does not cap how many requests you send per minute, it caps how many can be in flight at the same time. The limit applies to all traffic from your organization, across every API key, and it is set automatically from your organization’s tier and balance. When the limit is exceeded the request is rejected with HTTP 429. Wait for one of your requests to finish, increase your organization’s balance to raise the limit, or contact [email protected]. Nothing is queued and nothing is retried on your behalf, so a 429 is safe to retry once one of your in-flight requests completes.
If you are on an Enterprise plan, or you pay via monthly invoicing, your organization has custom in-flight rate limits and the balance-based table below does not apply to you. Contact [email protected] to review or change them.

The limit for all other customers

Your organization’s in-flight limit is your current balance multiplied by a multiplier, and the multiplier depends on your tier. Your tier comes from your lifetime spend with Requesty: Tiers are based on lifetime totals and only ever move up, so refunds or a spent down balance never take a tier away. Crossing a threshold promotes your organization within the hour. This is the multiplier for every combination of tier and current balance: The highest balance row you clear is the one that applies, and the balance is rounded up to the next whole dollar before it is multiplied. Examples: Because the limit is derived from your balance, it moves with your balance: topping up raises it, and spending down lowers it. Enabling auto top-up in billing settings keeps your limit stable.
Your auto top-up threshold is also your rate limit floor, because your balance never stays below it for long. Set the threshold at an amount whose multiplier still covers your peak concurrency. On tier 0, a threshold of $80 keeps you on the > $50 row and never below 2,000 in-flight requests, while a threshold of $20 leaves you at 200.

Best practices

  • Bound your own concurrency. Cap the number of parallel requests in your client below your organization limit instead of relying on 429s to throttle you.
  • Retry on 429. Retry with a short backoff. In-flight slots free up as your other requests complete, so a retry usually succeeds quickly.
  • Keep a buffer in your balance. A low balance is also a low concurrency limit, auto top-up avoids surprise throttling.
  • Split noisy workloads. Give batch or evaluation workloads their own service account, so interactive traffic is easier to keep separate.
Need higher in-flight rate limits than the table above gives you? Talk to us at [email protected] and we will review your limits.

Resources

Last modified on August 28, 2026