ThinkFacility

Plan limits

What are the GPT-6 Astra rate limits on the OpenAI API?

On the OpenAI API, GPT-6 Astra runs from 500 requests and 500,000 tokens a minute at Tier 1 to 15,000 requests and 40,000,000 tokens a minute at Tier 5. The free tier can't use it.

Checked against the lab's own pages on September 22, 2026. Limits change often, so the date matters.

GPT-6 Astra's API rate limits go by usage tier, and OpenAI prints the whole table on the model's own page. There's no free-tier access to this model at all. Tier 1 opens at 500 requests a minute, and Tier 5 tops out at 15,000.

Free
Not supported
Tier 1
500 RPM, 500,000 TPM, 1,500,000 batch queue limit
Tier 2
5,000 RPM, 1,000,000 TPM, 3,000,000 batch queue limit
Tier 3
5,000 RPM, 2,000,000 TPM, 100,000,000 batch queue limit
Tier 4
10,000 RPM, 4,000,000 TPM, 200,000,000 batch queue limit
Tier 5
15,000 RPM, 40,000,000 TPM, 15,000,000,000 batch queue limit

How you get to a higher tier

Tiers come off what you've already paid. Tier 1 needs $5 paid, Tier 2 $50, Tier 3 $100, Tier 4 $250 and Tier 5 $1,000, and OpenAI graduates you automatically as your spend goes up. Each tier carries an approved monthly usage limit of its own: $100 a month at Tier 1, $500 at Tier 2, $1,000 at Tier 3, $5,000 at Tier 4 and $200,000 at Tier 5.

One step in that table isn't really a step. Tier 2 to Tier 3 leaves requests per minute where they were, at 5,000 either side. Tokens per minute double, though, and the batch queue limit jumps from 3,000,000 to 100,000,000 (more than thirty times over, on the row nobody reads).

When the window resets

RPM and TPM are per-minute limits, and you don't have to guess where you are inside one. Every response carries headers with what's left and when it comes back, including x-ratelimit-remaining-requests, x-ratelimit-remaining-tokens and a reset time for each. Batch queue limits work differently: they count the input tokens sitting in pending batch jobs, and a job's tokens stop counting once it finishes.

So the meter runs on what you might use. OpenAI sizes each request off your max_tokens or its own estimate from the character count, whichever comes out higher.

Your rate limit is calculated as the maximum of max_tokens and the estimated number of tokens based on the character count of your request

From Rate limits | OpenAI API

What you see when you hit one

Going over returns a 429, and the response can include a Retry-After header giving the minimum seconds to wait. Treat that as a floor, wait at least that long, and add a small random delay so every client doesn't come back at the same moment. Failed requests still count against your per-minute limit, so hammering it doesn't help.

Two different 429s can come back, and they want different things from you. The slow_down code means your request rate increased too quickly, and OpenAI says it can fire even while you're inside your RPM and TPM. The rule of thumb on the same page: past 1 million input tokens a minute, raise traffic by no more than 50% every 15 minutes. A 503 with server_is_overloaded is the other one, and that's the model being busy rather than you.

If you don't need answers back immediately, the Batch API runs outside the synchronous request limits, which is what the batch queue column in the table is for.