How many DeepSeek V4 Flash tokens do you get on Ollama Cloud Pro?
Ollama Cloud Pro is $20 a month with $60 of usage credits included. At deepseek-v4-flash's off-peak rates that's about 270 million input tokens or 90 million output tokens, and half as many during peak hours.
Checked against the lab's own pages on September 15, 2026. Limits change often, so the date matters.
Ollama Cloud Pro is $20 a month, or $200 a year billed annually, and the plan includes "$60 of usage credits per month". No token allowance comes with it. "Usage is measured in tokens at each model's rates", so how far $60 goes depends on which DeepSeek you call and when you call it.
- Pro plan
- $20 / mo. or $200/yr billed annually
- Included
- $60 of usage credits per month
- deepseek-v4-flash
- $0.22 input, $0.007 cached input, $0.66 output
- Peak rates
- $0.44 input, $0.014 cached input, $1.32 output
- Concurrency
- Free 1, Pro 3, Max and Team 10
- Reset
- monthly on your subscription start date, no rollover
What $60 buys on deepseek-v4-flash
Ollama's model pricing table puts deepseek-v4-flash at $0.22 per million input tokens, $0.007 per million cached input tokens and $0.66 per million output tokens. Spend the whole $60 on input and that's about 270 million tokens. Spend it all on output and it's about 90 million.
Real work sits between the two. At three parts input to one part output, $60 comes to roughly 180 million tokens in a month, and heavy cache hits stretch it well past that, since cached input is $0.007. Granted, that ratio is my assumption (Ollama publishes no token figure for any plan). The rates and the credit are Ollama's. The division is yours.
Peak hours cost double
"Peak pricing applies between 12:00 and 18:00 UTC, Monday to Friday." Inside that window deepseek-v4-flash runs at $0.44 input and $1.32 output. Double, to the cent. If everything you do lands in it, the same $60 covers about 136 million input tokens or 45 million output tokens.
The newer Flash is cheaper
deepseek-v4.1-flash sits below it, at $0.15 input and $0.60 output off-peak, and $0.30 and $1.20 at peak. On the off-peak rate that's about 400 million input tokens for your $60 (and cached input drops to $0.003). If the model choice is open to you, that's the cheaper line in the same family.
When it resets, and what happens when the credit runs out
On Pro, Max, and Team plans, included usage resets monthly on the same day of the month your subscription started, including on annual plans.
Unused credit doesn't roll over; the included amount refreshes at each reset. Ollama emails a reminder at 90% of your included monthly usage, and you can turn that off in settings.
Past the $60, Ollama uses included plan credits first and then draws from any usage credits you've bought. Running models on your own hardware is always unlimited.
How many requests you can run at once
Pro gets 3 concurrent requests. Free gets 1, and Max and Team get 10. Anything over your limit is queued, and once the queue hits its own fixed limit (Ollama doesn't say what that is), the request is rejected until one of your slots opens.