Models & pricing
Every active model, across nine providers, with its real per-token price kept up to date.
Providers
The catalogue spans nine providers:
- Anthropic (Claude)
- OpenAI (ChatGPT)
- xAI (Grok)
- Google (Gemini)
- Perplexity
- DeepSeek
- Alibaba (Qwen)
- Mistral
- Moonshot (Kimi)
The full, current list is always on the Models page inside the app - prices and which models are available change over time as providers release and retire them, so that page is the source of truth, not a number written down here.
Filtering the catalogue
Wherever you pick models - a new test, a chat, a workflow step - you can filter by Flagship (a provider's best model), Cheap, or Fast, or search by name. Each model's row shows its provider, whether it can read images, and whether it's a premium-tier model before you pick it, so you're not guessing at capability or cost after the fact.
A model that's currently failing real calls is marked unavailable right in the picker and can't be selected, rather than letting you build a test around a model that would just fail.
New models
When a new model becomes available, it's surfaced as a suggestion - on a test, a workflow or an alert that could use it - rather than added to anything automatically. A saved workflow or alert never changes what it runs on its own, since that would change what an unattended scheduled run costs without you choosing it.
Retired models
Providers sometimes retire a model id. Where one has a clear successor, saved alerts, workflow steps and chat conversations built on the old id are moved to the replacement automatically so they keep running instead of silently breaking. Starting a brand-new test with a retired model by name is still refused.
Reasoning effort
Some models can "think" before answering, at a cost in extra tokens. Where a model supports it, you can choose Low, Medium or High, or leave it on Auto and let the model decide for itself - Auto is not the same as Low, and how much a given model reasons on Auto varies a lot model to model. A few models only support switching thinking on or off rather than a graded scale, and a few don't support the setting at all - the form only offers what a model actually accepts.
Pricing that changes by time or by document length
Most models charge a flat rate per token. A few don't, and the platform accounts for it rather than showing a misleading flat number:
- DeepSeek prices certain off-peak hours at half its normal rate - the Models page shows the off-peak price, and a run during a cheaper window is billed at it.
- Perplexity's Sonar models charge a flat fee per request on top of token costs, and its Deep Research model also bills per search it runs - reflected in the estimate before you run, not just in the final bill.
- A handful of models charge a different rate once a conversation gets very long (a "long-context" tier) - priced automatically if a prompt ever crosses that threshold.