Compare test

One prompt, every model you pick, run in parallel and shown side by side.

The three steps

Setting up a compare test walks through Prompt, Task and Models. You can jump back to any step you've already completed, but not ahead of one you haven't.

  • Prompt - what you want every model to answer. You can type it directly, generate an example to start from, or have it tidied up by a model before you run it.
  • Task - what kind of task this is (classification, extraction, summarization, generation, reasoning, and more). This helps route you to models that suit the task and feeds the platform's community benchmark data.
  • Models - which models to run the prompt against. You'll see an estimated cost in credits before you run anything.

Writing the prompt faster

Two small helpers sit above the prompt box. "Generate" drops in a random example prompt to start from when you're not sure what to test. "Tidy up" sends whatever you've already typed to a model and cleans it up - it's always available once you've written something, and defaults to a cheap model for the cleanup itself, though you can pick a different one.

You can also pull text straight out of a .docx, .txt or .md file and have it inserted into the prompt box, or type it by hand.

Text or image

Most tests are text in, text out. You can also attach an image or a PDF (each page is read as an image) and ask a vision-capable model about it - the model picker narrows automatically to models that can actually read an image, so you can't accidentally pick one that would just fail.

A PDF over 20 pages shows a range picker so you can choose which pages to include, rather than sending the whole document to every model by default - keeping the document the same length for every model in the comparison, and the cost predictable.

Reasoning effort

For models that support it, you can ask for Low, Medium or High reasoning effort, or leave it on Auto (the default) and let each model use its own default behaviour. Not every model supports every level, and the form will tell you if a choice isn't available for a model you've picked.

Cost

Before you run anything, you'll see an estimated cost in credits - the platform's own worst-case estimate for what the run could cost, with no markup. What you're actually charged after the run completes is usually lower, since the estimate assumes the longest possible answer from every model.

If your current balance is below the estimate, you'll see a note before you run so a test never stops partway through as a surprise.

After it runs

A finished test opens on four tabs: Setup (what you ran, so you can re-run it with changes), Playground (every model's answer, side by side, with cost/speed/tokens on each card), Summary (a sortable table across every result) and Winner (the result you've picked as best, if any).

On the Playground tab, an AI summary banner can read the results and recommend a model for you - one real model call, available once at least two results have finished. On the Winner tab, a "help me choose" assistant can talk through the trade-offs with you before you decide; that conversation isn't saved anywhere once you close it.

Picking a winner doesn't re-run or re-bill anything - it just marks which result you preferred. Winners are aggregated anonymously (no accounts, no names) on the platform's public Community Picks page, showing which models people actually choose for each kind of task.

Reliability runs

A single answer from a model is one sample - it doesn't tell you how much that model's output, speed or cost actually varies from one call to the next. A reliability run repeats an already-finished test's exact prompt and settings 2, 5 or 10 times (you choose), across whichever of its models you want to check, so you can see that variation directly instead of guessing at it.

Want to run this test again on a schedule, or chain it with others? See Smart alerts and Workflows.