Chat test

A multi-turn conversation, run independently against every model you pick.

How it's different from a compare test

A compare test is one prompt, one round of answers. A chat test is an ongoing back-and-forth: each message you send goes to every model you picked, and each model keeps its own separate conversation history. One model never sees another model's replies - you're having N independent conversations side by side, not one shared thread.

Because there's no single upfront prompt to write, a chat test skips straight to picking models - there's no Prompt or Task step. You can give the test a name, put it in a project, and set a system prompt that applies to every model's conversation before you start.

What you see while chatting

Since there's no Setup tab to show full settings in, the chat view keeps a compact row of the test's key settings (model, temperature, top P, max tokens) visible at the top instead. Each message you send appears once, followed by a row of reply cards - one per model, so you can read every model's answer to the same message side by side before deciding what to ask next.

Cost as you go

Each message shows a live estimated cost before you send it, the same way the cost preview on a compare test works - there's no confirmation popup in front of every message, since sending a chat message is the test's ordinary purpose, not an extra action layered on top of one.

If a model gets retired mid-conversation

Providers occasionally retire a model id while you're still mid-conversation with it. Where the platform has confirmed a live successor for that model, your next message continues with the replacement model instead, inheriting the earlier replies as its own history - so the conversation carries on rather than breaking.

Reasoning effort and a few other per-test settings available on a compare test aren't offered on chat tests yet - each message uses the model's default behaviour.