Workflows

A workflow is a suite of independent tests you can save, re-run, and schedule as a group.

What a workflow is

Each step in a workflow is its own test, with its own prompt, task, models and settings - a workflow doesn't chain steps together or pass output from one into the next, it just runs a set of tests as one unit and keeps their results together. A step can be a brand-new test you write right there, or a pointer at a test that already ran - reusing an existing test never re-runs or re-bills it, it just links to the historical result. A workflow can hold up to 10 steps.

Build and Runs

A workflow page has two tabs. Build is where you add, edit and remove steps and set the spend cap. Runs lists every time the workflow has actually been executed, each run showing every step's result together.

Spend cap

A workflow can carry its own maximum credits per run. Before a run starts, the whole workflow is estimated up front; if that estimate is over the cap, nothing runs at all rather than billing the first few steps and stopping partway through.

Comparing runs

Running the same workflow again shows how each step changed since the last run - cost, speed and whether the answer looks meaningfully different. Small token-level wobble on a short prompt is filtered out, so what you see is a real difference, not noise. It compares step by step, by position, and clearly skips (with a reason) any step whose prompt changed since the last run or that reuses an existing test, since there's nothing new to compare there.

Importing from n8n

If you already have model-calling steps built in an n8n workflow, you can import an exported n8n file when creating a new workflow. The file is read entirely in your browser - it's never uploaded or stored - and only the step prompts, models and structure are pulled out; credentials and anything that looks like a secret are never read, and nothing in the file is ever executed. Each imported step lands in the ordinary workflow form, so you can review and edit every one before saving.

  • Chain/agent nodes with a separate model node, and plain HTTP calls straight to a provider, both convert cleanly.
  • A prompt assembled in a Set or Code node one hop upstream, or built from a JavaScript expression, is converted on a best-effort basis - if it can't be reconstructed cleanly, you'll paste the prompt in yourself instead.
Want this workflow to re-run itself on a schedule and tell you when something changes? See Smart alerts.