Quality check
Judge your results against a rule you write in plain language, instead of reading every answer yourself.
How it works
On a completed test, you write a pass criterion in your own words - for example, "answers in under 50 words and doesn't mention competitors." Each result is then checked against that rule by other AI models acting as judges, and comes back marked pass, fail, or uncertain.
How the judging actually runs
Two cheap judge models look at an answer first. If they agree, that's the verdict. If they disagree, or either one gives an answer that can't be parsed into a clear verdict, a third judge is brought in automatically and the majority decides - you don't see the back-and-forth, just the final result. A model is never used to judge an answer that came from its own provider, so a result is never marked by something with a reason to be generous to itself.
Occasionally there's no majority - a genuine three-way split, or a two-judge disagreement with no third opinion available for that particular answer. Rather than guess, that result is marked uncertain instead of being forced into a pass or fail it didn't earn.
Cost
Running a quality check holds credits up front for every judge call that could happen across every result - the same worst-case-first approach a test run uses - and settles against what it actually cost once it's done. Most answers only need the first two, cheaper judges; the third only runs when there's a genuine disagreement to resolve.
What it doesn't cover yet
- Quality check runs on a test's own results. A reliability run's repeats are their own separate tests and aren't covered by one quality check on the original.