Configure and run a test in one screen. Select your test data, pick a judge, choose which models to compare, and hit run.
- Select 2 models or 20 from the full catalog (GPT, Claude, Gemini, open-source, any AI Gateway model)
- Set repetitions per question (1/3/5/10) for statistical reliability
- See a live cost estimate before running
- Automated design review flags issues: too few questions, judge bias, budget too tight, low statistical power
- Traffic-light quality checks so you know the results will be meaningful