Testing
Test AI models on your own questions, see which performs best, and route work to the right one.
Where to Find It
- Sidebar section: Design My AI
- Sidebar label: Testing
- URL path:
/model-testing
The Testing area opens on an Overview with counts of your test data, evaluators, test runs and smart routers, plus your most recent runs and a "Run a test" button.
Across the top is a tab bar you can use to move between the steps:
- Overview — your summary and recent runs (
/model-testing) - New Test Run — pick a test, an evaluator and the models to compare, then run it
- Test Data — the questions your models will be tested on
- Evaluation — the evaluators (judges) that score answers
- History — every past run and its results
- Smart Router — build a router that sends each request to the best model, based on your results
Step 1 — Create your test data
Go to the Test Data tab and click + Test Data. Choose one of three options:
- Create test manually — write your questions and answers yourself.
- Generate new test from file — upload a document and we write the questions for you.
- Import existing test — bring in a test you already have as JSON.
Generate from a file
Upload a document, then choose the kind of test:
- Knowledge test — checks what a model already knows about a topic from its own training (for example, upload a tax code and we create questions the model must answer from memory).
- Prediction test — checks a model's ability to predict an outcome (for example, upload A/B marketing case studies and the model is shown both variants and asked to predict which won).
We generate the questions with answers already filled in, each backed by a quote from your document. You confirm the ones you want (confirm at least the minimum shown), then save. The model being tested never sees your document.
Supported files: TXT, MD, CSV, JSON, PDF and DOCX.
Import from JSON
The Import screen shows the exact JSON format, a one-line instruction you can paste to any AI to convert your existing test, and a downloadable example file. Upload a .json file or paste the JSON, and we import the questions. If a file is the wrong type or the JSON isn't in the right format, you'll get a clear message explaining how to fix it.
Step 2 — Set up an evaluator
Go to the Evaluation tab and click + New Evaluator. Pick the model that will score answers and set the scoring instructions. When your test includes known correct answers, the evaluator judges each answer strictly against them, so a confident but wrong answer cannot score well.
Step 3 — Run a test
Go to New Test Run. Choose your test data, an evaluator, and the models you want to compare. Set:
- Repetitions per question — 1 (quick check), 3 (reasonable), 5 (solid), or 10 (rigorous). More repetitions give more reliable results.
- Budget cap (optional) — stop the run if the total cost would exceed this amount.
Before you start, an Estimated cost range is shown in credits, based on current model pricing and the size of your selected test. Then start the run and watch live progress — the test name and a progress bar are shown while it runs.
While a run is in progress you can click Cancel Run to stop it. If a run stops early (for example the servers are busy, or it reached your budget cap), open it from History and click Resume to continue from where it left off — you are not charged again for questions already completed.
Step 4 — Review results and build a router
- History lists every run; open one to see how each model scored.
- Smart Router lets you select one or more completed runs and build a router that automatically sends each request to the best model for the job, based on your results and your priorities (quality, speed, cost).
Requirements
- Plan: all
Common Issues
"Choose a test type" when generating from a file → Pick either Knowledge test or Prediction test before generating — the button stays disabled until you do.
"That isn't valid JSON" or "the wrong file type" when importing
→ Use the required format shown on the Import screen (download the example, or paste the example to your AI and ask it to convert your test). Upload a .json file.
Not enough confirmed questions to save a generated test → Confirm at least the minimum number of questions shown in the counter before saving.
Estimated cost looks high → Lower the number of models, reduce repetitions, or use a smaller test. You can also set a budget cap to stop a run automatically.
A run stopped with "servers are busy" or "budget cap reached" → Open the run from History and click Resume to finish it — completed questions are kept, so you are not charged for them again.