Model Testing

All AI benchmarks are useless. The only one that matters is yours.

Run tests in minutes that compare models from GPT, Claude, Gemini, open-source (or any model) with your specific use case. Measure quality, speed, and cost. No technical skills required.

results / composite scores · 20 tasksComplete
Quality: 70 Cost: 85 Speed: 45
Claude Sonnet · 4.6/5 · $0.003 · 1.2s
94
GPT-Sol · 4.2/5 · $0.002 · 0.9s
81
Gemini Flash · 3.8/5 · $0.0008 · 0.4s
73
Kimi K2 · 3.4/5 · $0.0003 · 0.6s
61

How model testing works.

Configure and run a test in one screen. Select your test data, pick a judge, choose which models to compare, and hit run.

  • Select 2 models or 20 from the full catalog (GPT, Claude, Gemini, open-source, any AI Gateway model)
  • Set repetitions per question (1/3/5/10) for statistical reliability
  • See a live cost estimate before running
  • Automated design review flags issues: too few questions, judge bias, budget too tight, low statistical power
  • Traffic-light quality checks so you know the results will be meaningful
new-test-run / configureReady
RequirementsContract review (20 tasks)
JudgeQuality Evaluator v2
Models4 selected
Repetitions5 per question
Est. cost$0.84 - $1.20

Define what your AI needs to be good at. Create test questions that mirror your actual work.

  • Use the 6-step "Design a Test" wizard: describe your work in English, the system generates relevant test questions
  • System detects capabilities from your description (drafting, compliance, summarization, reasoning, etc)
  • AI-generates questions per capability. Toggle individual questions on/off. Write your own manually.
  • Or start from a pre-built template for your industry
  • Set questions per category (3/5/7/10) and optional ground truth
design-test / step 1Your use case
DESCRIBE YOUR WORK:
"Client onboarding for an accounting practice. Drafting engagement letters, checking compliance."
CAPABILITIES DETECTED:
Drafting Compliance Summarization Accuracy

Create AI judges that score responses against your criteria. Each judge is a model paired with your scoring prompt.

  • Define what "good" looks like for your use case
  • Create multiple judges with different evaluation perspectives (accuracy, tone, completeness)
  • Judge calibration system verifies scoring consistency before you rely on results
  • Calibration badges: Gold (r > 0.85), Silver (r > 0.7), Bronze (r > 0.5)
judges / your evaluators2 active
Quality Evaluator v2Gold
Compliance CheckerSilver

Every run stored with full drill-down. Search, filter, compare, and export.

  • Search runs by title, test data, or description
  • Filter by tags and status (completed, running, failed)
  • Click into any run to see individual model responses and judge scores per task
  • Compare runs over time as models improve or your requirements change
  • Export full raw results as JSON
history / recent runs3 runs
Contract review (4 models)$1.04
Email drafting (3 models)$0.42
Data analysis (5 models)$2.18

Turn your test results into an intelligent router. One click. OpenAI-compatible API endpoint.

  • Set your priority weights for quality, cost, and speed
  • The router automatically dispatches each task to the winning model from your tests
  • Generates an OpenAI-compatible endpoint you can use anywhere
  • Use it in IIMAGINE chat, in agents, or via API in any external tool
  • Usage stats: total requests, total cost, per-model breakdown
  • One active router at a time. Switch instantly.
smart-router / activeLive
ENDPOINT:
https://api.iimagine.ai/smart-router/{key}/chat/completions
WeightsQ:70 / C:85 / S:45
Pool4 models
Requests12,847
Cost$4.22

Use your Smart Router everywhere.

Once created, your router works in three places. The same test results, the same weights, the same optimized model selection.

In IIMAGINE Chat

Every message automatically uses the best model for that type of question. No manual switching.

In Agents

Agents use the winning model for each step of their workflow. Complex reasoning gets the best model. Simple lookups get the cheapest.

Via API

Drop the OpenAI-compatible endpoint into any external tool. Cursor, VS Code, custom apps. Anything that supports a custom AI endpoint.

Stop paying for a brand name. Pay for what works.

Free account. No credit card. Run your first test in minutes.

Start free