# Evaluator: Make Sure the Tests Grading Your AI Are Right

> The Evaluator is the LLM evaluation checker in the Ai1 platform by MyZone AI that reviews the tests grading your AI prompts and agents, feeds them deliberately broken examples and gives a pass, flag or fail verdict on whether the grading can be trusted, while a person approves the decisions that matter.

Canonical page: https://myzone.ai/pages/agents/ai-evaluator-agent
Part of Ai1, by MyZone AI. Book an Evaluator walkthrough: https://calendly.com/d/ct6h-tcy-8qf/ai1-demo?a1=Evaluator&utm_source=myzone.ai&utm_medium=agent-page&utm_content=ai-evaluator-agent-final
Last updated: 2026-10-05. Reviewed by the MyZone AI team.

Check the test that grades your AI before you trust its scores.

## At a glance

- **What it checks:** The tests that grade your AI prompts and agents
- **How it tests them:** Feeds the grader deliberately broken examples to see what it misses
- **The verdict:** Pass, flag or fail, with the evidence and the fixes needed
- **Before paid runs:** A free dry check comes first
- **After a grading change:** Re-runs a fixed reference set and shows where results moved
- **A person approves:** Safety labels and pass bars, benchmark lock-ins, live adoption and published results

## What the Evaluator does

- **Reviews whether a test can be trusted:** Looks at the test setup, the reference answers and recent results, checks the data is split properly and the grader is independent, then gives a pass, flag or fail with evidence.
- **Plants mistakes to test the grader:** Feeds the grader examples that are wrong on purpose. Any it fails to catch are explained and turned into fixes.
- **Spots drift after a change:** When the grading instructions, grading model, scoring rules or data change, it re-runs a fixed reference set and shows where results moved.
- **Traces wrong results to their cause:** Works out whether a bad result came from the prompt, the test, the reference answers or the grader, logs it and says who should fix it.
- **Writes a short evidence report:** Shows how often the grader agreed with the right answer, which planted mistakes it caught, its known limits and the next step.

## How the Evaluator differs from a dashboard of test scores

A dashboard of test scores shows you the numbers and assumes the test behind them is sound. The Evaluator checks that test first: it reviews the data splits and the grader's independence, plants deliberate mistakes to see what the grader misses and re-runs a fixed reference set when the grading changes. It gives a pass, flag or fail with evidence, and a person approves safety labels, pass bars and published results.

How this differs from the QA Orchestrator (https://myzone.ai/pages/agents/ai-qa-agent): the QA Orchestrator runs general quality checks on apps and releases, while the Evaluator checks whether the tests grading your AI prompts and agents can be trusted.

## How it works

1. **A test setup comes in** (Before you rely on a test): You, or another agent, ask whether an evaluation setup is ready to trust.
2. **Review the setup** (Before any paid run): It checks the data splits, the automatic checks and whether the grader is independent of what it grades.
3. **Stress test the grader** (After the setup review): It runs deliberately broken examples through the grader to see which ones it catches.
4. **Give a verdict** (When the checks are done): Pass, flag or fail, with the evidence and the fixes needed before the setup can be relied on.
5. **A person approves what matters** (you approve) (Before anything important changes): Changing critical safety labels or pass bars, locking in a benchmark, adopting a change live or publishing results all wait for a person's approval.

## When to use it

- **You are about to pay for a large test run:** It runs a free dry check first and tells you whether the setup is ready or what to fix.
- **You switched the model that does the grading:** It compares new results against a fixed reference set so you know whether old and new scores are comparable.
- **A test passed an answer that was clearly wrong:** It traces the false pass to its cause, logs it and routes the fix to whoever owns that part.
- **You want to compare two prompt versions fairly:** It confirms the prompt evaluation can reliably tell the two versions apart before you pick a winner.

## What you get

- A readiness verdict of pass, flag or fail, with the evidence behind it
- Results from planted mistakes showing what the grader caught and missed
- Drift checks against a fixed reference set after any grading change
- A log of wrong results with their cause and who should fix them
- A short report covering the grader's limits and the next step

## Example: Support-reply grader readiness review

Example with a fictional company. Names, people and figures are invented to show the agent's output. Any resemblance to a real company or person is unintended.

What it was asked: Before our grader decides whether a cheaper model can answer our customer emails, tell us whether we can trust its scores.

I compared the grader's pass or fail on 120 past replies against the support team's own marks, planted 12 deliberate mistakes to see which it caught, re-judged 50 pairs in swapped order, and checked the test set for overlap with the assistant's instructions. I did not look at live customer traffic, the assistant's own reply quality or model costs, and I changed nothing: a person on the team approves the pass bars and makes every fix. Run date: 30 September 2026.

### What it found

- The grader lets through 9 of 30 replies the support team marked as bad, mostly refund and date mistakes, even though overall accuracy looks high at 87.5%.
- A single judge decides pass or fail, and it changed its pick in 11 of 50 pairs when the two replies were shown in the opposite order.
- 14 of the 120 test cases also appear in the assistant's own instructions, which lifts the score; on the other 106 cases accuracy is 85.8%.

## Guardrails

- It only reads the prompts, agents and tests it is checking. It does not change them.
- Critical safety labels and thresholds, benchmark lock-ins, live adoption of a change and published results all need a person's approval. It prepares the evidence.
- It reports what it measured on your own test data and states the limits of each review, rather than quoting accuracy figures it has not checked.

## Frequently asked questions

### What is AI evaluation testing?

It means grading the answers of your AI prompts or agents with an automated test. The catch is that the test can be wrong too. The Evaluator in Ai1 by MyZone AI checks the checker before you rely on its scores. A flawed test may pass bad answers, fail good ones, change its mind when answers are shown in a different order, or score higher because test cases leaked into the prompt. The Evaluator looks for exactly these problems.

### How do you compare two prompt versions fairly?

First make sure the test can tell them apart. The Evaluator checks that the data is split properly and the grader is independent, looks for test cases that leaked into the prompt, and confirms the prompt evaluation reliably separates the two versions before you pick a winner. It does not write the prompts. It tells the Prompt Engineer what the tests need and where they fall short, and the Prompt Engineer writes them.

### How can I tell whether an AI test is reliable?

Plant mistakes and see whether it catches them, compare its marks with your team's own, and check it does not change its mind when answers are shown in a different order. The Evaluator runs these checks and reports what it measured on your own test data, rather than quoting accuracy figures it has not checked. It only reads the prompts and agents it is testing, and a person on your team makes any fix.

### Is the Evaluator available today, and can it approve changes on its own?

Yes. The Evaluator is ready to use in Ai1 by MyZone AI. It cannot approve changes. Critical safety labels and thresholds, locking in a benchmark, adopting a change live and publishing results all need a person's approval; it supplies the evidence for that decision.

## About Ai1

Ai1 is the AI operations platform by MyZone AI, where each client runs on its own private server. The Evaluator is included on every Ai1 level, including Developer Core, with no per-agent charge. It works alongside the other Ai1 agents on your account.

Pricing: https://myzone.ai/pages/services/ai1-pricing. Security: https://myzone.ai/pages/security.
