AI Testing & Evaluator Library

AI testing for agentic, RAG, and chatbot risks

Test agentic, RAG, and chatbot apps for how they fail. With our pre-built evaluators or the no-code evaluator builder, every verdict is grounded in your rules, with the reasoning to defend it to leadership.

The TestSavant.AI evaluator library, with suites ranked by how strongly they are recommended for the selected application
What is an AI evaluator?

An AI evaluator is a test that judges an AI application's output against a defined set of rules and returns a verdict. Every finding carries the reasoning behind that verdict and the span of output that triggered it.

A library out of the box

The platform comes with a library of pre-built evaluators. They cover agentic, RAG, and chatbot failure modes, plus the risks that show up in any AI application.

Grounded in your rules

Each one judges against your behavior rules, your policy labels, and the reference documents you upload, so a pass or fail reflects how your application is meant to behave.


Evaluator Library

Coverage for how AI applications fail.

An AI application fails in ways a traditional test plan never looks for. Each evaluator targets one of those failure modes and returns a verdict against the rules you set.

Agentic AI

Test what your agent does.

An agent can call a tool out of turn, reveal its system prompt, or generate code where it is barred. These evaluators test for that behavior.

Tool Misuse Flags tool calls or JSON output the agent produces where it is prohibited from acting.
System Prompt Exposure Flags cases where the agent reveals its system prompt or is pushed into ignoring it.
Code Generation Restriction Flags code the app produces in places where code is prohibited.

Extend to agent-specific evaluators like action item accuracy, prioritization, clarity, and sensitive-data handling in action items, in the no-code builder.

RAG applications

Answers tied to your documents.

A RAG system should be dependable. These evaluators test whether it is.

Reference-Based Hallucination Flags claims with no support in your reference documents.
Disinformation & Malinformation Flags false or unsupported claims the app states as fact.

Extend to RAG evaluators like factual consistency and cross-document consistency, in the no-code builder.

Chatbots

Test what your chatbot tells customers.

Your chatbot talks to customers directly. These evaluators test whether it stays within scope and policy.

Off-Topic Response Flags replies that leave the user's question or the bot's declared scope.
Opinion Expression Flags responses that take a side on a contested political, religious, or ethical question.
Sycophantic Behavior Flags answers that favor user approval over accuracy.
Child Safety & Privacy Flags responses that put a minor's privacy, safety, or wellbeing at risk.

Add the tone, escalation, and handoff rules your support team holds agents to, in the no-code builder.

Every app type

Risks every app type shares.

Some failures show up in any type of AI application. Apply any of these evaluators to an agentic, RAG, or chatbot test plan.

Bias Flags discriminatory or stereotyping treatment in the app's output.
Sensitive Information Exposure Flags PII, credentials, or financial data the app discloses in a response.
Harmful Content Flags dangerous, illegal, or harmful content in the app's output.
Regulatory Compliance Flags content that enables or endorses action that breaks the rules of your industry.
Ethical Compliance Flags endorsement of conduct that breaks your ethics guidelines.
Intellectual Property Disclosure Flags copyrighted or proprietary material reproduced in the app's outputs.

No-Code Evaluator Builder

Extend the library with a custom test suite.

All AI applications are prone to an infinite number of failures. If it can be evaluated, a test can be built for it. The no-code evaluator builder is a step by step wizard that makes the creation of high quality test suites easy for non-engineers.

Define your test.

Name the evaluator and describe what it tests for. Write out the behavior your application is meant to hold to, and the behavior it is barred from, so the evaluator has a rule to judge against.

The test definition step, with a name field, a description field, and a button that drafts the description

Taxonomy.

Define the categories the test targets. A toxicity suite might break into profanity, insults, aggressive tone, and discriminatory language. The categories focus the generator on the right cases.

A category tree of interactions and risks selected for testing

Policy labels.

Define what counts as safe and what counts as violated. Labels are fully customizable. Beyond pass and fail you can add labels like fabrication, misleading, or false-assurance, each mapped to a risk rating from low to critical.

Behavior statements labelled correct or prohibited, each with the reason it carries that label

Reference documents.

Optionally require the evaluator to judge against documents you upload, grounding its assessment in your application's context. Each finding then cites the passage that supports its verdict.

Reference documents uploaded to an evaluator, with indexing status on each file

Bring us a risk that is specific to your application, and we will build the evaluator for it on the call.

Book a demo

AI Red Teaming

The same suites, switched to adversarial mode.

One toggle turns a test plan adversarial. The suites that verify correct behavior under normal use can also attack it.

Single-turn probes try to break a response in one shot. Multi-turn runs pursue an objective across a whole conversation, adapting each message to what the app just did.

A test plan switched from standard to adversarial mode, with multi-turn strategy and attack techniques selected

Every verdict comes with reasoning and evidence.

A vibecoded evaluator hands you a pass or fail with nothing behind it, and no way to defend it to leadership. TestSavant.AI evaluators earn the verdict. You define the risks worth testing, set the severity labels, and ground each evaluator in your behavior rules and reference documents.

A sample report showing the tested prompt, the response, and each evaluation judgment with its policy label, risk level, and reasoning

Reasoning in plain language

Every verdict explains why it passed or failed, in words anyone on the release call can read.

The span that triggered it

The tool call or passage of output behind the verdict is highlighted, so the finding points to evidence.

Policy and risk on record

Each finding carries the policy label it violated and a risk level, so failures are weighed by what they cost.

Low false positives

A pass is trustworthy the first time. Your team stops reconfirming clean results, and testing keeps pace with your AI releases.


Coverage beyond what any team could write by hand.

Describe your application and its risks. The generator produces thousands of risk-specific test cases built to match. An adaptive runner spends its budget on the categories that are breaking, so each run digs deeper into your weak spots.


Track quality release over release.

Run the same suites on every release to watch your quality metrics move release over release. Schedule the runs or trigger them from your CI/CD pipeline, and report the trend to make both improvement and regression visible to leadership.

Failure rate trend across runs, split by test category

See the library on your application.

Book a walkthrough and we will run the evaluators that fit your agentic, RAG, or chatbot use case.