# Test and evaluate your agent

A test scenario is a call you want the agent to handle right: who calls, what they want and what a good outcome is. Voice Logica runs it as a **real phone call**. A test caller phones your agent and plays the scenario, your agent answers with its real voice, speech recognition, tools, transfers and workflows, and a judge grades the call against what you expected.

You can do everything on the agent's **Tests** tab, or from Claude or ChatGPT ([connect your assistant](https://docs.voicelogica.ai/integrations/claude-chatgpt/)). Each step below names the tool.

> **Test calls are real calls:** - Each live run uses call minutes from the company plan, about the length of the call.
> - Tools run for real unless you mock them (step 3). A test that books, orders or opens a ticket does it for real.
> - Transfers are graded but never connect, so nobody's phone rings.
## Write the scenarios

1. **Generate drafts** - `scenarios_query` (action `"generate_agent_scenarios"`) with a focus: `prompt`, `transfer`, `tools` or `knowledge`. Drafts are not saved; keep the ones that match your real calls.

2. **Save the ones that matter** - `scenarios_execute` (action `"create_agent_scenario"`). Cover:
   - each of the top reasons people call;
   - an out-of-hours call;
   - a question the agent cannot answer;
   - a caller who spells a name or reads out a number;
   - a caller who asks for a person;
   - an off-topic caller.

   For each scenario fill in:
   - **Scenario** - the situation the caller creates.
   - **Ideal outcome** - what a good call ends with.
   - **Success criteria** - a checklist the judge grades one by one, e.g. "The agent asks for the caller's name". All must pass.
   - **Expected tool calls** - the tools that must fire, optionally with the values they must receive.
   - **Expected transfer** - whether the agent should transfer, and where.
   - **Caller** - persona, the facts they can give when asked, and behaviour: cooperative, confused, interrupts, changes mind, silent or noisy.
   - **Caller IDs** - the number the test call comes from. Enter your own number (or a customer's number from your CRM) and the agent recognises the caller exactly as on a real call: contact, integrations and earlier calls. Leave it empty to call from the default test line. Several numbers make one call per number. From a chat this is the scenario's `callerId` / `callerIds`.

3. **Mock the tools that change real data** - list them in the scenario's **mocked tools** with the response they should return, or turn off **use real tools** to block every tool you did not list.
   - Appointment booking tools cannot be mocked yet and book for real; use a test calendar or cancel the booking afterwards.
   - Set **test date/time** (`frozenNow`) when the right answer depends on the date, such as available slots or office hours.

## Run and evaluate

1. **Run one scenario** - **Run** on the Tests tab, or `run_agent_scenario`. Live is the default. One request makes at most 3 repetitions. To test a staged change, pass its `agentVersionId`.

2. **Read the verdict** - `scenarios_query` (action `"get_agent_scenario_results"`). Each result has pass or fail, a score, the result of every success criterion, the tools that fired with their results, and suggested improvements. Open the call itself to listen to the recording.

3. **Fix and run again** - change the prompt with `edit_agent_prompt` or the knowledge, or apply the suggestions with `scenarios_execute` (action `"apply_scenario_improvements"`; prompt changes go to staging). Run the same scenario again.

4. **Compare versions** - `scenarios_execute` (action `"run_agent_scenario_matrix"`) runs a scenario across several versions, caller numbers and repetitions, so you can compare a staged change with the live agent or measure a pass rate.

## Run the whole suite before you publish

1. **Run all** - **Run all** on the Tests tab, or `scenarios_execute` (action `"run_agent_scenario_suite"`) with the staged version. It runs every scenario of the agent, one live call each.

2. **Approve the cost** - from a chat, the first call only returns the estimated minutes; nothing runs until you approve and it is called again with `confirmCost: true`.

3. **Publish what passed** - `versions_execute` (action `"publish_agent_version"`) with the suite's `testSuiteId`.

4. **Make passing tests required (optional)** - on the Tests tab turn on **Require passing tests to publish**. A version that has never been live is then published only after a test run that started after its last change passed completely. Rolling back to a version that was live before is never blocked.

## Quick checks without a call

`mode: "text"` simulates the conversation without a phone call. It shows what the agent says and decides, but tools, transfers and workflows do not run, so a "missed tool call" there is not a finding. Use it to try prompt wording quickly; use live runs to decide what to publish.

To talk to the agent yourself, use `test_in_browser` (voice or chat in the browser) or `call_me` (the agent calls your verified mobile).

## Keep the suite growing

When a real call goes wrong, read it with `get_calls` (include `"aiDialogue"` shows the prompt, tool calls and tool responses of that call), fix the cause, and save that call as a new scenario. The suite then guards against it coming back.

The full build checklist is in [Build a reliable phone agent](https://docs.voicelogica.ai/agent-configuration/build-a-reliable-agent/).
