Skip to content
EN
Open app

Test and evaluate your agent

4 min read
Copy as Markdown
View as MarkdownOpen the raw .md file.md

A test scenario is a call you want the agent to handle right: who calls, what they want and what a good outcome is. Voice Logica runs it as a real phone call. A test caller phones your agent and plays the scenario, your agent answers with its real voice, speech recognition, tools, transfers and workflows, and a judge grades the call against what you expected.

You can do everything on the agent’s Tests tab, or from Claude or ChatGPT (connect your assistant). Each step below names the tool.

  1. Generate drafts - scenarios_query (action "generate_agent_scenarios") with a focus: prompt, transfer, tools or knowledge. Drafts are not saved; keep the ones that match your real calls.

  2. Save the ones that matter - scenarios_execute (action "create_agent_scenario"). Cover:

    • each of the top reasons people call;
    • an out-of-hours call;
    • a question the agent cannot answer;
    • a caller who spells a name or reads out a number;
    • a caller who asks for a person;
    • an off-topic caller.

    For each scenario fill in:

    • Scenario - the situation the caller creates.
    • Ideal outcome - what a good call ends with.
    • Success criteria - a checklist the judge grades one by one, e.g. “The agent asks for the caller’s name”. All must pass.
    • Expected tool calls - the tools that must fire, optionally with the values they must receive.
    • Expected transfer - whether the agent should transfer, and where.
    • Caller - persona, the facts they can give when asked, and behaviour: cooperative, confused, interrupts, changes mind, silent or noisy.
    • Caller IDs - the number the test call comes from. Enter your own number (or a customer’s number from your CRM) and the agent recognises the caller exactly as on a real call: contact, integrations and earlier calls. Leave it empty to call from the default test line. Several numbers make one call per number. From a chat this is the scenario’s callerId / callerIds.
  3. Mock the tools that change real data - list them in the scenario’s mocked tools with the response they should return, or turn off use real tools to block every tool you did not list.

    • Appointment booking tools cannot be mocked yet and book for real; use a test calendar or cancel the booking afterwards.
    • Set test date/time (frozenNow) when the right answer depends on the date, such as available slots or office hours.
  1. Run one scenario - Run on the Tests tab, or run_agent_scenario. Live is the default. One request makes at most 3 repetitions. To test a staged change, pass its agentVersionId.

  2. Read the verdict - scenarios_query (action "get_agent_scenario_results"). Each result has pass or fail, a score, the result of every success criterion, the tools that fired with their results, and suggested improvements. Open the call itself to listen to the recording.

  3. Fix and run again - change the prompt with edit_agent_prompt or the knowledge, or apply the suggestions with scenarios_execute (action "apply_scenario_improvements"; prompt changes go to staging). Run the same scenario again.

  4. Compare versions - scenarios_execute (action "run_agent_scenario_matrix") runs a scenario across several versions, caller numbers and repetitions, so you can compare a staged change with the live agent or measure a pass rate.

  1. Run all - Run all on the Tests tab, or scenarios_execute (action "run_agent_scenario_suite") with the staged version. It runs every scenario of the agent, one live call each.

  2. Approve the cost - from a chat, the first call only returns the estimated minutes; nothing runs until you approve and it is called again with confirmCost: true.

  3. Publish what passed - versions_execute (action "publish_agent_version") with the suite’s testSuiteId.

  4. Make passing tests required (optional) - on the Tests tab turn on Require passing tests to publish. A version that has never been live is then published only after a test run that started after its last change passed completely. Rolling back to a version that was live before is never blocked.

mode: "text" simulates the conversation without a phone call. It shows what the agent says and decides, but tools, transfers and workflows do not run, so a “missed tool call” there is not a finding. Use it to try prompt wording quickly; use live runs to decide what to publish.

To talk to the agent yourself, use test_in_browser (voice or chat in the browser) or call_me (the agent calls your verified mobile).

When a real call goes wrong, read it with get_calls (include "aiDialogue" shows the prompt, tool calls and tool responses of that call), fix the cause, and save that call as a new scenario. The suite then guards against it coming back.

The full build checklist is in Build a reliable phone agent.