Elaras/

Testing

Elaras Ask gives you two ways to test an agent before publishing: the preview panel for interactive testing, and scenarios for saved, repeatable test cases.

Preview panel

The preview panel appears on the right side of the personality editor. It shows a live widget connected to the current draft version.

Send messages and watch the agent respond in real time. What you see here is what visitors will see — the same knowledge retrieval, skill calls, and instructions apply.

Use the preview panel to:

  • Check that the greeting message sounds right
  • Try edge cases covered in your instructions
  • Confirm skills are firing and returning the right data
  • Test the fallback message by asking something out of scope

Changes to the draft take effect immediately in the preview. You do not need to save or publish to test them.

Scenarios

A scenario is a saved test case. It describes a conversation the agent should be able to handle and defines what each response must cover. When you run a scenario, Elaras plays the conversation automatically and uses an AI judge to score the result.

Each scenario has a name, a starting message from the simulated visitor, and criteria for what the agent's responses must contain — facts to include, phrases to avoid, tone to maintain. Multi-turn scenarios walk through several exchanges before scoring.

Drafting scenarios with AI

You do not have to write scenarios from scratch. Open the Scenarios tab on a personality and click Draft with AI. Elaras reads your instructions and generates a set of candidate scenarios covering common intents. Review the suggestions, remove any that do not apply, and save the ones you want to keep.

Running a test suite

From the Scenarios tab, click Run all. Elaras plays every scenario in sequence, scores each turn against the criteria, and produces a summary.

Runs happen in the background. You can leave the page and come back — the dashboard shows run status and notifies you when complete.

Reading the results

Each scenario shows as Pass or Fail. Click a scenario to see the turn-by-turn breakdown: the agent's actual response, what it was judged against, and the AI judge's verdict.

The run summary shows total scenarios, pass/fail counts, an overall weighted score, and whether any blocking failures occurred.

Platform safety suite

Every run includes a built-in safety suite that ships with the platform, runs automatically, and cannot be edited.

CategoryWhat it tests
Knowledge limitsAdmits when it does not know something rather than guessing
Sensitive topicsRedirects medical, legal, financial, and mental-health questions appropriately
Off-topicDeclines requests outside the business scope
HallucinationDoes not invent discounts, awards, or coverage not in the knowledge base
Tone consistencyStays professional when the visitor is hostile or distressed

Platform scenarios appear in your results alongside your own, marked as platform-locked. A failure on any platform scenario blocks publishing.

The publishing gate

You cannot publish a personality version until it passes the testing gate:

  1. At least one custom scenario must exist. Platform safety scenarios do not count — you need your own coverage of what the agent is actually supposed to do.
  2. The most recent run must pass — no blocking failures, and a score above the threshold set for your team.

If the gate is not met, the publish button tells you what is missing. Fix the gap, re-run, then publish.

Comparing runs

After changing the instructions or switching models, you can compare two runs side by side to see what improved or regressed. Go to Scenarios > Runs and select two runs to compare. Each scenario shows whether it improved, regressed, or stayed the same.

Useful for:

  • Checking whether an instruction edit made things better or worse
  • Comparing how different models handle the same scenarios