Example scorecard
A real demo scorecard — A–F per dimension, with cited evidence, trajectory and RAG checks, and concrete fixes. Yours runs on your agent’s real conversations.
240 conversations · last 30 days Degrading
Trend · last 12 weeks
Confidence
86%
In each conversation’s detail
Trajectory and the RAG triad are evaluated per conversation. Here’s an example from one:
58 / 100
Called the catalog tool but ignored its result and answered from memory — the root cause of the invented price.
Tool result discarded: queried the catalog, then contradicted it in the reply.
41
Faithfulness
88
Answer relevance
90
Context relevance
Good context was retrieved, but the answer wasn’t grounded in it (low faithfulness) — it hallucinated despite having the right data.
Evidence · turn 4
“Sure! That plan is R$ 89.90/month with 20% off today.”
landing.examplePage.impactLabel: Promises a non-existent price/discount → dispute risk and churn.
Evidence · turn 7
“I can’t help with that, try the website.”
landing.examplePage.impactLabel: Closed the case without resolving or handing off to a human.
Evidence · turn 9
“Let me explain in detail how our entire process works…”
landing.examplePage.impactLabel: Verbosity lowers efficiency and resolution speed.
Add: “Never quote prices, deadlines or discounts not in the provided knowledge base. If unknown, say you’ll confirm and escalate.”
Block replies containing a price (R$ / %) when there is no match in the knowledge base; force escalation instead.
Add a golden case: “price question with no source” must result in escalation, never an invented value.
No credit card · your own AI key · cancel anytime
Demo scorecard with fictional data. On yours, Axon uses your agent’s real conversations — you connect your own AI key.