The rubric is your policy
The dimensions evaluated are written by hand with your risk team, versioned, and it is against them that each conversation is scored. No generic vendor criteria.
Axon evaluates your AI agents' real conversations against your own internal policy and delivers a defensible dossier: a score per dimension with a margin of error, every finding backed by the literal excerpt that supports it, and a conclusion signed by a human being.
“The customer asked for a second copy of the invoice 3×, and the agent stated the charge had been cancelled, which the system does not show.”
What it was evaluated against, what backs each finding, and who signed the conclusion.
The dimensions evaluated are written by hand with your risk team, versioned, and it is against them that each conversation is scored. No generic vendor criteria.
The judge has to point to the literal excerpt that supports the problem. Axon checks whether that text really exists in the source conversation; a quote that doesn't match brings down the confidence of the evaluation.
“…that charge has already been cancelled, sir, you can ignore the invoice.”
Your reviewers score part of the sample blind. We publish the agreement between them and the judge (Cohen's kappa) alongside the result: if the judge doesn't agree with your team, it shows.
62 conversations labelled by 2 reviewers · every disagreement listed one by one in the dossier.
A piece of work with a beginning, a middle and an end — not a subscription. At the end you hold a dossier that supports an answer to internal audit, the ombudsman or a regulator.
The dimensions evaluated come from your internal risk and compliance policy, not from a template of ours.
Stratified by channel, agent and time window, with the confidence level and margin of error stated.
Every conversation in the sample goes through the LLM judge running on the key (BYOK) of the provider you already approved.
Your reviewers score part of the same sample without seeing the judge, and we measure the agreement with Cohen's kappa.
Sessions with risk and compliance to write the dimensions.
Transcripts come in with PII redacted and the sample is drawn.
The judge runs on your key and your reviewers label blind.
Delivery of the document and a joint read-through of the findings.
What we need from you: the transcripts for the period and two hours of one reviewer for the calibration.
A stated interval on each dimension, not a bare number.
Checked against the turn of the conversation it came from.
Sampling, judge model version and rubric version, so the result can be re-run.
The kappa, and every disagreement listed one by one.
By a person on your team. The model does not decide that.
Every score in the diagnostic stays navigable on the platform: your risk team signs in, drills from the agent's score down to the turn that produced the finding, and checks it without depending on us.
Current score, 7- and 30-day trend, and which agents are getting worse, by channel.
From any score you drill down to the conversation and see the exact excerpt that produced it, already verified.
Your reviewers label blind on the screen itself, and the agreement with the judge comes out calculated.
Every change to a dimension or weight creates a version, and each evaluation points to the version that judged it.
Transcripts by upload, REST API or typed SDK. Re-sending the same batch duplicates nothing.
Tax IDs, card numbers, phones, emails and names are stripped before anything is stored and before the judge sees the text.
That your AI agents do, or do not, behave according to your internal policy — with evidence. For every conversation evaluated there is a score per rubric dimension and, in each finding, the literal excerpt of the conversation that supports it. Observability shows latency, cost and volume; Axon evaluates the content of the interaction itself.
A person on your team, always. The LLM judge produces the score and the evidence; the compliance reading is human and is recorded as such in the dossier, with the name of whoever signed it. No model decides whether something is compliant.
Yours (BYOK). The judge runs on your institution's Anthropic, OpenAI or compatible key — the provider your team has already approved. Your data never passes through a key managed by Axon, never feeds third-party training, and the inference cost stays visible to you.
PII (emails, tax IDs, card numbers, phones, names) is redacted before anything is stored and before any text reaches the judge. The model only ever sees redacted text.
The retention window is set by you and applied automatically: conversations older than the limit are purged. At any moment the institution can request the erasure of all its data, as provided by Brazil's data protection law (LGPD).
The transcripts for the period (JSON, JSONL or CSV) — by upload, through the REST API or the typed SDK — and two hours of one reviewer from your team for the calibration. Ingestion is idempotent, so re-sending the same batch duplicates nothing.
Three mechanisms: every quote is checked against the turn it came from; the confidence of the evaluation reflects the agreement between independent samples of the judge; and the agreement with your human reviewers is measured by Cohen's kappa and published alongside the result. If the judge disagrees with your team, the dossier shows it.
A 30-minute conversation to define the scope. If it makes sense, the diagnostic starts the following week.
Already a customer? Sign in