Institutions regulated by Brazil's central bank (BACEN)

Compliance and quality of AI agents in regulated companies

Axon evaluates your AI agents' real conversations against your own internal policy and delivers a defensible dossier: a score per dimension with a margin of error, every finding backed by the literal excerpt that supports it, and a conclusion signed by a human being.

4-week diagnostic · evaluated on your own AI key (BYOK) · PII redacted before any processing
B+
Customer service · collections
Sample of 384 conversations · 5% margin of error
κ 0,74
Policy adherence
92
Duty to inform
88
Anti-hallucination
74
Evidence · turn 14

“The customer asked for a second copy of the invoice 3×, and the agent stated the charge had been cancelled, which the system does not show.”

The judge runs on your key, at the provider your team already approved
  • Anthropic
  • OpenAI
  • Google
  • Meta Llama
  • Mistral
  • DeepSeek
How it works

Auditable on the three points a regulator asks about

What it was evaluated against, what backs each finding, and who signed the conclusion.

  1. The rubric is your policyThe dimensions evaluated come from your internal policy, versioned.
  2. Every quote is verifiedThe literal excerpt is checked against the turn it came from.
  3. The conclusion is humanThe kappa between your reviewers and the judge ships with the result.

The rubric is your policy

The dimensions evaluated are written by hand with your risk team, versioned, and it is against them that each conversation is scored. No generic vendor criteria.

ADisclosed the total effective cost before closing#4821
CFailed to log the service ticket number#4822
FClaimed a fee waiver that does not exist#4823

Every finding cites the turn, and the quote is verified

The judge has to point to the literal excerpt that supports the problem. Axon checks whether that text really exists in the source conversation; a quote that doesn't match brings down the confidence of the evaluation.

Quote verified in turn 14

“…that charge has already been cancelled, sir, you can ignore the invoice.”

The conclusion is human, and the machine is measured

Your reviewers score part of the sample blind. We publish the agreement between them and the judge (Cohen's kappa) alongside the result: if the judge doesn't agree with your team, it shows.

κ 0.74substantial agreement

62 conversations labelled by 2 reviewers · every disagreement listed one by one in the dossier.

The diagnostic

A defensible snapshot of what your agents do today

A piece of work with a beginning, a middle and an end — not a subscription. At the end you hold a dossier that supports an answer to internal audit, the ombudsman or a regulator.

4 fronts

Scope

  • Rubric written with your team

    The dimensions evaluated come from your internal risk and compliance policy, not from a template of ours.

  • Statistical sample of the period

    Stratified by channel, agent and time window, with the confidence level and margin of error stated.

  • Evaluation on your AI key

    Every conversation in the sample goes through the LLM judge running on the key (BYOK) of the provider you already approved.

  • Blind calibration

    Your reviewers score part of the same sample without seeing the judge, and we measure the agreement with Cohen's kappa.

4 weeks

Timeline

  • Scope and rubric

    Sessions with risk and compliance to write the dimensions.

  • Ingestion and sampling

    Transcripts come in with PII redacted and the sample is drawn.

  • Evaluation and calibration

    The judge runs on your key and your reviewers label blind.

  • Dossier and presentation

    Delivery of the document and a joint read-through of the findings.

What we need from you: the transcripts for the period and two hours of one reviewer for the calibration.

5 parts

Deliverable

  • A score per dimension with a margin of error

    A stated interval on each dimension, not a bare number.

  • The literal excerpt in every finding

    Checked against the turn of the conversation it came from.

  • Run parameters on the record

    Sampling, judge model version and rubric version, so the result can be re-run.

  • Human-versus-judge calibration result

    The kappa, and every disagreement listed one by one.

  • A signed compliance conclusion

    By a person on your team. The model does not decide that.

See a full sample dossier

The platform

The dossier doesn't end in a PDF

Every score in the diagnostic stays navigable on the platform: your risk team signs in, drills from the agent's score down to the turn that produced the finding, and checks it without depending on us.

Agent health

Customer service · collections+1,2
Account opening+0,4
Debt renegotiation−2,8
Sample of 384 conversationsκ 0.74rubric v4
  • Health dashboard per agent

    Current score, 7- and 30-day trend, and which agents are getting worse, by channel.

  • Finding linked to the turn

    From any score you drill down to the conversation and see the exact excerpt that produced it, already verified.

  • Calibration and kappa

    Your reviewers label blind on the screen itself, and the agreement with the judge comes out calculated.

  • Versioned rubric

    Every change to a dimension or weight creates a version, and each evaluation points to the version that judged it.

  • Idempotent ingestion

    Transcripts by upload, REST API or typed SDK. Re-sending the same batch duplicates nothing.

  • PII redacted on the way in

    Tax IDs, card numbers, phones, emails and names are stripped before anything is stored and before the judge sees the text.

FAQ

Questions, answered.

What exactly does Axon prove?

That your AI agents do, or do not, behave according to your internal policy — with evidence. For every conversation evaluated there is a score per rubric dimension and, in each finding, the literal excerpt of the conversation that supports it. Observability shows latency, cost and volume; Axon evaluates the content of the interaction itself.

Who signs the compliance conclusion?

A person on your team, always. The LLM judge produces the score and the evidence; the compliance reading is human and is recorded as such in the dossier, with the name of whoever signed it. No model decides whether something is compliant.

Which AI key does the evaluation use?

Yours (BYOK). The judge runs on your institution's Anthropic, OpenAI or compatible key — the provider your team has already approved. Your data never passes through a key managed by Axon, never feeds third-party training, and the inference cost stays visible to you.

How is conversation data handled?

PII (emails, tax IDs, card numbers, phones, names) is redacted before anything is stored and before any text reaches the judge. The model only ever sees redacted text.

How long is data retained? (LGPD)

The retention window is set by you and applied automatically: conversations older than the limit are purged. At any moment the institution can request the erasure of all its data, as provided by Brazil's data protection law (LGPD).

What do we need to hand over to start?

The transcripts for the period (JSON, JSONL or CSV) — by upload, through the REST API or the typed SDK — and two hours of one reviewer from your team for the calibration. Ingestion is idempotent, so re-sending the same batch duplicates nothing.

How do we know the judge's score is trustworthy?

Three mechanisms: every quote is checked against the turn it came from; the confidence of the evaluation reflects the agreement between independent samples of the judge; and the agreement with your human reviewers is measured by Cohen's kappa and published alongside the result. If the judge disagrees with your team, the dossier shows it.

Find out what your agents do when nobody is watching.

A 30-minute conversation to define the scope. If it makes sense, the diagnostic starts the following week.

Already a customer? Sign in