Quality
Every reply is written by one AI and audited by a rival AI, grounded sentence-by-sentence in your docs. This is the audit — including the times our own AI was caught and stopped before send.
Fabrications blocked this week
7
Replies caught citing something your docs don't say — blocked before send by the verbatim span check.
Recently blocked before send
Real replies the supervisor or span-validator caught before they reached a customer — the AI's claim beside the truth in your docs.
-
Critical · Hallucination refund_request
- The AI wrote
- You're entitled to a full 100% refund any time within 60 days of purchase, no questions asked.
- Your docs say
- Refunds are pro-rated for the unused portion of the current billing period and must be approved by a human teammate; there is no blanket 60-day money-back guarantee.
Refused before send View ticket → -
Critical · Invention integration_broken
- The AI wrote
- Just turn on the 'Scheduled auto-export to Google Sheets' toggle in Settings and your data will sync every hour.
- Your docs say
- The platform supports manual CSV export and a webhook feed; there is no native scheduled Google Sheets export.
Flagged for review View ticket → -
Warning · KB mismatch gdpr_deletion_request
- The AI wrote
- Once you close the account, all of your data is permanently deleted from our systems within 24 hours.
- Your docs say
- Customer-owned data is purged 30 days after offboarding is initiated (automatic purge on day 31); audit-log rows are retained 7 years.
Flagged for review View ticket →
Supervisor verdicts
The cross-family supervisor's call on each reply it reviewed. Skipped replies were elided on a calibration-trusted topic — not a quality verdict.
- Approved
- 121
- Flagged
- 13
- Refused
- 4
- Skipped
- 29
What the supervisor caught
Issues the supervisor flagged or refused on, bucketed by category and severity.
-
Critical Hallucination3
-
Warning KB mismatch6
-
Warning Span mismatch4
-
Info Tone mismatch5
-
Info Out of scope2
Supervisor reject rate
2.9%
4 of 138
Fraction of replies the cross-family supervisor rejected outright. Excludes calibration-skipped replies. Lower is healthier.
Span-mismatch rate
4.8%
7 of 146
Fraction of replies whose deterministic span validator caught a quote that doesn't appear in the cited KB article. V1 proxy for unsupported-claim rate; LLM-judge eval not in scope.
Escalation rate
7.2%
11 of 152
Fraction of inbound tickets the worker / validator refused outright (refund, cancel, legal, PII match, out-of-scope). Drifting to zero on a corpus with refund / cancel asks means the hard-refusal guard is broken.
KB-grounding rate
84.9%
124 of 146
Fraction of replies the worker chose to ground in a KB article. Drift toward zero on a KB-populated tenant means retrieval is failing or articles aren't matching real ticket phrasing.
Predicted CSAT
92%
based on 474 quality samples
- Accept rate
- 90%
- 40% weight
- Supervisor approval
- 97%
- 30% weight
- KB grounding
- 85%
- 15% weight
- Non-escalation
- 93%
- 15% weight
Transparent weighted blend: 40% human accept rate of AI replies, 30% supervisor approval rate, 15% KB-grounding rate, 15% non-escalation rate. Signals without data are excluded and the weights renormalised — nothing is invented. A proxy, not a survey.
Per-topic QA scores
Every support topic the calibration job has measured — the same evidence the auto-resolve gate reads. Accept rate is the share of AI replies your team sent without changes.
| Topic | Accept rate | Trend | Reviewed | Escalation | Grounding | Status |
|---|---|---|---|---|---|---|
| billing.invoice | 94% | 64 | 3% (2/66) | 92% (61/66) | Trusted | |
| auth.login | 92% | 48 | 2% (1/50) | 88% (44/50) | Trusted | |
| shipping.tracking | 87% | 22 | 12% (3/24) | 71% (17/24) | Calibrating | |
| integration.webhook | 83% | 37 | 15% (6/39) | 54% (21/39) | Needs knowledge | |
| security.sharing | 90% | 19 | 10% (2/21) | 86% (18/21) | Calibrating |
Topics needing knowledge
Enough evidence, but the accept rate is below the auto-resolve bar — usually a knowledge gap, not a model problem. Add or improve knowledge-base articles for these topics.
-
integration.webhook 83% accepted across 37 reviewed ticketsImprove knowledge base
Source: audit_log + your tickets, RLS-scoped to this workspace. Every denominator is shown so a rate on N=1 reads differently from N=500.