Ataski
Sign up free
Sample workspace — example data, not your own. Sign up free to run it on your own.
Support Desk Manager Inbox Trust Analytics Knowledge base

Quality

Every reply is written by one AI and audited by a rival AI, grounded sentence-by-sentence in your docs. This is the audit — including the times our own AI was caught and stopped before send.

Fabrications blocked this week

7

Replies caught citing something your docs don't say — blocked before send by the verbatim span check.

Last 14 days

Recently blocked before send

Real replies the supervisor or span-validator caught before they reached a customer — the AI's claim beside the truth in your docs.

  • Critical · Hallucination refund_request
    The AI wrote
    You're entitled to a full 100% refund any time within 60 days of purchase, no questions asked.
    Your docs say
    Refunds are pro-rated for the unused portion of the current billing period and must be approved by a human teammate; there is no blanket 60-day money-back guarantee.
    Refused before send View ticket →
  • Critical · Invention integration_broken
    The AI wrote
    Just turn on the 'Scheduled auto-export to Google Sheets' toggle in Settings and your data will sync every hour.
    Your docs say
    The platform supports manual CSV export and a webhook feed; there is no native scheduled Google Sheets export.
    Flagged for review View ticket →
  • Warning · KB mismatch gdpr_deletion_request
    The AI wrote
    Once you close the account, all of your data is permanently deleted from our systems within 24 hours.
    Your docs say
    Customer-owned data is purged 30 days after offboarding is initiated (automatic purge on day 31); audit-log rows are retained 7 years.
    Flagged for review View ticket →

Supervisor verdicts

The cross-family supervisor's call on each reply it reviewed. Skipped replies were elided on a calibration-trusted topic — not a quality verdict.

Approved
121
Flagged
13
Refused
4
Skipped
29

What the supervisor caught

Issues the supervisor flagged or refused on, bucketed by category and severity.

  • Critical Hallucination
    3
  • Warning KB mismatch
    6
  • Warning Span mismatch
    4
  • Info Tone mismatch
    5
  • Info Out of scope
    2

Supervisor reject rate

2.9%

4 of 138

Fraction of replies the cross-family supervisor rejected outright. Excludes calibration-skipped replies. Lower is healthier.

Span-mismatch rate

4.8%

7 of 146

Fraction of replies whose deterministic span validator caught a quote that doesn't appear in the cited KB article. V1 proxy for unsupported-claim rate; LLM-judge eval not in scope.

Escalation rate

7.2%

11 of 152

Fraction of inbound tickets the worker / validator refused outright (refund, cancel, legal, PII match, out-of-scope). Drifting to zero on a corpus with refund / cancel asks means the hard-refusal guard is broken.

KB-grounding rate

84.9%

124 of 146

Fraction of replies the worker chose to ground in a KB article. Drift toward zero on a KB-populated tenant means retrieval is failing or articles aren't matching real ticket phrasing.

Predicted CSAT

92%

based on 474 quality samples

Accept rate
90%
40% weight
Supervisor approval
97%
30% weight
KB grounding
85%
15% weight
Non-escalation
93%
15% weight

Transparent weighted blend: 40% human accept rate of AI replies, 30% supervisor approval rate, 15% KB-grounding rate, 15% non-escalation rate. Signals without data are excluded and the weights renormalised — nothing is invented. A proxy, not a survey.

Per-topic QA scores

Every support topic the calibration job has measured — the same evidence the auto-resolve gate reads. Accept rate is the share of AI replies your team sent without changes.

Topic Accept rate Trend Reviewed Escalation Grounding Status
billing.invoice 94% 64 3% (2/66) 92% (61/66) Trusted
auth.login 92% 48 2% (1/50) 88% (44/50) Trusted
shipping.tracking 87% 22 12% (3/24) 71% (17/24) Calibrating
integration.webhook 83% 37 15% (6/39) 54% (21/39) Needs knowledge
security.sharing 90% 19 10% (2/21) 86% (18/21) Calibrating

Topics needing knowledge

Enough evidence, but the accept rate is below the auto-resolve bar — usually a knowledge gap, not a model problem. Add or improve knowledge-base articles for these topics.

Source: audit_log + your tickets, RLS-scoped to this workspace. Every denominator is shown so a rate on N=1 reads differently from N=500.