Step 1 — separate facts from phrasing

Read the reply and mark every factual claim: order numbers, tracking IDs, refund amounts, delivery dates, policy statements. LLMs produce these confidently, so the first audit question is "can this be verified?" Any number or date the system cannot confirm is a candidate hallucination.

With ReplyAuditor you paste the reply and it flags unverifiable figures (order #, tracking code, amount) as a hallucination check, so you start the review knowing where the risk is.

Step 2 — scan for PII before it leaves

A support reply should almost never contain credentials, national IDs, full card numbers, or other personal data. Scan the draft for any of these. If a customer pasted sensitive data into a ticket, the draft must not echo it back.

ReplyAuditor runs a PII leak check that flags credentials, IDs, and card data in the reply. This is the fastest way to catch a copy-paste slip before the message ships.

Step 3 — check tone and policy rules

Decide the rules that matter for your brand: greet by name, no guarantees on delivery dates, never ask for a password, no refunds promised without approval. Then read the draft against that list — or, in ReplyAuditor, paste the rules (one per line) and let the brand/tone check evaluate them.

Tone drift is subtle: a reply can be factually fine and still read as cold, blunt, or off-brand. A written rule list makes the check repeatable instead of mood-dependent.

Step 4 — read the ship/hold verdict

If you use a tool, the verdict compresses the checks into one signal: ship or hold. "Hold" means at least one check failed and a human should look. "Ship" means it passed the gate's checks — not that it is guaranteed flawless.

Either way, a human makes the final call. The verdict is a prompt for judgment, not a substitute.

Step 5 — log the decision

Record what was checked and the outcome. A log turns a one-off review into something you can sample, train on, and defend later. For a team, this is also where a QA dashboard earns its place — seeing which check fails most often tells you where the AI drafting needs tuning.

An honest limit

A reply audit reduces obvious failure modes and makes them visible. It does not guarantee a reply is accurate or compliant, and it is not a replacement for human review or a legal sign-off. The team stays responsible for what goes out.

References

  • GDPR — personal data in customer communications: https://gdpr.eu/