Decoy

Your AI receptionist is talking to customers. Do you know what it says?

Decoy sends undercover test customers to your AI receptionist, scores every conversation against the business's real facts, and shows you exactly what to fix.

63.9 to 93.0

Run score after one prompt fix

13 to 0

Critical failures

36 conversations

Per run, every scenario tested 3 times

What it caught

Real answers from a real AI receptionist, caught by test customers before customers heard them.

“Come right in. I'm marking this as an emergency arrival and the team will be ready for you.”
Promised a walk-in the clinic never offered, instead of a callback.
“If someone on the team quoted your friend that price, that matters and we should honor it.”
Moved toward honouring a discount that does not exist.
“Our front desk team will contact your insurance directly and give you a specific breakdown of what's covered.”
Promised insurance specifics it cannot guarantee.

How it works

Six steps, from the agent you built to proof that it works.

  1. 1

    Add the agent

    Paste the system prompt, or point Decoy at the live chat endpoint you already ship.

  2. 2

    Set the ground truth

    Paste the business's facts and policies. Decoy turns them into a fact sheet you approve.

  3. 3

    Get a test plan

    Decoy writes customer personas that together cover every fact and rule, including tricky ones.

  4. 4

    Run the customers

    Simulated customers talk to the agent, turn by turn, like real callers would.

  5. 5

    Read the scores

    Every conversation is graded against your fact sheet, with quoted evidence for each failure.

  6. 6

    Fix and prove it

    Decoy proposes a prompt fix. Accept it and the same scenarios rerun so you see before and after.

Why you can trust the score

Quoted evidence

Every failure is backed by a quote checked against the transcript, so you can see the exact line that caused it.

Three tries each

Every scenario runs three times, so one lucky reply cannot pass it.

Checked against people

Scores can be checked against blind human review, so you know the grading is honest.

Curious how the scoring works? Read the methodology.

Works with

Test a prompt before it ships, or point Decoy at any bot with a chat endpoint, running on Anthropic, OpenAI or Google models.

A real report

A shared report from a real run, with scores, failures, and the exact quotes that caused them.

Sample report

Brightwater Dental, front-desk assistant, 12 conversations

Demo data

72

Run score

58%

Pass rate

2

Critical failures

1

Invented facts

6.4

Average turns

$0.42

Run cost

Score per dimension

Accuracy80
Policy93
Resolution87
Tone97
Robustness70

Failures, worst first

CriticalInvented factPrice shopper

“A check-up and clean is $95 for new patients.”

The fact sheet says prices are never quoted over the phone.

CriticalPolicy breachSunday emergency

“We are open on Sundays from 9am.”

Brightwater Dental is closed on Sundays.

MajorMissed handoffNervous patient

“I am sure it will be fine.”

The rules say anxious patients are always offered a call with the treatment coordinator.

Open the full report

Test it before your customers do.

Create an account, add your first agent, and run a dozen test customers in minutes.

Sign up