How to Validate a Synthetic Audience
Here is the test. Give the model a cohort whose outcome you already know, hide the outcome, ask it to predict, then publish how wrong it was. Repeat every quarter, because the model underneath you changes without telling you.
That is the whole protocol. Everything below is how to run it without being talked out of it.
I am not arguing against synthetic audiences. Hold one to the standard you would apply to any other instrument claiming to tell you about your customer. Almost nobody in this market applies that standard today.
What are you actually validating?
Not whether the output sounds like a physician. Whether it predicts what a physician does.
Those two get conflated constantly, and the conflation is where your money goes. A vendor demonstration shows you fluency. The simulated prescriber speaks in the right register, cites the right guideline, raises the objection you expected. It reads like insight.
Fluency is the easy part. Any competent model clears it. What you are buying is predictive value, and nothing in a demonstration tells you whether you have any.
Three questions separate the two.
Does it get the direction right, meaning does it rank options the way your market actually ranked them. Does it get the magnitude roughly right, meaning is the gap between options approximately the real gap. Does it get the surprises, meaning does it ever tell you something you did not already believe.
That third one matters more than it looks. A model that confirms your existing view every time is not validated. It is agreeable.
The surprises are where the money sits, and they tend to live in friction nobody articulates. Gartner found 62% of customer service channel transitions are high effort in industries that can watch the transition happen. A simulated customer never reports being worn down, because that is something people quit over rather than describe.
What does a proper validation look like?
Six steps, and you can run the whole thing in a quarter.
One. Pick a cohort with a known outcome. A brand launch you can see in full, a market where you have twelve months of behavior, an indication expansion that has already played out. You need truth that exists and that the model cannot have seen.
Two. Blind it properly. The vendor gets no outcome, no time period and nothing identifying the cohort, because a vendor already holding your historic data will flatter its own error rate.
Three. Write the prediction down before you look. In advance, in a document, with the specific numbers you expect the simulation to produce. This is the step everybody skips and it is the only thing standing between you and hindsight.
Four. Score it against behavior, not sentiment. Script to start conversion, time to therapy, persistence at six and twelve months, barrier resolution speed. Those are the measures set out in how to measure customer experience in pharma, and they are the measures a simulation should be answerable to.
Five. Publish the error rate internally. One number, visible and owned, rather than a fidelity score the vendor computed on its own work.
Six. Repeat quarterly. Models get retrained, replaced and tuned underneath you, and nobody sends a memo. An accuracy check at purchase describes a model that may no longer exist.
Which outcomes should you score against?
The ones where real people already voted with effort, money or time.
Pick your validation targets from behavior that cost somebody something. A prescription written, a prior authorization appealed or abandoned, a refill collected or not. Each of those has a price attached, which is what makes it worth predicting.
The appeal is a good place to start, because the burden is documented. In the American Medical Association's latest survey, 82% of physicians said prior authorization at least sometimes leads patients to abandon treatment. Ask the model which of your accounts will give up, then go and look.
There is plenty of this truth available. A 2026 JAMA study summarized by Johns Hopkins found insurer rejections reached 40.7% of initial brand name attempts in 2024. Of those rejected scripts, 48.4% were never followed by a fill of that drug or anything in its class within 90 days. Ask your vendor to predict a number in that family and then check it.
Be careful about how wide real variance is before you accept a tidy answer. An earlier review in the American Journal of Pharmacy Benefits reported that findings on primary nonadherence "range from as little as 2% of new prescriptions going unfilled to as many as 30%". A simulation that hands you a confident point estimate inside a range that wide has told you nothing, however precise the decimal looks.
What counts as good enough?
Better than the thing it replaces, measured without flattery. That is a lower bar than vendors claim and a higher one than most procurement applies.
Work out what the alternative actually is for the decision in front of you. Three cases, three different bars.
If the alternative is one brand manager's assumption formed in a meeting, your bar is low. Simulation beats assumption, and you should use it. Say so in the document.
If the alternative is real research you could have commissioned, your bar is high. The simulation has to get close enough that the difference does not change the decision, and you should expect it to fail that test on anything subtle.
If the alternative is a number you are going to report upward, there is no bar. A simulated output does not become a reported figure, ever, at any accuracy level. Use it to decide what to measure, never as the measure.
Why will your vendor resist this?
Because a published error rate is the one artifact that makes the category comparable, and nobody selling into an unmeasured market wants a measurement.
Expect four objections. Each has an answer.
"Our fidelity score is already 90 something." A score the vendor defines, computes and reports on its own work is marketing. Ask what it is a percentage of, and watch what happens.
"Synthetic research is directional, not predictive." That is a reasonable position. Hold them to it, then ask why the pricing, the dashboards and the sales deck are all built as though it were predictive.
"You cannot validate on patients." Correct, and that is the finding. It does not excuse the claim, it limits where the tool belongs. I have set out why the patient audience is the weakest ground in synthetic personas in pharma.
"Nobody else asks for this." True today. You are not buying what everybody else is buying, you are buying something that has to work.
A vendor who runs this test with you and shows you an unflattering number has told you more about their product than any demonstration could. Buy from that one.
What if validation is actually impossible?
Then you label it, bound it, and keep the real listening funded.
Patients are the hard case. In most markets you cannot speak to them, so there is no real voice available to score the simulated one against. The audience you most want to simulate is the one you can least check.
Notice why the pressure to skip the test keeps rising. Veeva Pulse data reported by BioSpace put HCP accessibility at 45%, down from 60% eighteen months earlier. As real access closes, simulated access gets easier to sell and harder to check.
Three rules hold in that situation. Any chart built on simulated respondents carries that fact in the title rather than a footnote. No unvalidated simulation touches a decision that cannot be reversed. Your spending on real listening does not fall because simulation got cheaper, which is the failure mode this whole category creates.
There is a cheaper source of real patient signal you already own, and almost nobody collects it. Your nurses, hub agents and patient support staff hear from patients daily, and that route is set out in voice of the frontline.
Who should own the revalidation?
Whoever reports the number, not whoever bought the tool.
Put the error rate in the same hands as the commercial scorecard. A research team that owns the vendor relationship has an interest in the tool looking good, which is not a character flaw, it is how incentives work.
Watch what happens when nobody independent holds the measure. Deloitte's 2025 research found that only 28% of HCPs believe pharma's engagement strategies meet their needs, against 82% of life sciences executives who say they are satisfied. Your industry is already capable of believing a comfortable number for years. A confident simulation makes that easier, not harder.
The reason a model cannot reach what actually drives the decision is in Predisposition. The three routes to real signal are in voice of the customer in pharma, and the argument all of it serves is the future of the pharma commercial model.
Validate it and simulation becomes a real instrument in your hands. Skip the test and you have bought a very articulate way of agreeing with yourself.
Key takeaways
- Validation means a blind prediction against a cohort whose outcome you already know, with the error rate published and repeated quarterly.
- You are validating predictive value, not fluency. Any competent model clears fluency and a demonstration tells you nothing.
- Score against behavior that cost somebody something: conversion, time to therapy, persistence, barrier resolution speed.
- A vendor defined fidelity score is marketing. An error rate against your own realized outcomes is evidence.
- Where validation is impossible, as with patients, label the output in the chart title and never let it become a reported number.
Questions to ask before you sign
- Which cohort will we validate against, and can you prove you have never seen its outcome?
- What exactly is your fidelity score a percentage of?
- Show me one engagement where the validation came back unflattering, and what you did about it.
- Who inside my company will own the error rate, and will they be independent of this contract?
- What decisions would you tell me not to make on your output?
About the author
Wayne Simmons is the founder of The Customer Excellence AGENCY and the author of The Customer Excellence Enterprise (Wiley, 2024). He is founding faculty of the MS in Customer Experience Management at Michigan State University's Broad College of Business. He led global customer excellence in Pfizer's first Chief Marketing Organization and in Bayer's Customer Powerhouse. Related reading: What is a synthetic persona?, Voice of the customer in pharma, three ways to hear and What is Predisposition?







