What Is a Synthetic Persona? The Definition for Pharma
A synthetic persona is a model generated stand in for a customer, built from existing research and public data, used to produce simulated answers to questions you would otherwise put to a real person.
That is the whole thing. Everything else is implementation detail, and most of the confusion in this market comes from vendors describing the detail instead of the definition.
This page sets out what they are, how they are built, and what they are worth in each of your three pharma audiences. It also covers where they fail, and the test any of them should pass before touching a commercial decision.
What is a synthetic persona, precisely?
Three things have to be true, or what you are being sold is something else.
It stands in for a specific customer type rather than a general audience. It generates responses rather than retrieving them. Its outputs are then treated as if they were research, which is the part that creates all the risk.
The related terms are used loosely, so here is the vocabulary as it is actually used.
A synthetic respondent is one simulated individual answering a question. A synthetic panel or synthetic audience is many of them queried together to produce something resembling survey results. An AI persona usually means the same thing with the emphasis on characterization rather than sampling. A digital twin is a different idea borrowed from engineering and applied loosely here, normally meaning a simulated individual built from that specific person's own data rather than from a segment. In the research literature the practice of using a language model to stand in for survey respondents is sometimes called silicon sampling.
What none of them are is a measurement. A synthetic persona produces an estimate of what a category of person might say. That is a hypothesis, and treating it as a finding is the single most common error you will encounter in this field.
How are they actually built?
Three layers, and the quality of the whole thing is set by the first.
Source material. Your prior survey research, claims or behavioral data, public literature, transcripts, and whatever the vendor has in its own library. This defines what the persona can know.
The model. A large language model that generates plausible responses in character. This defines how the persona talks.
Calibration. Adjusting outputs so they resemble known distributions, usually by benchmarking against the client's historic data or syndicated category research. This defines how confident everybody feels, which is not the same as how accurate it is.
Notice the dependency. A synthetic persona cannot know anything your source material does not contain. It can only recombine, extrapolate and phrase. When a vendor says a persona is grounded in your research, read that as inheriting your research, including everything your research got wrong.
Are they equally useful across your three audiences?
Not remotely, and this is the distinction nobody selling them makes.
Your market has three customers. A physician chooses, a payer covers, a patient uses. Simulation fidelity differs so much across the three that treating them as one capability is the first mistake.
Physicians, medium fidelity and checkable. There is abundant public material in a clinician's professional register, from literature to guidelines to conference commentary. A model can produce a credible HCP voice. More importantly you can check it, because you have real prescribing behavior to score the simulation against. This is the audience where synthetic work is most defensible.
Patients, low fidelity and unverifiable. This is where the pitch is strongest and the ground is weakest. You usually cannot speak to patients, so there is no real voice available to validate the simulated one. The audience you most want to simulate is the one you can least check, and a confident simulation you cannot falsify is not evidence. It is a well written assumption.
The behavior you would want it to predict is also severe. A 2026 JAMA study summarized by Johns Hopkins found insurer rejections reached 40.7% of initial brand name attempts in 2024. Of those rejected scripts, 48.4% were never followed by a fill of that drug or anything in its class within 90 days. No simulated patient told anyone that was coming.
Payers, lowest fidelity of the three. Formulary decisions run on contracts, rebate economics, net price and committee politics. Almost none of that is in the text a model learned from, and the parts that matter most are confidential by design. A simulated payer will produce an articulate account of clinical value and miss the actual decision rule entirely.
So the honest ranking is physician first, patient a distant second with disclosure, payer last and mostly not worth doing.
Worth remembering why the temptation is strong. Veeva Pulse data reported by BioSpace put HCP accessibility at 45%, down from 60% eighteen months earlier. As real access closes, simulated access gets easier to sell.
What are they legitimately good for?
Four jobs, all of them upstream of real research rather than instead of it.
Replacing an assumption. At the design stage you are not choosing between a synthetic persona and real data. You are choosing between a persona and one brand manager's opinion, formed in a meeting, unchallenged. Simulation beats opinion. That is a genuine gain and the strongest case for the whole category.
Generating hypotheses to test. Run a hundred simulated reactions, find the three objections you had not anticipated, then go and check those three with real people. You have used the simulation to aim the expensive instrument.
Pressure testing a design before you build it. Walk a simulated patient through your access workflow and you will find the step where the instructions make no sense. That is a usability review, and it is a reasonable thing to automate.
Reaching what you cannot reach in time. A rare disease population, a market where recruitment takes four months, a decision due in three weeks. Simulation here is a considered compromise rather than a shortcut, provided you label it as one.
Where do they fail?
Four failure modes, and the first is structural rather than fixable.
They model the rationale, not the decision. A language model learns how people explain choices, not how people make them. Your physician decides on context, habit, trust and the difficult patient she saw on Tuesday, then produces a tidy reason when asked. The model reproduces the reason. What you receive is a fluent account of why someone would do something they will not actually do.
That gap matters more in your market than in most, because predisposition rather than stated preference is what moves prescribing. Simulated reasoning is a good model of the part of the customer that does not decide.
The calibration trap. Ground a persona in your own survey history and you have encoded your existing blind spots at speed. If your sentiment data was not predicting commercial outcome before, a model trained on it will reproduce that failure faster and with more confidence.
Homogenization. Models regress toward the center of their training distribution. Your outliers, your frustrated minority, your early signal, all the people who actually tell you something new, are precisely the voices a simulation smooths away.
Confident error. A real respondent hesitates, contradicts herself and says she does not know. A simulated one answers everything, immediately, in complete sentences. Fluency reads as reliability to a human audience and the two are unrelated.
Effort research outside pharma shows how much is hiding in the friction a tidy answer skips over. Gartner found 62% of customer service channel transitions are high effort, in industries where the company can watch it happen. A simulated customer never reports being worn down, because being worn down is not a thing you articulate, it is a thing you quit over.
How should you validate one?
Three tests, and anything that cannot pass them does not belong in a decision.
Provenance. Can the signal be traced to something a human did, rather than something a model inferred about what a human might say? Know which you are holding.
Predisposition or rationale. Does it capture what moved the person, or the story told afterwards? A synthetic respondent can only produce the latter, which caps what it can ever be worth.
Predictive hold. Run the simulation against a cohort whose outcome you already know, blind. Compare its prediction to realized script to start conversion, time to therapy and persistence. Publish the error rate. Repeat quarterly, because model behavior drifts.
That third test is the one the industry is skipping. A fidelity score a vendor defines, computes and reports on its own work is a marketing number. An error rate against your own outcomes is evidence. Insist on the second before you buy.
What rules should govern its use?
Four, and they are not onerous.
Never let a synthetic output become a reported number. Use it to decide what to measure, never as the measure.
Label it everywhere it appears. Any chart built on simulated respondents carries that fact in the title, not in a footnote, so that a decision made on it is made knowingly.
Validate before adoption and on a schedule after it. One accuracy check at purchase tells you about a model that no longer exists. Models are updated, retrained and replaced underneath you, and nobody sends a memo.
Keep real listening funded. The strongest argument for simulation is that it is cheaper than research, which is also how a company talks itself out of the research it needs.
Deloitte's 2025 research found that only 28% of HCPs believe pharma's engagement strategies meet their needs, against 82% of life sciences executives who say they are satisfied. You will not simulate your way across a 54 point perception gap.
The routes that produce real signal are set out in voice of the customer in pharma. The cheapest real signal you already own is covered in voice of the frontline.
So should you use them?
Yes, for the four jobs above, with the three tests applied and the four rules in place.
The position is simple. Simulation beats assumption and loses to evidence. Deploy it where the alternative is a guess. Keep it away from anything you report. Hold it to the same standard as any other instrument claiming to tell you about your customer.
That standard is the point. The measures it should ultimately be scored against are in how to measure customer experience in pharma, and the discipline that connects what you hear to what you realize is Customer Excellence.
Key takeaways
- A synthetic persona is a model generated stand in for a customer, producing simulated answers from existing research and public data.
- Fidelity differs enormously by audience. Physicians are checkable, patients are unverifiable, payers run on economics a text model never saw.
- They model the rationale rather than the decision, which caps their value in a market driven by predisposition.
- Grounding a persona in your own survey history encodes your existing blind spots and reproduces them faster.
- A vendor defined fidelity score is marketing. An error rate against your own realized outcomes is evidence.
Questions to ask before you buy one
- Which of my three audiences is this for, and what is your accuracy evidence for that audience specifically?
- What source material grounds it, and what happens to the output if that material was biased?
- Show me a blind holdout test against real outcomes, with the error rate published.
- How often is it revalidated, and who pays for that?
- What decisions would you tell me not to make on this?
About the author
Wayne Simmons is the founder of The Customer Excellence AGENCY and the author of The Customer Excellence Enterprise (Wiley, 2024). He is founding faculty of the MS in Customer Experience Management at Michigan State University's Broad College of Business. He led global customer excellence in Pfizer's first Chief Marketing Organization and in Bayer's Customer Powerhouse. Related reading: Voice of the customer in pharma, three ways to hear, Voice of the frontline and What is value leakage in pharma?







