The Calibration Trap in Synthetic Research
When a vendor tells you a synthetic persona is calibrated against your own research, hear what has actually been said. Your blind spots have been encoded, and they will now be reproduced faster, at scale, in complete sentences.
That is the calibration trap. It is the least discussed risk in this category and the most likely to cost you a decision.
The uncomfortable part is that calibration is the step that makes everyone comfortable. It is the word that gets the model past your procurement team and into a brand plan.
What is calibration, and why does it feel like rigor?
Adjusting a model's outputs until they resemble distributions you already have.
In practice that means your survey history, your tracking study, your segmentation, your syndicated category research. The vendor tunes the simulation until it produces numbers in the neighborhood of numbers you recognize. Then they show you the match.
Watch what that demonstration actually proves. It proves the model can reproduce your existing data. That is a statement about agreement, not about accuracy, and the two look identical on a slide.
Here is the circularity. You validate the model against your research. The model was tuned to match your research. Of course it matches.
Agreement with your own history is the one result the process guarantees in advance. Treating it as evidence is like marking your own exam with the answer sheet you wrote.
What exactly gets inherited?
Four things, and none of them are visible in the output.
Your sampling bias. If your research reaches the physicians who agree to be reached, the model learns that population and presents it as the market. The clinicians who never answer your surveys do not become visible because a model started generating them.
Your question design. Every survey encodes assumptions about what matters, so a model trained on your instrument inherits your framing and cannot surface the thing you never thought to ask about.
Your measurement error. If your satisfaction scores were not predicting commercial outcome before, a model grounded in them will produce confident numbers with the same predictive failure built in. Precision goes up, accuracy does not move.
Your institutional preference for comfortable answers. This one is cultural and it is the worst of the four. Research that contradicted leadership tended not to survive to the archive, so the archive leans optimistic, and now the model does too.
Each of these existed before you bought anything. What simulation changes is the speed and the confidence, which is a meaningful difference. A wrong assumption in a quarterly tracker is slow. A wrong assumption in a system that answers any question in four seconds is everywhere by Friday.
Why is pharma's historic data an especially bad anchor?
Because a lot of what you have been collecting was never predictive in the first place.
Look at the instrument that produced most of your archive. Bain's own published material says net promoter differences "explain anywhere from 10% to 70% of the variation in subsequent revenue growth rates". A range that wide is not a calibration target. It is an admission that the relationship varies enormously by context, and your context is the hardest one.
Bain made the limit explicit elsewhere, finding that goodwill is "a necessary but insufficient condition for generating revenue growth." Calibrate a model to the necessary but insufficient part and you get a necessary but insufficient model.
Your market makes it worse in three specific ways. Your customer is three people with different interests, so an aggregate sentiment score averages parties who disagree. Your patient voice is mostly missing by law, so the archive is thin exactly where the risk is highest. Your value event happens months after the data you collected, so the historic file rarely contains the outcome worth predicting.
I have set out why borrowed instruments break here in does NPS work with physicians and why consumer CX does not translate to pharma.
How would you know it has happened to you?
Four signs, and you can check all four this week.
The model never surprises you. Everything it returns is something somebody in the room already believed. Agreement is the symptom, not the proof.
It has no opinion about your known failures. Ask it where patients abandon your therapy and see whether it finds the seam your own field force complains about. If it cannot name a problem you already know is real, it does not know your market.
It is confident about things your data never covered. A model grounded in physician research will still answer a question about payer committee politics, fluently, with nothing underneath it.
Its error rate has never been published. Not a fidelity score the vendor computed, a comparison against realized outcomes. That protocol is set out in how to validate a synthetic audience.
What should you calibrate against instead?
Behavior that cost somebody something, and outcomes that exist independently of your research function.
Replace sentiment anchors with progression anchors wherever you can.
Script to start conversion. Time to therapy. Persistence at six and twelve months. Barrier resolution speed. Those are observable, they have a price attached, and nobody filled in a form to produce them.
There is plenty of external truth to anchor against. A 2026 JAMA study summarized by Johns Hopkins found insurer rejections reached 40.7% of initial brand name attempts in 2024. Of those rejected scripts, 48.4% were never followed by any fill in that class within 90 days. Calibrate to that and you are anchored to something the market did, not something it said.
Use your frontline as the correction layer. Your field force, hub agents and patient support staff observe failures that never reach a survey, which makes them the cheapest independent check you own. That route is in voice of the frontline.
Then hold the whole thing against the measures in how to measure customer experience in pharma. A model calibrated to a scorecard that does not move money will produce advice that does not move money.
Does this mean you should not calibrate?
No. It means calibrate to outcomes rather than to opinions, and keep a record of what you calibrated to.
Three rules make it workable.
Write down the anchor. Whatever the model was tuned against goes in a document that travels with every output, so a reader knows which archive is speaking.
Weight observed data above stated data in the anchor set. When both exist, behavior wins, and your vendor should be able to tell you the ratio.
Hold back a cohort the model has never seen and never will. One clean holdout, protected, used only for scoring. The moment it becomes training material you have lost your only independent read.
Do all three and calibration goes back to being what the word implies, which is an instrument checked against reality rather than against your own filing cabinet.
Why does this matter more than it sounds?
Because your industry has already demonstrated it can believe a comfortable number for years.
Deloitte's 2025 research found that only 28% of HCPs believe pharma's engagement strategies meet their needs, against 82% of life sciences executives who say they are satisfied with those same strategies. That 54 point spread was produced by human beings reading their own research.
Now imagine that archive turned into a system that answers instantly and never hesitates. The 82% view gets automated, and the 28% view has no route in at all. Nobody notices, because the output still sounds like research.
The layer that would have told you the truth is the one simulation structurally cannot reach, which is Predisposition. The whole argument this sits inside is the future of the pharma commercial model, and the discipline that keeps listening honest is Customer Excellence.
Calibration is not the enemy. Calibrating to yourself is.
Key takeaways
- Calibrating a synthetic persona against your own research guarantees agreement with your archive, which is not the same as accuracy.
- Four things get inherited invisibly: your sampling bias, your question design, your measurement error and your preference for comfortable answers.
- Pharma's archive is a poor anchor because much of it came from sentiment instruments that never predicted commercial outcome.
- If the model never surprises you and has no view on failures you already know are real, it has learned your file rather than your market.
- Anchor to behavior with a cost attached, record the anchor in writing, and protect one cohort the model will never see.
Questions to ask your research team
- What exactly was our model calibrated against, and is that written down anywhere a reader of the output would find it?
- In the anchor set, what is the ratio of observed behavior to stated opinion?
- Which cohort have we protected from the model, and who is enforcing that?
- Name one thing the simulation has told us that nobody in the room already believed.
- If our old satisfaction data was not predicting revenue, what are we expecting a model trained on it to do?
About the author
Wayne Simmons is the founder of The Customer Excellence AGENCY and the author of The Customer Excellence Enterprise (Wiley, 2024). He is founding faculty of the MS in Customer Experience Management at Michigan State University's Broad College of Business. He led global customer excellence in Pfizer's first Chief Marketing Organization and in Bayer's Customer Powerhouse. Related reading: How to validate a synthetic audience, What is a synthetic persona? and Voice of the customer in pharma, three ways to hear







