Whitepaper

Introducing consumer twins:
a reliable decision simulation infrastructure

How interview-grounded twins reached ~82% behavioral accuracy in held-out testing, and where teams can trust them today.

Abstract

Research, brand, product, and strategy teams are making thousands of deliberate choices today at an unprecedented pace. A pace, which cannot support the depth it needs. This depth is the backbone of great products that we build. Depth, which is produced by understanding the whys behind the whats of your consumers. 

However, these methods carry persistent constraints: access to the right participants is slow and expensive; niche populations can be essentially unreachable; and many questions, especially those involving second‑order effects or network dynamics, are simply not testable in the wild without unacceptable cost or risk. Experiments often take months, face operational or ethical limits, or require extrapolation from contexts that only imperfectly resemble the decision at hand. As a result, even the best operators are forced to make pivotal calls with partial views of their customers and markets.

Consumer twins are designed to solve this. They are digital twins of your target consumers, grounded in real interviews, diaries, and related data, then tested against held-out answers. This allows operators to access their consumers 24/7, and run multiple “what if” scenarios. The scenarios can range from concept testing, messaging, journey optimization, pricing, to anything in the realm of innovation.   

These simulations can therefore be treated as a new modality to support decision-making capability. Something like a reusable library of consumer twins, which calibrates periodically.

This paper offers a pragmatic playbook for adopting that capability. It begins with a plain-language overview of what consumer twins are and how they differ from traditional research artifacts and synthetic personas. It then summarizes held-out validation results that help leaders calibrate expectations about accuracy, strength, and limits. The heart of the piece is an operating guide: how to deploy twins today, where human oversight remains essential, and how the capability can mature over time through longitudinal simulation and twin-to-twin interaction.

Metric

Metric

Simulation accuracy against actual responses

Simulation accuracy against actual responses

LLM judge

LLM judge

81.6%

81.6%

BERTScore F1

BERTScore F1

83.3%

83.3%

Qualitative Alignment Score (QAS)

Qualitative Alignment Score (QAS)

71.2%

71.2%

Exact-match on available discrete / ordered probes

Exact-match on available discrete / ordered probes

100% where applicable

100% where applicable

What is a consumer twin?

A consumer twin is a digital twin, calibrated on a target consumer’s qualitative data, in the form of in-depth interviews, diary studies, ethnographies, surveys, and other behavioral data.

These twins are designed to preserve the behavioral actions of the said participants across their choices, decisions, and thoughts. These actions are classified into categories such as decision-making, economics, identity, culture, social influence, and more. It essentially mirrors the behavior of the participant, and is validated against held-out responses.

So, when you simulate a scenario, the twins respond under these circumstances and behavioral aspects.

Where consumer twins help the most

Modern research, product, brand, and strategy teams operate under a familiar constraint. Every primary question comes up with the next most important follow-on questions often immediately:

  • What if packaging changed?

  • What if we forced a convenience-versus-price tradeoff?

  • What if the offer was framed differently for different sub-segments?

  • What if we stress-tested this message?

Most of those questions are too important for improvisation and too narrow to justify another research cycle. It becomes more prohibitive as decision cycles compress and audiences become harder to reach consistently. Consumer simulations address these questions with high reliability.

Why synthetic personas fall short

Synthetic personas and related methods solve for speed. They make it easier to hypothesize and validate concept ideas, explore narrative versions, and generate early direction. However, they come with a lack of control over training data, and therefore introduce generic biases. They also limit validation of research. Without participant-level grounding, it is difficult to know where the answer is coming from, and how it sits in the question.

What teams need is a brand-specific and governed layer: a way to keep using real consumer data, with enough grounding, calibration, and measurement, to support decision-grade insights.

Capability

Capability

Primary Research

Primary Research

Synthetic Personas

Synthetic Personas

Consumer Twins

Consumer Twins

Continuity of research

Continuity of research

low

low

high

high

high

high

Controlled calibration data

Controlled calibration data

high

high

low

low

high

high

Speed of research

Speed of research

low

low

high

high

high

high

Emotional nuance

Emotional nuance

high

high

low

low

medium

medium

Behavioral depth

Behavioral depth

high

high

low

low

medium-high

medium-high

Ability to validate

Ability to validate

high

high

low

low

high

high

Grounding, fidelity, and validation of consumer twins

Grounding happens with real participant data

Participant data (relevant to the brand and use case), is collected in the form of in-depth interviews, diary studies, ethnographies, and surveys. They are then encoded into behavioral frameworks, specific to market research and brand’s context.

This grounding provides specificity over stereotyping. For example, in this paper’s study, the difference between participant responses on packaging vs convenience is as follows:

  1. Participant A sorts hard plastics from soft plastics using local council guidance.

  2. Participant B fluctuates plastic usage under travel constraints.

  3. Participant C prefers plastic when food waste feels worse than plastic harm.

  4. Participant D's uncertainty and lack of experience take over

These are all variations (and real) of a specific target consumer’s many behaviors, also observed in traditional research methods.

Behavioral fidelity over plausible narratives 

The goal of consumer twins is to not sound human, in an agreeable, and nice way. The goal is to duplicate actual behavior and scale it with statistical significance. 

These twins always trace their answers back to reasoning, alongside evidence to support the same. This structure introduces reliability, otherwise missing in synthetics and other similar research modalities. Also, since they are tied to real participants at the root, they can be validated across methods such as held-out testing, making it measurable and improvable. 

In practice, this means:

  1. Teams can introduce contextual data in a controlled, governed way

  2. Operators can ask for evidence, confidence intervals, and validation measurements.

  3. These simulations can be customized, and improved as they evolve.

Consumer behavior simulation: a case study

We adapted the evaluation methodology in this research paper called Interview-Informed Generative Agents for Product Discovery: A Validation Study (Wang & Siu, 2026)

The core question: given a research study with in-depth interviews, diary studies, ethnographies, and secondary data, can the calibrated consumer twin predict the target persona’s behaviour?

Study setting: an overview

Factor

Factor

Details

Details

Primary asset

Primary asset

A household study to understand plastic packaging preferences and usage for food and groceries

A household study to understand plastic packaging preferences and usage for food and groceries

Primary asset

Primary asset

24 participants from different households (demographic details are undisclosed for privacy) 

24 participants from different households (demographic details are undisclosed for privacy) 

Reuse mechanism

Reuse mechanism

In-depth interviews, diary studies, field notes

In-depth interviews, diary studies, field notes

Reuse mechanism

Reuse mechanism

In-depth interviews, diary studies, field notes (except final interview)

In-depth interviews, diary studies, field notes (except final interview)

Reuse mechanism

Reuse mechanism

Final interview 

Final interview 

Simulation evaluation: metrics we used

Method

Method

Metric

Metric

What it contains

What it contains

LLM judge

LLM judge

a prompt based accuracy analysis (qualitatively)

a prompt based accuracy analysis (qualitatively)

Did the twin recover the person behaviorally?

Did the twin recover the person behaviorally?

Discreet agreement

Discreet agreement

Exact match (yes/no)

Exact match (yes/no)

Did the twin make the same core choice?

Did the twin make the same core choice?

Semantic similarity

Semantic similarity

BERTScore F1

BERTScore F1

Are these answers paraphrased closely? 

Are these answers paraphrased closely? 

Open-ended fidelity

Open-ended fidelity

Qualitative Alignment Score (QAS)

Qualitative Alignment Score (QAS)

Which behavioral dimensions are transferred? 

Which behavioral dimensions are transferred? 

Definitions:

  1. Behavioral dimensions applied: constraint navigation, trade-offs, emotional intensity, comparative judgement, social influence. 

  2. LLM judge: a prompt based accuracy analysis of the qualitative match of the simulated vs actual answer

  3. BERTScore F1: a semantic similarity metric which looks at paraphrasing similarity. However, it misses the behavioral dimension applied by a real human while making these decisions. 

  4. QAS: a LLM-based qualitative alignment score, across these metrics - stance alignment, explanation alignment, topic coverage, voice / tone similarity, uncertainty calibration, and specificity preservation.

Finally, the study treats “LLM judge” as the primary evaluation metric.

The results: how the consumer simulations fared

Below is a table with illustrative examples of twins generated from participant A, B, C, and D, and their results from the held-out interview testing.

Twin

Twin

Probes

Probes

LLM judge

LLM judge

BERTScore F1

BERTScore F1

QAS

QAS

Strongest segment

Strongest segment

Weakest segment

Weakest segment

Twin A

Twin A

9

9

86%

86%

84.6%

84.6%

74%

74%

trade-offs

trade-offs

comparative judgement

comparative judgement

Twin B

Twin B

8

8

82%

82%

82.2%

82.2%

73.3%

73.3%

constraint navigation

constraint navigation

social influence

social influence

Twin C

Twin C

5

5

81%

81%

83.1%

83.1%

76.5%

76.5%

emotional intensity

emotional intensity

trade-offs

trade-offs

Twin D

Twin D

9

9

71%

71%

83.1%

83.1%

60.8%

60.8%

trade-offs

trade-offs

comparative judgement

comparative judgement

Cohort

Cohort

81.6%

81.6%

83.3%

83.3%

71.2%

71.2%

discussed below

discussed below

discussed below

discussed below

Following patterns emerged:

  1. The twins usually talk about the right topic: BERTScore is measuring whether the twin’s answer and the real person’s answer are about the same things. A score above 80 means the twin rarely goes completely off-topic. If someone is talking about packing, recycling, or food waste, the twin usually answers in that same territory. 

  2. How well the twin recovers the actual person varies more: LLM judge and QAS are closer to human judgment. They ask: did the twin get the stance, the reasoning, the tone, the uncertainty, and the detail right? These scores swing more, between 71% to 86%, proving that twins preserve the individuality of research.

  3. Twins are better at some kinds of questions than others: qualitative answers perform better on constraint navigation and tradeoffs over emotional and social influence. These can however be further improved with second-order simulations.

Behavioral segment level analysis

Below is a table with a more focused breakdown of how the twins perform on a behavioral level. It is important to measure these, as they form the backbone of decision making.

Twin

Twin

LLM judge

LLM judge

BERTScore F1

BERTScore F1

QAS

QAS

Remark

Remark

Constraint navigation

Constraint navigation

83.1%

83.1%

85.5%

85.5%

81.5%

81.5%

Good at simulating decision frameworks

Good at simulating decision frameworks

Trade-offs

Trade-offs

89.6%

89.6%

82.6%

82.6%

79.9%

79.9%

Reliable at choice questions 

Reliable at choice questions 

Emotional intensity

Emotional intensity

83.5%

83.5%

81.9%

81.9%

74.2%

74.2%

Gets that something matters, not how strongly they felt it

Gets that something matters, not how strongly they felt it

Comparative judgement 

Comparative judgement 

70.1%

70.1%

82.4%

82.4%

63.0%

63.0%

Can only give directional ranking

Can only give directional ranking

Social influence 

Social influence 

82.4%

82.4%

82.8%

82.8%

68.8%

68.8%

The number of instances were limited for inference 

The number of instances were limited for inference 

What the simulations do well:

  1. They reliably predict how a target consumer cohort usually decides, such as sorting rules, travel compromises, food-waste norms, and preferences for fresh food over reheated leftovers, in this research study’s case. 

  2. They also handle everyday tradeoffs well, including convenience versus sustainability and effort versus payoff.

  3. Twins often get comparative preferences right, such as which option feels better or which compromise is least bad. The accuracy is especially good when the answer is a clear yes/no or an ordered stance.

  4. They are also able to say no, and do not exhibit agreeable behavior, a common occurrence in synthetic personas.

These strengths make twins useful for scenarios, such as concept variant choices, claim comparisons, and messaging triage.

Simulations are less reliable in the following cases:

  1. They are slightly more confident than an actual human. For example, in some cases, a preference is simulated strongly, which actually came with hesitation. 

  2. They can soften emotion into mild concern, which can hamper the intensity of choices and actions.

  3. They cannot (and should not) simulate lived-in experiences. They can tell you how a consumer would feel about a taste change, but cannot experience it.

Consumer behavior simulation vs synthetic personas

Below is the comparison of the same held-out interview held on grounded consumer twins vs synthetic personas generated by demographic LLM prompt:

Metric

Metric

Grounded twin

Grounded twin

Synthetic persona

Synthetic persona

LLM judge

LLM judge

82.1%

82.1%

66.5%

66.5%

BERTScore F1

BERTScore F1

82.2%

82.2%

82.9%

82.9%

QAS

QAS

73.3%

73.3%

59.5%

59.5%

Results show that grounded twins show stronger qualitative accuracy, while semantic similarity alone can make synthetic answers look closer than they are. 

How KikiLabs deploys consumer twins

Below is an overview of how we calibrate consumer twins for reliable simulations:

  1. We collect behavioral and use case specific data via in-depth interviews, run using our AI moderator. The moderator is trained on behavioral frameworks, to extract deeper insights. 

  2. These interviews are encoded into our proprietary behavioral framework engine. This encoding forms the base of the twin's grounding. 

  3. Then, simulations are run across follow-on scenarios, and validated against actual human responses. 

  4. Every simulated scenario comes with reasoning, evidence, and confidence intervals.

  5. Every twin’s calibration is updated as market metrics such as category, product, or audience shifts.

In an overview, grounded simulations can reliably help pressure-test packaging, claims, and concepts earlier, research hard-to-reach respondents, and surface decision logic with more precision. The best fit is structured, high-signal questions where the team needs a fast directional answer to funnel research topics before committing to downstream methods.

Looking ahead

We are working on improving the accuracy of these twins across all relevant behavioral segments. Extensions such as, including longitudinal effects and twin-to-twin interaction, remain on the roadmap and should follow as we build ahead.

How we handle data governance

Because consumer twins are built from real participant material, governance is part of the method. Here’s how we ensure it:

  1. Use only consented data, with clear permission for research usage.

  2. Mask or remove PII that is not required for the research purpose.

  3. Do not use participant or client materials for general AI training.

  4. Keep brand data isolated, time-bound, access-controlled, and auditable.

Conclusion

Consumer behavior simulation is an emerging modality for reliable research. 

They are grounded in real participant data, already strong on behavioral prediction, and improvable through measurable validation. They enable using the brand’s contextual historical and evolving data.

For organizations that need to learn faster than the market changes, it shifts both the timeline and cost curve. The question for leaders is no longer whether to explore this capability. It is where to start, what to trust today, and how to build a program that lasts.