Skip to content

How it works

The short version: we build a panel of believable people, put your product in front of them, and write up what they say. The longer version is below, including where we ladder past the surface answer and how we keep the whole thing honest.

01 · The panel

People who could argue with each other

A panel is only useful if the people in it actually differ. Each participant gets a real shape: an age and a job, a few lifestyle anchors in their own words, what they buy today, and the deeper value the category serves for them. One scores strictly, another gives the benefit of the doubt. That spread is the point. If everyone says four out of five, you have learned nothing.

A participant the system might recruit
Akiko, 32 · investment analyst, Tokyo

Reads the ingredient list before the marketing copy. Trusts a clinical trial more than an influencer. Keeps an anti-aging routine that earns her every step. Underneath all that, what she wants from skincare is to age on her own terms.

02 · The conversation

We keep asking why

In a focus group, a first answer is rarely the real one. Someone says they like the ten percent vitamin C. Fine, but why does that matter? Because it fades dark spots faster. And why does that matter? The moderator keeps pulling that thread until it reaches something a person actually cares about. Researchers call this laddering. It turns a feature into a reason you can act on.

One ladder from the vitamin-C study
What she noticed
"It's a high-strength vitamin C, ten percent."
What it does for her
"My dark spots fade faster than with the gentle stuff."
How it makes her feel
"I can skip foundation and not think about my skin all day."
What she actually values
Feeling in control of how she shows up.

We also borrow a projective trick or two. Asking "if this product were a person, who would it be?" gets past the polite answer to the gut one. And we map where the room divides, because the disagreement usually hides the most interesting decision.

03 · The numbers

Turning words into scores you can trust

People answer in sentences, not on a scale, so each answer gets read and placed on a one-to-five rating. Three things go into that placement: what the words mean, how spread out the room is, and whether the score fits the kind of rater the person is. A strict reviewer who suddenly gushes is a flag, not a data point.

Every number then carries its own margin of error, computed by resampling the responses hundreds of times to see how much the average wobbles. On a survey, if that range comes out too wide to be useful, the read-out says so instead of quietly shipping a shaky figure.

Purchase intent, from a real run
Target serum 2.6 (range 2–3.2)
Competitor A 2.7 (range 2–3.5)
Competitor B 2.9 (range 2.2–3.6)

Three serums sit close on the average. The target's range is the widest, which tells you opinion on it was genuinely split. That is the kind of thing a single number would have hidden.

04 · Keeping it honest

Checks at every step

Real differences, not mush

A focus group has to disagree. If the room converges too neatly, the panel is rebuilt until it reflects a real spread of attitudes.

In character, start to finish

Each answer has to sound like the person who gave it. A response that drifts off the participant's anchors gets redrawn.

The right questions, all of them

A focus group covers the full guide and a survey covers every dimension you asked for. Missing pieces get filled before the read-out.

Backed by the transcript

Every theme, gap, and recommendation points to the responses behind it. If it cannot cite the room, it does not ship.

When a check fails, the step does not just error out. It gets told what went wrong and tries again, up to twice, before the study stops and tells you why. You never get a polished read-out built on a step that quietly broke.

05 · Comparing to the real thing

What every benchmark discloses

A claim that synthetic answers track real ones is only worth what the comparison discloses. Each benchmark case puts the source FGI or survey next to the SSR run question-by-question — and lists every setting on both sides so the comparison can be redone.

Recruitment screener (verbatim)

Who was invited and on what criteria, in the source's own wording.

Sample weighting & demographics

Final N, demographic mix, weighting scheme if any.

Fieldwork dates, country, language

When and where the source data was collected.

Original question text & translations

The exact question wording shown to respondents, in every language used.

Likert / scale labelling

The labels for every point on the scale — not just "1 to 5".

FGI moderator guide & coding rules

For focus groups, the discussion guide and how answers were coded into scores.

Source license

Where the dataset came from and the terms under which it is reused.

On the SSR side every benchmark also publishes the run's template version, the LLM seed, the bootstrap seed, and which model wrote the answers — so an outside reader can request the same run again and get the same numbers.

Headline metrics — Pearson r, Top-2 box agreement, Bland-Altman bias — are reported once at least 5 paired studies and 300 question pairs have been collected. Until then every case is shown, but no aggregate claim is made.

06 · What the source data already reveals

Findings the benchmarks already settle

Before the SSR re-runs anything, the source data already shows what good answers should look like. Each finding below is a gradient an honest re-run has to reproduce — and a unit test for any system that claims to stand in for survey respondents.

Wording effect — same policy, different label

Δ 1.19

"Welfare" reads as 3.04. "Assistance to the poor" reads as 4.23 — a 1.19-point gap on the same policy, same 2024 wave.

Source: GSS 2024 NATSPEND / NATSPENDY split-ballot.

Scenario gradient — traumatic vs elective

Δ 1.59

Woman's life endangered reads as 4.65. Unmarried by choice reads as 3.06 — the same respondents shift 1.59 points across the conditions.

Source: GSS 2024 abortion battery (AB × 7 split-ballot).

Target gradient — in-group vs out-group

Δ 1.53

Free speech for an anti-religion atheist reads as 4.18. For an anti-American Muslim cleric, 2.65 — same action, different target, a 1.53-point gap.

Source: GSS 2024 SPK battery (4 target categories, split-ballot).

Action gradient — book vs speech

Δ LIB > SPK

Across every target category, removing a book from the library is a higher bar than blocking a speech. Same wave, same respondents.

Source: GSS 2024 SPK + LIB batteries (4 paired cells).

Shape mismatch — same mean, different distribution

Δ |skew| > 0.5

homosex reads as SSR 3.46 on average, but the distribution is sharply bimodal — 33.6% "always wrong" and 55.7% "not wrong at all", with only 10.6% in the two middle cells. A re-run that lands the right mean from a unimodal 3.46 cluster is wrong about what Americans actually believe. SD and skew are the acceptance test the mean alone cannot run.

Source: GSS 2024 homosex item (SD ≈ 1.85, skew ≈ −0.48).

Catalogue composition — shape mismatch is the rule, not the exception

Δ 16% unimodal

Across the full 98 paired-question catalogue, only 16 questions (16%) classify as unimodal. The other 84% break down as 34 multimodal-symmetric, 31 bimodal-asymmetric, 14 left-skewed, and 3 right-skewed. A mean-only acceptance test passes the wrong shape on most source items — the shape badge is what catches it.

Source: catalogue shape mix as of this page-load (18 fixtures, 98 paired questions, classifier thresholds SD ≥ 1.2 / |skew| ≥ 0.5).

Every gradient above is published as part of a benchmark fixture, with its weighted means, CIs, and raw response distributions. A synthetic re-run is useful to the degree it reproduces these gradients on the same wording and the same target.

07 · Download schema reference

What the CSV and JSON downloads contain

Every benchmark detail page exposes the underlying paired-question data as either a CSV (one row per question) or a fixture-shaped JSON. The same columns appear in the catalogue-wide download and the per-topic-shelf download. Below is what each column means and how the SSR 1-5 axis is derived from the source codebook.

Column Meaning
study_slug Catalogue and topic downloads only — fixture slug (e.g. gss-abortion-2024). Matches the URL on the detail page.
source_label Catalogue and topic downloads only — human-readable source citation (e.g. "GSS (2024) — Abortion attitudes battery").
topic_key Catalogue and topic downloads only — shelf key from the BenchmarkTopicGrouper (e.g. "abortion", "hedonic-cpg").
question_template_slotStable identifier for the paired question. Survives renaming of the prompt_text.
prompt_text The exact wording shown to a synthetic respondent in the SSR run — usually the source codebook verbatim.
ground_truth_mean Weighted source mean on the SSR 1-5 axis. Higher = more of the dimension named in the disclosure (more permissive / more confidence / more tolerance / more support).
ground_truth_ci_lower / _upper95% confidence interval on the source mean, after the same SSR 1-5 rescale.
ground_truth_sd SSR-axis weighted standard deviation: sqrt(Σ w · (x − μ)² / Σ w). Same canvas as the mean/CI.
ground_truth_skewness Fisher-Pearson skewness on the SSR axis: (Σ w · (x − μ)³ / Σ w) / SD³. Negative = long left tail, positive = long right tail.
ground_truth_shape Shape badge: unimodal / multimodal-symmetric / right-skewed / left-skewed / bimodal-asymmetric. Dispatch on (SD, |skew|) with thresholds 1.2 / 0.5. A re-run must reproduce this label, not just the mean.
n_respondents Per-question source n. Subgroup-restricted items (e.g. married-only HAPMAR) carry their subgroup n, not the wave total.
distribution_json The raw weighted distribution as inline JSON. Keys carry the SSR anchor (1, 3, 5, or 1..9) so the same dispatch can recover the SSR mean across all five distribution shapes.
weighting_notes Per-row prose — the weight variable used, the rescale formula applied, and any subgroup / split-ballot caveat for this specific question.

How the SSR 1-5 axis is derived

Five distribution-key shapes cover every fixture in the catalogue. The dispatch in BenchmarkFixtureConsistencyTest is the canonical reference — any new fixture has to round-trip its published mean through one of these formulas to within ±0.05.

Distribution shape SSR formula Example fixtures
9-pt hedonic — keys "1".."9", counts 1 + (raw − 1) · 4 / 8Visalli, iFishIENCi
3-pt labelled — leading ints 1/2/3 7 − 2 · raw GSS CONFIDENCE, GSS NATSPEND, GSS HAPPY
4-pt labelled — leading ints 1/2/3/4 1 + (4 − raw) · 4 / 3GSS SATJOB
Dichotomous — leading ints 1 and 5 leading_int GSS abortion, free speech, policy trio
Trust-style — leading ints 1, 3, 5 with "depends" middleleading_int GSS TRUST + FAIR + HELPFUL

Every fixture in the catalogue publishes its weighting_notes column with the exact rescale formula applied — the table above is the canonical map, not the only source of truth. The downloads themselves remain reproducible from the JSON fixture files in the repository.

For a deeper look

Each score looks at three practical signals: what the participant said, how much the room agreed or split, and whether the reaction fits that participant’s usual rating style. A focus group highlights disagreement. A survey looks for a stable average.

Survey reports include a confidence range so you can see whether a number is strong enough to use in a decision, or whether the sample needs to be larger.

Every study keeps the original brief and settings, so a saved study can be rerun later and compared against a revised product, message, or price.

The approach builds on published research into estimating purchase intent from customer-style responses.

Want the deeper guide?

The research guide explains how InsightForge turns customer reactions into scores, where the results are strongest, and where to use them with care.

Read the guide →

See it on a real product

The case studies put our output next to published benchmarks.