How it works
The short version: we build a panel of believable people, put your product in front of them, and write up what they say. The longer version is below, including where we ladder past the surface answer and how we keep the whole thing honest.
People who could argue with each other
A panel is only useful if the people in it actually differ. Each participant gets a real shape: an age and a job, a few lifestyle anchors in their own words, what they buy today, and the deeper value the category serves for them. One scores strictly, another gives the benefit of the doubt. That spread is the point. If everyone says four out of five, you have learned nothing.
Reads the ingredient list before the marketing copy. Trusts a clinical trial more than an influencer. Keeps an anti-aging routine that earns her every step. Underneath all that, what she wants from skincare is to age on her own terms.
We keep asking why
In a focus group, a first answer is rarely the real one. Someone says they like the ten percent vitamin C. Fine, but why does that matter? Because it fades dark spots faster. And why does that matter? The moderator keeps pulling that thread until it reaches something a person actually cares about. Researchers call this laddering. It turns a feature into a reason you can act on.
We also borrow a projective trick or two. Asking "if this product were a person, who would it be?" gets past the polite answer to the gut one. And we map where the room divides, because the disagreement usually hides the most interesting decision.
Turning words into scores you can trust
People answer in sentences, not on a scale, so each answer gets read and placed on a one-to-five rating. Three things go into that placement: what the words mean, how spread out the room is, and whether the score fits the kind of rater the person is. A strict reviewer who suddenly gushes is a flag, not a data point.
Every number then carries its own margin of error, computed by resampling the responses hundreds of times to see how much the average wobbles. On a survey, if that range comes out too wide to be useful, the read-out says so instead of quietly shipping a shaky figure.
Three serums sit close on the average. The target's range is the widest, which tells you opinion on it was genuinely split. That is the kind of thing a single number would have hidden.
Checks at every step
Real differences, not mush
A focus group has to disagree. If the room converges too neatly, the panel is rebuilt until it reflects a real spread of attitudes.
In character, start to finish
Each answer has to sound like the person who gave it. A response that drifts off the participant's anchors gets redrawn.
The right questions, all of them
A focus group covers the full guide and a survey covers every dimension you asked for. Missing pieces get filled before the read-out.
Backed by the transcript
Every theme, gap, and recommendation points to the responses behind it. If it cannot cite the room, it does not ship.
When a check fails, the step does not just error out. It gets told what went wrong and tries again, up to twice, before the study stops and tells you why. You never get a polished read-out built on a step that quietly broke.
What every benchmark discloses
A claim that synthetic answers track real ones is only worth what the comparison discloses. Each benchmark case puts the source FGI or survey next to the SSR run question-by-question — and lists every setting on both sides so the comparison can be redone.
Recruitment screener (verbatim)
Who was invited and on what criteria, in the source's own wording.
Sample weighting & demographics
Final N, demographic mix, weighting scheme if any.
Fieldwork dates, country, language
When and where the source data was collected.
Original question text & translations
The exact question wording shown to respondents, in every language used.
Likert / scale labelling
The labels for every point on the scale — not just "1 to 5".
FGI moderator guide & coding rules
For focus groups, the discussion guide and how answers were coded into scores.
Source license
Where the dataset came from and the terms under which it is reused.
On the SSR side every benchmark also publishes the run's template version, the LLM seed, the bootstrap seed, and which model wrote the answers — so an outside reader can request the same run again and get the same numbers.
Headline metrics — Pearson r, Top-2 box agreement, Bland-Altman bias — are reported once at least 5 paired studies and 300 question pairs have been collected. Until then every case is shown, but no aggregate claim is made.
Findings the benchmarks already settle
Before the SSR re-runs anything, the source data already shows what good answers should look like. Each finding below is a gradient an honest re-run has to reproduce — and a unit test for any system that claims to stand in for survey respondents.
Wording effect — same policy, different label
Δ 1.19"Welfare" reads as 3.04. "Assistance to the poor" reads as 4.23 — a 1.19-point gap on the same policy, same 2024 wave.
Source: GSS 2024 NATSPEND / NATSPENDY split-ballot.
Scenario gradient — traumatic vs elective
Δ 1.59Woman's life endangered reads as 4.65. Unmarried by choice reads as 3.06 — the same respondents shift 1.59 points across the conditions.
Source: GSS 2024 abortion battery (AB × 7 split-ballot).
Target gradient — in-group vs out-group
Δ 1.53Free speech for an anti-religion atheist reads as 4.18. For an anti-American Muslim cleric, 2.65 — same action, different target, a 1.53-point gap.
Source: GSS 2024 SPK battery (4 target categories, split-ballot).
Action gradient — book vs speech
Δ LIB > SPKAcross every target category, removing a book from the library is a higher bar than blocking a speech. Same wave, same respondents.
Source: GSS 2024 SPK + LIB batteries (4 paired cells).
Shape mismatch — same mean, different distribution
Δ |skew| > 0.5homosex reads as SSR 3.46 on average, but the distribution is sharply bimodal — 33.6% "always wrong" and 55.7% "not wrong at all", with only 10.6% in the two middle cells. A re-run that lands the right mean from a unimodal 3.46 cluster is wrong about what Americans actually believe. SD and skew are the acceptance test the mean alone cannot run.
Source: GSS 2024 homosex item (SD ≈ 1.85, skew ≈ −0.48).
Catalogue composition — shape mismatch is the rule, not the exception
Δ 16% unimodalAcross the full 98 paired-question catalogue, only 16 questions (16%) classify as unimodal. The other 84% break down as 34 multimodal-symmetric, 31 bimodal-asymmetric, 14 left-skewed, and 3 right-skewed. A mean-only acceptance test passes the wrong shape on most source items — the shape badge is what catches it.
Source: catalogue shape mix as of this page-load (18 fixtures, 98 paired questions, classifier thresholds SD ≥ 1.2 / |skew| ≥ 0.5).
Every gradient above is published as part of a benchmark fixture, with its weighted means, CIs, and raw response distributions. A synthetic re-run is useful to the degree it reproduces these gradients on the same wording and the same target.
Per-shelf top gradients
auto-derived from the current benchmark setBelow: the biggest mean-to-mean spread inside each topic shelf, computed live from the benchmark fixtures. New fixtures land here automatically as soon as they are seeded.
Abortion attitudes
Δ 1.594.65 () → 3.06 ()
1 paired study on this shelf →
Well-being (HAPPY + SAT*)
Δ 1.314.15 () → 2.83 ()
1 paired study on this shelf →
Generalised social trust
Δ 0.712.91 () → 2.20 ()
1 paired study on this shelf →
Political-policy trio
Δ 0.253.74 () → 3.48 ()
1 paired study on this shelf →
Sexual-morality trio
Δ 2.344.00 () → 1.66 ()
1 paired study on this shelf →
Political identity (polviews + partyid)
Δ 0.113.12 () → 3.01 ()
1 paired study on this shelf →
Religion belief (postlife + god + pray)
Δ 1.024.29 () → 3.27 ()
1 paired study on this shelf →
Institutional confidence
Δ 1.583.63 () → 2.05 ()
1 paired study on this shelf →
Federal spending priorities
Δ 2.534.38 () → 1.85 ()
1 paired study on this shelf →
Civil liberties (SPK + LIB)
Δ 1.534.18 () → 2.65 ()
1 paired study on this shelf →
Hedonic taste tests (CPG)
Δ 1.653.63 () → 1.97 ()
4 paired studies on this shelf →
Aquaculture sensory panels
Δ 0.604.10 () → 3.50 ()
4 paired studies on this shelf →
What the CSV and JSON downloads contain
Every benchmark detail page exposes the underlying paired-question data as either a CSV (one row per question) or a fixture-shaped JSON. The same columns appear in the catalogue-wide download and the per-topic-shelf download. Below is what each column means and how the SSR 1-5 axis is derived from the source codebook.
| Column | Meaning |
|---|---|
| study_slug | Catalogue and topic downloads only — fixture slug (e.g. gss-abortion-2024). Matches the URL on the detail page. |
| source_label | Catalogue and topic downloads only — human-readable source citation (e.g. "GSS (2024) — Abortion attitudes battery"). |
| topic_key | Catalogue and topic downloads only — shelf key from the BenchmarkTopicGrouper (e.g. "abortion", "hedonic-cpg"). |
| question_template_slot | Stable identifier for the paired question. Survives renaming of the prompt_text. |
| prompt_text | The exact wording shown to a synthetic respondent in the SSR run — usually the source codebook verbatim. |
| ground_truth_mean | Weighted source mean on the SSR 1-5 axis. Higher = more of the dimension named in the disclosure (more permissive / more confidence / more tolerance / more support). |
| ground_truth_ci_lower / _upper | 95% confidence interval on the source mean, after the same SSR 1-5 rescale. |
| ground_truth_sd | SSR-axis weighted standard deviation: sqrt(Σ w · (x − μ)² / Σ w). Same canvas as the mean/CI. |
| ground_truth_skewness | Fisher-Pearson skewness on the SSR axis: (Σ w · (x − μ)³ / Σ w) / SD³. Negative = long left tail, positive = long right tail. |
| ground_truth_shape | Shape badge: unimodal / multimodal-symmetric / right-skewed / left-skewed / bimodal-asymmetric. Dispatch on (SD, |skew|) with thresholds 1.2 / 0.5. A re-run must reproduce this label, not just the mean. |
| n_respondents | Per-question source n. Subgroup-restricted items (e.g. married-only HAPMAR) carry their subgroup n, not the wave total. |
| distribution_json | The raw weighted distribution as inline JSON. Keys carry the SSR anchor (1, 3, 5, or 1..9) so the same dispatch can recover the SSR mean across all five distribution shapes. |
| weighting_notes | Per-row prose — the weight variable used, the rescale formula applied, and any subgroup / split-ballot caveat for this specific question. |
How the SSR 1-5 axis is derived
Five distribution-key shapes cover every fixture in the catalogue. The dispatch in BenchmarkFixtureConsistencyTest is the canonical reference — any new fixture has to round-trip its published mean through one of these formulas to within ±0.05.
| Distribution shape | SSR formula | Example fixtures |
|---|---|---|
| 9-pt hedonic — keys "1".."9", counts | 1 + (raw − 1) · 4 / 8 | Visalli, iFishIENCi |
| 3-pt labelled — leading ints 1/2/3 | 7 − 2 · raw | GSS CONFIDENCE, GSS NATSPEND, GSS HAPPY |
| 4-pt labelled — leading ints 1/2/3/4 | 1 + (4 − raw) · 4 / 3 | GSS SATJOB |
| Dichotomous — leading ints 1 and 5 | leading_int | GSS abortion, free speech, policy trio |
| Trust-style — leading ints 1, 3, 5 with "depends" middle | leading_int | GSS TRUST + FAIR + HELPFUL |
Every fixture in the catalogue publishes its weighting_notes column with the exact rescale formula applied — the table above is the canonical map, not the only source of truth. The downloads themselves remain reproducible from the JSON fixture files in the repository.
For a deeper look
Each score looks at three practical signals: what the participant said, how much the room agreed or split, and whether the reaction fits that participant’s usual rating style. A focus group highlights disagreement. A survey looks for a stable average.
Survey reports include a confidence range so you can see whether a number is strong enough to use in a decision, or whether the sample needs to be larger.
Every study keeps the original brief and settings, so a saved study can be rerun later and compared against a revised product, message, or price.
The approach builds on published research into estimating purchase intent from customer-style responses.
Want the deeper guide?
The research guide explains how InsightForge turns customer reactions into scores, where the results are strongest, and where to use them with care.
See it on a real product
The case studies put our output next to published benchmarks.