Skip to content

Visalli et al. (2022) — Iced tea panel

https://data.mendeley.com/datasets/fshtbhffth/1 →

benchmark:visalli-icetea-2022

Reuse this data: CSV JSON headers match the published fixture; distributions JSON-encoded inline
Shape mix: 2 unimodal 1 left-skewed 1 multimodal-symmetric
Δ 1.03 3.46 () → 2.44 () within-fixture gradient
Acceptance rubric

An SSR re-run acceptably reproduces this benchmark when it clears every check below on the same 4 paired questions the source publishes. The chip badges name the checks the round-38 → round-46 shape-mismatch finding argued are necessary but not sufficient in isolation.

CI overlap 4/4 Shape match 4/4 SSR mean inside CI · shape badge reproduces the source label.
What this benchmark discloses
Recruitment screener (verbatim)
Recruited by gender, age, and ice-tea consumption frequency. Excluded: allergy, restrictive diet, vulnerable persons, pregnant women. Source: Visalli et al. (2022), PMC9679663.
Sample weighting & demographics
Unweighted hedonic. 9-point discrete scale, single rating per product per consumer. Anchor labels taken from questionnaire.pdf p.18 — only end-points (1 and 9) are named.
Fieldwork dates, country, language
Dijon and vicinity, France. Specific collection dates not published in PMC9679663. Mendeley release: 2022-10-03.
Original question text & translations
On a scale of 1 to 9, how do you like this iced tea? — single hedonic rating per ice tea per consumer. Source: questionnaire.pdf p.18 (Mendeley fshtbhffth/1). Screen [7] reused for ice tea per q.txt:99 ("Screens 4 to 8 repeated for all ice tea samples").
Likert / scale labelling
9-point discrete checkbox scale, end-anchored only. Anchor 1 = "I did not like it at all". Anchor 9 = "I liked it very much". Points 2-8 unlabelled (equal-interval assumption).
FGI moderator guide & coding rules
Hedonic survey, not FGI. No moderator guide. The temporal methods (TDS, TCATA, AEF, FC-AEF) in the same paper are out of scope for this paired comparison.
Data source & license
CC-BY 4.0. Attribution: Visalli M. et al. (2022). https://data.mendeley.com/datasets/fshtbhffth/1

Anything not disclosed here was not available in the source dataset and is recorded as such.

SSR run — reproducibility keys
template_version
3.0.0
llm_seed
4203
bootstrap_seed
8406
llm_model
deepseek-v4-flash
embedding_model
nvidia/nv-embed-v1

SSR score vs ground truth, by question

Each row holds one question. The teal mark and range are SSR. The sky mark and range are the source FGI/survey.

§ How much do you like this iced tea? (IT1 — black tea, white peach, 4.7 g sugar)
SSR
Source 3.46 (3.38–3.55)
Source distribution SD 0.90 skew -0.66 9 cells left-skewed
1.02.03.04.05.0
§ How much do you like this iced tea? (IT2 — white tea, peach and rosemary, 4.5 g sugar)
SSR
Source 2.44 (2.34–2.53)
Source distribution SD 1.00 skew 0.50 9 cells unimodal
1.02.03.04.05.0
§ How much do you like this iced tea? (IT3 — black tea, peach, sweeteners, 4.3 g sugar)
SSR
Source 2.99 (2.88–3.10)
Source distribution SD 1.16 skew -0.13 9 cells unimodal
1.02.03.04.05.0
§ How much do you like this iced tea? (IT4 — black tea, peach, sweeteners only, 0 g sugar)
SSR
Source 2.55 (2.43–2.67)
Source distribution SD 1.25 skew 0.34 9 cells multimodal-symmetric
1.02.03.04.05.0
More on this shelf

Hedonic taste tests (CPG)

9-pt hedonic panel scored on four SKU per category in the Visalli et al. open dataset. Same panel, multiple products — the gradient is product-vs-product within one consumer.