Visalli et al. (2022) — Guacamole panel
https://data.mendeley.com/datasets/fshtbhffth/1 →benchmark:visalli-guacamole-2022
Shape mix:
2
unimodal
1
multimodal-symmetric
1
left-skewed
Δ 1.09
3.51
()
→ 2.42
()
within-fixture gradient
Acceptance rubric
An SSR re-run acceptably reproduces this benchmark when it clears every check below on the same 4 paired questions the source publishes. The chip badges name the checks the round-38 → round-46 shape-mismatch finding argued are necessary but not sufficient in isolation.
CI overlap
4/4
Shape match
4/4
SSR mean inside CI · shape badge reproduces the source label.
What this benchmark discloses
Recruitment screener (verbatim)
Recruited by gender, age, and guacamole consumption frequency. Excluded: allergy, restrictive diet, vulnerable persons, pregnant women. Source: Visalli et al. (2022), PMC9679663.
Sample weighting & demographics
Unweighted hedonic. 9-point discrete scale, single rating per product per consumer. Anchor labels taken from questionnaire.pdf p.18 — only end-points (1 and 9) are named.
Fieldwork dates, country, language
Dijon and vicinity, France. Specific collection dates not published in PMC9679663. Mendeley release: 2022-10-03.
Original question text & translations
On a scale of 1 to 9, how do you like this guacamole? — single hedonic rating per guacamole per consumer. Source: questionnaire.pdf p.18 (Mendeley fshtbhffth/1). Screen [7] reused for guacamole per q.txt:140-143 ("Same as screens 4 to 12 with guacamole").
Likert / scale labelling
9-point discrete checkbox scale, end-anchored only. Anchor 1 = "I did not like it at all". Anchor 9 = "I liked it very much". Points 2-8 unlabelled (equal-interval assumption).
FGI moderator guide & coding rules
Hedonic survey, not FGI. No moderator guide. The temporal methods (TDS, TCATA, AEF, FC-AEF) in the same paper are out of scope for this paired comparison.
Data source & license
CC-BY 4.0. Attribution: Visalli M. et al. (2022). https://data.mendeley.com/datasets/fshtbhffth/1
Anything not disclosed here was not available in the source dataset and is recorded as such.
SSR run — reproducibility keys
- template_version
- 3.0.0
- llm_seed
- 4204
- bootstrap_seed
- 8408
- llm_model
- deepseek-v4-flash
- embedding_model
- nvidia/nv-embed-v1
SSR score vs ground truth, by question
Each row holds one question. The teal mark and range are SSR. The sky mark and range are the source FGI/survey.
§
How much do you like this guacamole? (G1 — 92% avocado, 16 g fat)
SSR
—
Source
3.13 (3.02–3.24)
Source distribution
SD 1.16
skew -0.26
9 cells
unimodal
1.02.03.04.05.0
§
How much do you like this guacamole? (G2 — 13% avocado, 9.5 g fat)
SSR
—
Source
3.07 (2.96–3.19)
Source distribution
SD 1.25
skew -0.14
9 cells
multimodal-symmetric
1.02.03.04.05.0
§
How much do you like this guacamole? (G3 — 90% avocado, 14.6 g fat)
SSR
—
Source
2.42 (2.32–2.52)
Source distribution
SD 1.08
skew 0.41
9 cells
unimodal
1.02.03.04.05.0
§
How much do you like this guacamole? (G4 — 95% avocado, 18 g fat)
SSR
—
Source
3.51 (3.41–3.61)
Source distribution
SD 1.08
skew -0.53
9 cells
left-skewed
1.02.03.04.05.0
More on this shelf
Hedonic taste tests (CPG)
9-pt hedonic panel scored on four SKU per category in the Visalli et al. open dataset. Same panel, multiple products — the gradient is product-vs-product within one consumer.