Design, Sensory Science, Art, and Health
The last two chapters asked where preferences come from and what a set of comparisons can tell a model. This chapter turns to the disciplines that ask people about designed things for a living: designers and engineers, architects and planners, food scientists who run consumer panels, researchers of music and visual aesthetics, and health researchers who elicit patients' priorities before a treatment decision. Many of them have faced, at scale and for decades, the measurement problems that preferential Bayesian optimization (PBO) meets in every user study. Their most transferable answer is the placebo pair of sensory science (Section 44.5.1); the chapter ends by gathering what PBO can adopt from this part of the book without new theory (Section 44.9).
44.1 Design research #
Design research has long described creativity as the co-evolution of problem and solution, with designers who discover what they want by designing (Dorst and Cross, 2001). But "co-evolution" names a family of different concepts, some meaning that problem and solution change each other, others only that attention alternates between them, and parallel processes appear across creative work, so it looks like a general feature of creative work rather than something specific to design (Crilly, 2021a; Crilly, 2021b). Product-innovation methods fare worse. In a working paper presented at a practitioner conference, not peer reviewed, Chapman and Callegaro (2022) had 1,501 people answer the same item of the Kano model, which sorts product attributes into "must-be" attributes and "delighters", twice within seconds: only 39% gave the same "expect" answer again.
Fixation is the clearest new evidence. Design fixation is the tendency to reuse features of examples one has seen, even when they are flawed or the task asks for something new. In a large experiment with novice engineers, Leahy et al. (2020) showed half of them a given example and let the other half generate their own first concept. The students who saw the example fixated on it less than the control students fixated on their own first concept, which revises the classic experiment of Jansson and Smith (1991), whose control condition had been assumed to be free of fixation. In a between-subjects experiment (), Wadinambiarachchi et al. (2024) found that participants who used an AI image generator during ideation fixated more on the initial example and produced fewer, less varied, and less original ideas, with effects that depended on how participants prompted and ideated. And in a study of designers working with an optimizer (Mo et al., 2024), 12 of 18 participants preferred a mode in which designer and optimizer collaborate, and the mode led by the optimizer gave the lowest sense of control (compare Section 32.2).
What it means for PBO. A PBO gallery is a sequence of examples supplied by the system; used for ideation rather than for tuning parameters, it should be expected to produce the fixation and loss of diversity that Wadinambiarachchi et al. observed (inference). The incumbent, the design the user chose last, plays the role of a designer's own first concept and may anchor later comparisons (inference), the more so because choosing changes value: in a fitted revaluation model of repeated food choices, each choice raised the value of the chosen snack by about 0.18 US dollars and lowered that of the rejected one by as much (Zylberberg et al., 2024). Comparing periodically against diverse designs not derived from the incumbent is a practical check (inference). A Kano "must-be" attribute is a utility with a threshold, better represented as a constraint or a hinge-shaped component (Section 14.4), but Kano classifications are too unreliable to serve as a hard prior (inference). We found no study that measures fixation within a PBO session, for example against a control group shown a random gallery.
Sources cited in Section 44.1 9
- Dorst and Cross (2001) Creativity in the Design Process: Co-Evolution of Problem-Solution
- Crilly (2021a) The Evolution of “Co-evolution” (Part I): Problem Solving, Problem Finding, and Their Interaction in Design and Other Creative Practices
- Crilly (2021b) The Evolution of “Co-evolution” (Part II): The Biological Analogy, Different Kinds of Co-evolution, and Proposals for Conceptual Expansion
- Chapman and Callegaro (2022) Kano Analysis: A Critical Survey Science Review
- Leahy et al. (2020) Design Fixation From Initial Examples: Provided Versus Self-Generated Ideas
- Jansson and Smith (1991) Design Fixation
- Wadinambiarachchi et al. (2024) The Effects of Generative AI on Design Fixation and Divergent Thinking
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
44.2 Architecture #
Architecture is often credited with universal spatial preferences, which has prompted the suggestion to encode a universal good as a prior separate from personal preference. Jay Appleton's prospect-refuge theory, which holds that people prefer places where they can see without being seen, has inconsistent quantitative support. A meta-analysis of 34 quantitative studies, published slightly before the period this part covers (Dosen and Ostwald, 2016), found prospect supported in 19 of its 31 tests (61%) and refuge in only 8 of its 27 tests (30%); the 53% and 22% in the paper's own summary are these factors' shares of the 36 supporting findings, not rates of support. Refuge found support mainly in studies of natural environments (5 of 8 tests in landscapes, against 2 of 17 in interiors), and of the 29 studies with survey results, 16 had 20 or fewer participants. Perceptual dimensions fare better: Coburn et al. (2020) had 798 participants rate 200 images of interiors on 16 scales, and three components, coherence, fascination, and hominess, explained 90% of the variance in the ratings; the structure replicated in an independent sample ().
Shared and individual taste. The key result since 2017 is Vessel et al. (2018). Agreement between individuals was high for faces and landscapes and low for building exteriors, building interiors, and artworks. Within participants, agreement was significantly higher for landscapes than for building exteriors, while the reliability of each person's own ratings did not differ between the two. The authors proposed that taste is shared for natural kinds and individual for cultural artifacts. Expertise splits taste further: young architects and matched laypeople rating Czech detached houses had the same overall means (5.08 against 5.09), but the architects rated modern and wooden houses higher, and catalog houses and "McMansions" lower (Šafárová et al., 2019).
What it means for PBO. Low agreement combined with equal reliability means that a population-average prior gives little guidance about one person's architectural taste: architecture calls mainly for learning per person, while landscapes or faces can borrow more from population data, so the choice between informative population priors and weak priors depends on the domain (inference; Section 31.2). Defining the surrogate on a few perceptual dimensions, rather than on dozens of raw geometric parameters, can reduce the effective dimension (inference; Chapter 30). An architect who runs PBO on a client's behalf optimizes the architect's utility unless the client is in the loop (inference). Given the low agreement on buildings, a universal architectural good should not be encoded as a prior (inference).
Sources cited in Section 44.2 4
- Dosen and Ostwald (2016) Evidence for prospect-refuge theory: a meta-analysis of the findings of environmental preference research
- Coburn et al. (2020) Psychological and neural responses to architectural interiors
- Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
- Šafárová et al. (2019) Differences between young architects' and non-architects' aesthetic evaluation of buildings
44.3 Urban planning #
Surveys of urban perception are the largest practical application of pairwise visual comparison outside machine learning, built on Place Pulse 2.0 (Dubey et al., 2016), which collected comparisons of street-view images along dimensions such as "safer" or "livelier". Quintana et al. (2025) (with methodological details in the arXiv version) had 1,000 demographically balanced participants from five countries each make 50 pairwise comparisons, with the option "Both are the same to me". They scored images with the Strength of Schedule method, which credits a win by the strength of the images it was won against, because it needs fewer comparisons per image than the TrueSkill rating system (4 against 22 to 29). Gender produced the most differences between groups, especially for "safe", and a vision transformer trained on Place Pulse overestimated positive indicators such as "lively" relative to the human ratings. In participatory budgeting, where residents decide how to spend part of a public budget, Benade et al. (2017) (with the journal version Benadè et al., 2021) found that threshold approval voting, in which each voter approves the projects whose value to them exceeds a threshold, was qualitatively better than knapsack voting and value rankings on distortion (how far the welfare of the chosen outcome falls short of the best possible when only ordinal information is available) and on regret.
What it means for PBO. Aggregating pairwise judgments from a heterogeneous population biases the learned scores and hides differences between subgroups, so group or civic PBO should model respondents' covariates or keep a surrogate per group (inference; Section 20.5). Urban surveys treat an explicit indifference option as standard, and a forced-choice interface discards that signal (inference; Section 20.4). With several stakeholders, the query format is a design variable with social-choice consequences (inference; Section 40.7).
Sources cited in Section 44.3 4
- Dubey et al. (2016) Deep Learning the City: Quantifying Urban Perception at a Global Scale
- Quintana et al. (2025) Global urban visual perception varies across demographics and personalities
- Benade et al. (2017) Preference Elicitation For Participatory Budgeting
- Benadè et al. (2021) Preference Elicitation for Participatory Budgeting
44.4 Materials and engineering design #
Engineering design produced a direct precursor of PBO: Ren and Papalambros (2011) cast design preference elicitation as an optimization problem, using a support vector machine with efficient global optimization in user tests of car exterior styling. Since 2017, Kanarik et al. (2023) compared human engineers with Bayesian optimization algorithms on the design of a semiconductor plasma etch process, in a controlled virtual process game. Humans did well early, algorithms were far more cost-efficient close to tight tolerances, and a "human first, computer last" strategy halved the cost of reaching the target compared with humans alone. Nandy and Goucher-Lambert (2025) optimized the perceived comfort of parametric mugs with interactive Bayesian optimization per participant (), started from scratch or from the data of an earlier group (), and interactive models aligned intent and generated designs better than non-interactive ones. A 2025 paper (Huber et al., 2025) built a Bayesian model of a decision maker's utility over a Pareto set from pairwise comparisons, tested it on problems with up to nine objectives, and released its code.
What it means for PBO. A workflow should plan a handoff from the person to the algorithm instead of using one mode throughout; but Kanarik et al.'s evidence concerns expert search with an objective metric, and whether it transfers to design driven by taste is uncertain (inference; compare Section 23.4). Pairwise acquisition on a precomputed Pareto front is a practical way to bring a designer's preference in after multi-objective optimization (Section 14.5, Section 40.11).
Sources cited in Section 44.4 4
- Ren and Papalambros (2011) A Design Preference Elicitation Query as an Optimization Process
- Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development
- Nandy and Goucher-Lambert (2025) Exploring the Effectiveness of Interactive Preference Learning for Adapting Designs to Abstract Semantic Attributes
- Huber et al. (2025) Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization
44.5 Food and sensory science #
Food scientists have long run paired preference tests with consumers (O'Mahony and Wichchukit, 2017), and they have learned uncomfortable things about what people say when asked "which do you prefer?". One famous result needs care first. In Morrot et al. (2001), white wine colored red with an odorless dye was described in olfactory terms as red wine by 54 tasters, oenology students in Bordeaux. One reading says that color completely overrides chemosensory information, with devastating consequences for preference data. The effect is real and has been replicated conceptually, but the task was choosing descriptors, not stating a preference or a liking. In Wang and Spence (2019), participants with wine experience judged a white wine dyed to look like rosé far more similar to the rosé than to the identical white wine, yet the dyed wine was liked less than either real wine, and participants felt it was somehow different. Spence (2020) reviewed how contextual factors, from the color of ambient light to background music, have profound and sometimes predictable effects on tasting.
44.5.1 Paired preference tests and placebo pairs #
The most useful lesson for PBO is the methodology of paired preference tests. O'Mahony and Wichchukit (2017) reviewed its history. Forced choice was adopted first because its binomial statistics are simple, but it gave consumers no way to say "no preference". Consumers then turned out to report preferences between stimuli that should have been identical, so placebo pairs, pairs made of two identical samples, became the control condition. The foundational study carries the finding in its title, "Consumers report preferences when they should not" (Marchisano et al., 2003). Frequency measures (how many consumers chose each sample) were later supplemented by the discrimination index of signal detection theory, the perceived distance between two stimuli in units of the standard deviation of perceptual noise, which measures the strength of a preference more accurately.
Sensory science reads such data with Thurstonian models, the same model PBO uses (Section 16.3). A paired preference test with a "no preference" option is technically identical to the 2-AC protocol ("two alternatives, with a no-difference option"), whose Thurstonian model has two parameters: the distance between the two stimuli and a decision threshold below which the person reports no preference; it is closely related to a cumulative probit model, the ordinal likelihood of Section 27.2 (Christensen et al., 2012).
- In the Thurstonian model of Section 16.3, each of the two options is perceived with independent Gaussian noise of standard deviation around its utility, so the perceived difference is , where is the true difference.
- In the 2-AC model, the person reports "no preference" when and otherwise prefers the option that favors.
- A placebo pair has . The probability that the person nevertheless states a preference is , by the symmetry of the Gaussian.
- Solving for the threshold gives . A false-preference rate of means ; a forced-choice interface has and therefore by construction.
- The placebo pairs thus identify the person's tie threshold relative to their noise, which is exactly the quantity a PBO likelihood with ties needs and cannot learn from forced choices.
Figure 44.1 lets you take a short paired preference test with placebo pairs. Answer as you would in a study, and only then read the results.
Things to notice:
- Consumers in paired tests often state preferences between identical samples (Marchisano et al., 2003); check whether you did. Every preference you stated on an identical pair is a "false preference": in a PBO session it would enter the likelihood as evidence that one design beats the other.
- If your false preferences went mostly to one side, your answers carried a position bias, which a likelihood can absorb with a side term (Equation (43.1)) if the sides are balanced.
- A forced-choice interface, the default in most PBO studies, would have recorded every identical pair as a preference.
Several studies refine the method. Halim et al. (2020) found that placing the placebo pair after the target pair raised the share of "no preference" answers, though not always significantly. Xia et al. (2020) found that for some products the effect is weaker because samples of the same product vary a lot, so the "identical" pairs were not really identical. On reliability, Nijman et al. (2022) had 62 participants rate their liking of one beer, and ten emotions it evoked, in two sessions each in a bar and in a central-location test. The intraclass correlation coefficient (ICC), here the share of rating variance due to stable differences between participants rather than to noise between sessions, averaged 0.66 (±0.1) in the bar and 0.60 (±0.15) in the central-location test across the eleven ratings, and for liking alone lay between 0.58 and 0.76.
What it means for PBO. Occasionally presenting an identical pair, the same design twice, estimates the user's false-preference rate; allowing a tie answer and fitting the noise or lapse term of the probit likelihood with the tie and placebo data calibrates the model. Because the Thurstonian model of sensory science is PBO's probit likelihood, the transfer is mathematically direct, and is simply the utility difference in units of the noise standard deviation (inference). An "identical" pair must be truly identical, with the same random seed and the same rendering, or the placebo is contaminated (inference). Hedonic test-retest correlations of about 0.6 to 0.7 give a reasonable expectation of consistency across sessions for sensory stimuli, and a PBO posterior far more confident than that has probably overfitted a single session (inference).
Sources cited in Section 44.5 9
- O'Mahony and Wichchukit (2017) The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection
- Morrot et al. (2001) The Color of Odors
- Wang and Spence (2019) Drinking through rosé-coloured glasses: Influence of wine colour on the perception of aroma and flavour in wine experts and novices
- Spence (2020) Wine psychology: basic & applied
- Marchisano et al. (2003) Consumers report preferences when they should not: a cross-cultural study
- Christensen et al. (2012) Estimation of the Thurstonian Model for the 2-AC Protocol
- Halim et al. (2020) Paired preference tests and placebo placement: 1. Should placebo pairs be placed before or after the target pair?
- Xia et al. (2020) Paired preference tests and placebo placement: 2. Unraveling the effects of stimulus variance
- Nijman et al. (2022) The stability of self-reported emotional response and liking of beer in context
44.6 Empirical aesthetics and music #
Empirical aesthetics has well-known principles: Berlyne's Wundt curve says liking peaks at intermediate complexity or novelty, and processing fluency says that the easier a stimulus is to process, the more positively it is judged. The inverted U is well supported at the level of groups: of 57 music studies reviewed by Chmiel and Schubert (2017), 50 (87.7%) were consistent with a (segmented) inverted U between complexity or arousal and liking, though the authors note that the model may not match Berlyne's arousal mechanism. But Güçlütürk et al. (2016) showed, with 30 participants rating grayscale images, that the group's inverted U was composed of two clusters of people: in one, liking fell with complexity; in the other, it rose. Figure 44.2 shows how two kinds of monotone individual produce an inverted U that nobody has.
At an even split, the population peaks near complexity 0.5, and a design placed there gives each group only about 70% of what a design aimed at it would give. Moving the slider to a 25% minority shows the minority served at only about a third of its best.
Fluency is not uniformly positive either. Graf and Landwehr (2017) found that for chairs and lamps, the effect of fluency on attractiveness ran through pleasure, especially under automatic processing, while under controlled processing disfluent designs could gain attractiveness through interest. And pleasure depends on what came before. Cheung et al. (2019) used a machine learning model to quantify the uncertainty and surprise of 80,000 chords in US Billboard pop songs: chords were most pleasurable when they were surprising in a context of low uncertainty, or unsurprising in a context of high uncertainty. On repetition the evidence conflicts: Gold et al. (2019) found that seven repetitions of a stimulus lowered liking but did not destroy the preference for intermediate complexity, while in Madison and Schiölde (2017), liking for 40 unfamiliar pieces heard 28 times each over about four weeks rose monotonically at every level of complexity. The timescale and whether listening is voluntary may explain the difference, but no study has resolved it.
Since 2017, a computational theory of aesthetic value. Brielmann and Dayan (2022) proposed a model in which the observer's sensory-cognitive state is a generative model of stimuli, and aesthetic value has two parts: an immediate sensory reward, the fluency of the stimulus measured as its likelihood under the current state, and the change in expected future reward as that state moves toward or away from the distribution of stimuli expected in the long run. Brielmann et al. (2024) had 59 participants rate 55 morphed images of dogs. Individual models using features from the deep convolutional network VGG-16 captured trial-by-trial liking with a median correlation of 0.65, against 0.01 for the population average, and the learning component explained on average 17% more variance for the actual order of presentation than for simulated random orders.
What it means for PBO. In the Brielmann-Dayan model every stimulus seen updates the observer's state, so the latent utility is a function of the stimulus and of a state driven by the sequence of queries: in preference elicitation every query is also a write (inference; Section 42.2). The order of a gallery is part of the stimulus, and showing the incumbent many times may lower or raise its value, depending on the timescale (inference). For individualized content, individual models beat the population average by a wide margin (0.65 against 0.01), so per-user surrogates are necessary, and deep network features are a usable input space (inference). A population inverted U can be a mixture of opposite monotone preferences, so a proposal that the acquisition function should target a zone of optimal novelty assumes that each person has an inverted U, which is not established (inference). Fast gallery-style judgments may favor fluent, typical designs and slower comparisons interesting ones (inference). An observer-state model such as Brielmann and Dayan's could serve as the non-stationarity of a PBO surrogate (Section 46.5) (inference).
Sources cited in Section 44.6 8
- Chmiel and Schubert (2017) Back to the inverted-U for music preference: A review of the literature
- Güçlütürk et al. (2016) Liking versus Complexity: Decomposing the Inverted U-curve
- Graf and Landwehr (2017) Aesthetic Pleasure versus Aesthetic Interest: The Two Routes to Aesthetic Liking
- Cheung et al. (2019) Uncertainty and Surprise Jointly Predict Musical Pleasure and Amygdala, Hippocampus, and Auditory Cortex Activity
- Gold et al. (2019) Predictability and Uncertainty in the Pleasure of Music: A Reward for Learning?
- Madison and Schiölde (2017) Repeated Listening Increases the Liking for Music Regardless of Its Complexity: Implications for the Appreciation and Aesthetics of Music
- Brielmann and Dayan (2022) A computational model of aesthetic value
- Brielmann et al. (2024) Modelling individual aesthetic judgements over time
44.7 Health sciences and pharmacology #
Health research elicits patients' preferences before treatment decisions, for regulators approving medical products, and for health technology assessment. Its main stated-preference method, the discrete choice experiment (DCE), in which respondents choose between hypothetical alternatives described by their attributes (Section 40.9), is close to PBO but developed independently; from 2018 to 2023, 1,279 health DCEs were published (Nouwens et al., 2025). Pharmacology supplies a favorite argument against stable preferences, that drugs can change choices; the substance survives, though its citations circulate with errors (Table 44.1). Manipulating serotonin changed moral judgments and responses to unfair offers (Crockett et al., 2008; Crockett et al., 2010), in the 2010 study more so in highly empathic individuals, and impulse control disorders in Parkinson's disease are linked mainly to dopamine agonists (Voon et al., 2006; Weintraub et al., 2010).
Do stated choices predict real ones? On external validity, Quaife et al. (2018) pooled 8 studies (6 in the meta-analysis): sensitivity 88% (the share of actual choices to take up an option that the DCE predicted), specificity 34% (the share of actual refusals it predicted), and an area under the ROC curve of 0.60, where 0.5 is chance. Zhang et al. (2025b) updated this to 14 studies (10 in the meta-analysis): sensitivity 89%, specificity 52%, area under the curve 0.81, with very high heterogeneity between studies ( of 95% to 97%, the share of variation not due to chance). Prediction was better for preventive, opt-in decisions. Figure 44.3 turns the two pooled estimates into counts of people.
Things to try:
- At the defaults, with Quaife et al.'s estimates and half the people taking the option up, the experiment predicts take-up for 77 people, and 44 of them (57%) do. The 33 filled magenta dots are refusals it called take-ups.
- Switch to Zhang et al. The higher specificity cuts the false take-ups to 24: 69 predicted, 45 of them real (65%).
- Lower the share who take it up to 20%. With Quaife et al.'s estimates, 71 people are predicted to take it up and only 18 of them (25%) do; with Zhang et al.'s, 56 and 18 (32%). When most people would refuse, a stated "yes" says little.
Reliability and dependence on method. Retesting 162 people after two weeks, Xie et al. (2022) found that 76.4% of DCE choices were identical (Cohen's kappa 0.528, agreement corrected for chance), while the time trade-off method, which asks how many years in full health a person would trade for a longer life in a worse state, had an intraclass correlation of 0.958 with 59.3% of values identical. In Whichello et al. (2023), 459 Dutch adults with diabetes completed a DCE and swing weighting (a direct method that asks how much moving each attribute from its worst to its best level matters; Section 40.10) in balanced order. Both ranked cost and precision highest, but the DCE's weights for the most and least important attributes differed 14.9-fold against 1.4-fold for swing weighting; among 307 lung cancer patients, Veldwijk et al. (2024) likewise found DCE weights more spread out.
Adaptive elicitation and decision aids. The "patient preference diagnostic" of Gonzalez Sepulveda et al. (2023) uses latent preference classes from an earlier survey as a prior and asks adaptive choice questions to place a patient in one of the known preference phenotypes. For first-time anterior shoulder dislocation with four classes, the posterior class probabilities reached 87% to 89% after a sequence of two questions; these are simulation results only. A Cochrane review of 209 randomized trials with 107,698 participants (Stacey et al., 2024) found that decision aids probably improve the congruence between informed values and the choice made (relative risk 1.75, 95% confidence interval 1.44 to 2.13; 21 studies; moderate certainty), with no difference in decision regret. And measured preferences can drift for reasons other than a change in what is measured: response shift is observed change that is not fully explained by change in the target construct (Vanier et al., 2021).
What it means for PBO. This is one of the chapter's strongest links; the following are inferences. Population latent classes as a prior and two adaptive questions translate into a mixture prior over utility functions learned from earlier users: early queries identify the class, later ones refine within it (Section 20.5), though the accuracy with real patients is unknown. Stated choices predict actual uptake well but refusal poorly, so the "winner" of a PBO session should be validated in real use before it is treated as the user's real choice. Pairwise PBO is a choice-based method, so the trade-offs it learns may be more extreme than direct ratings would give (14.9 against 1.4). About 76% identical binary health choices after two weeks gives a realistic noise floor for repeated pairwise judgments. Values congruence and decision regret can evaluate PBO as a decision aid, beside simulated regret (Section 46.7). In long-term PBO for assistive devices (Chapter 33), drift may be response shift rather than a change of preference, which a time-varying kernel alone cannot tell apart. And clinical PBO should record medication status and timing as covariates. We found no health DCE or shared decision-making study that uses Gaussian process PBO or an acquisition function.
Sources cited in Section 44.7 13
- Nouwens et al. (2025) The Evolving Landscape of Discrete Choice Experiments in Health Economics: A Systematic Review
- Crockett et al. (2008) Serotonin Modulates Behavioral Reactions to Unfairness
- Crockett et al. (2010) Serotonin selectively influences moral judgment and behavior through effects on harm aversion
- Voon et al. (2006) Prevalence of repetitive and reward-seeking behaviors in Parkinson disease
- Weintraub et al. (2010) Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients
- Quaife et al. (2018) How well do discrete choice experiments predict health choices? A systematic review and meta-analysis of external validity
- Zhang et al. (2025b) Prediction accuracy of discrete choice experiments in health-related research: a systematic review and meta-analysis
- Xie et al. (2022) Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation
- Whichello et al. (2023) Discrete choice experiment versus swing-weighting: A head-to-head comparison of diabetic patient preferences for glucose-monitoring devices
- Veldwijk et al. (2024) Comparing Discrete Choice Experiment with Swing Weighting to Estimate Attribute Relative Importance: A Case Study in Lung Cancer Patient Preferences
- Gonzalez Sepulveda et al. (2023) Patient-Preference Diagnostics: Adapting Stated-Preference Methods to Inform Effective Shared Decision Making
- Stacey et al. (2024) Decision aids for people facing health treatment or screening decisions
- Vanier et al. (2021) Response shift in patient-reported outcomes: definition, theory, and a revised model
44.8 Common claims, checked #
Table 44.1 collects the claims from design, the senses, art, and health that were checked in this chapter.
| Claim | What the sources say | Section |
|---|---|---|
| Color completely overrides taste and smell in wine, with devastating consequences for preference data | color biases verbal description and replicates; the task was choosing descriptors, and tasters still noticed a difference | Section 44.5 |
| Prospect-refuge theory is established | prospect supported in 19 of 31 tests (61%), refuge in 8 of 27 (30%); refuge supported mainly in landscapes | Section 44.2 |
| The Kano model shows that deep needs are stable | only 39% of "expect" answers repeated within seconds | Section 44.1 |
| The more fluent, the more liked | true for fast judgments; under deliberate processing, disfluent designs can win through interest | Section 44.6 |
| Liking follows an inverted U in complexity | at group level; in one data set it is a mixture of rising and falling individuals | Section 44.6 |
| SSRIs change moral preferences (Crockett et al., Science 2010) | PNAS 2010 (citalopram); Science 2008 is tryptophan depletion; the substance holds | Section 44.7 |
| 13.6% of levodopa patients develop impulse control disorders (Voon et al., Annals of Neurology 2006) | Voon et al. is Neurology: 13.7% on dopamine agonists; 13.6% is Weintraub et al. 2010; mainly agonists, with levodopa use also independently associated | Section 44.7 |
44.9 Methods PBO can adopt directly #
Table 44.2 gathers what in this chapter and the previous one is solid enough to use now, without new theory, in the order of building a PBO system; Chapter 42 adds controls for evaluation and deployment, and Chapter 46 turns all of it into advice. The third column is the book's inference.
| Method | Evidence | What changes in PBO (inference) |
|---|---|---|
| Placebo pairs: the same option twice (Section 44.5.1) | O'Mahony and Wichchukit (2017); Halim et al. (2020); Xia et al. (2020) | calibrate the noise or lapse term with the answers to identical pairs |
| A "no preference" option | O'Mahony and Wichchukit (2017); Quintana et al. (2025) | add a tie likelihood (Section 20.4) instead of reading indifference as a random choice |
| "Cannot decide" modeled apart from "about the same" (Section 43.4.2) | Ok and Tserenjigmid (2022); Cettolin and Riedl (2019) | distinguish indifference, indecision, and experimentation |
| Noise scale fitted per feedback type (Section 43.5) | Ghosal et al. (2023) | estimate noise separately for pairs, sliders, and ratings |
| Position, default, and choice-history terms (Equation (43.1)) | Matějka and McKay (2015); Enisman et al. (2021) | turn tilts near indifference into an estimable signal; balance sides and order |
| Diverse comparisons not derived from the incumbent (Section 44.1) | Wadinambiarachchi et al. (2024); Leahy et al. (2020) | add diverse candidates to galleries, and measure fixation |
| Latent-class prior with a short adaptive diagnosis (Section 44.7) | Gonzalez Sepulveda et al. (2023) (simulation) | spend early queries on identifying the user's preference class |
| Per-person models where taste is individual (Section 44.2) | Vessel et al. (2018); Brielmann et al. (2024) | per-person surrogates for buildings and art; population data for faces and landscapes |
| Comparison-graph and Hodge diagnostics (Section 43.3.1, Section 43.3.2) | Hendrickx et al. (2019); Jiang et al. (2011) | check how well utilities are anchored, and the cyclic share against chance |
| Test-retest benchmarks | Nijman et al. (2022); Xie et al. (2022) | flag posteriors far more confident than an ICC of 0.6 to 0.7 or 76% repeated choices |
| Values congruence, decision regret, and validation in use | Stacey et al. (2024); Quaife et al. (2018); Zhang et al. (2025b) | report human outcomes beside simulated regret; validate the "winner" in real use |
If a study can afford only a few changes, five have the best ratio of evidence to cost, and none needs new software beyond a likelihood with ties (inference):
- Mix identical pairs into the session, a few percent of the queries, and report the false-preference rate.
- Offer "no preference" (and, where it matters, "cannot decide") and model it.
- Fit the noise scale to this person and this interface, rather than fixing it.
- Randomize a control group to a random or balanced query order, and retest some early pairs at the end or a few days later.
- Report a human outcome, such as endorsement at retest, values congruence, or real use, next to regret.
Sources cited in Section 44.9 21
- O'Mahony and Wichchukit (2017) The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection
- Halim et al. (2020) Paired preference tests and placebo placement: 1. Should placebo pairs be placed before or after the target pair?
- Xia et al. (2020) Paired preference tests and placebo placement: 2. Unraveling the effects of stimulus variance
- Quintana et al. (2025) Global urban visual perception varies across demographics and personalities
- Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
- Cettolin and Riedl (2019) Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization
- Ghosal et al. (2023) The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
- Matějka and McKay (2015) Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model
- Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
- Wadinambiarachchi et al. (2024) The Effects of Generative AI on Design Fixation and Divergent Thinking
- Leahy et al. (2020) Design Fixation From Initial Examples: Provided Versus Self-Generated Ideas
- Gonzalez Sepulveda et al. (2023) Patient-Preference Diagnostics: Adapting Stated-Preference Methods to Inform Effective Shared Decision Making
- Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
- Brielmann et al. (2024) Modelling individual aesthetic judgements over time
- Hendrickx et al. (2019) Graph Resistance and Learning from Pairwise Comparisons
- Jiang et al. (2011) Statistical ranking and combinatorial Hodge theory
- Nijman et al. (2022) The stability of self-reported emotional response and liking of beer in context
- Xie et al. (2022) Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation
- Stacey et al. (2024) Decision aids for people facing health treatment or screening decisions
- Quaife et al. (2018) How well do discrete choice experiments predict health choices? A systematic review and meta-analysis of external validity
- Zhang et al. (2025b) Prediction accuracy of discrete choice experiments in health-related research: a systematic review and meta-analysis
44.10 Settled, contested, missing #
Settled. Consumers state preferences between identical samples, which is why sensory science uses placebo pairs and a "no preference" option. System-supplied examples, including the outputs of AI image generators, increase fixation and reduce the diversity of ideas. Taste for faces and landscapes is shared, taste for buildings and artworks individual, with equally reliable ratings where the two were compared. Individual models of liking far outperform population averages for individualized content (0.65 against 0.01). Choice-based and direct weighting give different trade-offs for the same people (14.9 against 1.4). Decision aids probably improve the congruence of choices with values. Color biases the verbal description of wine, replicably.
Contested. Whether the inverted U describes individuals or only groups. Whether repetition raises or lowers liking. Whether a handoff from humans to algorithms helps in design driven by taste, as it does in expert search with an objective metric. How well stated choices predict real refusals (specificity 34% to 52%). Whether co-evolution is specific to design.
Missing. A measurement of fixation within a PBO session, against a random gallery. Test-retest consistency of designers' pairwise preferences across sessions. PBO with human sensory panels for food or flavor. A measurement of how much presentation changes preference rather than description. A pairwise loop with residents driven by an acquisition function for public space. A health DCE or shared decision-making study that uses Gaussian process PBO. An observer-state model, such as Brielmann and Dayan's, used as the non-stationarity of a PBO surrogate.
44.11 Exercises #
In a pilot study, participants state a preference on 30% of placebo pairs. Using the derivation in Section 44.5.1, what tie threshold does a Thurstonian model infer, in units of the noise ? A second interface lowers the false-preference rate to 10%. What changed in the model's terms, and what might have changed in the interface?
Solution
With , . With , . In the model, either the threshold for reporting no preference rose or the perceptual noise fell, relative to each other; placebo pairs alone identify only the ratio. In the interface, the "no preference" option may have become more prominent, or the order changed (Halim et al. found that placing the placebo pair after the target pair raises "no preference" answers). To separate the two explanations, a study also needs pairs that differ by known amounts.
A PBO system renders each design with a random texture seed, so the same parameters never look exactly the same twice. What happens to placebo pairs in this system, and what does Xia et al. (2020) suggest about the consequence for the estimated false-preference rate?
Solution
Two renderings of the same parameters are no longer identical stimuli, so a "placebo" pair contains a real, if small, perceptual difference. Answers to it mix false preferences with genuine discriminations of the texture, and the estimated false-preference rate is biased; Xia et al. found exactly this when samples of the same food product varied. The fix is to render placebo pairs with the same seed, or, if variation between renderings is part of the product, to measure it separately with pairs of renderings of the same design.
Sources cited in Section 44.11 1
- Xia et al. (2020) Paired preference tests and placebo placement: 2. Unraveling the effects of stimulus variance
Further reading #
- O'Mahony and Wichchukit (2017) tells the history of paired preference testing, from forced choice to placebo pairs and , the most directly transferable method in this part of the book.
- Christensen et al. (2012) shows how the Thurstonian model of a preference test with a "no preference" option becomes a cumulative probit model.
- Vessel et al. (2018) separates shared from individual taste across domains.
- Brielmann and Dayan (2022) is the computational theory of aesthetic value that ties liking to the observer's changing state.
- Stacey et al. (2024) is the Cochrane review of decision aids, whose outcome measures PBO evaluations can borrow.
References
- (2017). Preference Elicitation For Participatory Budgeting. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §44.3
- (2021). Preference Elicitation for Participatory Budgeting. Management Science. Cited in §44.3
- (2022). A computational model of aesthetic value. Psychological Review. Cited in §44.6
- (2024). Modelling individual aesthetic judgements over time. Philosophical Transactions of the Royal Society B: Biological Sciences. Cited in §44.6 §44.9
- (2019). Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization. Journal of Economic Theory. Cited in §44.9
- (2022). Kano Analysis: A Critical Survey Science Review. Sawtooth Software Conference Proceedings (practitioner conference, not peer-reviewed). working paper Cited in §44.1
- (2019). Uncertainty and Surprise Jointly Predict Musical Pleasure and Amygdala, Hippocampus, and Auditory Cortex Activity. Current Biology. Cited in §44.6
- (2017). Back to the inverted-U for music preference: A review of the literature. Psychology of Music. Cited in §44.6
- (2012). Estimation of the Thurstonian Model for the 2-AC Protocol. Food Quality and Preference. Cited in §44.5
- (2020). Psychological and neural responses to architectural interiors. Cortex. Cited in §44.2
- (2021a). The Evolution of “Co-evolution” (Part I): Problem Solving, Problem Finding, and Their Interaction in Design and Other Creative Practices. She Ji: The Journal of Design, Economics, and Innovation. Cited in §44.1
- (2021b). The Evolution of “Co-evolution” (Part II): The Biological Analogy, Different Kinds of Co-evolution, and Proposals for Conceptual Expansion. She Ji: The Journal of Design, Economics, and Innovation. Cited in §44.1
- (2008). Serotonin Modulates Behavioral Reactions to Unfairness. Science. Cited in §44.7
- (2010). Serotonin selectively influences moral judgment and behavior through effects on harm aversion. Proceedings of the National Academy of Sciences. Cited in §44.7
- (2001). Creativity in the Design Process: Co-Evolution of Problem-Solution. Design Studies. Cited in §44.1
- (2016). Evidence for prospect-refuge theory: a meta-analysis of the findings of environmental preference research. City, Territory and Architecture. Cited in §44.2
- (2016). Deep Learning the City: Quantifying Urban Perception at a Global Scale. Computer Vision – ECCV 2016. Cited in §44.3
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §44.9
- (2023). The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. AAAI. Cited in §44.9
- (2019). Predictability and Uncertainty in the Pleasure of Music: A Reward for Learning? The Journal of Neuroscience. Cited in §44.6
- (2023). Patient-Preference Diagnostics: Adapting Stated-Preference Methods to Inform Effective Shared Decision Making. Medical Decision Making. Cited in §44.7 §44.9
- (2017). Aesthetic Pleasure versus Aesthetic Interest: The Two Routes to Aesthetic Liking. Frontiers in Psychology. Cited in §44.6
- (2016). Liking versus Complexity: Decomposing the Inverted U-curve. Frontiers in Human Neuroscience. Cited in §44.6
- (2020). Paired preference tests and placebo placement: 1. Should placebo pairs be placed before or after the target pair? Food Research International. Cited in §44.5 §44.9
- (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Cited in §44.9
- (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Cited in §44.4
- (1991). Design Fixation. Design Studies. Cited in §44.1
- (2011). Statistical ranking and combinatorial Hodge theory. Mathematical Programming. Cited in §44.9
- (2023). Human–machine collaboration for improving semiconductor process development. Nature. Cited in §44.4
- (2020). Design Fixation From Initial Examples: Provided Versus Self-Generated Ideas. Journal of Mechanical Design. Cited in §44.1 §44.9
- (2017). Repeated Listening Increases the Liking for Music Regardless of Its Complexity: Implications for the Appreciation and Aesthetics of Music. Frontiers in Neuroscience. Cited in §44.6
- (2003). Consumers report preferences when they should not: a cross-cultural study. Journal of Sensory Studies. Cited in §44.5
- (2015). Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model. American Economic Review. Cited in §44.9
- (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. Cited in §44.1
- (2001). The Color of Odors. Brain and Language. Cited in §44.5
- (2025). Exploring the Effectiveness of Interactive Preference Learning for Adapting Designs to Abstract Semantic Attributes. Journal of Mechanical Design. Cited in §44.4
- (2022). The stability of self-reported emotional response and liking of beer in context. Food Quality and Preference. doi:10.1016/j.foodqual.2022.104603. Cited in §44.5 §44.9
- (2025). The Evolving Landscape of Discrete Choice Experiments in Health Economics: A Systematic Review. PharmacoEconomics. Cited in §44.7
- (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Cited in §44.5 §44.9
- (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §44.9
- (2018). How well do discrete choice experiments predict health choices? A systematic review and meta-analysis of external validity. The European Journal of Health Economics. Cited in §44.7 §44.9
- (2025). Global urban visual perception varies across demographics and personalities. Nature Cities. Cited in §44.3 §44.9
- (2011). A Design Preference Elicitation Query as an Optimization Process. Journal of Mechanical Design. Cited in §44.4
- (2019). Differences between young architects' and non-architects' aesthetic evaluation of buildings. Frontiers of Architectural Research. Cited in §44.2
- (2020). Wine psychology: basic & applied. Cognitive Research: Principles and Implications. Cited in §44.5
- (2024). Decision aids for people facing health treatment or screening decisions. Cochrane Database of Systematic Reviews. Cited in §44.7 §44.9
- (2021). Response shift in patient-reported outcomes: definition, theory, and a revised model. Quality of Life Research. Cited in §44.7
- (2024). Comparing Discrete Choice Experiment with Swing Weighting to Estimate Attribute Relative Importance: A Case Study in Lung Cancer Patient Preferences. Medical Decision Making. Cited in §44.7
- (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §44.2 §44.9
- (2006). Prevalence of repetitive and reward-seeking behaviors in Parkinson disease. Neurology. Cited in §44.7
- (2024). The Effects of Generative AI on Design Fixation and Divergent Thinking. Proceedings of the CHI Conference on Human Factors in Computing Systems. Cited in §44.1 §44.9
- (2019). Drinking through rosé-coloured glasses: Influence of wine colour on the perception of aroma and flavour in wine experts and novices. Food Research International. Cited in §44.5
- (2010). Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients. Archives of Neurology. Cited in §44.7
- (2023). Discrete choice experiment versus swing-weighting: A head-to-head comparison of diabetic patient preferences for glucose-monitoring devices. PLOS ONE. Cited in §44.7
- (2020). Paired preference tests and placebo placement: 2. Unraveling the effects of stimulus variance. Food Research International. Cited in §44.5 §44.9 §44.11
- (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. Cited in §44.7 §44.9
- (2025b). Prediction accuracy of discrete choice experiments in health-related research: a systematic review and meta-analysis. eClinicalMedicine. Cited in §44.7 §44.9
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §44.1