Social, Affective, and Developmental Psychology
Chapter 37 examined the seconds in which a person decides between two options. This chapter turns to the branches of psychology that study people over longer spans and in richer settings: under the eyes of others, in pursuit of goals, as consumers, in moods, facing moral trade-offs, across the lifespan, and in rooms and cities. Each has a view on whether a person carries a stable preference that a sequence of comparisons could find.
Three findings run through them. Several famous effects once used to argue that answers depend on one another did not replicate in multi-lab tests, or shrank to near zero after correction for publication bias. Broad dispositions that people report about themselves are quite stable, while incentivized choices and single judgments are the least stable measures. And the evidence closest to preferential Bayesian optimization (PBO) comes from moral preference elicitation: when people answer the same pairwise question in different sessions, 6% to 20% of the answers flip, and in simulations in which preferences are unstable or the model class is wrong, actively chosen queries can do no better than random ones. Section 38.12 assembles the observation model these results point to; the yardsticks used throughout are glossed in Section 37.1.
38.1 Social psychology #
Social psychology contributes claims about choice, about attitudes people cannot report, and about the influence of others. Cognitive dissonance theory holds that people feel discomfort when their actions and attitudes conflict and reduce it by changing the attitude, and it has long been used to explain why the chosen option gains after a choice (Section 37.2.2). The dual attitudes model holds that people can have an implicit and an explicit attitude toward the same object at once (Wilson et al., 2000), so a user might have "two true preferences". And Asch's conformity experiments, usually summarized as showing that about 75% of participants went along with a wrong majority at least once (Asch, 1956), and information cascades, in which people copy earlier choices instead of using their own information (Bikhchandani et al., 1992), suggest that showing a user other people's ratings may create a false consensus.
38.1.1 Choice changes preference, from infancy #
Beyond the artifact-free meta-analysis with , 95% confidence interval (Section 37.2.2), Silver et al. (2020) found in 7 experiments () that preverbal infants avoid the option they did not choose, ruling out novelty and prior attitudes: choice shapes preference before a mature self-concept exists to be defended. The other classic route to dissonance did not replicate. In the induced-compliance paradigm, people write an essay against their own view, freely or under instruction, and the theory predicts more attitude change when they feel they chose to. Across 39 laboratories in 19 countries (), Vaidis et al. (2024) found no difference in attitudes between the high-choice and low-choice conditions, although writing a counter-attitudinal essay did change attitudes relative to writing a neutral one. This favors sequential sampling and revaluation (Section 37.2.2) over classic dissonance as the mechanism.
38.1.2 Implicit attitudes and conformity #
Implicit measures, such as the Implicit Association Test, which infers attitudes from how quickly people pair concepts, predict poorly for individuals. Across 217 reports (), implicit and explicit measures each made a unique contribution to predicting intergroup behavior, and , with highly heterogeneous correlations (Kurdi et al., 2019). A network meta-analysis of 492 studies found that implicit measures can be changed, but only weakly (), and implicit change did not mediate explicit or behavioral change (Forscher et al., 2019). The temporal stability of implicit measures (weighted mean ) was lower than that of explicit ones () (Gawronski et al., 2017). Conformity replicated: in an Asch replication with five confederates (), the error rate on the standard line-judgment task was 33%, and monetary incentives lowered it to 25% (Franzen and Mader, 2023).
38.1.3 Social influence as belief updating #
Social influence now reads as belief updating weighted by uncertainty. In a large sample of people aged 14 to 24, the tendency to adopt others' preferences in delay discounting (choosing between smaller rewards sooner and larger ones later) came from greater uncertainty about one's own preferences, and this uncertainty, and with it susceptibility to influence, declined over 1.5 years (Reiter et al., 2021). Artificial intelligence is now one of the sources: while people interacted with a biased AI system, the share of arrays of faces they classified as more sad than happy rose from 49.9% at baseline to 56.3% () (Glickman and Sharot, 2025).
38.1.4 What this means for the observation model #
- Keep a choice-history term, with a revaluation mechanism. It can predict the size of the drift from the difficulty of the choice and the response time, both of which an interface can record (inference).
- Do not show others' choices early. Other people's choices, popularity, or a system-endorsed default should not appear early in a session, when the user's own uncertainty, and with it susceptibility, is highest (inference). Section 20.5 discusses population models that pool people without showing one person's answers to another.
- Preference uncertainty is measurable. Inconsistency on repeated pairs estimates it, as a noise scale and as a warning that the person is open to influence (inference).
- Implicit measures are not a second preference channel. Response time is the better-supported auxiliary signal (inference), as Section 39.3 reports.
Sources cited in Section 38.1 11
- Wilson et al. (2000) A model of dual attitudes
- Asch (1956) Studies of independence and conformity: I. A minority of one against a unanimous majority
- Bikhchandani et al. (1992) A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades
- Silver et al. (2020) When Not Choosing Leads to Not Liking: Choice-Induced Preference in Infancy
- Vaidis et al. (2024) A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance
- Kurdi et al. (2019) Relationship between the Implicit Association Test and intergroup behavior: A meta-analysis
- Forscher et al. (2019) A meta-analysis of procedures to change implicit measures
- Gawronski et al. (2017) Temporal Stability of Implicit and Explicit Measures: A Longitudinal Analysis
- Franzen and Mader (2023) The power of social influence: A replication and extension of the Asch experiment
- Reiter et al. (2021) Preference uncertainty accounts for developmental effects on susceptibility to peer influence in adolescence
- Glickman and Sharot (2025) How human–AI feedback loops alter human perceptual, emotional and social judgements
38.2 Motivation and goals #
A hierarchical view of preference holds that ultimate preferences are stable while instrumental ones are constructed in context. In control theory, behavior is a hierarchy of goals, with higher goals more abstract and slower to change (Carver and Scheier, 1982); in self-determination theory, needs for autonomy, competence, and relatedness drive motivation (Deci and Ryan, 2000); and terminal values are said to be stable (Schwartz, 1992). If ultimate goals are stable, there is something stable for PBO to find.
38.2.1 Habits, needs, and values #
How habits relate to goals remains unresolved: Wood et al. (2022) argue that habits operate independently of goals, yet in human laboratories de Wit et al. (2018) report five failed attempts to induce habits by overtraining. Value hierarchies are stable for most people: among early adolescents over two years, for 75% of people the value hierarchy at the two times correlated at least .85, and for 5% at most .12 (Vecchione et al., 2020). But value is relative to a goal. When participants had to choose the worst option instead of the best, behavior and the neural activity usually read as value were dominated by goal congruency, how well an option served the current goal, rather than expected reward (Frömer et al., 2019).
38.2.2 What this means for the observation model #
- The utility is relative to the task framing. Choosing what you like, what suits a client, or which to rule out are different goals, so instructions should be held constant and recorded, and "which is worse?" should not share a likelihood with "which is better?" without a test (inference).
- Response habits are nuisance, not preference. Always choosing the left option or keeping the incumbent belongs in person-specific nuisance terms (inference).
- Autonomy gives a reason to bound influence, a concern Section 41.2 develops (inference).
Sources cited in Section 38.2 7
- Carver and Scheier (1982) Control theory: A useful conceptual framework for personality–social, clinical, and health psychology
- Deci and Ryan (2000) The "What" and "Why" of Goal Pursuits: Human Needs and the Self-Determination of Behavior
- Schwartz (1992) Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries
- Wood et al. (2022) Habits and Goals in Human Behavior: Separate but Interacting Systems
- de Wit et al. (2018) Shifting the balance between goals and habits: Five failures in experimental habit induction
- Vecchione et al. (2020) Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study
- Frömer et al. (2019) Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making
38.3 Consumer psychology #
Consumer research holds that making choices, and choosing repeatedly, make preferences more stable (Hoeffler and Ariely, 1999), and that early experiences strongly predict final preferences (Hoeffler et al., 2006), so the initial candidates of a session would decide what crystallizes. Choice overload is the claim that too many options harm choice: in a famous field study, 3% of shoppers bought jam when 24 varieties were on display against 30% when there were 6 (Iyengar and Lepper, 2000), but a meta-analysis of 63 conditions found a mean effect near zero with large variance (Scheibehenne et al., 2010).
38.3.1 Choice overload and crystallization #
Choice overload is real but highly conditional: a multilevel multivariate meta-analysis showed that the effect varies greatly across six dependent variables and four moderators, which interact (McShane and Böckenholt, 2018). When people chose from 6, 12, or 24 items, activity in the striatum and anterior cingulate cortex followed an inverted U with its peak at 12, the number participants also judged about right, and the pattern vanished when people only browsed without choosing (Reutskaja et al., 2018). We found no direct replication of preference crystallization. In a repeated discrete choice experiment with eye tracking, the latent preferences of unstable respondents did not differ from those of stable ones, and instability reflected the difficulty of choosing when utilities were close (Fraser et al., 2021).
38.3.2 What this means for the observation model #
- Inconsistency is a noise scale, not a user type. It should rise near indifference and for complex options (inference).
- Galleries may have a best size, but not a universal one. A multi-option interface (Section 20.3) may have an interior optimum that depends on whether the user must choose or may browse; 12 is a result from one domain (inference).
- Randomize and record the initial designs. Early experience and choice both shape where a session ends, so randomizing the starting designs across users, as the starting image settings in the photo-enhancement case study (Chapter 25) could be, separates what the optimizer did from what the user preferred (inference).
Sources cited in Section 38.3 7
- Hoeffler and Ariely (1999) Constructing Stable Preferences: A Look Into Dimensions of Experience and Their Impact on Preference Stability
- Hoeffler et al. (2006) Path dependent preferences: The role of early experience and biased search in preference development
- Iyengar and Lepper (2000) When choice is demotivating: Can one desire too much of a good thing?
- Scheibehenne et al. (2010) Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload
- McShane and Böckenholt (2018) Multilevel Multivariate Meta-analysis with Application to Choice Overload
- Reutskaja et al. (2018) Choice overload reduces neural signatures of choice set value in dorsal striatum and anterior cingulate cortex
- Fraser et al. (2021) Preference stability in discrete choice experiments. Some evidence using eye-tracking
38.4 Emotion and hedonic psychology #
Research on emotion supplies three claims. Hedonic adaptation: lottery winners were famously found to be no happier than others (Brickman et al., 1978), so early preference pairs might expire. Feelings as information: people consult their current mood when judging, as when ratings of life satisfaction depended on the weather (Schwarz and Clore, 1983), so a session might need an "emotional cool-down". And dual-process theories hold that fast forced choices capture intuitive preferences.
38.4.1 Three revisions #
All three need revision. A preregistered analysis of Swedish lottery players 5 to 22 years after winning found that life satisfaction rose and stayed higher for more than a decade with no sign of fading, while the effects on happiness and mental health were significantly smaller (Lindqvist et al., 2020); across 18 life events, affect adapted to every positive event within two years, but evaluative well-being kept lasting benefits of financial gains and retirement (Kettlewell et al., 2020). In nine direct and conceptual replications with larger samples ( from 118 to 401), the effect of mood on judgments of life satisfaction was mostly not significant, and much smaller than earlier results when it was (Yap et al., 2017). And in 21 preregistered replications pooled, the difference in contributions to a common project between people who had to decide quickly and people forced to wait was percentage points in an intention-to-treat analysis (which compares groups as assigned), against 8.6 percentage points in the original data, and 65.9% of the time-pressure group did not meet the time limit (Bouwmeester et al., 2017).
38.4.2 What this means for the observation model #
- Incidental mood is a weak noise source. Emotional cool-downs and mood-timed exploration lack support (inference).
- Two forces act on a repeated option, in opposite directions. Choice-induced revaluation raises the incumbent that keeps winning, while the decline in enjoyment with repeated consumption (Galak and Redden, 2018) lowers it; the sign of a drift term should be estimated, not fixed (inference).
- Fast forced choices do not reveal an intuitive self (inference).
- Satisfaction at the end of a session measures early affect. Liking now and fitness for purpose diverge over time, so a delayed follow-up is needed (inference), as Section 46.7 recommends.
Sources cited in Section 38.4 7
- Brickman et al. (1978) Lottery winners and accident victims: Is happiness relative?
- Schwarz and Clore (1983) Mood, misattribution, and judgments of well-being: Informative and directive functions of affective states
- Lindqvist et al. (2020) Long-Run Effects of Lottery Wealth on Psychological Well-Being
- Kettlewell et al. (2020) The differential impact of major life events on cognitive and affective wellbeing
- Yap et al. (2017) The effect of mood on judgments of subjective well-being: Nine tests of the judgment model
- Bouwmeester et al. (2017) Registered Replication Report: Rand, Greene, and Nowak (2012)
- Galak and Redden (2018) The Properties and Antecedents of Hedonic Decline
38.5 Moral psychology #
Moral psychology offers claims about the structure of preference and the most direct evidence on repeated pairwise elicitation. Sacred values resist trade-offs: a proposal to trade a sacred value for money provokes anger (Tetlock et al., 2000), from which comes the argument that sacred values should be hard constraints in PBO. A second argument holds that successive moral answers compensate for one another. In moral licensing, a good deed licenses a later lapse (Monin and Miller, 2001); in self-concept maintenance, honest people cheat only as much as lets them keep seeing themselves as honest, and reminders of morality reduce cheating (Mazar et al., 2008).
38.5.1 Honesty, licensing, and retractions #
In 25 direct replications of self-concept maintenance (), the primary meta-analysis (19 replications, ) found against an original (Verschuere et al., 2018). The Journal of Marketing Research issued an expression of concern about the 2008 paper in 2024 (Journal of Marketing Research, 2024), and a 2012 paper from the same group of authors, on signing an honesty pledge at the top of a form, has been retracted (Shu et al., 2012).
Moral licensing does not hold in the within-person form that the compensation argument needs. Rotella et al. (2026) pooled 115 experiments (). The uncorrected multilevel estimate was Hedges' ( is Cohen's with a small-sample correction), but a bias-corrected robust Bayesian meta-analysis gave an estimate near zero ( between and ), with very strong evidence of publication bias. The effect depended on being watched: when participants were explicitly observed, against when they were not, which the authors read as an interpersonal, reputational effect. A registered replication of the original moral credentials study () did not support a consistent effect (Xiao et al., 2024). Figure 38.1 sets these results beside the choice-induced change that held.
Things to try:
- At the original moral-reminder effect, , the distributions overlap 81% and a person from the higher group outscores one from the lower group 63% of the time. At the replication's they overlap 98%.
- Licensing under observation, , overlaps 75%; without it, , 95%.
38.5.2 Sacred values and dilemmas #
Sacred values are neither universal nor fixed. Seeing a sacred value used for profit lowered its sacredness and resistance to trade-offs (7 studies, ) (Ruttan and Nordgren, 2021). When a penalty for taboo trade-offs was added to a random utility model in discrete choice experiments, a latent class analysis showed some groups treating the trade-offs as taboo and others not, and ignoring the aversion overestimated the willingness to pay for saving lives by a factor of about 3.5 (Smeele et al., 2025). And hypothetical moral judgments do not predict real behavior: answers to a trolley-style dilemma did not predict whether people would actually deliver an electric shock to one mouse to spare five (Bostyn et al., 2018).
38.5.3 How stable are moral answers? #
Aggregated utilitarian judgments are highly consistent over time and across eight contexts (Helzer et al., 2017). Single judgments are not: between two waves 6 to 8 days apart, an average of 49% of participants changed their rating of a given sacrificial dilemma, and only about 17% of the changes came from people who said they had changed their mind (Rehren and Sinnott-Armstrong, 2022). Two studies asked the same pairwise question again and again. Asked ten times over two weeks which of two patients should receive one available kidney, people changed their answers on controversial scenarios about 10% to 18% of the time, more often for slower and harder choices (Boerstler et al., 2024). In a larger study, more than 400 participants answered pairwise kidney-allocation comparisons in 3 to 5 sessions; on average 6% to 20% of their answers to the same scenario changed, and the predictive performance of simple models fell with instability and over time (Keswani et al., 2026). Take a short version of the test in Figure 38.2.
The result most relevant to PBO concerns active learning, choosing each next question to be as informative as possible, as acquisition functions do (Section 19.3). Keswani et al. (2024) identify its three premises: preferences are stable and unaffected by question order, the hypothesis class is right, and noise is limited. In simulations that violate them, active learning did as well as or worse than random questions in some settings, and stayed worthwhile only when instability and noise were small and the preferences could be approximately represented by the hypothesis class.
38.5.4 Noise or change? #
A flip rate of 6% to 20% can come from a fixed preference answered with noise, or from a preference that moves between sessions, a possible "moral change" in the words of Keswani et al. (2026). Noise should be averaged away; change should be tracked. Let the long-run utility difference between two options be , shifted in each session by , with probit response noise of scale per option (Section 16.3). Two answers in one session share ; answers in two sessions do not. With , the probabilities that the two answers differ are
since two independent sessions each say "yes" with average probability . When , answers differ more often across sessions than within one, as Figure 38.3 shows.
With and the curves coincide, and the average flip rate, about 16%, is inside the reported range: pure noise is enough. With and the average across sessions is again about 15%, but within a session it falls to about 3%. Only comparing a retest within the session with one in a later session tells the two apart, which is why Section 37.4.4 recommends retesting either immediately or much later.
38.5.5 What this means for the observation model #
- Test acquisition functions against random queries. In value-sensitive domains such as allocation, policy, and safety trade-offs, compare them with random queries on held-out repeated pairs, and keep a fixed share of random or repeated queries in every session (inference); the comparisons in Section 28.9 rarely include such a control.
- Expect a floor on consistency. A flip rate of 6% to 20% across sessions calls for session-level random effects, and convergence within one session should not be extrapolated across sessions (inference).
- Replace hard constraints with a person-specific penalty on trading a protected attribute, placed in a mixture over classes of users (inference).
- Drop licensing, and check the evidence before citing it. Compensation based on moral licensing does not belong in private sessions, since licensing appears only under observation; who can see the answers is then an experimental factor. Papers that justify a design choice with a behavioral effect should check its replication and retraction status (inference).
- Hypothetical validation is not real validation (inference). A PBO study in a moral domain, one of the gaps in Section 38.13, should validate against real decisions, not only against further hypothetical answers.
Sources cited in Section 38.5 16
- Tetlock et al. (2000) The psychology of the unthinkable: Taboo trade-offs, forbidden base rates, and heretical counterfactuals
- Monin and Miller (2001) Moral credentials and the expression of prejudice
- Mazar et al. (2008) The Dishonesty of Honest People: A Theory of Self-Concept Maintenance
- Verschuere et al. (2018) Registered Replication Report on Mazar, Amir, and Ariely (2008)
- Journal of Marketing Research (2024) Expression of Concern: “The Dishonesty of Honest People: A Theory of Self-Concept Maintenance”
- Shu et al. (2012) RETRACTED: Signing at the beginning makes ethics salient and decreases dishonest self-reports in comparison to signing at the end
- Rotella et al. (2026) Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms
- Xiao et al. (2024) Licensing via Credentials: Replication Registered Report of Monin and Miller (2001) with Extensions Investigating the Domain-Specificity of Moral Credentials and the Association Between the Credential Effect and Trait Reputational Concern
- Ruttan and Nordgren (2021) Instrumental use erodes sacred values
- Smeele et al. (2025) Taboo trade-off aversion in choice behaviors: A discrete choice model and application to health-related decisions
- Bostyn et al. (2018) Of Mice, Men, and Trolleys: Hypothetical Judgment Versus Real-Life Behavior in Trolley-Style Moral Dilemmas
- Helzer et al. (2017) Once a Utilitarian, Consistently a Utilitarian? Examining Principledness in Moral Judgment via the Robustness of Individual Differences
- Rehren and Sinnott-Armstrong (2022) How Stable are Moral Judgments?
- Boerstler et al. (2024) On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
- Keswani et al. (2026) Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
- Keswani et al. (2024) On the Pros and Cons of Active Learning for Moral Preference Elicitation
38.6 Personality and lifespan development #
A widely cited meta-analysis found that the rank-order stability of personality traits rises from .31 in childhood to .64 at age 30 and reaches a plateau around .74 between ages 50 and 70 (Roberts and DelVecchio, 2000), from which comes the suggestion to treat data from middle-aged users as more reliable.
38.6.1 What is stable, and from when #
A meta-analysis of longitudinal studies published since 2005, covering rank-order stability (189 studies, ) and mean-level change (276 studies, ), found that rank-order stability reaches a plateau in early adulthood, with little evidence of further increase after age 25 (Bleidorn et al., 2022). Stated, trait-like preferences are quite stable: in 1,507 adults who completed 39 measures of risk preference, stated and behavioral measures correlated only weakly, but the stated measures formed a general factor that was highly reliable over six months (Frey et al., 2017). The Global Preferences Survey of 80,000 people in 76 countries found more heterogeneity within countries than between them (Falk et al., 2018).
38.6.2 What this means for the observation model #
- Population priors help, but each user still has to be learned, since heterogeneity within countries is large (inference); Section 32.5 reports how population priors have fared in interactive systems.
- Pair a few stated items with the comparisons. Pairwise choices belong to the less reliable, behavioral class of measures (Section 37.1). Stated items can carry the stable component, for example through a Gaussian process prior mean that depends on them (inference; the implementation comes from the Gaussian process framework).
- Report retest agreement across sessions, since convergence within one session is weak evidence of a stable preference (inference).
Sources cited in Section 38.6 4
- Roberts and DelVecchio (2000) The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies
- Bleidorn et al. (2022) Personality stability and change: A meta-analysis of longitudinal studies
- Frey et al. (2017) Risk preference shares the psychometric structure of major psychological traits
- Falk et al. (2018) Global Evidence on Economic Preferences*
38.7 Evolutionary psychology #
Sex differences in mate preferences, first reported across 37 cultures (Buss, 1989), replicated in 45 countries () (Walter et al., 2020), but they say little about individuals. A registered report (, 43 countries) found that stated ideal-partner preferences matched evaluations of partners in the overall pattern (), but the trait-specific effects averaged only (Eastwick et al., 2025). A prior initialized from the importance a user states for each attribute should therefore be discounted, or corrected from revealed comparisons (inference).
Sources cited in Section 38.7 3
- Buss (1989) Sex differences in human mate preferences: Evolutionary hypotheses tested in 37 cultures
- Walter et al. (2020) Sex Differences in Mate Preferences Across 45 Countries: A Large-Scale Replication
- Eastwick et al. (2025) A worldwide test of the predictive validity of ideal partner preference matching
38.8 Developmental psychology and behavioral genetics #
Heritability is the share of a trait's variation across a population that is associated with genetic differences. In 9,169 Swedish twins, genetic effects explained up to 54% of the variance in musical reward sensitivity (Bignardi et al., 2025), while individual preferences for faces come mostly from each person's unique environment (Germine et al., 2015). Heritability does not by itself set the strength of a prior; what a model needs is the split between a component shared across people and a person-specific one, and the weight on the shared part can itself be a person-level parameter (inference). Since choice-induced change exists from infancy (Section 38.1.1), it cannot be avoided by recruiting more reflective users (inference).
Sources cited in Section 38.8 2
- Bignardi et al. (2025) Twin modelling reveals partly distinct genetic pathways to music enjoyment
- Germine et al. (2015) Individual Aesthetic Preferences for Faces Are Shaped Mostly by Environments, Not Genes
38.9 Sleep and circadian rhythms #
A morning morality effect, in which people are more honest in the morning (Kouchaki and Smith, 2014), is cited to suggest eliciting preferences at a person's circadian peak. A conceptual replication with found an odds ratio of 1.04 (95% confidence interval ), and a meta-analysis found (Zickfeld et al., 2024), a 2024 preprint. Self-reported sleep is better used as a covariate for the noise scale and the lapse rate than as a signed bias on utility (inference).
Sources cited in Section 38.9 2
- Kouchaki and Smith (2014) The Morning Morality Effect
- Zickfeld et al. (2024) Investigating the Morning Morality Effect and its Mediating and Moderating Factors
38.10 Neurodiversity #
In a preregistered adversarial collaboration, people who reported an autism diagnosis learned much less about features irrelevant to the outcome, and the reduction scaled with autistic traits across the whole sample, a dimensional rather than categorical pattern (Ben-Artzi et al., 2026). Nuisance biases such as position, order, and irrelevant visual features therefore call for person-level nuisance parameters estimated from the interaction itself, which adapts to each user without collecting any diagnostic information (inference).
Sources cited in Section 38.10 1
- Ben-Artzi et al. (2026) Autism-associated learning patterns show reduced credit assignment to outcome-irrelevant features
38.11 Environmental psychology #
The quantitative evidence for prospect-refuge theory, which says people prefer places that offer a view while affording shelter, is inconsistent (Dosen and Ostwald, 2016). When 798 participants rated 200 interior spaces, three components, coherence, fascination, and hominess, explained 90% of the variance, and the structure replicated in an independent sample () (Coburn et al., 2020); they could serve as interpretable objectives in a multi-objective formulation (Section 14.5) (inference). 81,630 volunteers made 1.17 million pairwise comparisons of street-view images from 56 cities (Dubey et al., 2016), data that could supply population priors for urban design if kept specific to a culture or city (inference). Section 44.2 and Section 44.3 take up design for buildings and cities.
Sources cited in Section 38.11 3
- Dosen and Ostwald (2016) Evidence for prospect-refuge theory: a meta-analysis of the findings of environmental preference research
- Coburn et al. (2020) Psychological and neural responses to architectural interiors
- Dubey et al. (2016) Deep Learning the City: Quantifying Urban Perception at a Global Scale
38.12 A common pattern #
Across the branches the same pattern recurs (Table 38.1): broad self-reports are the most stable measures of preference, single choices and incentivized behavior the least, and population regularities predict little of one person's specific shape.
| Domain | Broad, self-reported | Single choices, behavior, or specific matches |
|---|---|---|
| Risk | stated propensity: mean reliability over time 0.61 | behavioral measures such as lottery choices: 0.25 |
| Morality | aggregated utilitarian judgments consistent over time | 49% change a dilemma rating within days; 6% to 20% of pairwise answers flip across sessions |
| Partners | stated ideals match evaluations in overall pattern, | trait-specific matching |
| Attitudes | explicit measures: stability | implicit measures: stability |
Taken together with Chapter 37, these results suggest an observation model of the following shape (inference). Let a person's utility in session be
where is a population mean, informed by population data or by a few stated preferences; is the person's stable component, the Gaussian process of Chapter 18; and is a session-level deviation with a small variance, which lets the utility move between sessions without moving within one. A comparison at step then follows
where is a lapse rate (the probability of an answer unrelated to the options), codes which option was on the left, is a person-specific position bias, and is a noise scale that may depend on the person's uncertainty and on how close the options are. Repeated pairs inform and , counterbalanced positions inform , and repeats in later sessions inform the variance of . No preferential optimization method we found implements this model, and its pieces have been tested separately, mostly outside optimization (inference). Two practices complete it: a share of random queries as a control for the acquisition function, and a delayed evaluation of the result, separate from the session that produced it (inference).
38.13 Settled, contested, missing #
Settled. Choice changes preference, from infancy on (Silver et al., 2020). The induced-compliance route to dissonance did not replicate across 39 labs (Vaidis et al., 2024). Moral reminders do not reduce cheating ( in the primary analysis of 25 direct replications) (Verschuere et al., 2018), and moral licensing appears only under observation (Rotella et al., 2026). Conformity on simple judgments is still about one in three (Franzen and Mader, 2023). Broad stated dispositions are stable while single choices are not, and personality stability plateaus at about 25 (Bleidorn et al., 2022). The same moral pairwise question flips 6% to 20% of the time across sessions (Keswani et al., 2026). Mood effects on judgments of well-being are much smaller than once reported (Yap et al., 2017), and time pressure does not reveal an intuitive preference (Bouwmeester et al., 2017).
Contested. Whether habits run independently of goals. Whether choice overload exists outside particular moderator combinations. Whether flips across sessions are noise or genuine change in preference, which the flip rate alone cannot decide (Figure 38.3). Whether stated and behavioral preferences measure one construct.
Missing. A direct replication of preference crystallization. A test of ultimate-goal stability and instrumental construction in the same design task. Any use of PBO in a moral domain. A PBO study that compares its acquisition function with random queries on held-out repeated pairs, models session-level deviations, or estimates person-specific nuisance parameters. Delayed evaluations of designs chosen by PBO, separate from the session.
Sources cited in Section 38.13 9
- Silver et al. (2020) When Not Choosing Leads to Not Liking: Choice-Induced Preference in Infancy
- Vaidis et al. (2024) A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance
- Verschuere et al. (2018) Registered Replication Report on Mazar, Amir, and Ariely (2008)
- Rotella et al. (2026) Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms
- Franzen and Mader (2023) The power of social influence: A replication and extension of the Asch experiment
- Bleidorn et al. (2022) Personality stability and change: A meta-analysis of longitudinal studies
- Keswani et al. (2026) Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
- Yap et al. (2017) The effect of mood on judgments of subjective well-being: Nine tests of the judgment model
- Bouwmeester et al. (2017) Registered Replication Report: Rand, Greene, and Nowak (2012)
Further reading #
- Keswani et al. (2024) is a short, clear statement of the assumptions behind active preference elicitation and what happens when they fail; read it with the repeated-elicitation data of Boerstler et al. (2024) and Keswani et al. (2026).
- Rotella et al. (2026) is a model of how a bias-corrected meta-analysis can turn a famous effect into a conditional one.
- Frey et al. (2017) explains why stated and behavioral measures of the same preference disagree, and which one is stable.
- Eastwick et al. (2025) tests whether stated preferences predict evaluations of real people, a direct analogue of initializing a prior from stated attribute importance.
References
- (1956). Studies of independence and conformity: I. A minority of one against a unanimous majority. Psychological Monographs: General and Applied. Cited in §38.1
- (2026). Autism-associated learning patterns show reduced credit assignment to outcome-irrelevant features. Translational Psychiatry. Cited in §38.10
- (2025). Twin modelling reveals partly distinct genetic pathways to music enjoyment. Nature Communications. Cited in §38.8
- (1992). A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades. Journal of Political Economy. Cited in §38.1
- (2022). Personality stability and change: A meta-analysis of longitudinal studies. Psychological Bulletin. Cited in §38.6 §38.13
- (2024). On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods. AIES. Cited in §38.5
- (2018). Of Mice, Men, and Trolleys: Hypothetical Judgment Versus Real-Life Behavior in Trolley-Style Moral Dilemmas. Psychological Science. Cited in §38.5
- (2017). Registered Replication Report: Rand, Greene, and Nowak (2012). Perspectives on Psychological Science. Cited in §38.4 §38.13
- (1978). Lottery winners and accident victims: Is happiness relative? Journal of Personality and Social Psychology. Cited in §38.4
- (1989). Sex differences in human mate preferences: Evolutionary hypotheses tested in 37 cultures. Behavioral and Brain Sciences. Cited in §38.7
- (1982). Control theory: A useful conceptual framework for personality–social, clinical, and health psychology. Psychological Bulletin. Cited in §38.2
- (2020). Psychological and neural responses to architectural interiors. Cortex. Cited in §38.11
- (2018). Shifting the balance between goals and habits: Five failures in experimental habit induction. Journal of Experimental Psychology: General. Cited in §38.2
- (2000). The "What" and "Why" of Goal Pursuits: Human Needs and the Self-Determination of Behavior. Psychological Inquiry. Cited in §38.2
- (2016). Evidence for prospect-refuge theory: a meta-analysis of the findings of environmental preference research. City, Territory and Architecture. Cited in §38.11
- (2016). Deep Learning the City: Quantifying Urban Perception at a Global Scale. Computer Vision – ECCV 2016. Cited in §38.11
- (2025). A worldwide test of the predictive validity of ideal partner preference matching. Journal of Personality and Social Psychology. Cited in §38.7
- (2018). Global Evidence on Economic Preferences*. The Quarterly Journal of Economics. Cited in §38.6
- (2019). A meta-analysis of procedures to change implicit measures. Journal of Personality and Social Psychology. Cited in §38.1
- (2023). The power of social influence: A replication and extension of the Asch experiment. PLOS ONE. Cited in §38.1 §38.13
- (2021). Preference stability in discrete choice experiments. Some evidence using eye-tracking. Journal of Behavioral and Experimental Economics. Cited in §38.3
- (2017). Risk preference shares the psychometric structure of major psychological traits. Science Advances. Cited in §38.6
- (2019). Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making. Nature Communications. Cited in §38.2
- (2018). The Properties and Antecedents of Hedonic Decline. Annual Review of Psychology. Cited in §38.4
- (2017). Temporal Stability of Implicit and Explicit Measures: A Longitudinal Analysis. Personality and Social Psychology Bulletin. Cited in §38.1
- (2015). Individual Aesthetic Preferences for Faces Are Shaped Mostly by Environments, Not Genes. Current Biology. Cited in §38.8
- (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. doi:10.1038/s41562-024-02077-2. Cited in §38.1
- (2017). Once a Utilitarian, Consistently a Utilitarian? Examining Principledness in Moral Judgment via the Robustness of Individual Differences. Journal of Personality. Cited in §38.5
- (1999). Constructing Stable Preferences: A Look Into Dimensions of Experience and Their Impact on Preference Stability. Journal of Consumer Psychology. Cited in §38.3
- (2006). Path dependent preferences: The role of early experience and biased search in preference development. Organizational Behavior and Human Decision Processes. Cited in §38.3
- (2000). When choice is demotivating: Can one desire too much of a good thing? Journal of Personality and Social Psychology. Cited in §38.3
- (2024). Expression of Concern: “The Dishonesty of Honest People: A Theory of Self-Concept Maintenance”. Journal of Marketing Research. Cited in §38.5
- (2024). On the Pros and Cons of Active Learning for Moral Preference Elicitation. AIES. Cited in §38.5
- (2026). Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback. AAAI. Cited in §38.5 §38.13
- (2020). The differential impact of major life events on cognitive and affective wellbeing. SSM - Population Health. Cited in §38.4
- (2014). The Morning Morality Effect. Psychological Science. Cited in §38.9
- (2019). Relationship between the Implicit Association Test and intergroup behavior: A meta-analysis. American Psychologist. Cited in §38.1
- (2020). Long-Run Effects of Lottery Wealth on Psychological Well-Being. The Review of Economic Studies. Cited in §38.4
- (2008). The Dishonesty of Honest People: A Theory of Self-Concept Maintenance. Journal of Marketing Research. Cited in §38.5
- (2018). Multilevel Multivariate Meta-analysis with Application to Choice Overload. Psychometrika. Cited in §38.3
- (2001). Moral credentials and the expression of prejudice. Journal of Personality and Social Psychology. Cited in §38.5
- (2022). How Stable are Moral Judgments? Review of Philosophy and Psychology. doi:10.1007/s13164-022-00649-7. Cited in §38.5
- (2021). Preference uncertainty accounts for developmental effects on susceptibility to peer influence in adolescence. Nature Communications. Cited in §38.1
- (2018). Choice overload reduces neural signatures of choice set value in dorsal striatum and anterior cingulate cortex. Nature Human Behaviour. Cited in §38.3
- (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin. Cited in §38.6
- (2026). Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms. Personality and Social Psychology Bulletin. Cited in §38.5 §38.13
- (2021). Instrumental use erodes sacred values. Journal of Personality and Social Psychology. Cited in §38.5
- (2010). Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload. Journal of Consumer Research. Cited in §38.3
- (1992). Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries. Advances in Experimental Social Psychology. Cited in §38.2
- (1983). Mood, misattribution, and judgments of well-being: Informative and directive functions of affective states. Journal of Personality and Social Psychology. Cited in §38.4
- (2012). RETRACTED: Signing at the beginning makes ethics salient and decreases dishonest self-reports in comparison to signing at the end. Proceedings of the National Academy of Sciences. Cited in §38.5
- (2020). When Not Choosing Leads to Not Liking: Choice-Induced Preference in Infancy. Psychological Science. Cited in §38.1 §38.13
- (2025). Taboo trade-off aversion in choice behaviors: A discrete choice model and application to health-related decisions. Social Science & Medicine. Cited in §38.5
- (2000). The psychology of the unthinkable: Taboo trade-offs, forbidden base rates, and heretical counterfactuals. Journal of Personality and Social Psychology. Cited in §38.5
- (2024). A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance. Advances in Methods and Practices in Psychological Science. Cited in §38.1 §38.13
- (2020). Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study. Journal of Personality. Cited in §38.2
- (2018). Registered Replication Report on Mazar, Amir, and Ariely (2008). Advances in Methods and Practices in Psychological Science. Cited in §38.5 §38.13
- (2020). Sex Differences in Mate Preferences Across 45 Countries: A Large-Scale Replication. Psychological Science. Cited in §38.7
- (2000). A model of dual attitudes. Psychological Review. Cited in §38.1
- (2022). Habits and Goals in Human Behavior: Separate but Interacting Systems. Perspectives on Psychological Science. Cited in §38.2
- (2024). Licensing via Credentials: Replication Registered Report of Monin and Miller (2001) with Extensions Investigating the Domain-Specificity of Moral Credentials and the Association Between the Credential Effect and Trait Reputational Concern. International Review of Social Psychology. Cited in §38.5
- (2017). The effect of mood on judgments of subjective well-being: Nine tests of the judgment model. Journal of Personality and Social Psychology. Cited in §38.4 §38.13
- (2024). Investigating the Morning Morality Effect and its Mediating and Moderating Factors. PsyArXiv. preprint Cited in §38.9