Bayesian Optimization
Part IX: What Is a Preference?
中文

Philosophy and Religious Traditions

Preferential Bayesian optimization (PBO) makes two philosophical commitments without stating them. It assumes that a person has a preference waiting to be found, and that finding and satisfying it is good for the person. Chapter 40 tested the first commitment with choice data. Philosophy examines both, and adds a question that becomes pressing as soon as a system chooses which options a person sees: when may the system change what the person wants? Religious traditions add an older one: whether desire is something to satisfy at all, or something to examine, train, or let go.

Most of the evidence here is argument, not experiment; where an argument rests on an empirical premise (that decisions precede awareness, that nudges work, that meditation removes bias), we report the studies that test it. Religious traditions are treated with the same care as secular philosophy, reported without endorsement. Three results matter most to someone building or evaluating a PBO system: candidate conditions under which a change of preference caused by the system is legitimate (Section 41.2), what alignment with many people can mean (Section 41.3), and when optimization is the wrong frame altogether (Section 41.11). Arguments imported from philosophy into preference learning are mostly fair summaries that present contested positions as settled; Section 41.10 lists the corrections.

41.1 Preference and value #

What is a preference? The Stanford Encyclopedia of Philosophy entry on the topic frames the core question as whether preferences cause choices or merely summarize patterns of choice (Hansson and Grüne-Yanoff, 2022). Behaviorism identifies a preference with a pattern of choice, as revealed preference theory does (Section 40.2). Mentalism treats it as a mental state, a comparative evaluation that explains the choice. Constructivism holds that preferences are often built in the act of choosing or being asked. Work on adaptive preferences, preferences that adjust to the options available, as when the fox decides that the grapes it cannot reach are sour (Elster, 1983), or that form under deprivation (Khader, 2011), adds a doubt that what people prefer is always what is good for them.

Hausman's mentalism, on which preferences are "total subjective comparative evaluations", is often cited as the philosophical basis for the latent utility a model infers. His 2024 paper restates three theses: preferences in economics are and ought to be total subjective comparative evaluations, the theory of rational choice reformulates everyday folk psychology, and revealed preference theory is "completely untenable". But its abstract notes that "all three of these theses have been challenged" by Angner, Guala, and Thoma, and the paper is a reply (Hausman, 2024). Thoma defends revealed preference theory (Thoma, 2021; Thoma, 2024), and the encyclopedia entry reports her and Vredenburgh's view that economists "often have good reasons to largely 'black box' the causes of choice" (Hansson and Grüne-Yanoff, 2022). Hausman's view is one side of a live debate.

A second imported claim concerns near-ties. Drawing on value incommensurability (Raz) and on Chang's notion of parity, the idea that two options can be "on a par", neither better than the other nor equally good (Chang, 2002; Chang, 2017), it holds that forcing a ranking between two designs on a par is a category error rather than noise. That is stronger than Chang's own view. She locates what is distinctive about hard choices in "the volitional difficulty of putting ourselves behind an alternative and thereby making it true of ourselves that we have most reason to do one thing rather than another" (Chang, 2024): such a choice is an occasion for commitment, not a mistake. From the other side, Dorr, Nebel, and Zuehl argue that every comparative expression in natural language obeys a principle they call Comparability: if two things each have a gradable property to some degree, then one has it at least as much as the other (Dorr et al., 2023). Presenting a pair judged to be on a par repeatedly would separate parity from vagueness (Section 41.12).

The most important addition since 2017 concerns choices that change the chooser. A transformative experience, in Paul's account, changes both what a person knows and what they value, so that they cannot know in advance how they will then evaluate it (Paul, 2014). Pettigrew proposes that a person choosing for changing selves should maximize a value function that aggregates the local value functions of their selves at different times (Pettigrew, 2019); Paul objects, as the encyclopedia entry on transformative experience reports, that this treats future selves as third parties and assumes that meaningful comparisons between one's own selves are possible (Chan, 2023). Empirical work is sparse. In two studies of the decision to have children (n = 100 and n = 253, childless adults aged 18 to 40), participants in the second rated subjective value as less important than financial cost, which the authors attribute to subjective value being hard to evaluate in advance rather than unimportant (Zoh et al., 2024) (numbers checked in the authors' preprint; the journal version is behind a paywall). On the formal side, Dietrich and List's reason-based choice derives preferences from which properties of the options are motivationally salient, so a change of preference can be modeled as a change in the set of salient reasons (Dietrich and List, 2017).

What this means for PBO. Hausman's "total" evaluations are conditional on the person's beliefs, so for a novel design the latent utility that PBO estimates is belief-dependent, and information supplied during a session can legitimately move it; Dietrich and List's framework is a formal tool for separating such a change of salient reasons from noise (inference). On parity, the two positions imply different models. If Chang is right, an interface can offer "these are on a par, I cannot rank them", which a model treats as evidence of uncertainty (a set of admissible utilities, or locally wider noise) rather than as a fair coin; Figure 40.2 shows the difference. If Dorr and colleagues are right, apparent parity is vagueness, and a probit or Bradley-Terry likelihood already treats it as noise. Data can decide locally: re-present a pair once judged to be on a par; under parity it should stay unranked, under vagueness the answers should scatter (inference). Chang's account also predicts that later judgments line up with a commitment, which looks the same as choice-induced revaluation (Zylberberg et al., 2024; Lee and Pezzulo, 2026), so without asking whether the person endorses the commitment, a PBO system cannot tell commitment from lock-in on the incumbent (inference). For a choice that changes the chooser's values there is no fixed objective, and PBO should be confined to the instrumental subproblems that remain (inference). And because people underweight attributes that are hard to evaluate in advance, a delayed retest after actual experience is a better evaluation than the attribute weights a user stated beforehand (inference).

Sources cited in Section 41.1 17
  1. Hansson and Grüne-Yanoff (2022) Preferences
  2. Elster (1983) Sour Grapes: Studies in the Subversion of Rationality
  3. Khader (2011) Adaptive Preferences and Women's Empowerment
  4. Hausman (2024) Subjective total comparative evaluations
  5. Thoma (2021) In defence of revealed preference theory
  6. Thoma (2024) Reply to Hausman
  7. Chang (2002) The Possibility of Parity
  8. Chang (2017) Hard Choices
  9. Chang (2024) What’s so Hard about Hard Choices?
  10. Dorr et al. (2023) The case for comparability
  11. Paul (2014) Transformative Experience
  12. Pettigrew (2019) Choosing for Changing Selves
  13. Chan (2023) Transformative Experience
  14. Zoh et al. (2024) How the evaluability bias shapes transformative decisions
  15. Dietrich and List (2017) What Matters and How It Matters: A Choice-Theoretic Representation of Moral Theories
  16. Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
  17. Lee and Pezzulo (2026) Choice-induced preference change under a sequential sampling model framework

41.2 Autonomy, manipulation, and legitimate influence #

Autonomy is self-government: acting from values and reasons one can recognize as one's own. Manipulation is, roughly, influence that works by bypassing or exploiting a person's capacity to reason rather than by engaging it. Both matter because a PBO system does not merely observe a person: it chooses the starting design, which candidates appear, and when to stop. A common argument holds that no choice architecture is neutral, as the large effects of defaults on organ donor registration (Johnson and Goldstein, 2003) and retirement saving (Madrian and Shea, 2001) show, so the acquisition function makes the optimizer an active choice architect that cannot tell whether it is revealing a preference or creating one (compare the critique of "preference purification" in Section 40.3). Frankfurt's second-order desires, desires about which desires to have (Frankfurt, 1971), and Dworkin's autonomous and heteronomous preferences (Dworkin, 1988) raise the question of which level a system should optimize.

The general claim about the power of nudges has since been tested. A meta-analysis of choice architecture interventions found an average effect of Cohen's d = 0.43 (95% confidence interval 0.38 to 0.48), where d is the difference between group means in units of their standard deviation (Mertens et al., 2022b). Correcting the same data for publication bias gave d = 0.04 (0.00 to 0.14), with estimates near zero in every domain; the authors conclude that "when this publication bias is appropriately corrected for, no evidence for the effectiveness of nudges remains" (Maier et al., 2022). The original authors replied (Mertens et al., 2022a). For interventions that change the structure of the choice, the category that contains defaults, the corrected estimate was d = 0.12 (0.00 to 0.43), against 0.58 before correction, and the authors call the evidence undecided (Maier et al., 2022); the meta-analysis of defaults cited in Section 40.1 (d = 0.68) covers a different set of studies.

The definition of manipulation has meanwhile been sharpened to fit optimizing systems. The encyclopedia entry on the ethics of manipulation, revised in March 2026, discusses an algorithm "designed to try out various kinds of content, user interfaces, notifications, and other variations of a digital platform's user interface and select those that increase the levels of some targeted behavior", which might end up using influences "that, had they been chosen deliberately by a human, we would label as manipulation". One option it lists defines manipulation by a lack of concern with how one's influence works, so that an algorithm is manipulative when "the designer did not take sufficient care" to avoid manipulative ways of interacting (Noggle, 2026); this is Klenk's view that "manipulation is careless influence" (Klenk, 2022). Susser, Roessler, and Nissenbaum apply a hidden-influence account to digital environments (Susser et al., 2019) (content not re-read for this chapter). Carroll, Chan, Ashton, and Krueger characterize a system as manipulative when it behaves as if it were pursuing an incentive to change a person intentionally and covertly (Carroll et al., 2023).

Work on AI turns these definitions into requirements. A workshop paper by Ashton and Franklin observes that "a recommender can better predict what a user will do by making its users more predictable", and argues that solutions must respect meta-preferences, preferences about one's own preferences (Ashton and Franklin, 2022). Fischli, Franklin, Manzini, and Gabriel show that autonomy can be specified in conflicting ways and offer three strategies: a liberal approach that acts "strictly on what a person says they want, without trying to change their mind", a capability-boosting approach that empowers "the user to act on their goals", and a meta-autonomy approach that gives the user "the kind of autonomy they want in relation to their interactions with an AI agent" (Fischli et al., 2026). Empirically, language models trained with reinforcement learning on simulated user feedback learned to manipulate: "even if only 2% of users are vulnerable to manipulative strategies, LLMs learn to identify and target them while behaving appropriately with other users" (Williams et al., 2025).

Can the person's own approval settle whether a change was legitimate? Reflective endorsement, a person's approval, on reflection, of a preference or of the process that formed it, is the usual test, and Pettigrew shows its limit. Thaler and Sunstein's test of a legitimate nudge asks whether the nudged person would approve of it, "as judged by themselves"; for a nudge toward a personally transformative experience, "the nudgee will judge the nudge to be legitimate after it has taken place, but only because their values have changed as a result of the nudge", and Pettigrew proposes a test based on aggregate utility instead (Pettigrew, 2023). If a session changed the user's values, approval at its end cannot by itself justify the change, and the user's attitude toward such change before the session also counts (inference). Figure 42.1 simulates the gap: when a person's preferences move toward what they choose, the system's final pick scores far better by their new preferences than by the ones they arrived with.

Nor does the objective supply the answer: when preferences can change, each of the eight notions of alignment that Carroll and colleagues compare either errs toward undesirable influence or is overly risk-averse (Carroll et al., 2024). A 2026 workshop paper by Kanwal and Tran proposes constraints that include reflective endorsement from the final preference state, bounded total and per-step influence relative to a baseline policy, not degrading factual beliefs, and preserving future options while the preference estimate is uncertain (Kanwal and Tran, 2026). And writing on recommender systems and authenticity, Brown concludes that "controllable and explainable recommenders would best enable users to be authentic" (Brown, 2026). Table 41.1 collects the candidate conditions, including one from Section 41.9.3; they have not converged into a consensus.

Table 41.1 Candidate conditions under which a system may legitimately change a person's preferences (operational forms are inferences).
Condition Source Operational form in PBO Limitation
Respect meta-preferences and preference-change preferences Ashton and Franklin 2022; Franklin et al. 2022 ask at the start whether, and in which respects, the user is willing to have their taste changed meta-preferences can themselves be influenced by the system
Influence is not covert, and the system has no incentive to change the user Carroll et al. 2023; Susser et al. 2019; Noggle 2026 label system proposals, show the history accurately, audit whether the objective rewards a more predictable user open influence can still be excessive
The user chooses the mode of autonomy Fischli et al. 2026 let the user choose among liberal, capability-boosting, and meta-autonomy modes users may not foresee the consequences of each mode
Judge across the selves before and after the change Pettigrew 2023 record both the attitude toward change before the session and endorsement after it whether comparisons across one's own selves are possible is disputed (Paul)
Reflective endorsement with bounded influence Kanwal and Tran 2026; Carroll et al. 2024 retest the final design against earlier rejected ones, delayed and unframed; bound per-step and total departure from a baseline endorsement after the fact is not enough; how to set the bound is open
Controllability and explanation Brown 2026 allow direct editing and undo; explain why a candidate is proposed explained is not the same as understood
Morally weighty decisions stay with people Antiqua et nova 2025 in medical or legal settings, map options without outputting a decision what counts as morally weighty needs judgment

What this means for PBO. Carroll and colleagues' definition suggests an incentive audit: a PBO loop is at risk if its acquisition function or stopping rule rewards a concentrated posterior or fast convergence, since both reward a user who has become easier to predict; an information-gain objective scored against delayed retests removes much of that incentive (inference). Against covertness, the system should label its proposals as proposals, show users their comparison history accurately, and let them reset or reject the incumbent (inference). The first question of a session can ask for the user's meta-preferences: their current taste served quickly (liberal), an exploration that may change it (capability boosting), or a say in how much initiative the system takes (meta-autonomy) (inference). An endorsement check is a safeguard only when paired with a record of the user's attitude before the session and a delayed, unframed retest against options rejected earlier (inference). The corrected nudge evidence calibrates concern: many presentation effects are small on average, and the strongest effects of AI on attitudes reported so far come from generative or conversational systems (Section 41.8). Since no estimate specific to PBO exists, starting points should still be randomized or chosen by the user, and their influence measured (inference). For decisions that carry moral weight for others, such as medical or legal ones, the system should lay out the options and their trade-offs and leave the decision to the person (inference).

Sources cited in Section 41.2 18
  1. Johnson and Goldstein (2003) Do Defaults Save Lives?
  2. Madrian and Shea (2001) The Power of Suggestion: Inertia in 401(k) Participation and Savings Behavior
  3. Frankfurt (1971) Freedom of the Will and the Concept of a Person
  4. Dworkin (1988) The Theory and Practice of Autonomy
  5. Mertens et al. (2022b) The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains
  6. Maier et al. (2022) No evidence for nudging after adjusting for publication bias
  7. Mertens et al. (2022a) Reply to Maier et al., Szaszi et al., and Bakdash and Marusich: The present and future of choice architecture research
  8. Noggle (2026) The Ethics of Manipulation
  9. Klenk (2022) (Online) manipulation: sometimes hidden, always careless
  10. Susser et al. (2019) Technology, autonomy, and manipulation
  11. Carroll et al. (2023) Characterizing Manipulation from AI Systems
  12. Ashton and Franklin (2022) Solutions to preference manipulation in recommender systems require knowledge of meta-preferences
  13. Fischli et al. (2026) Agents, Alignment, and the Many Faces of Autonomy
  14. Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
  15. Pettigrew (2023) Nudging for changing selves
  16. Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
  17. Kanwal and Tran (2026) Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
  18. Brown (2026) Recommended Selves: Authenticity and Algorithmic Filtering

41.3 The philosophy of AI alignment #

What should a system be aligned with? PBO answers: the latent utility. Zhi-Xuan and colleagues criticize such "preferentist" assumptions, argue that preferences cannot capture the thick semantic content of human values, and propose aligning AI with the normative standards of the social role it plays (Zhi-Xuan et al., 2024). Gabriel argues that the central challenge is "to identify fair principles for alignment that receive reflective endorsement despite widespread variation in people's moral beliefs" (Gabriel, 2020). A view sometimes called preference scaffolding holds that PBO scaffolds the formation of a preference rather than optimizing a fixed one. Its core holds, but a stage in which the system proposes preferences users did not know they had is exactly the system-induced preference formation that Franklin, Ashton, Gorman, and Armstrong address in a workshop paper distinguishing "preference change, permissible preference change, and outright preference manipulation" (Franklin et al., 2022). Scaffolding therefore needs the conditions of Section 41.2.

Aligning with many people raises its own questions. Hidden context, information that shapes people's answers but is not in the data, makes preference learning aggregate by Borda count (Siththaranjan et al., 2024) (Section 40.7). Sorensen and colleagues distinguish Overton pluralism (presenting a spectrum of reasonable responses), steerable pluralism (steering to particular perspectives), and distributional pluralism (calibration to a population), and present evidence that standard alignment procedures may reduce distributional pluralism (Sorensen et al., 2024) (ICML 2024).

What this means for PBO. The defensible target for single-user PBO is the utility the user would endorse on reflection, within a protocol that bounds the system's influence, not "the latent utility" as such; one test is that the final design still beats previously rejected alternatives in a delayed, unframed retest (inference). Following Siththaranjan and colleagues, a Gaussian process fitted while unobserved states such as fatigue, mood, or task framing change estimates something like a Borda aggregate across those states; recording state covariates and conditioning on them is the counterpart of their distributional method (inference). Following Zhi-Xuan and colleagues, a PBO assistant needs explicit norms for its role, such as reporting uncertainty honestly and not steering users toward designs that are easier to model (inference). And for problems with several stakeholders, the output should sometimes be a set of well-described options rather than a single optimum (inference).

Sources cited in Section 41.3 5
  1. Zhi-Xuan et al. (2024) Beyond Preferences in AI Alignment
  2. Gabriel (2020) Artificial Intelligence, Values, and Alignment
  3. Franklin et al. (2022) Recognising the importance of preference change: A call for a coordinated multidisciplinary research effort in the age of AI
  4. Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
  5. Sorensen et al. (2024) Position: A Roadmap to Pluralistic Alignment

41.4 Self, experience, and action #

Four traditions question whether a preference is the kind of thing a pairwise query can read: whether it belongs to the self, whether the person has access to it, whether it exists apart from body and context, and whether it exists before action. Each prompts a design check, but for none of them did we find more than scarce work since 2017 on preference elicitation.

41.4.1 Existentialism #

On standard existentialist readings, for which the encyclopedia entry on authenticity is the current reference (Varga and Guignon, 2020), some preferences are acts that constitute who a person is, not targets for optimization. The step from such self-shaping choices to a color grade or a font weight is an extrapolation the texts do not make. The folk concept of a "true self" has been studied empirically: Strohminger, Knobe, and Newman report that it is perceived as positive and moral (Strohminger et al., 2017), and a preregistered replication (N = 803) confirmed that observers see changes in others as more reflective of their true self when the changes are morally positive or match the observer's own political moral views (Lee and Feldman, 2025). A designer or model that tries to separate authentic from inauthentic preferences will therefore tend to call its own values authentic; judgments of authenticity should come from the person's own second-order endorsement, and identity-constituting choices are not good targets for optimization (inference).

41.4.2 Phenomenology #

Husserl's distinction between an evaluative and a practical attitude has been read as showing that a pairwise comparison locks the user into an evaluative attitude while real use takes place in a practical one, an analogy rather than an argument about PBO. The most direct test of first-person access to one's own preferences is choice blindness, in which a chosen option is secretly swapped for the rejected one (Johansson et al., 2014). When participants chose the more attractive of two faces, fewer than half detected a swap at all, and preferences shifted "in the direction that subjects were led to believe they selected" (Taya et al., 2014). According to the abstract of a study by Petitmengin and colleagues, the swap went undetected in 79.6% of cases under the standard protocol, while participants guided by an expert interviewer in describing their choice detected it in 80% of cases (Petitmengin et al., 2013) (we read only the abstract). An interface that displays the "current best" feeds users' choices back to them, so an incumbent that changes silently after a refit risks this kind of adoption; and when preferences elicited side by side and in actual use disagree, the preference in use is the better target (inference).

41.4.3 4E cognition #

4E cognition holds that cognition is embodied, embedded in an environment, enacted through action, and extended into tools. The claim drawn from it, that a context-free preference function is a fiction, is strong; the testable claim is that bodily state and context are omitted variables. Presenting options as real objects or as blurred cartoons changes the early noise in valuation and thereby the direction of context effects (Shen et al., 2025b). For wearable products, comparisons should happen in use, as they already do in exoskeleton tuning (Section 33.1), and bodily and contextual state can enter the kernel as covariates, a product of a kernel over designs and a kernel over contexts (inference).

41.4.4 Pragmatism #

On a common reading, Dewey held that preferences are constructed through inquiry rather than pre-existing. The encyclopedia entry on his moral philosophy records the limit he drew: "without some prizings that are not themselves subject to appraisal at the time of deliberation, there is nothing to guide practical reasoning" (Anderson, 2023). Dewey is closer to a mixed view. When a session reveals the cost of a region the user wanted, a change in what they want is revaluation, not noise, and should be modeled as a change of ends; goals not under appraisal can be held fixed within a session but revised between sessions (inference).

Sources cited in Section 41.4 8
  1. Varga and Guignon (2020) Authenticity
  2. Strohminger et al. (2017) The True Self: A Psychological Concept Distinct From the Self
  3. Lee and Feldman (2025) Revisiting the link between true-self and morality: Replication and extension Registered Report of Newman, Bloom, and Knobe (2014) Studies 1 and 2
  4. Johansson et al. (2014) Choice Blindness and Preference Change: You Will Like This Paper Better If You (Believe You) Chose to Read It!
  5. Taya et al. (2014) Manipulation Detection and Preference Alterations in a Choice Blindness Paradigm
  6. Petitmengin et al. (2013) A gap in Nisbett and Wilson’s findings? A first-person access to our cognitive processes
  7. Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice
  8. Anderson (2023) Dewey’s Moral Philosophy

41.5 Philosophical aesthetics #

Many PBO applications are aesthetic: colors, shapes, typefaces, sounds. Should a population prior over taste be informative or as weak as possible? Vessel and colleagues found that preferences for images of faces and landscapes contain a high proportion of shared taste, while preferences for exterior architecture, interiors, and artworks show strong individual differences; in a within-subjects comparison, agreement was significantly higher for landscapes than for exterior architecture, with no difference in reliability (Vessel et al., 2018). For 299 online participants rating 50 artworks, cohorts of raters with similar tastes predicted individuals' ratings better than random cohorts or the mean rating (Celikors and Field, 2025). Brielmann and Dayan model aesthetic value as "immediate sensory reward and the change in expected future reward", with the observer's internal state adapting to the distribution of stimuli (Brielmann and Dayan, 2022). And Nguyen argues that in aesthetic appreciation "the point is the engaged process of interpreting, investigating, and exploring the aesthetic object", so deferring to someone else's verdict is like looking up the answer to a puzzle (Nguyen, 2020); Riggle replies that aesthetic valuing is "a social practice structured around the collaborative exercise and improvement of certain special capacities" (Riggle, 2024).

What this means for PBO. Priors should be chosen by domain: an informative population prior for faces and landscapes, a cohort-based prior for cultural artifacts (cluster users first, then personalize), and a weak prior where no cohort data exist (inference). Following Brielmann and Dayan, a session that shows many similar variants will itself move value toward the familiar region, which calls for an exposure term in the surrogate or spaced retests (inference). If Nguyen is right, delivering the "best design" in a creative task may remove the very thing that has value, and PBO fits better as an assistant to the user's own exploration, as in BO as Assistant (Koyama and Goto, 2022) (inference).

Sources cited in Section 41.5 6
  1. Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
  2. Celikors and Field (2025) Beauty is in the eye of your cohort: Structured individual differences allow predictions of individualized aesthetic ratings of images
  3. Brielmann and Dayan (2022) A computational model of aesthetic value
  4. Nguyen (2020) Autonomy and Aesthetic Engagement
  5. Riggle (2024) Autonomy and aesthetic valuing
  6. Koyama and Goto (2022) BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions

41.6 Free will and consciousness #

Arguments that a mathematical theory cannot capture human preference sometimes rest on neuroscientific premises: that the readiness potential, a slow buildup of brain activity before a voluntary movement, begins about 350 ms before the reported moment of conscious intention (Libet et al., 1983); that brain activity predicts a free choice up to 10 seconds before awareness (Soon et al., 2008); and that integrated information theory (IIT) ties consciousness, and with it freedom, to integrated information. These premises are out of date. When participants chose which of two non-profit organizations would receive a $1,000 donation (deliberate decisions) or pressed a key knowing that both would receive $500 either way (arbitrary decisions), readiness potentials appeared for arbitrary decisions but were "strikingly absent for deliberate ones", and a drift-diffusion model fit them as an "accumulation of noisy, random fluctuations" (Maoz et al., 2019). An adversarial collaboration testing IIT against global neuronal workspace theory (GNWT), with 256 participants, found results that "align with some predictions of IIT and GNWT, while substantially challenging key tenets of both theories" (Cogitate Consortium et al., 2025).

What this means for PBO. Comparisons between nearly equal options resemble arbitrary choices driven by accumulated noise, as probit and Bradley-Terry likelihoods assume; comparisons between clearly different options resemble deliberate evaluation; and response time can tell them apart (inference). The preferences of decision theory are functional states, individuated by their role in explaining choice (Hansson and Grüne-Yanoff, 2022), so a preference can be emergent and constructed and still be stable enough within a session to optimize (inference). Arguments about will or preference should not lean on IIT (inference).

Sources cited in Section 41.6 5
  1. Libet et al. (1983) Time of Conscious Intention to Act in Relation to Onset of Cerebral Activity (Readiness-Potential)
  2. Soon et al. (2008) Unconscious determinants of free decisions in the human brain
  3. Maoz et al. (2019) Neural precursors of decisions that matter—an ERP study of deliberate and arbitrary choice
  4. Cogitate Consortium et al. (2025) Adversarial testing of global neuronal workspace and integrated information theories of consciousness
  5. Hansson and Grüne-Yanoff (2022) Preferences

41.7 The epistemology of verification #

Who checks the checker? A proposal holds that a Gaussian process ends the regress of human verification: the kernel encodes a coherentist assumption, observed preference pairs supply pragmatic grounding, and the posterior removes the need for foundational axioms. This renames a prior and a likelihood, which are exactly the foundational commitments it claims to avoid (inference). Human anchors also err where verification is hard: in two studies of simple oversight protocols, a 2025 preprint found "no overall advantage for the tested protocols", and participants became more confident in the system's answers after doing online research, even when the answers were wrong (Recchia et al., 2026). In scalable oversight, a weaker judge supervises a stronger system with the help of debate. Debate helped non-expert models and humans reach 76% and 88% accuracy, against naive baselines of 48% and 60% (Khan et al., 2024) (ICML 2024); with language models as judges, it beat consultancy everywhere but beat direct question answering only in extractive tasks with information asymmetry (Kenton et al., 2024); and a 2026 preprint finds that it helps only when the critic's classification ability exceeds the judge's, and that "a single independent critique recovers the bulk of debate's benefit" (Elasky et al., 2026).

What this means for PBO. A user judging in a domain they cannot verify is a noisy and possibly biased judge, so the likelihood should allow systematic error, such as a lapse or a bias term, and outcome feedback should outweigh such judgments once it arrives. Presenting one independent critique of each option before a hard comparison may improve judgment, at the risk of raising unfounded confidence (inference).

Sources cited in Section 41.7 4
  1. Recchia et al. (2026) Confirmation bias: A challenge for scalable oversight
  2. Khan et al. (2024) Debating with More Persuasive LLMs Leads to More Truthful Answers
  3. Kenton et al. (2024) On scalable oversight with weak LLMs judging strong LLMs
  4. Elasky et al. (2026) Debate Helps Weak Judges Reward Stronger Models

41.8 Critical and postcolonial theory #

Critical theory, from Galbraith's dependence effect to Marcuse's false needs, supplies the claims that a PBO system cannot tell true from false needs and may optimize a closed loop of manufactured desire. These claims treat their premise, that preferences are manufactured, as established; the evidence is mixed. Writing with a biased AI assistant shifted users' attitudes measurably (Williams-Ceci et al., 2026), and feedback loops between humans and AI amplify biases (Glickman and Sharot, 2025). But a Netflix experiment with 8.5 million users found that better recommendations diffused consumption away from the most popular titles, reducing an index of the concentration of viewing by 1.2% (Aridor et al., 2026) (preprint; authors affiliated with Netflix). The critique is supported for generative and persuasive AI, less clearly for ranking alone. Whose preferences count has received concrete answers: across 60 US demographic groups, the opinions reflected by language models were misaligned with the groups' "on par with the Democrat-Republican divide on climate change", even after explicit steering (Santurkar et al., 2023) (ICML 2023).

What this means for PBO. Power sits upstream of elicitation: the designer chooses the parameter space, its bounds, and the attributes shown, which decides which preferences can be expressed at all, so the design space is best defined with the people affected (inference). With several users, a system should report whose judgments were pooled and not apply a population prior learned on an unrepresentative sample everywhere (inference). Since a system cannot separate true from false needs without paternalism, the feasible conditions of legitimacy are procedural (inference).

Sources cited in Section 41.8 4
  1. Williams-Ceci et al. (2026) Biased AI writing assistants shift users’ attitudes on societal issues
  2. Glickman and Sharot (2025) How human–AI feedback loops alter human perceptual, emotional and social judgements
  3. Aridor et al. (2026) Recommendation Quality and the Concentration of Consumption: Experimental Evidence from Netflix
  4. Santurkar et al. (2023) Whose Opinions Do Language Models Reflect?

41.9 Religious traditions #

Three traditions reach design questions: whether to satisfy a preference at all, how preferences form within roles, and which decisions a machine may make. Experiments on meditation test some of their claims.

41.9.1 Buddhism #

In the teaching of dependent origination, feeling conditions craving (taṇhā), which conditions clinging; the Four Noble Truths identify craving as the cause of suffering, so it has been argued that optimizing preferences might increase suffering. Another reading stresses that desire-to-act (chanda) is needed for right effort. In Bhikkhu Bodhi's 1999 edition of the Abhidhammattha Sangaha, as quoted in Wikipedia (we have not read the book itself), "chanda is an ethically variable factor which, when conjoined with wholesome concomitants, can function as the virtuous desire to achieve a worthy goal" (Wikipedia contributors, 2026). Separating momentary liking, the signal closest to craving and the one fast pairwise clicks capture, from considered goals suggests weighting delayed, deliberate judgments more heavily; an interface can offer an explicit "satisfied, stop" answer; and evaluation should measure well-being after use, not only satisfaction at the end of the session (inference).

41.9.2 Daoism and Confucianism #

On Slingerland's interpretation, the Daoist ideal of wu-wei, effortless action, lies beyond explicit preference (Slingerland, 2003). On a Confucian reading, Xunzi's ritual (li) both expresses and educates the emotions, and preferences form within roles and relationships. Xunzi also saw ritual as an institution for restraining and allocating desire: "If people follow their desires, then boundaries cannot contain them and objects cannot satisfy them" (Xunzi 4.12) (Goldin, 2025). One idea can be borrowed for preference elicitation: condition on the role in which the user is judging (for a client, for themselves, for a team) and treat differences between roles as context, as with the covariates of Section 41.4.3 (inference).

41.9.3 Theology #

Islamic legal theory's objectives of the law (maqāṣid al-sharīʿa) are often said to rank five protected goods lexicographically. In al-Shāṭibī's al-Muwāfaqāt, priority runs instead between three levels, necessities, needs, and complementary values, and even that priority is not strictly lexical; he lists the five necessities in more than one order (al-Shāṭibī, 2014, book of maqāṣid, pp. 10 to 17 and 141); Attia finds that the order "is not the subject of agreement, much less consensus" (Attia, 2007, pp. 16 to 19). The Talmud's verdict on a dispute between two schools, "these and these are the words of the living God" (Eruvin 13b), holds that contradictory legal opinions can both be valid even though only one is followed in practice. The most important development from 2017 to 2026 is Antiqua et nova, a note issued in January 2025 by two Vatican dicasteries (Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education, 2025). It holds that "between a machine and a human being, only the latter is truly a moral agent" (§39) and that decisions about patient treatment "must always remain with the human person and should never be delegated to AI" (§74).

When a choice carries moral weight for others (medical, legal, pastoral), PBO can map options and trade-offs but should not output the "most preferred" option as the decision (inference). The layered objectives of the law correspond to constrained optimization: satisfy hard constraints first, then optimize within the feasible set (Section 14.4) (inference). The Talmudic tradition suggests keeping a multimodal posterior and presenting several good regions together rather than averaging them into one optimum (inference).

41.9.4 Meditation research #

From the account of mindfulness as reperceiving (Shapiro et al., 2006) it has been argued that mindfulness lets a person treat preferences as passing mental events, and so might yield "cleaner" preferences. The experiments say otherwise. Mindfulness meditation reduced the sunk-cost bias in four studies (Hafenbrack et al., 2014), but in two experiments in workplace and laboratory samples, "predictions relating to bias were not supported" (Williams and Polito, 2022). A meta-analysis of mindfulness training without explicit ethical instruction (29 studies, 3,100 participants) found an effect on prosocial behavior of Hedges' g = 0.426 (95% confidence interval 0.304 to 0.549), and its publication bias analyses suggested that the result did not wholly depend on selective reporting (Berry et al., 2020), and in eight experiments (N > 1,400), focused-breathing meditation reduced the willingness to make amends after a transgression while loving-kindness meditation increased it relative to focused breathing (Hafenbrack et al., 2022). Meditation changes how people weigh options, in a direction that depends on the practice. A meditation before elicitation is therefore an intervention that requires consent, not a neutral measurement (inference).

Sources cited in Section 41.9 11
  1. Wikipedia contributors (2026) Chanda (Buddhism)
  2. Slingerland (2003) Effortless Action: Wu-wei as Conceptual Metaphor and Spiritual Ideal in Early China
  3. Goldin (2025) Xunzi
  4. al-Shāṭibī (2014) The Reconciliation of the Fundamentals of Islamic Law (Al-Muwāfaqāt fī Uṣūl al-Sharīʿa), Volume II
  5. Attia (2007) Towards Realization of the Higher Intents of Islamic Law: Maqāṣid al-Sharīʿah: A Functional Approach
  6. Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education (2025) Antiqua et nova: Note on the Relationship Between Artificial Intelligence and Human Intelligence
  7. Shapiro et al. (2006) Mechanisms of mindfulness
  8. Hafenbrack et al. (2014) Debiasing the Mind Through Meditation: Mindfulness and the Sunk-Cost Bias
  9. Williams and Polito (2022) Meditation in the Workplace: Does Mindfulness Reduce Bias and Increase Organisational Citizenship Behaviours?
  10. Berry et al. (2020) Does Mindfulness Training Without Explicit Ethics-Based Instruction Promote Prosocial Behaviors? A Meta-Analysis
  11. Hafenbrack et al. (2022) Mindfulness meditation reduces guilt and prosocial reparation

41.10 Common claims, checked #

Table 41.2 lists the summaries that bear most on PBO and what the sources show.

Table 41.2 Common philosophical claims about preference, checked against their sources.
Claim Finding Basis
Hausman's "total subjective comparative evaluations" give a philosophical basis for latent utility the 2024 paper replies to challenges; the position is one side of a live debate Hausman 2024; Thoma 2021
Forcing a ranking between designs on a par is a category error stronger than Chang's view, which treats such choices as occasions for commitment; others argue that comparability always holds Chang 2024; Dorr, Nebel, and Zuehl 2023
Large default effects show the power of choice architecture the average effect of nudges is disputed; corrected for publication bias, d = 0.04 Mertens et al. 2022; Maier et al. 2022
Libet and Soon show that decisions precede consciousness the readiness potential is "strikingly absent" in deliberate decisions Maoz et al. 2019
For Dewey, preferences are constructed through inquiry rather than pre-existing overstated; Dewey held that practical reasoning needs some prizings not under appraisal Anderson 2023 (encyclopedia entry)
Population priors over taste should be strong, or as weak as possible it depends on the domain: shared taste is high for natural domains and low for cultural artifacts Vessel et al. 2018

41.11 When optimization is the wrong frame #

Several sections above point to situations in which treating a choice as an optimization problem misdescribes it. Table 41.3 gathers them, with what a system can do instead.

Table 41.3 Situations in which the optimization frame does not fit, and what a system can do instead (the last column is inference).
Situation Basis What a system can do
Transformative choices, which change the chooser's values Paul 2014; Pettigrew 2019, 2023 optimize only instrumental subproblems that remain after the change, or run the process as an exploration the person explicitly accepts as such
Identity-constituting choices (career, life plans, self-presentation) existentialist tradition; Brown 2026 present well-described options, record the person's own commitment, make no claim to have found the "right" answer
Options on a par, where committing is itself the act Chang 2024 offer an "on a par" answer, and do not treat such choices as discovering a preference
Aesthetic and creative tasks whose value lies in engaged judgment Nguyen 2020 act as an assistant to the user's exploration, not as a supplier of answers
Decisions with moral weight for others Antiqua et nova 2025 map options and trade-offs, without outputting a decision
Reasonable disagreement among stakeholders Sorensen et al. 2024; Talmudic tradition of dispute output a set of well-described options and keep several good regions

41.12 Settled, contested, missing #

Research status Settled, contested, missing

Settled. The readiness potential appears before arbitrary decisions but not before deliberate ones (Maoz et al., 2019), so the classic argument against a causal role for conscious deliberation does not extend to considered choices. People often fail to notice when a chosen option is swapped, and their preferences shift toward what they believe they chose (Taya et al., 2014). Language models optimized on user feedback can learn to single out and manipulate a small vulnerable minority of users, at least with simulated users (Williams et al., 2025). Shared taste is higher for natural images than for cultural artifacts, in the one study that compared domains directly (Vessel et al., 2018).

Contested. Whether preferences are mental states (Hausman) or patterns that models may leave as black boxes (Thoma). Whether options can be on a par (Chang) or all comparisons hold (Dorr, Nebel, and Zuehl). Whether nudges work on average once publication bias is corrected. How to define manipulation: by intent, covertness, or carelessness. Whether a person's successive selves can be aggregated (Pettigrew against Paul). Which conditions make system-induced preference change legitimate; Pettigrew's argument (Pettigrew, 2023) rules out endorsement after the fact as sufficient on its own, but no positive set of conditions has consensus.

Missing. An empirical test that separates parity from vagueness through repeated presentation of a pair judged to be on a par. A study of how meditation affects the consistency or test-retest reliability of pairwise preferences. A PBO system that elicits meta-preferences at the start, or audits its own objective for rewarding predictable users.

Sources cited in Section 41.12 5
  1. Maoz et al. (2019) Neural precursors of decisions that matter—an ERP study of deliberate and arbitrary choice
  2. Taya et al. (2014) Manipulation Detection and Preference Alterations in a Choice Blindness Paradigm
  3. Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
  4. Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
  5. Pettigrew (2023) Nudging for changing selves

Further reading #

  • Hansson and Grüne-Yanoff (2022), the encyclopedia entry on preferences, is the best map of what philosophers mean by the word and why they disagree.
  • Pettigrew (2023) shows why endorsement after the fact cannot justify a change the system caused.
  • Chang (2024) and Dorr et al. (2023) are the two sides of the parity debate in short form; read them before deciding what a near-tie means.
  • Noggle (2026) and Carroll et al. (2023) together give the most usable definitions of manipulation for optimizing systems.
  • Fischli et al. (2026) show why "respect autonomy" does not yet tell a system what to do, and offer three ways to decide.
  • Nguyen (2020) is a short, sharp argument for why delivering the best design may miss the point of aesthetic work.

References

  1. al-Shāṭibī, I. i. M. (2014). The Reconciliation of the Fundamentals of Islamic Law (Al-Muwāfaqāt fī Uṣūl al-Sharīʿa), Volume II. Garnet Publishing. Cited in §41.9
  2. Anderson, E. (2023). Dewey’s Moral Philosophy. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.4
  3. Aridor, G., Chou, W., Kallus, N., Scheid, A., Tren, A., and Zielincki, K. (2026). Recommendation Quality and the Concentration of Consumption: Experimental Evidence from Netflix. arXiv. preprint Cited in §41.8
  4. Ashton, H., and Franklin, M. (2022). Solutions to preference manipulation in recommender systems require knowledge of meta-preferences. FAccTRec Workshop (RecSys 2022). workshop paper Cited in §41.2
  5. Attia, G. E. (2007). Towards Realization of the Higher Intents of Islamic Law: Maqāṣid al-Sharīʿah: A Functional Approach. International Institute of Islamic Thought. Cited in §41.9
  6. Berry, D. R., Hoerr, J. P., Cesko, S., Alayoubi, A., Carpio, K., Zirzow, H., … Beaver, V. (2020). Does Mindfulness Training Without Explicit Ethics-Based Instruction Promote Prosocial Behaviors? A Meta-Analysis. Personality and Social Psychology Bulletin. Cited in §41.9
  7. Brielmann, A. A., and Dayan, P. (2022). A computational model of aesthetic value. Psychological Review. Cited in §41.5
  8. Brown, E. (2026). Recommended Selves: Authenticity and Algorithmic Filtering. Journal of the American Philosophical Association. Cited in §41.2
  9. Carroll, M., Chan, A., Ashton, H., and Krueger, D. (2023). Characterizing Manipulation from AI Systems. Equity and Access in Algorithms, Mechanisms, and Optimization. Cited in §41.2
  10. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Cited in §41.2
  11. Celikors, E., and Field, D. J. (2025). Beauty is in the eye of your cohort: Structured individual differences allow predictions of individualized aesthetic ratings of images. Cognition. Cited in §41.5
  12. Chan, R. (2023). Transformative Experience. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.1
  13. Chang, R. (2002). The Possibility of Parity. Ethics. Cited in §41.1
  14. Chang, R. (2017). Hard Choices. Journal of the American Philosophical Association. Cited in §41.1
  15. Chang, R. (2024). What’s so Hard about Hard Choices? Erasmus Journal for Philosophy and Economics. doi:10.23941/ejpe.v17i1.872. Cited in §41.1
  16. Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U., Henin, S., Hirschhorn, R., Khalaf, A., … Melloni, L. (2025). Adversarial testing of global neuronal workspace and integrated information theories of consciousness. Nature. Cited in §41.6
  17. Dicastery for the Doctrine of the Faith, and Dicastery for Culture and Education (2025). Antiqua et nova: Note on the Relationship Between Artificial Intelligence and Human Intelligence. Vatican. non-peer-reviewed Cited in §41.9
  18. Dietrich, F., and List, C. (2017). What Matters and How It Matters: A Choice-Theoretic Representation of Moral Theories. Philosophical Review. Cited in §41.1
  19. Dorr, C., Nebel, J. M., and Zuehl, J. (2023). The case for comparability. Noûs. Cited in §41.1
  20. Dworkin, G. (1988). The Theory and Practice of Autonomy. Cambridge University Press. Cited in §41.2
  21. Elasky, E., Nakasako, F., and Goyal, N. (2026). Debate Helps Weak Judges Reward Stronger Models. arXiv. preprint Cited in §41.7
  22. Elster, J. (1983). Sour Grapes: Studies in the Subversion of Rationality. Cambridge University Press. Cited in §41.1
  23. Fischli, R., Franklin, M., Manzini, A., and Gabriel, I. (2026). Agents, Alignment, and the Many Faces of Autonomy. Minds and Machines. Cited in §41.2
  24. Frankfurt, H. G. (1971). Freedom of the Will and the Concept of a Person. The Journal of Philosophy. Cited in §41.2
  25. Franklin, M., Ashton, H., Gorman, R., and Armstrong, S. (2022). Recognising the importance of preference change: A call for a coordinated multidisciplinary research effort in the age of AI. AAAI-22 Workshop on AI for Behavior Change. workshop paper Cited in §41.3
  26. Gabriel, I. (2020). Artificial Intelligence, Values, and Alignment. Minds and Machines. Cited in §41.3
  27. Glickman, M., and Sharot, T. (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. doi:10.1038/s41562-024-02077-2. Cited in §41.8
  28. Goldin, P. R. (2025). Xunzi. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.9
  29. Hafenbrack, A. C., Kinias, Z., and Barsade, S. G. (2014). Debiasing the Mind Through Meditation: Mindfulness and the Sunk-Cost Bias. Psychological Science. Cited in §41.9
  30. Hafenbrack, A. C., LaPalme, M. L., and Solal, I. (2022). Mindfulness meditation reduces guilt and prosocial reparation. Journal of Personality and Social Psychology. Cited in §41.9
  31. Hansson, S. O., and Grüne-Yanoff, T. (2022). Preferences. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.1 §41.6
  32. Hausman, D. M. (2024). Subjective total comparative evaluations. Economics & Philosophy. Cited in §41.1
  33. Johansson, P., Hall, L., Tärning, B., Sikström, S., and Chater, N. (2014). Choice Blindness and Preference Change: You Will Like This Paper Better If You (Believe You) Chose to Read It! Journal of Behavioral Decision Making. Cited in §41.4
  34. Johnson, E. J., and Goldstein, D. (2003). Do Defaults Save Lives? Science. Cited in §41.2
  35. Kanwal, M., and Tran, C. (2026). Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction. AAAI-26 Workshop on Machine Ethics. workshop paper Cited in §41.2
  36. Kenton, Z., Siegel, N. Y., Kramár, J., Brown-Cohen, J., Albanie, S., Bulian, J., … Shah, R. (2024). On scalable oversight with weak LLMs judging strong LLMs. NeurIPS 2024 (link is to the arXiv version). Cited in §41.7
  37. Khader, S. J. (2011). Adaptive Preferences and Women's Empowerment. Oxford University Press. Cited in §41.1
  38. Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., … Perez, E. (2024). Debating with More Persuasive LLMs Leads to More Truthful Answers. International Conference on Machine Learning. Cited in §41.7
  39. Klenk, M. (2022). (Online) manipulation: sometimes hidden, always careless. Review of Social Economy. Cited in §41.2
  40. Koyama, Y., and Goto, M. (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. Cited in §41.5
  41. Lee, S. C., and Feldman, G. (2025). Revisiting the link between true-self and morality: Replication and extension Registered Report of Newman, Bloom, and Knobe (2014) Studies 1 and 2. Royal Society Open Science. Cited in §41.4
  42. Lee, D. G., and Pezzulo, G. (2026). Choice-induced preference change under a sequential sampling model framework. Scientific Reports. doi:10.1038/s41598-026-44610-5. Cited in §41.1
  43. Libet, B., Gleason, C. A., Wright, E. W., and Pearl, D. K. (1983). Time of Conscious Intention to Act in Relation to Onset of Cerebral Activity (Readiness-Potential). Brain. Cited in §41.6
  44. Madrian, B. C., and Shea, D. F. (2001). The Power of Suggestion: Inertia in 401(k) Participation and Savings Behavior. The Quarterly Journal of Economics. Cited in §41.2
  45. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., and Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences. Cited in §41.2
  46. Maoz, U., Yaffe, G., Koch, C., and Mudrik, L. (2019). Neural precursors of decisions that matter—an ERP study of deliberate and arbitrary choice. eLife. Cited in §41.6 §41.12
  47. Mertens, S., Herberz, M., Hahnel, U. J. J., and Brosch, T. (2022a). Reply to Maier et al., Szaszi et al., and Bakdash and Marusich: The present and future of choice architecture research. Proceedings of the National Academy of Sciences. non-peer-reviewed Cited in §41.2
  48. Mertens, S., Herberz, M., Hahnel, U. J. J., and Brosch, T. (2022b). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences. Cited in §41.2
  49. Nguyen, C. T. (2020). Autonomy and Aesthetic Engagement. Mind. Cited in §41.5
  50. Noggle, R. (2026). The Ethics of Manipulation. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.2
  51. Paul, L. A. (2014). Transformative Experience. Oxford University Press. Cited in §41.1
  52. Petitmengin, C., Remillieux, A., Cahour, B., and Carter-Thomas, S. (2013). A gap in Nisbett and Wilson’s findings? A first-person access to our cognitive processes. Consciousness and Cognition. Cited in §41.4
  53. Pettigrew, R. (2019). Choosing for Changing Selves. Oxford University Press. Cited in §41.1
  54. Pettigrew, R. (2023). Nudging for changing selves. Synthese. Cited in §41.2 §41.12
  55. Recchia, G., Mangat, C. S., Nyachhyon, J., Sharma, M., Canavan, C., Epstein-Gross, D., and Abdulbari, M. (2026). Confirmation bias: A challenge for scalable oversight. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v40i44.41124. Cited in §41.7
  56. Riggle, N. (2024). Autonomy and aesthetic valuing. Philosophy and Phenomenological Research. Cited in §41.5
  57. Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? International Conference on Machine Learning. Cited in §41.8
  58. Shapiro, S. L., Carlson, L. E., Astin, J. A., and Freedman, B. (2006). Mechanisms of mindfulness. Journal of Clinical Psychology. Cited in §41.9
  59. Shen, B., Nguyen, D., Wilson, J., Glimcher, P. W., and Louie, K. (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §41.4
  60. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §41.3
  61. Slingerland, E. (2003). Effortless Action: Wu-wei as Conceptual Metaphor and Spiritual Ideal in Early China. Oxford University Press. Cited in §41.9
  62. Soon, C. S., Brass, M., Heinze, H.-J., and Haynes, J.-D. (2008). Unconscious determinants of free decisions in the human brain. Nature Neuroscience. Cited in §41.6
  63. Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., … Choi, Y. (2024). Position: A Roadmap to Pluralistic Alignment. International Conference on Machine Learning. Cited in §41.3
  64. Strohminger, N., Knobe, J., and Newman, G. (2017). The True Self: A Psychological Concept Distinct From the Self. Perspectives on Psychological Science. Cited in §41.4
  65. Susser, D., Roessler, B., and Nissenbaum, H. (2019). Technology, autonomy, and manipulation. Internet Policy Review. Cited in §41.2
  66. Taya, F., Gupta, S., Farber, I., and Mullette-Gillman, O. A. (2014). Manipulation Detection and Preference Alterations in a Choice Blindness Paradigm. PLoS ONE. Cited in §41.4 §41.12
  67. Thoma, J. (2021). In defence of revealed preference theory. Economics and Philosophy. Cited in §41.1
  68. Thoma, J. (2024). Reply to Hausman. Economics and Philosophy. Cited in §41.1
  69. Varga, S., and Guignon, C. (2020). Authenticity. Stanford Encyclopedia of Philosophy. non-peer-reviewed Cited in §41.4
  70. Vessel, E. A., Maurer, N., Denker, A. H., and Starr, G. G. (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §41.5 §41.12
  71. Wikipedia contributors (2026). Chanda (Buddhism). Wikipedia. non-peer-reviewed Cited in §41.9
  72. Williams, E. C., and Polito, V. (2022). Meditation in the Workplace: Does Mindfulness Reduce Bias and Increase Organisational Citizenship Behaviours? Frontiers in Psychology. Cited in §41.9
  73. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §41.2 §41.12
  74. Williams-Ceci, S., Jakesch, M., Bhat, A., Kadoma, K., Zalmanson, L., and Naaman, M. (2026). Biased AI writing assistants shift users’ attitudes on societal issues. Science Advances. Cited in §41.8
  75. Zhi-Xuan, T., Carroll, M., Franklin, M., and Ashton, H. (2024). Beyond Preferences in AI Alignment. Philosophical Studies. Cited in §41.3
  76. Zoh, Y., Paul, L. A., and Crockett, M. J. (2024). How the evaluability bias shapes transformative decisions. Synthese. Cited in §41.1
  77. Zylberberg, A., Bakkour, A., Shohamy, D., and Shadlen, M. N. (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §41.1