Interactive Design and Human-Computer Interaction
Part IV built preferential Bayesian optimization (PBO) around a simulated person: a fixed utility, constant noise, and unlimited patience, and Section 19.7 warned that real people are none of these things. This chapter collects what human-computer interaction (HCI), computer graphics, and interactive design learned between 2017 and September 2026 by putting real people in the loop: designers tuning photographs, artists simplifying 3D meshes, crowd workers moving sliders, drivers rating dashboards.
The chapter reads "preferential" broadly: besides pairwise duels, it includes sliders, galleries, and rankings, as the authors of one main system do when they use the term PBO "in a broader sense" (Koyama et al., 2020), and it includes optimization driven by ratings or by measured performance, because most of what is known about agency and population priors comes from those studies. Each study is labeled with its feedback type, preference (choices, rankings, comparisons), rating (a score on a scale), or performance (a measurement of what the person did, such as task time), because a result about one type does not automatically transfer to another. is the number of participants.
Three findings carry the chapter, and each bears on the book's argument that the hard part of PBO is now measurement, including what asking does to the person (Section 45.1). When the optimizer leads, outcomes improve but the person's sense of agency drops, and most of it returns when the person can steer. Feedback in real sessions is unstable in ways the standard noise model does not describe. And against strong comparisons, the usual result is the same outcome at lower cost, not a better outcome, from samples that in pure preference studies are mostly 6 to 60 people.
Sources cited in the introduction 1
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
32.1 Research lines #
Two papers mark the start: Brochu et al. (2007) used gallery-style active preference learning to design materials, and González et al. (2017) gave the formal definition based on duels in 2017. Since then, HCI systems have mostly changed the form of the query and rarely the acquisition function, the rule that picks the next question (Section 12.1), which is usually expected improvement (Section 12.3) or the upper confidence bound (Section 12.4) rewritten to fit the query (inference, from the system descriptions below). After 2024 the emphasis moved toward priors borrowed from other users and toward language and generative models. The subsections follow the research groups; Table 32.1 collects the user studies.
32.1.1 Koyama and collaborators: the query interface as the variable #
Yuki Koyama and colleagues (University of Tokyo, AIST, and OMRON SINIC X) treat the query interface as the main thing to study: from 2017 to 2025 each paper added one channel for feedback, from a single slider to a gallery, several sliders, implicit slider logs, references, constraints, and natural language (inference, from the designs below).
Sequential line search turns each query into a single-slider microtask along a line through the design space (Koyama et al., 2017) (Section 20.2; Chapter 25 runs it on a photograph). Its crowdsourced pipeline cost about $5.25 and 68 minutes per result (15 iterations of 7 microtasks at $0.05 each), and against the automatic enhancement of Photoshop and Lightroom the optimized versions of three photos received 32, 26, and 29 crowd votes, and each alternative 0 to 3 (Koyama and Igarashi, 2018). The same book chapter argues that absolute ratings generally work poorly because a rater must know the design space, whereas a pairwise comparison "is easy to answer even for non-experts".
Sequential Gallery replaces the slider with a zoomable grid (5 by 5 for photos) from which the person picks the best item (Koyama et al., 2020) (Section 20.3). Its plane search beat line search in simulation, and in a preliminary study with , participants were satisfied after 5.36 iterations on average and rated the grid as a source of inspiration at 6.50 on a 7-point scale. The authors list as limitations that the method suits only visual designs, handles neither discrete parameters nor prior knowledge, performs poorly above about 20 dimensions, and assumes that the user's perceptual function "does not change over time".
Systems from the same period search the latent space of a generative model, the compact input vector that a trained network decodes into an image, a melody, or a font. Chiu et al. (2020) built slider subspaces in it without Bayesian optimization; Chong et al. (2021) let people blend image candidates with several sliders and guide the search by painting; and Zhou, Koyama, Goto, and Igarashi had people pick the best of several generated melodies, first in a pilot study (Zhou et al., 2020), then with a slider that sets the exploration weight of the upper confidence bound (Zhou et al., 2021).
BO as Assistant removes the explicit query loop: it extracts preference pairs from ordinary slider edits and offers suggestions asynchronously, so the designer can "take the full initiative"; it had no formal user study, and its authors state that usability, usefulness, and agency were not evaluated (Koyama and Goto, 2022). Photographer-in-the-loop lighting is PBO with a gallery of four photographs, which in the authors' demonstrations found pleasing lighting in about 10 iterations (Yamamoto et al., 2022). FontCraft explores font styles from text, image, and font references, with a history whose choices can be withdrawn (Tatsukawa et al., 2025). And constrained PBO asks banner-ad designers only for pairwise preferences while a model of the click-through rate (the fraction of viewers who click an ad) acts as a constraint; its acquisition function, EUBOC, is the expected utility of the best option (EUBO, Section 19.4) with the constraint added (Iwai et al., 2025). The line then turned to language: designers intervening in a system-led optimization in natural language (Niwa et al., 2025), a language model turning vague preferences into a combinatorial problem (Kuroki et al., 2026), and, in a 2026 preprint, a feasibility-aware latent space of car-wheel designs (Owaki et al., 2026).
32.1.2 Cambridge, Aalto, and NYCU: multi-objective optimization with designers #
The second line, around Per Ola Kristensson, John Dudley, Antti Oulasvirta, Liwei Chan, George Mo, and Yi-Chi Liao (Cambridge, Aalto, and National Yang Ming Chiao Tung University), mostly uses multi-objective Bayesian optimization with performance objectives, not preference feedback. Multi-objective optimization (Section 14.5) measures several objectives at once, such as speed and accuracy, and returns the Pareto front, the designs that cannot be improved on one objective without getting worse on another.
Dudley et al. (2019) optimized a map search interface for task time, with 200 crowd workers in its first experiment. Chan et al. (2022) compared optimizer-led with designer-led design of a virtual reality touch technique (4 parameters) in a between-subjects experiment, in which each participant saw only one condition, with 40 novices; its results anchor Section 32.2. Shen et al. (2022) let 12 users pick their final gesture-keyboard design from their own Pareto front; Liao et al. (2023) had 8 designers work with the optimizer and with their own strategy; Mo et al. (2024) proposed cooperative optimization, in which designers mark designs and regions to guide the optimizer; and Liao et al. (2024b) pooled sessions with a "global Gaussian process" (the surrogate of Chapter 8) and started new users from a "warm-start Gaussian process". Later papers color-graded 360-degree images with two-level PBO (Yuan et al., 2025), built the cost of each prototype into the acquisition function (Langerak et al., 2026), and separated expert opinion, empirical studies, and simulators as feedback sources in MUSE (Fernandes Junior et al., 2026), accepted at ACM Transactions on Interactive Intelligent Systems (TiiS). A review by Dudley, Oulasvirta, Chan, and Kristensson appeared online in ACM Computing Surveys in August 2026; we could not obtain its text (Dudley et al., 2026).
32.1.3 ETH Zurich and Saarland: priors learned from other people #
Since 2024 a third line, around Yi-Chi Liao, Zhipeng Li, Christoph Gebhardt, Christian Holz, and Anna Maria Feit (ETH Zurich and Saarland University), has studied population priors: data from earlier users, so that a new user needs to be asked as little as possible. Liao et al. (2024a), a collaboration of Aalto University and Meta, built a population model from 14 participants and used it as the prior of a transfer acquisition function for 11 new ones, and continual human-in-the-loop optimization represents the whole population with a Bayesian neural network, one that keeps a distribution over its weights (Liao et al., 2025); both use performance feedback. Meta-PO meta-learns over the Gaussian process models of earlier users with the Sequential Gallery interface (Li et al., 2025a), and Song et al. (2025) inferred objective weights from users' manual edits of mixed-reality layouts. Three CHI 2026 papers followed: HOMI trained a neural acquisition function on simulated users before any real user arrived (Liao et al., 2026); APPO hides an image generator's text prompt and asks only for binary preferences while a language model rewrites the prompt (Li et al., 2026f); and AutoOptimization turns spoken preferences into the objectives of a layout optimizer (Li et al., 2026a).
32.1.4 Ulm: automotive interfaces rated by drivers and passengers #
The group of Pascal Jansen, Mark Colley, and Enrico Rukzio at Ulm University optimizes automotive interfaces on questionnaire ratings (trust, perceived safety, mental demand, predictability), with some of the largest samples in the field: the visualizations of an automated vehicle in OptiCarVis (, online, between-subjects) (Jansen et al., 2025), external vehicle displays () (Colley et al., 2025), an air-taxi simulation with and without a motion seat () (Meinhardt et al., 2025), a proactive in-car assistant in ProVoice (, within-subjects, meaning every participant experienced every condition) (Susak et al., 2026), blur in virtual-reality driving in BlurDriving (IMWUT, September 2026) (Li et al., 2026b), and a three-day study accepted at AutomotiveUI 2026 () that compared continued optimization with keeping the day-one design (Colley et al., 2026).
32.1.5 LMU Munich: expertise and field deployment #
Ou, Buschek, Mayer, and Butz built, with an industry partner, a PBO system for simplifying 3D meshes (9 parameters), in which artists' ratings were converted into pairwise relations; it ran for 3 months with two full-time technical artists and then in a lab study with (Ou et al., 2022). Ou et al. (2023) compared 60 users of different expertise on three tasks with an interface that asked for a ranking of 4 candidates, offered a "don't know" area, and ended when the user pressed a satisfaction button. These two studies supply much of Section 32.3 and Section 32.6.
32.1.6 Guiding generative models #
Four more systems put preference queries inside a generative model, often a diffusion model, the kind of generator behind most current text-to-image systems. BOgen builds 3D shapes part by part with PBO and a variational autoencoder (a network that decodes a short latent vector into an output), and was compared with a baseline by 30 designers (Lee et al., 2026). ROMBO personalizes music generation, evaluated in simulation and with 16 volunteers who scored each piece from 0 to 10 (Marcos et al., 2025). GimmBO runs PBO over the merging weights of 20 to 30 adapters, small add-on networks that each push a diffusion model toward a style (Liu et al., 2026b). MultiBO (ICML 2026) has the user pick, from 4 generated images, the one closest to the image they have in mind, and steers edits in the model's attention space for up to 50 rounds (Rajagopalan et al., 2026). It reports an image similarity of 0.9364 against 0.8365 for DiffusionDPO and won 70.82% of comparisons rated by 30 users, and its authors observe that the preference signal weakens as the images approach the target.
32.1.7 Other domains: touch, text entry, reading, architecture, engineering #
Catkin and Patoglu (2023) optimized the perceived realism of rendered springs and friction from pairwise comparisons, validated against a model tuned by experts, and in Zhang et al. (2026b) 13 participants made 40 rounds of comparisons with a 5-level confidence rating, for a held-out accuracy of 92.3% (range 85% to 100%). AdaptiFont optimized a generative font space for reading speed (performance, in the main study) (Kadner et al., 2021); Tanaka et al. (2026) had 40 design students design a pavilion with 6 parameters, by optimization driven by ratings from 0 to 100 or by sliders; and Nandy and Goucher-Lambert (2025) adapted a parametric mug to the attribute "comfortable" ( from scratch, from the previous group's data) and improved the perceived match between intent and design over a non-interactive model. McCourt and Dewancker (2019), a workshop paper, had to extend PBO to handle ties when coloring artwork, and Peng et al. (2026), a preprint, personalized generative user interfaces from pairwise judgments. For text entry the only study we found is the gesture keyboard of Shen et al. In visualization design, the applications of Bayesian optimization we found tune particular displays (OptiCarVis, a case in MUSE), and none of them is a PBO study.
32.1.8 The studies at a glance #
Table 32.1 lists the main user studies by year. The questionnaires named in it are explained in Section 32.2. Hypervolume measures how much of the objective space a Pareto front covers; larger is better.
| System or study | Venue | Participants | Feedback | Comparison | Main result |
|---|---|---|---|---|---|
| Sequential line search | SIGGRAPH 2017 | crowd workers, about 105 microtasks per result | preference: slider | Photoshop and Lightroom auto-enhancement | crowd votes 32, 26, 29 against 0 to 3 |
| Dudley et al. | CHI 2019 | 200 crowd workers (experiment 1) | performance: task time | baseline interface | no difference in batch 1 (); from batch 2, median time from about 34 s to about 21 s |
| Sequential Gallery | SIGGRAPH 2020 | 6, plus simulation | preference: best of a 2D gallery | line search, in simulation | plane search better in simulation; users satisfied after 5.36 iterations |
| Zhou et al. | IUI 2021 | 12 novice composers | preference: best of several | automatic expected improvement | Creativity Support Index higher; 11 of 12 preferred manual balancing |
| Chan et al. | CHI 2022 | 40 novices, between-subjects | performance | designer-led design | better spatial error; lower agency, ownership, expressiveness |
| Ou et al. | Mensch und Computer 2022 | 2 artists (3-month field), 20 (lab) | rating converted to pairs | none (observational) | 415 of 549 field sequences stopped at iteration 1; satisfied 16/134 (field), 97/200 (lab) |
| Photographer-in-the-loop | UIST 2022 | 12 | preference: best of 4 | manual lighting workflow | experience rated better on all six items; satisfaction with the lighting not significantly different |
| Shen et al. | ISMAR 2022 | 12 | performance; final pick by preference | baseline keyboard | speed +14.4%, accuracy +13.8% |
| Ou et al. | IUI 2023 | 60, three tasks | preference: rank 4, with "don't know" | levels of expertise | novices reach expert-level quality; experts iterate more, less satisfied |
| Liao et al. | CHI 2024 | 14 for the model, 11 for evaluation | performance | standard Bayesian optimization; manual calibration | absolute pointing improved 22.92% and 21.35% |
| Mo et al. | ACM TiiS 2024 | 18, within-subjects | performance; designers can mark | designer-led; optimizer-led | sense-of-control medians 6.0, 5.5, 1.5; 12 of 18 preferred cooperative |
| FontCraft | CHI 2025 | 10 non-experts | preference: slider, references, retractable history | single-slider Bayesian optimization | closer to target fonts (no inferential statistics seen) |
| Constrained PBO | IJCAI 2025 | 11 professional ad designers | preference: pairs | no control system | 4.2 s per choice; positive attitudes; in a pre-study, designer preference not positively correlated with click-through |
| OptiCarVis | CHI 2025 | 117, online | rating | no visualization; expert design; user-customized | personalized designs rated better on perceived safety, predictability, trust |
| Meta-PO | UIST 2025 | 36 (three groups of 12) | preference: best of a gallery | no transfer | iterations to satisfaction from 9.54 to 5.86 (same theme) and 7.41 (other theme) |
| Niwa et al. | UIST 2025 | 18 (study 1), 12 (study 2) | performance, plus natural language | designer-led; optimizer-led; explicit constraints | optimizer-led highest hypervolume, lowest agency; language lowers workload, gives less agency than constraints |
| Song et al. | UIST 2025 | 12 | weights inferred from manual edits | manual; ParetoSelect | fewer elements moved; layouts ranked above ParetoSelect's (); no difference in hypervolume or overall quality |
| GimmBO | SIGGRAPH North America 2026 | 12 (computing or machine-learning background) | preference: top- ranking; slider; gallery | slider; Sequential Gallery | ranking better on similarity and success rate; 50.5 s against 34.7 s per step |
| APPO | CHI 2026 | 16 | preference: binary, prompt hidden | PromptCharm; DSPy; clarifying questions | satisfied in fewer than 4 iterations (baselines more than 6); less expressive |
| Cost-aware Bayesian optimization | CHI 2026 | 12 | performance plus a comfort rating | Bayesian optimization ignoring cost | same performance at about 67% of the cost; final quality no different () |
| HOMI | CHI 2026 | 12 | performance | transfer acquisition; continual Bayesian optimization | better only at iterations 2 and 3; equal from iteration 6 |
| Tanaka et al. | CAADRIA 2026 | 40 design students | rating, 0 to 100 | sliders | 62.5% preferred the optimized results () |
| Owaki et al. | preprint 2026 | 40 | preference exploration | raw 9-dimensional parameters | higher shape similarity and more feasible suggestions in a 5-dimensional latent space |
Sources cited in Section 32.1 52
- Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
- González et al. (2017) Preferential Bayesian Optimization
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Koyama and Igarashi (2018) Computational Design with Crowds
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Chiu et al. (2020) Human-in-the-loop differential subspace search in high-dimensional latent space
- Chong et al. (2021) Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance
- Zhou et al. (2020) Generative Melody Composition with Human-in-the-Loop Bayesian Optimization
- Zhou et al. (2021) Interactive Exploration-Exploitation Balancing for Generative Melody Composition
- Koyama and Goto (2022) BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions
- Yamamoto et al. (2022) Photographic Lighting Design with Photographer-in-the-Loop Bayesian Optimization
- Tatsukawa et al. (2025) FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Kuroki et al. (2026) LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- Owaki et al. (2026) Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs
- Dudley et al. (2019) Crowdsourcing Interface Feature Design with Bayesian Optimization
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Shen et al. (2022) Personalization of a Mid-Air Gesture Keyboard using Multi-Objective Bayesian Optimization
- Liao et al. (2023) Interaction Design With Multi-Objective Bayesian Optimization
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Liao et al. (2024b) Practical approaches to group-level multi-objective Bayesian optimization in interaction technique design
- Yuan et al. (2025) Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality
- Langerak et al. (2026) Cost-Aware Bayesian Optimization for Prototyping Interactive Devices
- Fernandes Junior et al. (2026) Integrating Multi-Source Feedback in Computational Design
- Dudley et al. (2026) Putting the Human Back in the Loop: A Review of Interactive Bayesian Optimization
- Liao et al. (2024a) A Meta-Bayesian Approach for Rapid Online Parametric Optimization for Wrist-based Interactions
- Liao et al. (2025) Continual Human-in-the-Loop Optimization
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Song et al. (2025) Preference-Guided Multi-Objective UI Adaptation
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
- Li et al. (2026a) Automating UI Optimization through Multi-Agentic Reasoning
- Jansen et al. (2025) OptiCarVis: Improving Automated Vehicle Functionality Visualizations Using Bayesian Optimization to Enhance User Experience
- Colley et al. (2025) Improving External Communication of Automated Vehicles Using Bayesian Optimization
- Meinhardt et al. (2025) Fly Away: Evaluating the Impact of Motion Fidelity on Optimized User Interface Design via Bayesian Optimization in Automated Urban Air Mobility Simulations
- Susak et al. (2026) ProVoice: Designing Proactive Functionality for In-Vehicle Conversational Assistants using Multi-Objective Bayesian Optimization to Enhance Driver Experience
- Li et al. (2026b) BlurDriving: Investigating How Personalized Blur Techniques Impact Drivers' Performance in Virtual Reality
- Colley et al. (2026) Multi-Session User Experience Assessments of Computationally Optimized Automated Vehicle Functionality Visualizations
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Lee et al. (2026) Part-level 3D shape generation driven by user intention inference with preferential Bayesian optimization
- Marcos et al. (2025) Random rotational embedding Bayesian optimization for human-in-the-loop personalized music generation
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- Rajagopalan et al. (2026) Personalized Image Generation via Human-in-the-loop Bayesian Optimization
- Catkin and Patoglu (2023) Preference-Based Human-in-the-Loop Optimization for Perceived Realism of Haptic Rendering
- Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
- Kadner et al. (2021) AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization
- Tanaka et al. (2026) Human-in-the-Loop Bayesian Optimization Approach to Supporting Early-Stage Architectural Design
- Nandy and Goucher-Lambert (2025) Exploring the Effectiveness of Interactive Preference Learning for Adapting Designs to Abstract Semantic Attributes
- McCourt and Dewancker (2019) Sampling Humans for Optimizing Preferences in Coloring Artwork
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
32.2 Agency versus performance #
The most consistent result of the period concerns who leads the search. When the optimizer leads, outcome measures are better and mental demand is lower, but the person's agency (the sense of being in control of the process), ownership (the sense that the result is one's own), and expressiveness (the sense of being able to put one's own ideas into the result) fall clearly. When designers get ways to steer, most of the agency comes back at a small cost in outcome quality. Almost all of this evidence comes from multi-objective optimization with performance objectives.
Two questionnaires recur. The Creativity Support Index scores support for creative work on a 0 to 100 scale, from factors such as exploration, expressiveness, and enjoyment; the NASA Task Load Index (NASA-TLX) measures workload from six subscales, among them mental demand and effort. is the probability of a difference at least this large if the conditions truly did not differ, and is a statistic with 38 degrees of freedom whose sign says which group scored lower.
The only study large enough to isolate the effect. In the between-subjects experiment of Chan et al. (2022), with 40 novice designers and performance objectives, the optimizer-led group reported lower agency (, ) and ownership (, ), while satisfaction and confidence did not differ. The Creativity Support Index was 75.3 for designer-led and 65.4 for optimizer-led design (), a difference that came mainly from expressiveness (44.9 against 23.0, ). The NASA-TLX total did not differ (57.6 against 49.6, ), but the optimizer-led group reported lower mental demand (14.9 against 8.4, ) and effort (24.9 against 13.7, ). Optimizer-led sessions were longer, 78.0 minutes against 51.8, their designs had better spatial error (, ), and the group explored a larger share of the design space. In the words of the abstract, designers guided by an optimizer "reported lower mental effort but also felt less creative and less in charge of the progress", and the authors conclude that such optimization can support novice designers "in cases where agency is not critical".
Three levels of control in one study. Mo et al. (2024) compared designer-led, cooperative, and optimizer-led design within subjects (). For "I felt that I had control over searching different areas of the design space", the median ratings were 6.0, 5.5, and 1.5 on a 7-point scale. Twelve participants preferred the cooperative condition, 4 the designer-led one, and 2 the optimizer-led one, with medians for willingness to use the method again of 6.0, 5.0, and 3.0 in the same order. The cooperative condition's relative hypervolume did not differ significantly from the optimizer-led one, and it needed fewer formal evaluations but found fewer Pareto-optimal designs.
Smaller studies point the same way. Among the 8 designers of Liao et al. (2023), optimizer-found designs showed "very similar performance" to the designers' own at significantly lower workload, while "designers may become detached from the design process when typical aspects of their role are subsumed" by the optimizer. In the melody study of Zhou et al. (2021) (preference feedback, 12 novice composers), letting users set the balance between exploration and exploitation themselves raised the Creativity Support Index total () and satisfaction (), and 11 of 12 preferred balancing by hand. In study 1 of Niwa et al. (2025) (), the optimizer-led condition reached a higher relative hypervolume than designer-led design () and than cooperation in natural language () but had significantly lower agency than both (corrected ); in study 2 (), language and explicit constraints performed the same (), language gave lower task load (), and constraints gave more agency (). Some participants were unsure "how much their instructions were actually being followed".
Weaker evidence. FontCraft's participants (preference feedback, ) reported more agency with multimodal references than with the baseline, and 7 of the 10 mentioned the history view (Tatsukawa et al., 2025). In Colley et al. (2025), "I felt in control of the design process" averaged 5.81 and "the final design is mine" 5.30 on 7-point scales. Tanaka et al. (2026) found "process disclosure" (parameter plots, the range of evaluations, the direction of optimization) essential for agency and trust, even when participants' understanding stayed superficial, and Song et al. (2025) note that an unconstrained optimizer can move elements the user had placed by hand.
You can feel the trade-off in Figure 32.1: run the same budget three times, choosing every design yourself, letting the optimizer choose, and letting it suggest while you decide, and rate after each run how much you controlled where the search went.
Things to try:
- Do the You run first, without revealing the landscape. You will probably find the broad hill first; notice whether you spend the remaining tries polishing it.
- Switch to Optimizer and press Run to the end. Compare its best score with yours, then compare your two control ratings. Which gap is larger?
- Press New landscape and repeat. On the first landscape the optimizer's exploration finds the narrow, taller peak; on most of the next ones it does not reach that peak within its 10 tries, and your own search can do as well.
The map shows the whole design space at a glance. The design spaces in the studies were larger (4 parameters in Chan et al., 6 in the pavilion study, 9 in the mesh simplification of Ou et al.), and a person moving 9 sliders can no longer see which regions they have not tried, which is where an optimizer's systematic exploration should count for more (inference); Figure 19.4 shows from the other side how much less forty comparisons accomplish as parameters are added.
What it means for preferential optimization. Agency falls in the order designer-led, cooperative, user-guided, optimizer-led, and the differences in agency (sense-of-control medians of 6.0 against 1.5) are much larger than the differences in outcome between cooperative and optimizer-led design. For design tasks, this favors adding channels through which the person can steer over making the optimizer more autonomous (inference). Plain PBO, in which the person only answers the optimizer's questions, is the optimizer-led condition by construction (inference). The effect was isolated with performance objectives; among preference-feedback systems only Zhou et al. and FontCraft measured anything close to agency, with 12 and 10 participants (inference). A team building a preference tool should therefore measure agency in its own study, with an item such as Mo et al.'s, rather than assume the effect transfers (inference).
Sources cited in Section 32.2 9
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Liao et al. (2023) Interaction Design With Multi-Objective Bayesian Optimization
- Zhou et al. (2021) Interactive Exploration-Exploitation Balancing for Generative Melody Composition
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Tatsukawa et al. (2025) FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
- Colley et al. (2025) Improving External Communication of Automated Vehicles Using Bayesian Optimization
- Tanaka et al. (2026) Human-in-the-Loop Bayesian Optimization Approach to Supporting Early-Stage Architectural Design
- Song et al. (2025) Preference-Guided Multi-Objective UI Adaptation
32.3 Unstable feedback and non-convergence #
A PBO model assumes that each answer is a noisy reading of one fixed utility, with independent, identically distributed (i.i.d.) noise (Chapter 19). The longest observation of real use shows how far practice can be from that. In the deployment of Ou et al. (2022) with two professional 3D artists, 415 of 549 evaluation sequences stopped at the first iteration, without any optimization being requested. The remaining 134 sequences ran 4.1 iterations on average (standard deviation 4.2, range 1 to 23), and only 16 of them (11.9%) produced a satisfactory result. In the lab study, 97 of 200 sequences (48.5%) ended satisfied. Figure 32.2 draws every sequence.
Things to try:
- With Count against set to All sequences, compare the two percentages: 2.9% in the field, 48.5% in the lab.
- Switch to Sequences that asked for optimization. The 415 gray squares drop out of the count and the field figure rises to 11.9%. Both numbers are correct; Exercise 32.1 asks which one describes the tool in practice.
The authors write that participants, if not rating at random, "at least behave highly unstable and inconsistent in the rating process". Asked about specific inconsistent ratings, the experts pointed to anchoring on grids they had seen before, the availability and representativeness heuristics, loss aversion, and diminishing returns, the judgment biases of Section 37.2. The authors concluded that the assumptions of a stable and complete preference and of i.i.d. noise do not hold: "human judgment is a fragile function to optimize for". Reading this as direct evidence that preferences are constructed during the session is an interpretation: the paper documents instability and non-convergence, for which construction is one explanation (inference).
Scattered but consistent signs elsewhere. A Sequential Gallery participant selected designs "based on criteria that I didn't have at the beginning" (Koyama et al., 2020); AdaptiFont finds the best font at the time of use rather than a stable optimum, since the font that maximizes reading speed may depend on the text, fatigue, and display (Kadner et al., 2021); and Dudley et al. (2019) had to cope with noise from differences between users and tasks and from learning effects.
Real people against simulated ones. In personalizing a retinal implant with 17 sighted participants, only about 50% of choices agreed with the simulated agent's (Schoinas et al., 2025) (Section 33.5). In the preprint of Peng et al. (2026), 20 participants with varying interface-design experience judged the same 600 pairs of generated interfaces and agreed with each other at only 0.25 on a chance-corrected agreement statistic where 0 means chance and 1 perfect agreement (Krippendorff's ; Cohen's is the same to two decimals). The vibrotactile study notes that comparisons between nearby candidates are "dominated by response noise" and that its model assumes stationarity, with no account of adaptation, fatigue, or drift (Zhang et al., 2026b). Meta-PO, Song et al., and LAPPI list static preferences as a limitation without measuring change (Li et al., 2025a; Song et al., 2025; Kuroki et al., 2026).
From 2017 to 2026 no design study tested preference drift within a session with a repeated-measures design, for example by showing early pairs again at the end of a session; most papers list drift as a limitation and leave it there. A Gaussian process model built on a latent utility and i.i.d. noise treats anchoring, diminishing returns, and criteria that form during the session all as noise (inference). The test costs a few comparisons per session (Exercise 32.2), so a team running PBO with people can add it to every study (inference).
The models that could absorb such effects are discussed in Section 27.2 and Section 29.10, and the experiment that would separate noise from drift in Section 47.4.
Sources cited in Section 32.3 10
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Kadner et al. (2021) AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization
- Dudley et al. (2019) Crowdsourcing Interface Feature Design with Bayesian Optimization
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
- Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Song et al. (2025) Preference-Guided Multi-Objective UI Adaptation
- Kuroki et al. (2026) LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
32.4 Forms of feedback #
What should the system ask? Chapter 20 treats the query forms as methods; this section collects what was learned with people. Before 2026 each study varied one aspect of the query. Koyama and Igarashi (2018) argue that a query over a continuous space carries much more information than a comparison but is hard for crowd workers, so the single slider is their compromise. Chong et al. (2021) measured 17.2 seconds per iteration with one slider and 53.4 with four; four converged faster per iteration and were preferred, but . In virtual-reality color grading, users preferred comparing two options to four (Yuan et al., 2025). A "don't know" area let people express incomplete preferences (Ou et al., 2023); in the field study, rating 4 models instead of picking 1 did not remove inconsistency (Ou et al., 2022) (Section 20.4 covers ties and abstention). And Mikkola et al. (2020) argue that slider-like projective queries greatly reduce user workload.
The first comparisons on the same task, in 2026. GimmBO (, within-subjects, 20 iterations) found that ranking beat sliders in similarity to the target, in success rate, and in recovering the right sparse set of adapters; sliders switched on superfluous adapters and Sequential Gallery users got stuck in local minima, but ranking took longer per step (50.5 against 34.7 seconds) (Liu et al., 2026b). APPO, with binary preferences only, converged faster and with lower load but was less expressive than baselines that allowed text feedback or manual edits (Li et al., 2026f). MultiBO argues that choosing one of 4 candidates per round balances information against load better than a pair, from a design analysis rather than a controlled experiment (Rajagopalan et al., 2026). In an online study with 12 new users, Peng et al. (2026) (a preprint) found that personalization from 8 pairwise judgments beat every baseline, including shared design guidelines and the users' own written preferences, with an aggregate win rate of 60.35% against 50.9% for a judge conditioned on a persona. And Mo et al. (2024) criticize galleries because all options are still supplied by the optimizer, leaving the user nothing to express beyond liking or disliking them.
These results point the same way: lower the burden of each query or raise the information it carries. No study yet compares pairs, galleries, sliders, rankings, and ratings on one task with real users, so advice on format rests mainly on simulation, very small samples, and the authors' arguments (inference). Until it exists, the cheapest safeguard is a short pilot of two candidate formats on the real task, timing each answer and repeating a few queries (inference). Section 20.6 takes up why the interface is part of the model.
Sources cited in Section 32.4 11
- Koyama and Igarashi (2018) Computational Design with Crowds
- Chong et al. (2021) Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance
- Yuan et al. (2025) Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
- Rajagopalan et al. (2026) Personalized Image Generation via Human-in-the-loop Bayesian Optimization
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
32.5 Individual differences and population priors #
A crowdsourcing framework assumes one preference shared by the crowd, though crowds from different backgrounds "can have clearly different preferences" (Koyama and Igarashi, 2018). Population priors borrow other people's data to help each new person. They work in experiments, but their gains are concentrated in the first iterations.
The evidence for population priors. The global Gaussian process of Liao et al. (2024b) (performance) produced a design that cut completion time by 5.5% and spatial error by 48%, and the warm-start Gaussian process raised hypervolume by 38% for experienced and 18% for novice users, though a global design oriented toward speed was not better than the baseline in spatial error. Meta-Bayesian optimization of wrist interaction improved absolute pointing by 22.92% over standard Bayesian optimization and 21.35% over manual calibration, and relative pointing by 25.43% and 13.60% (Liao et al., 2024a). Meta-PO (preference feedback) cut the iterations to satisfaction from 9.54 without transfer (standard deviation 2.19) to 5.86 when transferring between users with the same theme (standard deviation 1.20, , ) and to 7.41 across themes (standard deviation 1.28, ); across themes needed more iterations than within a theme (), so the more the goals differ, the less transfer helps (Li et al., 2025a). HOMI was significantly better only at iterations 2 and 3 ( and ); from iteration 6 all methods converged to the same level, and NASA-TLX did not differ (Liao et al., 2026).
The evidence for individual differences. Every study that measured individual differences found them large: hearing-aid gain adjustments () (Søgaard Jensen et al., 2019) and preferred compression ratios (Baltzell et al., 2018) (Section 33.4), the font that maximized reading speed (Kadner et al., 2021), the trade-off between typing speed and accuracy (Shen et al., 2022), design preferences for air-taxi interfaces (Meinhardt et al., 2025), and preferred blur in driving (Li et al., 2026b). Even expert preference and measured outcome can disagree: in a pre-study, professional ad designers' preferences showed no positive correlation with actual click-through rates, which is why constrained PBO hands the click-through rate to the machine (Iwai et al., 2025).
A population prior shifts the unit of analysis from the person to the population, which sits uneasily with these individual differences. The studies show that population priors speed up early convergence; none shows that they do no harm when an individual departs from the population (inference). A system using such a prior should therefore let the person's own answers outweigh it quickly, and report results separately for the people farthest from the population (inference). Section 20.5 and Section 27.3 describe the models.
Sources cited in Section 32.5 12
- Koyama and Igarashi (2018) Computational Design with Crowds
- Liao et al. (2024b) Practical approaches to group-level multi-objective Bayesian optimization in interaction technique design
- Liao et al. (2024a) A Meta-Bayesian Approach for Rapid Online Parametric Optimization for Wrist-based Interactions
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Søgaard Jensen et al. (2019) Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference
- Baltzell et al. (2018) Efficient characterization of individual differences in compression ratio preference
- Kadner et al. (2021) AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization
- Shen et al. (2022) Personalization of a Mid-Air Gesture Keyboard using Multi-Objective Bayesian Optimization
- Meinhardt et al. (2025) Fly Away: Evaluating the Impact of Motion Fidelity on Optimized User Interface Design via Bayesian Optimization in Automated Urban Air Mobility Simulations
- Li et al. (2026b) BlurDriving: Investigating How Personalized Blur Techniques Impact Drivers' Performance in Virtual Reality
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
32.6 Novices and experts #
Who is being optimized for matters as much as how. In the study of Ou et al. (2023) (), "novices can achieve an expert level of quality performance, but participants with higher expertise led to more optimization iteration with more explicit preference while keeping satisfaction low. In contrast, novices were more easily satisfied and terminated faster." The authors' explanation is that "experts seek more diverse outcomes while the machine reaches optimal results". In the field study, too, lab participants were more easily satisfied than the expert artists, who applied quality criteria the lab participants ignored (Ou et al., 2022).
Experts repeatedly asked for direct control: to manipulate parameters outside the gallery (Koyama et al., 2020), or to edit glyphs, switch off style propagation, and see a fuller history (Tatsukawa et al., 2025). They were positive when the optimizer took over what they could not judge themselves: the 11 ad designers of constrained PBO, with 5.82 years of experience on average, endorsed letting the system manage click-through rate (Iwai et al., 2025). Materials-science experiments also found experts and non-experts behaving differently (Section 34.2).
Most samples, however, are novices or students: 40 novices in Chan et al., 18 participants aged 20 to 36 in Mo et al., 10 non-experts in FontCraft, participants all with a computing or machine-learning background in GimmBO, and participants aged 24 to 28 in APPO (Chan et al., 2022; Mo et al., 2024; Liu et al., 2026b; Li et al., 2026f). What is known about professional designers using PBO tools comes mainly from two technical artists and two interviews (inference).
Sources cited in Section 32.6 9
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Tatsukawa et al. (2025) FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
32.7 Explanation and trust #
Explanations improved task performance in experiments. In a between-subjects experiment with tuning 6 parameters of a simulated egg-boiling task, any explanation (a bar chart, rules, or text) improved task success, reduced the trials needed, and improved understanding and confidence, without adding workload (Chakraborty et al., 2025b). MOLONE's authors report that comparative explanations over inputs and outcomes gave significantly faster convergence in a user study whose size we could not obtain (Chakraborty et al., 2025a). ShapleyBO, a preprint, explains each proposal with Shapley values, which divide the proposal's expected value among the input parameters; in personalizing an exosuit, teams that saw the explanations had lower cumulative regret, the total shortfall from the best design over the session (Chapter 13) (Rodemann et al., 2024). CoExBO explains each candidate to foster trust and comes with a no-harm guarantee: even under adversarial input from the user, it converges like ordinary Bayesian optimization (Adachi et al., 2024). Participants of Niwa et al. (2025) found that explanations made unexpected suggestions more acceptable, while some found them too generic or too long.
Understanding the optimizer can also change the feedback. Colella et al. (2020) found that users who understood the optimizer gave strategically biased answers (21 people, a one-dimensional task, scalar feedback), unlike the truthful simulated users of most papers. Sandholtz et al. (2023) (Bayesian Analysis, 2023) found that many participants explored in ways standard acquisition functions do not capture, and a preprint by Weichert et al. (2025) documents an industrial case in which adding data and expert knowledge made Bayesian optimization worse.
The "trust" in the Ulm studies is trust in the vehicle, an objective being optimized, not trust in the optimizer (Jansen et al., 2025), and trust in the optimizer itself has almost never been measured directly in HCI studies of PBO (inference); a study that wants to claim it should ask about the optimizer by name. Making the search visible and editable (explicit constraints, process disclosure, a withdrawable history) appears to sustain agency better than text explanations alone (inference).
Sources cited in Section 32.7 9
- Chakraborty et al. (2025b) Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
- Chakraborty et al. (2025a) Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection
- Rodemann et al. (2024) Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration
- Adachi et al. (2024) Looping in the Human Collaborative and Explainable Bayesian Optimization
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Colella et al. (2020) Human Strategic Steering Improves Performance of Interactive Optimization
- Sandholtz et al. (2023) Inverse Bayesian Optimization: Learning Human Acquisition Functions in an Exploration vs Exploitation Search Task
- Weichert et al. (2025) When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge
- Jansen et al. (2025) OptiCarVis: Improving Automated Vehicle Functionality Visualizations Using Bayesian Optimization to Enhance User Experience
32.8 Against strong baselines #
A method can look good by beating a weak competitor. Table 32.2 sorts the main comparisons by how strong the comparison condition was. The grading is this chapter's judgment (inference): random queries, a single slider, and no visualization count as weak; designers' own design, manual parameter tuning, and settings from the literature as medium; similar optimizers, Bayesian optimization that ignores cost, and expert designs as strong.
| Study | Compared with | Strength | Result |
|---|---|---|---|
| ROMBO (2025) | random queries | weak | new favorite tracks found 40% more often and 16% faster, 18% less time on disliked tracks () |
| FontCraft (2025) | single-slider Bayesian optimization | weak | closer to target; no inferential statistics seen () |
| OptiCarVis (2025) | no visualization; user-customized; expert design | weak to strong | personalized designs beat expert and customized designs on several ratings |
| Sequential line search (2017) | commercial auto-enhancement | medium | large lead in crowd votes |
| Dudley et al. (2019) | baseline interface | medium | no difference in batch 1; shorter task times from batch 2 |
| Shen et al. (2022) | baseline keyboard | medium | speed +14.4% (), accuracy +13.8% (), learning effects ruled out |
| Liao et al. (2024) | standard Bayesian optimization; manual calibration | strong; medium | absolute pointing improved 22.92% and 21.35%, relative pointing 25.43% and 13.60% |
| Tanaka et al. (2026) | sliders | medium | 62.5% preferred the optimized results; 85% endorsed their diversity |
| GimmBO (2026) | sliders; gallery | medium | ranking better, but slower per step |
| Liao et al. (IEEE Pervasive Computing 2023) | designers' own strategies | medium | "very similar" design performance |
| Song et al. (2025) | manual; ParetoSelect | medium to strong | layouts ranked above ParetoSelect's (); no difference in hypervolume, overall quality, or experience |
| Multi-session vehicle visualization (AutomotiveUI 2026) | the day-one design kept on days two and three | medium to strong | continued optimization rated better on cognitive load, trust, predictability, and perceived safety (); the authors report shortcomings for subjective measures |
| Mo et al. (2024) | optimizer-led | strong | no difference in hypervolume; fewer Pareto designs |
| Niwa et al., study 2 (2025) | cooperation by explicit constraints | strong | no difference in performance () |
| Cost-aware Bayesian optimization (2026) | Bayesian optimization ignoring cost | strong | about 67% of the cost; final quality no different () |
| HOMI (2026) | transfer acquisition; continual Bayesian optimization | strong | no difference from iteration 6 on |
| BlurDriving (IMWUT 2026) | no blur | medium | no significant improvement in objective driving performance |
| Ou et al., field deployment (2022) | none | not applicable | 415 of 549 sequences never reached optimization |
Figure 32.3 places the studies on two axes, sample size and the strength of the comparison, colored by feedback type.
The pattern is clear (inference). The stronger positive results come from performance objectives: task time, typing speed, pointing error. Studies of pure preference feedback more often report satisfaction or iteration counts, or a win over a weak comparison such as a single slider (ROMBO's win over random queries was obtained with ratings). Against skilled manual work or a similar optimizer, final quality usually does not differ. What interactive Bayesian optimization delivered in design from 2017 to 2026 is best summarized as the same result at lower cost, in fewer iterations, evaluations, or minutes of effort, rather than a better result (inference). Section 46.7 turns this into advice on how to evaluate a new system.
Sources cited in Section 32.8 2
- Marcos et al. (2025) Random rotational embedding Bayesian optimization for human-in-the-loop personalized music generation
- Colley et al. (2026) Multi-Session User Experience Assessments of Computationally Optimized Automated Vehicle Functionality Visualizations
32.9 Critiques and alternatives from design research #
Between 2017 and 2025 the critiques came mainly from the authors of optimization papers and from neighboring HCI research, and they are of two kinds. The first concerns agency, with the evidence of Section 32.2: Chan et al. propose "to push the optimizer to the background, making its suggestions recommendations and not dictations" (Chan et al., 2022), and Mo et al. (2024) add that an optimizer typically cannot "leverage the designer's expertise in quickly identifying that a given 'bad' design is not worth" evaluating.
The second concerns the utility model itself. Ou et al. (2022) conclude that "we need to generally rethink basic assumptions and approaches in the design of HITL systems". Koyama and Igarashi (2018) ask "Whose preference?": crowd frameworks assume "a 'general' (or universal) preference shared among crowds", while in some domains "only experts can adequately assess the quality of designs". Koyama and Goto (2022) leave open whether BO-generated suggestions are creative, and Dudley et al. (2019) could not tell whether their optimizer truly optimized or only excluded poorly performing regions. In 2026, Langerak et al. (2026) wrote that a predefined parameter space limits early exploration while the design space is still evolving, and Owaki et al. (2026) argue that the representation shown to the user is part of interaction design; in their study () a 5-dimensional feasibility-aware latent space beat the raw 9-dimensional parameters.
Fixation. Design fixation is the tendency to stay close to examples one has seen, or to one's first idea. An AI image generator used during ideation led to more fixation, fewer ideas, and lower variety and originality () (Wadinambiarachchi et al., 2024), and students can fixate more on their own first idea than on given examples (Leahy et al., 2020). A PBO gallery is a stream of system-supplied examples with the current best as the first idea, so the same fixation can be expected when PBO is used for ideation (inference). Gmeiner et al. (2023) watched 14 trained designers struggle to understand and steer generative design tools.
Preferences formed in the interaction. Research on co-creation with generative AI in 2025 and 2026 argued that preferences form during the interaction, that premature convergence and fixation are the main failures, and that some friction may help. Switchable divergent and convergent modes (HAICo, a preprint) scored higher than ChatGPT on every factor of the Creativity Support Index (, ) (Wen et al., 2026); IdeaBlocks led designers to explore 2.13 times as many images, with 12.5% higher visual diversity (Choi et al., 2026); a workshop paper argues for keeping reflective friction (Avelino et al., 2026); a preprint found that elicitation surfaces preferences users had not yet formed (Kim et al., 2026); and Saracay et al. (2026) (COLM 2026) argue that agents should help users construct preferences, with a simulated-user benchmark whose user model a study with 25 people supports. Chapter 45 returns to what a comparison measures if preferences are partly built while answering.
The alternatives move toward mixed initiative, in which either side can take the lead: the systems of Section 32.2, a withdrawable history (Tatsukawa et al., 2025), constraints handled by the machine (Iwai et al., 2025), and the designer as a curator of constraints rather than a maker of prototypes (Jansen, 2025), a workshop paper. For ideation, Koch et al. (2019) used cooperative contextual bandits, a sequential recommender that adapts to feedback, to suggest inspirational material for mood boards instead of converging on an optimum; 14 of 16 professional designers preferred the tool.
Counter-evidence: optimization can support exploration. Optimizer-led designers explored more of the design space (Chan et al., 2022), cooperative designers moved farther between evaluations than designer-led ones (Mo et al., 2024), participants of Niwa et al. described suggestions that broke their fixation as "Oh, I see!" moments (Niwa et al., 2025), and 85% of the architecture students endorsed the diversity the optimizer brought () (Tanaka et al., 2026).
Where the two critiques stand. The agency critique already has tested remedies (cooperative control, explicit constraints, an exploration slider) that restored most of the agency in experiments. The utility-model critique has only proposed remedies inside PBO, such as discarding history, decaying old data, and withdrawing choices, and a PBO tool has yet to be compared with an ideation tool built for divergence on the same creative task (inference). Critics and optimization authors largely agree on the diagnosis (fixation, unclear goals, preferences that form during the interaction) and disagree on the remedy.
The evidence suggests a working rule for design tools (inference): put the optimizer in the role of an advisor; give the machine the measurable sub-goals, such as click-through rate, feasibility, or task time; leave judgments of appearance to comparisons or rankings; make the state of the search visible and editable; and evaluate against manual tuning or skilled designers, not only against a single slider or random queries. The rule comes from the empirical results of this chapter and does not depend on the Gaussian process framework.
Sources cited in Section 32.9 22
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Koyama and Igarashi (2018) Computational Design with Crowds
- Koyama and Goto (2022) BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions
- Dudley et al. (2019) Crowdsourcing Interface Feature Design with Bayesian Optimization
- Langerak et al. (2026) Cost-Aware Bayesian Optimization for Prototyping Interactive Devices
- Owaki et al. (2026) Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs
- Wadinambiarachchi et al. (2024) The Effects of Generative AI on Design Fixation and Divergent Thinking
- Leahy et al. (2020) Design Fixation From Initial Examples: Provided Versus Self-Generated Ideas
- Gmeiner et al. (2023) Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools
- Wen et al. (2026) Exploration vs. Fixation: Scaffolding Divergent and Convergent Thinking for Human-AI Co-Creation with Generative Models
- Choi et al. (2026) IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI
- Avelino et al. (2026) Creativity from Friction: Human-AI Interaction for Exploratory Structural Design
- Kim et al. (2026) Elicitive User Interfaces: Designing How Users Shape Generative Interfaces
- Saracay et al. (2026) Beyond expert users: agents should help users construct preferences, not just elicit them
- Tatsukawa et al. (2025) FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Jansen (2025) Human-in-the-Loop Optimization for Inclusive Design: Balancing Automation and Designer Expertise
- Koch et al. (2019) May AI? Design Ideation with Cooperative Contextual Bandits
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Tanaka et al. (2026) Human-in-the-Loop Bayesian Optimization Approach to Supporting Early-Stage Architectural Design
32.10 Common claims, checked #
Table 32.3 checks claims that circulate in secondary accounts against the primary sources.
| Claim | Verdict | What the sources show |
|---|---|---|
| Sequential line search converges in about 15 to 20 iterations on a 6-dimensional problem. | partly right | The photo task has 6 parameters, and the crowdsourced runs were fixed at 15 iterations; runs from different starts came together within the first 4 to 5 iterations. Nothing supports "15 to 20" (Koyama and Igarashi, 2018). |
| AdaptiFont, continual human-in-the-loop optimization, and Chan et al. are evidence about PBO. | wrong scope | AdaptiFont optimizes reading speed, Chan et al. speed and accuracy, and continual human-in-the-loop optimization uses performance feedback; none uses preference feedback (Kadner et al., 2021; Chan et al., 2022; Liao et al., 2025). |
| In Niwa et al., a language model turns design intent into constraints on the search space. | partly right | The language model lets designers intervene in a system-led optimization and explains its reasoning; the study compared it with a method that uses explicit constraints (Niwa et al., 2025). |
| Constrained PBO (Iwai et al. 2025) is the first PBO with inequality constraints. | the paper's own claim, too strong | The authors are Iwai, Kumagae, Koyama, Hamasaki, and Goto. "First" overlooks C-GLISp (IEEE TCST 2022) and StageOpt (ICML 2018) (Iwai et al., 2025; Zhu et al., 2022; Sui et al., 2018b). |
Two citation details are easy to get wrong. The 2020 SIGGRAPH paper is called Sequential Gallery; "sequential plane search" is the method inside it (Koyama et al., 2020). And the melody work is two papers, a 2020 pilot (Zhou et al., 2020) and the IUI 2021 study with 12 participants (Zhou et al., 2021).
Sources cited in Section 32.10 11
- Koyama and Igarashi (2018) Computational Design with Crowds
- Kadner et al. (2021) AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Liao et al. (2025) Continual Human-in-the-Loop Optimization
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Zhou et al. (2020) Generative Melody Composition with Human-in-the-Loop Bayesian Optimization
- Zhou et al. (2021) Interactive Exploration-Exploitation Balancing for Generative Melody Composition
32.11 Settled, contested, missing #
Settled. In design tasks with performance objectives, letting the optimizer lead lowers agency, ownership, and expressiveness by a large margin while improving outcome measures modestly, and channels for steering restore most of the lost agency at a small cost in outcome (Chan et al., 2022; Mo et al., 2024; Niwa et al., 2025). Population priors cut the iterations people need (Li et al., 2025a), and in the one study that tracked the whole session their advantage was gone by the sixth iteration (Liao et al., 2026). Individual differences are large wherever they were measured. Real sessions often end early: in the one long field deployment, three quarters of sequences never reached optimization (Ou et al., 2022).
Contested. Whether optimizers help or hinder creative exploration: the fixation literature points one way, measures of design-space coverage the other. Whether the gains of interactive Bayesian optimization survive strong comparisons, where final quality usually does not differ. Which query format is best: the first same-task comparisons appeared only in 2026, with 12 to 16 participants. Whether explanations build trust in the optimizer, or only improve task performance.
Missing. A repeated-measures test of preference drift within a session. A comparison of pairs, galleries, sliders, rankings, and ratings on one task with real users. A measurement of fixation inside a PBO session. A direct measure of trust in the optimizer. A study of agency with pure preference feedback and a sample comparable to Chan et al.'s. Studies of professional designers beyond two artists and two interviews. A comparison of a PBO tool with an ideation tool built for divergence on the same task. Evidence that population priors do no harm to people far from the population (inference).
Sources cited in Section 32.11 6
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
32.12 Exercises #
Figure 32.2 gives three success rates for the same tool: 2.9% of all field sequences, 11.9% of the field sequences that asked for optimization, and 48.5% of the lab sequences. Which of them describes how the tool performs in practice? Why might a paper that reports only the 134 sequences, or only the lab study, give a reader the wrong idea?
Solution
The 2.9% () does, against of the sequences that requested optimization and in the lab. The 415 sequences that stopped at the first iteration are part of how the tool was used: the artists looked at the first grid and stopped, perhaps because a candidate was good enough, perhaps because they gave up. Reporting only the sequences that entered optimization conditions on people having already chosen to keep going, and reporting only the lab study replaces professional users with participants who, as the authors found, are more easily satisfied. Both choices make the method look better than its field use.
No design study between 2017 and 2026 tested preference drift within a session. Design the simplest test you can add to an existing PBO session of 30 comparisons. What would you repeat, when, and what pattern in the answers would distinguish drift from ordinary answer noise?
Solution
Repeat some early pairs at two later points: for example, show pairs 2, 4, and 6 again right after they were first answered (short lag), and again at the end of the session (long lag), in random order among the regular queries. Answer noise alone predicts the same agreement rate at both lags, because each answer is an independent noisy reading of a fixed utility. Drift predicts lower agreement at the long lag than at the short lag, and in a consistent direction: if the person's criteria moved, the late answers should agree with the model fitted to late data better than with the model fitted to early data. With a handful of repeats per person the test is weak for one person and useful across a group; the cost is a few extra comparisons per session. Section 47.4 describes a fuller version of this experiment.
A new paper reports that its preference-based design tool reached a satisfactory result in 6 iterations against 11 for a single-slider baseline, with 12 participants. Using the grading of Table 32.2, how strong is this comparison? Name two comparison conditions that would make the result more informative, and say what each would rule out.
Solution
A single slider is a weak comparison in the chapter's grading. Two stronger ones: manual tuning of the same parameters by the same participants with the same time budget, which tests whether people would do as well on their own; and a similar optimizer with a different query form or acquisition function, which tests whether the gain comes from the new component rather than from optimization in general. Measuring final quality with judges who did not take part, not only iterations to satisfaction, would also separate "faster" from "better".
Further reading #
- Chan et al. (2022) is the clearest controlled study of agency against performance; read its discussion of mixed initiative alongside the results.
- Mo et al. (2024) show what a cooperative middle ground looks like and measure it against both extremes in one within-subjects study.
- Ou et al. (2022) is the only long field deployment of PBO with professionals, and the most detailed record of how real feedback departs from the model.
- Koyama et al. (2020) and the book chapter Koyama and Igarashi (2018) give the interface line's reasoning about query forms and its frank list of limitations.
- Niwa et al. (2025) measure outcome and agency together when the designer steers in natural language, and are the bridge to the language-model systems of Chapter 35.
References
- (2024). Looping in the Human Collaborative and Explainable Bayesian Optimization. AISTATS 2024. Cited in §32.7
- (2026). Creativity from Friction: Human-AI Interaction for Exploratory Structural Design. ICML 2026 Workshop on Human-AI Co-Creativity. workshop paper Cited in §32.9
- (2018). Efficient characterization of individual differences in compression ratio preference. The Journal of the Acoustical Society of America. doi:10.1121/1.5067390. Cited in §32.5
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §32.1
- (2023). Preference-Based Human-in-the-Loop Optimization for Perceived Realism of Haptic Rendering. IEEE Transactions on Haptics. doi:10.1109/toh.2023.3266726. Cited in §32.1
- (2025a). Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection. World Conference on eXplainable AI 2025. Cited in §32.7
- (2025b). Explanation format does not matter; but explanations do – An Eggsbert study on explaining Bayesian Optimisation tasks. Information Systems Frontiers. doi:10.1007/s10796-025-10671-6. Cited in §32.7
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Cited in §32.1 §32.2 §32.6 §32.9 §32.10 §32.11
- (2020). Human-in-the-loop differential subspace search in high-dimensional latent space. ACM Transactions on Graphics. doi:10.1145/3386569.3392409. Cited in §32.1
- (2026). IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI. Proceedings of the 2026 Designing Interactive Systems Conference. doi:10.1145/3800645.3813005. Cited in §32.9
- (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. Cited in §32.1 §32.4
- (2020). Human Strategic Steering Improves Performance of Interactive Optimization. UMAP 2020. Cited in §32.7
- (2025). Improving External Communication of Automated Vehicles Using Bayesian Optimization. CHI 2025. Cited in §32.1 §32.2
- (2026). Multi-Session User Experience Assessments of Computationally Optimized Automated Vehicle Functionality Visualizations. Proceedings of the 18th International Conference on Automotive User Interfaces and Interactive Vehicular Applications. doi:10.1145/3828157.3828785. Cited in §32.1 §32.8
- (2019). Crowdsourcing Interface Feature Design with Bayesian Optimization. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. doi:10.1145/3290605.3300482. Cited in §32.1 §32.3 §32.9
- (2026). Putting the Human Back in the Loop: A Review of Interactive Bayesian Optimization. ACM Computing Surveys. Cited in §32.1
- (2026). Integrating Multi-Source Feedback in Computational Design. Accepted at ACM TiiS. Cited in §32.1
- (2023). Exploring Challenges and Opportunities to Support Designers in Learning to Co-create with AI-based Manufacturing Design Tools. CHI 2023. Cited in §32.9
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §32.1
- (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Cited in §32.1 §32.5 §32.6 §32.9 §32.10
- (2025). Human-in-the-Loop Optimization for Inclusive Design: Balancing Automation and Designer Expertise. CHI 2025 Workshop Access InContext. workshop paper Cited in §32.9
- (2025). OptiCarVis: Improving Automated Vehicle Functionality Visualizations Using Bayesian Optimization to Enhance User Experience. CHI 2025. Cited in §32.1 §32.7
- (2021). AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization. CHI 2021. Cited in §32.1 §32.3 §32.5 §32.10
- (2026). Elicitive User Interfaces: Designing How Users Shape Generative Interfaces. arXiv. preprint Cited in §32.9
- (2019). May AI? Design Ideation with Cooperative Contextual Bandits. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. doi:10.1145/3290605.3300863. Cited in §32.9
- (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. Cited in §32.1 §32.9
- (2018). Computational Design with Crowds. Computational Interaction. Cited in §32.1 §32.4 §32.5 §32.9 §32.10
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §32.1
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §32.1 §32.3 §32.6 §32.10
- (2026). LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation. IEEE Access 14. Cited in §32.1 §32.3
- (2026). Cost-Aware Bayesian Optimization for Prototyping Interactive Devices. CHI 2026. Cited in §32.1 §32.9
- (2020). Design Fixation From Initial Examples: Provided Versus Self-Generated Ideas. Journal of Mechanical Design. Cited in §32.9
- (2026). Part-level 3D shape generation driven by user intention inference with preferential Bayesian optimization. Scientific Reports. doi:10.1038/s41598-026-38916-7. Cited in §32.1
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §32.1 §32.3 §32.5 §32.11
- (2026a). Automating UI Optimization through Multi-Agentic Reasoning. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI 2026). doi:10.1145/3772318.3791444. Cited in §32.1
- (2026b). BlurDriving: Investigating How Personalized Blur Techniques Impact Drivers' Performance in Virtual Reality. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies. doi:10.1145/3831646. Cited in §32.1 §32.5
- (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Cited in §32.1 §32.4 §32.6
- (2023). Interaction Design With Multi-Objective Bayesian Optimization. IEEE Pervasive Computing. doi:10.1109/mprv.2022.3230597. Cited in §32.1 §32.2
- (2024a). A Meta-Bayesian Approach for Rapid Online Parametric Optimization for Wrist-based Interactions. Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3613904.3642071. Cited in §32.1 §32.5
- (2024b). Practical approaches to group-level multi-objective Bayesian optimization in interaction technique design. Collective Intelligence. doi:10.1177/26339137241241313. Cited in §32.1 §32.5
- (2025). Continual Human-in-the-Loop Optimization. CHI 2025. Cited in §32.1 §32.10
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §32.1 §32.5 §32.11
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §32.1 §32.4 §32.6
- (2025). Random rotational embedding Bayesian optimization for human-in-the-loop personalized music generation. PLOS One. doi:10.1371/journal.pone.0335853. Cited in §32.1 §32.8
- (2019). Sampling Humans for Optimizing Preferences in Coloring Artwork. ICML 2019 Workshop on Human in the Loop Learning. workshop paper Cited in §32.1
- (2025). Fly Away: Evaluating the Impact of Motion Fidelity on Optimized User Interface Design via Bayesian Optimization in Automated Urban Air Mobility Simulations. CHI 2025. Cited in §32.1 §32.5
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §32.4
- (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. Cited in §32.1 §32.2 §32.4 §32.6 §32.9 §32.11
- (2025). Exploring the Effectiveness of Interactive Preference Learning for Adapting Designs to Abstract Semantic Attributes. Journal of Mechanical Design. Cited in §32.1
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §32.1 §32.2 §32.7 §32.9 §32.10 §32.11
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §32.1 §32.3 §32.4 §32.6 §32.9 §32.11
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Cited in §32.1 §32.4 §32.6
- (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. preprint Cited in §32.1 §32.9
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §32.1 §32.3 §32.4
- (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. Cited in §32.1 §32.4
- (2024). Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration. arXiv. preprint Cited in §32.7
- (2023). Inverse Bayesian Optimization: Learning Human Acquisition Functions in an Exploration vs Exploitation Search Task. Bayesian Analysis. doi:10.1214/21-BA1303. Cited in §32.7
- (2026). Beyond expert users: agents should help users construct preferences, not just elicit them. Conference on Language Modeling (COLM 2026). Cited in §32.9
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §32.3
- (2022). Personalization of a Mid-Air Gesture Keyboard using Multi-Objective Bayesian Optimization. 2022 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). doi:10.1109/ismar55827.2022.00088. Cited in §32.1 §32.5
- (2019). Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference. Trends in Hearing. doi:10.1177/2331216519847413. Cited in §32.5
- (2025). Preference-Guided Multi-Objective UI Adaptation. Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3746059.3747645. Cited in §32.1 §32.2 §32.3
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §32.10
- (2026). ProVoice: Designing Proactive Functionality for In-Vehicle Conversational Assistants using Multi-Objective Bayesian Optimization to Enhance Driver Experience. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. doi:10.1145/3772318.3791877. Cited in §32.1
- (2026). Human-in-the-Loop Bayesian Optimization Approach to Supporting Early-Stage Architectural Design. CAADRIA proceedings. doi:10.52842/conf.caadria.2026.1.347. Cited in §32.1 §32.2 §32.9
- (2025). FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization. CHI 2025. Cited in §32.1 §32.2 §32.6 §32.9
- (2024). The Effects of Generative AI on Design Fixation and Divergent Thinking. Proceedings of the CHI Conference on Human Factors in Computing Systems. Cited in §32.9
- (2025). When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge. arXiv. preprint Cited in §32.7
- (2026). Exploration vs. Fixation: Scaffolding Divergent and Convergent Thinking for Human-AI Co-Creation with Generative Models. arXiv. preprint Cited in §32.9
- (2022). Photographic Lighting Design with Photographer-in-the-Loop Bayesian Optimization. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545690. Cited in §32.1
- (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. Cited in §32.1 §32.4
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Cited in §32.1 §32.3
- (2020). Generative Melody Composition with Human-in-the-Loop Bayesian Optimization. CSMC-MuMe 2020. Cited in §32.1 §32.10
- (2021). Interactive Exploration-Exploitation Balancing for Generative Melody Composition. 26th International Conference on Intelligent User Interfaces. Cited in §32.1 §32.2 §32.10
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §32.10