Wearable Robots, Health, and Assistive Technology
In the design studies of Chapter 32, a person judges a picture on a screen. In this chapter the thing being judged is felt through the body: the push of an exoskeleton, the swing of a prosthetic knee, the sound of a hearing aid, the light pattern of a retinal implant, the gait of a walking robot. Each evaluation takes seconds to minutes, sessions are limited by fatigue and safety, and the people who would benefit most are the hardest to recruit. These are the settings Chapter 1 used to motivate the book: an expensive, noisy objective that only the person can report.
The chapter covers wearable robots, hearing aids, neural prostheses, and robot control, from 2017 to September 2026, and counts other methods aimed at the same problem (radial basis function surrogates, linear reward models, neural rankers) alongside Gaussian process models, because several of the most cited studies use them. The thread is measurement. Studies are small, and validation is almost always internal: the same person's later choices agree with the model. Where an objective outcome or a manual baseline was measured, the picture is mixed: people tuning a hip exoskeleton themselves reached a benefit of the same order as algorithm tuning, and preference tuning of hearing aids improved sound quality but not speech clarity.
33.1 Exoskeletons #
An exoskeleton is a wearable robot that applies torque at a joint (hip, knee, or ankle) to assist walking; a soft exosuit does the same through textile straps and cables. Its controller is a curve of torque over the gait cycle, one stride from heel strike to the next heel strike of the same foot, shaped by a few numbers: peak torque, the timing of the peak, and when assistance starts and stops. The right values differ from person to person.
33.1.1 Before preferences: optimizing measured effort #
The field's starting point optimized metabolic cost, the energy the body spends, estimated by indirect calorimetry: the wearer breathes through a mask while oxygen and carbon dioxide are measured, and each setting must be walked for minutes before the reading settles. This is called human-in-the-loop optimization. In Zhang et al. (2017), torque patterns optimized by an evolution strategy (CMA-ES, Section 15.8.1) reduced metabolic energy consumption by 24.2 ± 7.4% compared with no torque, for 11 participants with an ankle exoskeleton. Ding et al. (2018) used Bayesian optimization to tune two timing parameters of hip assistance from a soft exosuit for 8 participants; the optimum was found in 21.4 ± 1.0 minutes on average, a convergence time computed afterward from a 40-minute optimization, and reduced metabolic cost by 17.4 ± 3.2% compared with walking without the device (mean ± standard error). Slade et al. (2022) replaced the mask with a model that estimates the metabolic effect from wearable sensors: it reached the same parameters as metabolic optimization, within 5%, after 32 minutes of walking instead of 128 (9 participants), and after one hour of walking outdoors the optimized assistance reduced metabolic cost on a treadmill at 1.5 m/s by 23 ± 8% compared with normal shoes (10 participants).
Metabolic measurement is slow and noisy and leaves out comfort and the feeling of safety (Schäfer et al., 2026); a 2024 perspective (Slade et al., 2024) and a 2023 review (Ingraham et al., 2023) discuss subjective objectives and user preference as part of the design space. Adaptation is the other complication: in Poggensee and Collins (2021), customized assistance after training reduced metabolic rate by 39% compared with the exoskeleton turned off, but training contributed about half of that benefit and customization about a quarter, and becoming an expert user took about 109 minutes of assisted walking. A person whose response is still changing is a moving target for any optimizer (Section 39.7).
33.1.2 The Caltech line: comparisons, suggestions, and ordinal labels #
A group at Caltech, with Maegan Tucker as first author of its main papers, built the main Gaussian process line of preference-based gait tuning on Atalante, a self-balancing lower-body exoskeleton that walks for its user. CoSpar combines pairwise preferences with coactive feedback, improvements the user volunteers ("a slightly longer step"), and picks the next gait by Thompson sampling, drawing possible utility functions from the posterior and letting them compete (Section 12.5) (Tucker et al., 2020b). With three able-bodied participants and 20 gait trials, each participant's blind ranking of three gaits agreed with the posterior; because "users have difficulty remembering more than two trials at once", each trial was compared only with the one before.
CoSpar became infeasible at 5 or more parameters, so LineCoSpar restricts each iteration to a random line through the gait with the highest posterior mean, the idea of Section 20.2 (Tucker et al., 2020a). Six able-bodied newcomers optimized 6 gait parameters in 30 trials and agreed with the model on 4 validation preferences each at 75%, 100%, 100%, 25%, 100%, and 100%; utility functions differed between people. ROIAL adds ordinal labels, a rating of each gait from "very bad" to "good", and learns the landscape only inside a region of interest that excludes uncomfortable gaits (Li et al., 2021). With 3 participants, 4 parameters, and 40 trials, most ordinal predictions were within one level of the reported label, and the search explored less than 2% of the action space: comfortable for the wearer, and a landscape that covers only a small corner.
The only study with complete paraplegia. Maegan Tucker's doctoral thesis reports the only preference-based exoskeleton study we found with participants who have complete motor paraplegia (Tucker, 2023): two people classified ASIA A (no motor or sensory function below the injury), one with more than 300 hours of experience with Atalante and one with fewer than 30. Three gait parameters certified for use in the European Union were tuned, on day one with ROIAL for 15 iterations and on day two with LineCoSpar for 15 and 25, with a 5-minute break every 20 minutes to prevent pressure sores. Each person could complete only one evaluation trial, and both rated the posterior's best gait as "good". The thesis also reports 8 able-bodied participants.
Several objectives. Astudillo et al. (2025) proposed an asymptotically consistent method for multi-objective optimization when every objective is observed only through preferences, dueling scalarized Thompson sampling, evaluated on simulated exoskeleton and driving tasks only (arXiv 2024, TMLR 2025). In a September 2026 preprint, MO-HILBO optimized metabolic cost and ordinal comfort feedback together over 3 controller parameters of a hip-knee exoskeleton, with 3 participants who each spent 3 to 4 hours; the surrogate predicted 94% (17 of 18) of the pairwise rankings among Pareto-front controllers correctly (Janwani et al., 2026).
33.1.3 Other groups #
A neural ranker. Lee et al. (2023) used a neural network ranker, pretrained on earlier preference data, to score ankle-assistance settings (4 parameters) proposed by an evolutionary algorithm, while the wearer made forced choices between pairs. It converged on the wearer's preferred parameters with an accuracy of 88% on average when compared with randomly generated parameters, and the preferred setting stabilized after 43 ± 7 queries. It is not a Gaussian process method. We could read only the abstract: it does not name the evolutionary algorithm (a later preprint calls it CMA-ES (Liu et al., 2026d)), and the participant count of 14 given in a third-party summary could not be checked against the paper.
A radial basis function surrogate. WANDER, an omnidirectional walking-aid robot, tuned 2 admittance parameters (how heavy and how damped the robot feels when pushed) with GLISp, the active preference learning algorithm of Bemporad and Piga, which fits a radial basis function surrogate instead of a Gaussian process (Fortuna et al., 2024; Bemporad and Piga, 2021). Against two settings from the literature, with 12 healthy adults and up to 15 iterations, linear energy fell by 17.61% and 13.93% (both significant); angular energy fell by 13.26% against the first (significant) and rose by 3.57% against the second (not significant); and jerk fell by 1.12% (not significant) and 5.95% (significant). The preferred virtual mass correlated with body weight (, ).
A linear reward. Ramella et al. (2025) modeled preference over hip torque curves as a linear reward over 6 features, with a posterior sampled by the Metropolis-Hastings algorithm (Section 17.5). Eight healthy participants made 12 pairwise comparisons, walking 20 seconds with each curve. Against perturbed versions (±2 N·m, ±7% of the gait cycle) most kept their original preference, though for four of them the check failed when the rise time was perturbed. Preferred torque was synchronized with the wearer's movement and had lower negative device power: the device absorbed less energy from the wearer.
Perception thresholds in the acquisition function. Arens et al. (2025) optimized lifting and lowering assistance from a soft back exosuit with 15 healthy participants, in 3 blocks of 15 iterations (each with 3 hidden validation checks) over 2 parameters. The upper confidence bound (Section 12.4) accounted for the just noticeable difference (JND), the smallest change a person can reliably detect (Section 16.2), measured separately at 13.1 N for lifting and 9.8 N for lowering. The intraclass correlation coefficient (ICC), how consistently a person lands on the same setting (1 is perfect), was 0.80 for lowering and 0.91 for lifting. Preferred assistance rose by 15% and 29% from block 1 to block 3, a trend that was not significant ().
Pairwise preferences with metabolic follow-up. In a 2026 preprint, Liu et al. (2026d) optimized 6 parameters of hip assistance from pairwise comparisons with five healthy adults who knew the device. After 20 iterations they chose the optimized torque profile over a random one in 90.7 ± 1.3% of validation comparisons, and in follow-up measurements on 2 to 3 of them metabolic rate fell by 14.5% to 15.4% across three speeds relative to assistance switched off, and muscle activation by 6.7% to 31.5% relative to walking without the exoskeleton (Section 33.3 discusses the baselines).
Context, inverse models, and perception. In a retrospective analysis of data from 9 healthy adults, a 2026 preprint found that preference landscapes in neighboring operating conditions tended to be more similar, and in simulation sharing data across contexts helped when that continuity held and caused negative transfer when it was weak (Baek et al., 2026); Park and Collins (2026) infer an individual reward while learning a forward model of the person, in simulation only. Maberry and Martin (2026) found that a walking person notices a change in ankle exoskeleton stiffness only at about 42%, and in the rate of ankle angle control at about 49%, more than three times the values measured standing still. A 16-person preprint on dynamic Bayesian optimization handles a non-stationary objective, but that objective is a measured hip angle, not a preference (Kim and Sergi, 2025).
33.1.4 The wearable studies at a glance #
Table 33.1 lists the wearable studies. Most "validation" in the table is internal: the same person's later choices are compared with the model's predictions.
| Study | Venue | Participants | Parameters | Comparisons or trials | Validation | Objective outcome | Compared with manual or expert settings |
|---|---|---|---|---|---|---|---|
| CoSpar | ICRA 2020 | 3 able-bodied | 1 (also 2) | 20 gait trials | blind ranking of 3 gaits matched the posterior | not measured | no |
| LineCoSpar | IROS 2020 | 6 able-bodied | 6 | 30 plus 6 validation | 25% to 100% per person | preference correlated with gait dynamism | no |
| ROIAL | ICRA 2021 | 3 able-bodied | 4 | 40 (30 training, 10 validation) | most ordinal predictions within one level | not measured | no |
| Tucker thesis | Caltech 2023 | 2 with complete paraplegia, plus 8 able-bodied | 3 | 15; next day 15 and 25 | one evaluation each, both rated "good" | not measured | no |
| Lee et al. (neural ranker) | Science Robotics 2023 | not checked against the paper | 4 | stable after 43 ± 7 queries | 88% on random parameters | not covered | no |
| WANDER (GLISp) | ICRA 2024 | 12 healthy adults | 2 | up to 15 iterations | not reported | linear energy 13.93% to 17.61% lower; other measures partly significant | settings from the literature; partly better |
| Ramella et al. (linear reward) | ICRA 2025 | 8 healthy | 6 features | 12 pairwise comparisons | most kept their preference against perturbed versions | lower negative device power | no |
| Arens et al. | Science Advances 2025 | 15 healthy | 2 | 3 blocks of 15 | ICC 0.80 and 0.91 | preferred assistance rose 15% and 29% (not significant) | no |
| Liu et al. | preprint 2026 | 5 healthy adults | 6 | 20 | 90.7% | metabolic rate 14.5% to 15.4% lower than with assistance off (2 to 3 participants) | no |
| Taddei et al. | preprint 2026 | 2 amputees, 2 able-bodied | 4 | discrete 10.3 ± 2.5 (3 participants); continuous 35.0 ± 6.2 (1 participant) | 93% and 67% | gait asymmetry improved in one amputee | no |
| MO-HILBO | preprint 2026 | 3 | 3 | 3 to 4 hours each | 94% of Pareto rankings predicted | metabolic cost was one objective | no |
| Ingraham et al. (self-tuning) | Science Robotics 2022 | 24 able-bodied | 2 | about 105 s per tuning | repeatability SD 1.7 N·m and 1.5% of gait cycle | naive users drifted toward more torque | is itself self-tuning |
| Schäfer et al. (self-tuning) | npj Biomedical Innovations 2026 | 11 healthy adults | 4 | 30.5 settings in 10.9 min | not reported | metabolic cost 16.6% lower than with zero torque | is itself self-tuning |
| Díaz et al. (self-exploration) | JNER 2026 | 3 transfemoral amputees | 2 | 8 to 14 settings per trial | 2 of 3 inconsistent across trials | less intact-side muscle activity | is itself self-exploration |
Sources cited in Section 33.1 23
- Zhang et al. (2017) Human-in-the-loop optimization of exoskeleton assistance during walking
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Slade et al. (2022) Personalizing exoskeleton assistance while walking in the real world
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Slade et al. (2024) On human-in-the-loop optimization of human–robot interaction
- Ingraham et al. (2023) Leveraging user preference in the design and evaluation of lower-limb exoskeletons and prostheses
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Tucker (2023) Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Janwani et al. (2026) Multi-Objective Human-in-the-Loop Bayesian Optimization of a Lower-Limb Exoskeleton
- Lee et al. (2023) User preference optimization for control of ankle exoskeletons using sample efficient active learning
- Liu et al. (2026d) Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization
- Fortuna et al. (2024) A Personalizable Controller for the Walking Assistive omNi-Directional Exo-Robot (WANDER)
- Bemporad and Piga (2021) Global optimization based on active preference learning with radial basis functions
- Ramella et al. (2025) Rapid Online Learning of Hip Exoskeleton Assistance Preferences
- Arens et al. (2025) Preference-based assistance optimization for lifting and lowering with a soft back exosuit
- Baek et al. (2026) Context-Continuous Preference Learning for Exoskeleton Personalization
- Park and Collins (2026) Simultaneous Forward and Inverse Human-in-the-Loop Optimization
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Kim and Sergi (2025) Validation of Dynamic Bayesian Optimization for a Non-Stationary Human-in-the-Loop Optimization Problem
33.2 Prostheses #
A powered prosthesis adds a further constraint: the wearer depends on the device for every step, so a poor setting is risky, and few people can take part. Taddei et al. (2026), a 2026 preprint, tuned 4 parameters of an active prosthesis with 2 transfemoral amputees (amputation above the knee) and 2 able-bodied people walking with an adapter. A discrete version, which combined the expected utility of the best option (EUBO, Section 19.4) with LineCoSpar and was run with one amputee and the two able-bodied participants, converged in 10.3 ± 2.5 iterations (8.5 ± 3.6 minutes) with a validation identification rate of 93%; for the amputee, swing-phase asymmetry improved from -13.4% to -7.6% and stance-phase asymmetry from 6.2% to 3.9% compared with the person's own prosthesis. A continuous version, run with the other amputee alone, needed 35.0 ± 6.2 iterations and reached a validation rate of only 67%. The authors attribute growing variability after the 15th iteration of one trial to fatigue, and write that "it remains unclear which biomechanical variables drive these preferences".
Díaz et al. (2026) let 3 transfemoral amputees explore the timing and magnitude of a powered knee-ankle prosthesis themselves, blind to the values, on a touchscreen grid (8 to 14 settings per trial). Preferred settings went with lower muscle activity on the intact side and less variable muscle synergies, but for 2 of the 3 people the preferred setting was not consistent across trials of the same day. Two neighboring studies are not PBO: in Alili et al. (2023), amputees choose a preferred knee kinematics curve and a reinforcement learning tuner fits 12 control parameters to it, and a preliminary look at gait found no clear difference between the preferred control and control tuned to a normative gait; Sun et al. (2024) validated individualized prosthesis control only on a gait data set, with no online test with amputees.
Sources cited in Section 33.2 4
- Taddei et al. (2026) Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
- Díaz et al. (2026) User preference in the personalized control of an ankle prosthesis: a case study
- Alili et al. (2023) A Novel Framework to Facilitate User Preferred Tuning for a Robotic Knee Prosthesis
- Sun et al. (2024) An Individual Prosthesis Control Method with Human Subjective Choices
33.3 Against manual tuning #
Every method in this chapter is meant to save effort compared with tuning by hand, which makes manual tuning the comparison that matters most. It is also the one the field has run least.
How reliable are people at tuning themselves? In Ingraham et al. (2022), twelve naive and twelve knowledgeable able-bodied participants adjusted the magnitude and timing of ankle exoskeleton torque themselves, blind to the values. Preferred settings ranged from 7.9 to 19.4 N·m and from 54.1% to 59.2% of the gait cycle; the repeatability, the standard deviation of a person's settings across repeated tunings, was 1.7 N·m and 1.5% of the gait cycle; each tuning took about 105 seconds; naive users' preferences shifted toward more torque over the experiment; and knowledgeable users preferred more torque than naive ones. Exposure and knowledge change the preferred setting itself, not only the noise around it.
The most informative direct comparison. In Schäfer et al. (2026), 11 healthy adults tuned 4 timing parameters of a bilateral hip exoskeleton themselves with a thumbstick while walking on a treadmill. They took 10.9 ± 0.9 minutes on average (at most 16.2), testing a median of 30.5 settings (range 16 to 111), almost always changing one parameter at a time (97.5% of adjustments). Walking with the preferred assistance cost 16.6 ± 1.1% less metabolic energy than with the exoskeleton at zero torque. Shifting the timing by up to ±8% of the stride time did not change the reduction significantly, and preferred timings differed between people by up to 22.5% of the stride. A sense-of-agency score (0 to 1) fell from 0.80 ± 0.07 at zero torque to 0.49 ± 0.07 with the self-tuned assistance (): even assistance the wearers chose reduced their sense of controlling their own movement. The authors note that self-tuning took a quarter of the time Ding et al. (2018) needed for 2 parameters, and half of Ding et al.'s convergence time extracted after the fact.
Reading the two 2026 results side by side. Liu et al. (2026d) tuned 6 parameters by pairwise Bayesian optimization in sessions of 20.6 ± 4.6 minutes including validation, with reductions of 14.5% to 15.4%. Its abstract calls the baseline "unassisted walking", but its results section shows that it is the exoskeleton worn with assistance off, the same kind of baseline as Schäfer et al.'s zero torque; relative to walking without the exoskeleton, the reductions at the three speeds were 8.6%, 9.3%, and 13.0%. Devices, parameters, and speeds differ, so the studies cannot be compared directly; what they show is that for a problem with few parameters, tuning by hand already reaches the order of benefit that algorithm tuning reports. Figure 33.1 places these results, and the metabolic optimizations of Section 33.1.1, on one chart.
Things to look for:
- In Time and benefit, the self-tuning result sits at the left, inside the range of the optimizer-tuned results, not below it.
- The one result clearly above the others, Slade et al.'s 23%, is measured against normal shoes and comes after an hour of walking.
- In Who took part, almost every marker sits below 15 participants, and the hollow markers for clinical participants sit at 2 to 4.
What it means. The ±8% tolerance around the optimum, the inconsistency of self-chosen optima in Díaz et al., and the drop in validation rate when candidates are close (67% for the continuous version of Taddei et al., with one participant) all suggest that the preference landscape is flat near the optimum, so a fast coarse search fits these problems better than a precise one (inference). The perception data agree: if a walking person cannot feel a stiffness change smaller than about 42% (Maberry and Martin, 2026), comparisons between closer candidates carry almost no information, which gives a principled stopping rule for that parameter (inference; the rules in use, listed in Section 46.6, ignore perception).
No study has compared preferential Bayesian optimization with expert manual tuning on the same patients with preregistered endpoints. The experiment that would settle it is simple: the same participants and time budget, self-tuning against pairwise preference optimization in balanced order, with an objective endpoint and a delayed check of whether the person still prefers the result. Until it is run, a team tuning a device with 4 to 6 parameters should treat self-tuning as the baseline to beat (inference). Section 34.4 carries this into the book's recommendations, and Chapter 24 lets you take both roles, tuning a simulated walker by hand and letting an optimizer tune it.
Sources cited in Section 33.3 5
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Liu et al. (2026d) Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
33.4 Hearing aids and audio #
Hearing aids are one of the first areas in which preference learning reached a commercial product. A hearing aid amplifies frequency bands differently, and the gain in each band and the compression ratio, how strongly loud sounds are turned down relative to soft ones, are normally set by an audiologist from a prescription, a formula such as NAL-NL2 that maps a person's audiogram to settings. In 2015 Nielsen et al. (2015) already used a Gaussian process with active learning for perception-based personalization, and Baltzell et al. (2018), at the research center of the maker Starkey, found with preference-based Bayesian optimization that listeners differed in the compression ratios they preferred and tended to prefer the learned ratio over linear gain and over the NAL-NL2 prescription.
SoundSense Learn. Søgaard Jensen et al. (2019), whose authors are from the maker Widex and from FORCE Technology, evaluated Widex SoundSense Learn, which models the user's utility over the gain in 3 frequency bands with a Gaussian process prior and learns from paired comparisons in which the user marks a degree of preference on a slider (20 comparisons in the study). Twenty hearing-impaired participants each tuned 12 sound scenes, four for each of three attributes (basic audio quality, listening comfort, speech clarity), with the hearing aids on an acoustic manikin, and recordings of the tuned settings were rated double-blind against two settings from Widex's own prescription. Basic audio quality improved generally; listening comfort improved only in traffic noise, not in babble from several talkers; and speech clarity did not improve significantly. The size of the gain adjustments varied widely between people and did not predict who would benefit. Balling et al. (2021), also from the manufacturer, describe a Bayesian optimization mechanism that runs continuously on user input and summarize laboratory and field results.
Presets and dueling bandits. For over-the-counter hearing aids, Vyas et al. (2022) discretized a 24-dimensional configuration space into 15 presets and treated the choice as a dueling bandit problem, a bandit that learns from pairwise wins and losses (Chapter 21). Thirty-five people with mild-to-moderate hearing loss each compared every pair of presets four times; the algorithms were then compared in simulations that drew each answer from a person's recorded preferences, where the authors' algorithm found the best preset in a median of 25 comparisons, half as many as the best baseline.
Other work. Ignatenko et al. (2025) proposed an information-theoretic query design, validated in simulation and on one "real-life example" whose domain the abstract does not name (we could not read the full text), and Tasnim et al. (2024) review machine learning for hearing-aid personalization. The closest cochlear-implant study, Gilbert et al. (2022), had 14 MED-EL users rate music under fixed settings, not adaptive optimization. The HearClip trial of Bayesian preference elicitation for hearing aids was registered in Amsterdam in May 2008, and its record, last updated on 1 September 2025, still lists recruitment as "pending" with no results (Amsterdam UMC, 2008). A preprint review of preference learning in audio cites one music-generation study with 60% annotator agreement, against 75% for text summarization; these are that study's figures, not a pooled estimate (Broukhim et al., 2025).
The hearing-aid evidence is the strongest in this chapter by design, with a double-blind comparison against prescribed settings, and it shows a recurring pattern: preference tuning improves what the person judges directly (overall sound quality) and does not reliably improve the outcome that motivated the device, understanding speech, which was measured only as rated clarity and not with an intelligibility test (inference). The evaluations of the commercial system were run or co-authored by the manufacturer, and the compression-ratio study came from a manufacturer's research center; none of the controlled evaluations we found is independent of a manufacturer (inference, from the authors' affiliations). An independent evaluation with a speech intelligibility test is the study this area needs (inference).
Sources cited in Section 33.4 10
- Nielsen et al. (2015) Perception-based Personalization of Hearing Aids using Gaussian Processes and Active Learning
- Baltzell et al. (2018) Efficient characterization of individual differences in compression ratio preference
- Søgaard Jensen et al. (2019) Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference
- Balling et al. (2021) The Collaboration between Hearing Aid Users and Artificial Intelligence to Optimize Sound
- Vyas et al. (2022) Personalizing over-the-counter hearing aids using pairwise comparisons
- Ignatenko et al. (2025) On preference learning based on sequential Bayesian optimization with pairwise comparison
- Tasnim et al. (2024) A Review of Machine Learning Approaches for the Personalization of Amplification in Hearing Aids
- Gilbert et al. (2022) Cochlear Implant Compression Optimization for Musical Sound Quality in MED-EL Users
- Amsterdam UMC (2008) Personalization of Hearing Aids through Bayesian Preference Elicitation
- Broukhim et al. (2025) Preference-Based Learning in Audio Applications: A Systematic Analysis
33.5 Neuroprostheses and neurostimulation #
A neuroprosthesis stimulates the nervous system to replace a lost function. In a retinal implant, electrodes on the retina produce spots of light called phosphenes, and an encoder turns a camera image into stimulation patterns whose quality depends on parameters that differ between patients. In spinal cord stimulation, an electrode array over the spinal cord can restore some standing or grasping after injury, depending on which electrodes are active and at what frequency and amplitude.
Visual prostheses. Fauvel and Chalk (2022) tuned the encoder of simulated prosthetic vision for sighted viewers with PBO and report significant, robust improvements in perceived image quality that transferred to other tasks; the bioRxiv preprint reports 24 participants and runs of 60 duels, and we could not read the journal version to confirm them. Granley et al. (2023) tuned 13 patient-specific parameters of a deep encoder with duels on simulated patients only, reaching high-quality percepts after about 20 duels and close to an ideal encoder after about 75, robust to simulated response noise and to thresholds misspecified by up to 300%, where it converged to slightly worse encodings.
The follow-up of Schoinas et al. (2025), at the IEEE Engineering in Medicine and Biology Society conference (EMBC) in 2025, ran 60 duels per condition with 17 sighted undergraduates. Only about 50% of the human choices agreed with the simulated agent's, and the final loss in the main condition was 0.27 for humans against 0.07 in the earlier simulations; yet 16 of 17 participants preferred the optimized result in the main condition, all 17 with misspecified thresholds, and 14 in the out-of-distribution condition. All three studies used sighted people or simulated patients.
Spinal cord stimulation. CorrDuel, a correlated dueling bandit, chose electrode configurations by having clinicians compare a patient's standing under two configurations, in a live trial with two patients and 414 comparisons (Sui et al., 2017a). Clinical knowledge first cut the 16-channel array's roughly configurations to the order of to , and the comparison with the doctors' own choices was qualitative: their selections were a subset of what the algorithm found. StageOpt, stage-wise safe Bayesian optimization, explores only settings it is confident are safe before optimizing within them (Section 14.4), with preference feedback in its clinical part (Sui et al., 2018b). Over 10 weeks it ran 564 grasping therapy trials with one patient with tetraplegia; it never sampled an unsafe pattern, and after around 400 iterations a Gaussian process fit to its chosen patterns exceeded the physicians' best choice. The authors note that it assumes the patient's response does not change over time. Zhao et al. (2021) built a Bayesian preference model for the 5 participants of the E-STAND trial, with accuracy averaged over participants of 71.5% in cross-validation and 65.6% in prospective validation, both significantly above chance. A later paper with overlapping authors adds meta-learning; its abstract reports validation on synthetic data and mentions no patients, and we could not read the full text (Farooqi et al., 2025).
Two lessons stand out (inference). The gap between simulated and real people is large even for a simple task, so results on simulated patients cannot stand in for clinical evidence (Section 31.7). And the clinical successes, StageOpt and CorrDuel, are measured against physicians' choices only qualitatively or on a single patient.
Sources cited in Section 33.5 7
- Fauvel and Chalk (2022) Human-in-the-loop optimization of visual prosthetic stimulation
- Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Sui et al. (2017a) Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces
- Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
- Zhao et al. (2021) Optimization of Spinal Cord Stimulation Using Bayesian Preference Learning and Its Validation
- Farooqi et al. (2025) An augmented preference-based Bayesian approach for optimizing neuromodulation stimulation parameters using meta learning
33.6 Robot control and human-robot interaction #
Robots are tuned by people, too: an engineer adjusts gains and constraints until the robot moves "well", a judgment easier to make than to write down as a cost function. Section 15.4 covers Bayesian optimization in robotics with measured objectives; this section covers the cases where the objective is a person's judgment.
Legged robots. Tucker et al. (2021) used LineCoSpar to tune the constraints of a gait optimization based on hybrid zero dynamics (a standard method for stable biped gaits) on the AMBER-3M robot, producing robust walking in 30 iterations with rigid feet and 50 with springs the gait model ignored. Csomay-Shanklin et al. (2022) tuned controller gains on AMBER, where none of the initially sampled gains could walk and after 50 iterations 3 sets were rated very good, and on Cassie, with a domain expert and with a naive user. Cosner et al. (2022) tuned a safety filter on a Unitree A1 quadruped from pairwise preferences and safety labels, in 30 iterations in simulation and 7 on hardware indoors and 3 outdoors. In these studies the judge was one operator at a time. The same group released POLAR, a MATLAB toolbox (Tucker et al., 2022), a preprint.
Controller calibration, two schools. The GLISp family uses radial basis function surrogates: GLISp (Bemporad and Piga, 2021), GLISp-r with a convergence guarantee (Previtali et al., 2023), a robotic sealing planner whose abstract reports better deposition quality within 20 trials than programming by demonstration and manual tuning (Roveda et al., 2021) (we could read only the abstract), and collaborative spray painting validated with 15 PhD or master's students (Cella et al., 2026). The Gaussian process school includes a simulated proportional-integral controller tuned from real users' pairwise feedback, which came closer to the response users wanted than multi-objective alternatives (Coutinho et al., 2024), with a follow-up in Control Engineering Practice (Coutinho et al., 2026); coactive multi-objective optimization for personalized plasma medicine (Shao et al., 2024); CrashPBO, which lets the user report a crash as worse than every successful experiment, on a backflipping quadcopter (three lab members as judges, 15 trials each), a Furuta pendulum, and a unicycle robot (Menn et al., 2026b), the earliest preference tuning of a quadrotor we found; and a reaching demonstration on a Franka Panda arm at the 2024 International Conference on Ubiquitous Robots, which reports no user study (Feith and Rueckert, 2024).
An expert's cost function against the expert's choices. A 2025 preprint by De Witte et al. (2025) gives the most important negative finding in this area. A robot arm pushed a block to a target, and 4 time-scale parameters of a cascade of controllers were tuned; three preferential methods were run on hardware, where an expert chose between two pushes (2 random and 13 algorithm-driven comparisons, 30 experiments per trial). PBO tuned the system to the expert's satisfaction, but the cost function the expert designed to describe their own priorities was not in full agreement with their choices, even after its weights were refitted to the pushes the expert had rated highest. The expert rarely disagreed with it among the 10% of experiments with the lowest cost; beyond that, agreements and disagreements on its goal-reaching term were roughly balanced. Across iterations the expert's choices followed the total cost relatively consistently, but no repeatable pattern of shifting priorities could be modeled. The Gaussian process utility learned from the duels was more consistent with the expert's decisions than the expert's own cost function. The experiments refer to one expert; the paper gives no count.
Active preference-based reward learning. Robot learning offers better controlled user studies, mostly with linear rewards. Sadigh et al. (2017) introduced active preference-based reward learning from a person's choices between pairs of trajectories (Section 36.1), and Bıyık and Sadigh (2018) extended it to batch queries. DemPref combined demonstrations with preferences and was rated more successful than inverse reinforcement learning (learning a reward from demonstrations alone) by 15 users of a Fetch robot (Palan et al., 2019), with more user studies in its journal version (Bıyık et al., 2022); "Asking Easy Questions" accounts for how hard a question is for the person (Bıyık et al., 2019). With a Gaussian process reward, 10 users taught a Fetch robot a variant of minigolf and rated the learned behavior, on a 9-point scale, 6.9 ± 0.6 (mean ± standard error) for the active Gaussian process, 3.4 ± 0.7 for the active linear model, and 5.1 ± 0.7 for a Gaussian process with random queries () (Bıyık et al., 2020). Myers et al. (2021) learned rewards from rankings, the APReL library collects these algorithms (Bıyık et al., 2021), and Kwon et al. (2022) plated Japanese food from comparisons of rendered dishes, with simulated users and 4 real users.
What it means. Control papers present PBO as an alternative to trial-and-error tuning, but their baselines are a scoring function the engineers could not write (as in Section 34.1) or an expert-written cost function, not measured expert tuning time and quality, so the claimed saving in tuning effort is not yet quantified (inference), with the sealing study as a possible exception. De Witte et al. also suggest that synthetic decision makers defined by a hidden cost function may overstate how well a method works on real people (inference).
Sources cited in Section 33.6 23
- Tucker et al. (2021) Preference-Based Learning for User-Guided HZD Gait Generation on Bipedal Walking Robots
- Csomay-Shanklin et al. (2022) Learning Controller Gains on Bipedal Walking Robots via User Preferences
- Cosner et al. (2022) Safety-Aware Preference-Based Learning for Safety-Critical Control
- Tucker et al. (2022) POLAR: Preference Optimization and Learning Algorithms for Robotics
- Bemporad and Piga (2021) Global optimization based on active preference learning with radial basis functions
- Previtali et al. (2023) GLISp-r: a preference-based optimization algorithm with convergence guarantees
- Roveda et al. (2021) Pairwise Preferences-Based Optimization of a Path-Based Velocity Planner in Robotic Sealing Tasks
- Cella et al. (2026) Adaptive Human-Robot Collaborative Painting Combining Preference-Based Optimization and Dynamic Motion Primitives
- Coutinho et al. (2024) Human-in-the-loop controller tuning using Preferential Bayesian Optimization
- Coutinho et al. (2026) Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization
- Shao et al. (2024) Coactive Preference-Guided Multi-Objective Bayesian Optimization: An Application to Policy Learning in Personalized Plasma Medicine
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
- Feith and Rueckert (2024) Integrating Human Expertise in Continuous Spaces: A Novel Interactive Bayesian Optimization Framework with Preference Expected Improvement
- De Witte et al. (2025) How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation
- Sadigh et al. (2017) Active Preference-Based Learning of Reward Functions
- Bıyık and Sadigh (2018) Batch Active Preference-Based Learning of Reward Functions
- Palan et al. (2019) Learning Reward Functions by Integrating Human Demonstrations and Preferences
- Bıyık et al. (2022) Learning Reward Functions from Diverse Sources of Human Feedback: Optimally Integrating Demonstrations and Preferences
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Bıyık et al. (2020) Active Preference-Based Gaussian Process Regression for Reward Learning
- Myers et al. (2021) Learning Multimodal Rewards from Rankings
- Bıyık et al. (2021) APReL: A Library for Active Preference-based Reward Learning Algorithms
- Kwon et al. (2022) Physically Consistent Preferential Bayesian Optimization for Food Arrangement
33.7 Settled, contested, missing #
Settled. Preference-based tuning of exoskeletons works in the sense its studies test: people's later choices agree with the learned model at about 65% to 100% when averaged over a study's participants, within sessions of 12 to 50 comparisons. Self-tuning by the wearer is repeatable within a person but varies widely between people (Ingraham et al., 2022), and a thumb-controlled self-tuning of 4 hip parameters took about 11 minutes and reduced metabolic cost by 16.6% against zero torque (Schäfer et al., 2026). Hearing-aid preference tuning improved overall sound quality without improving speech clarity (Søgaard Jensen et al., 2019). Simulated users can differ sharply from real ones (Schoinas et al., 2025).
Contested. Whether algorithm tuning beats tuning by hand for problems with 4 to 6 parameters; the two 2026 results are of the same order, on different devices. How much the preferred setting is the person's and how much it is a product of adaptation and of the optimizer's own dynamics: preferences rose across blocks in Arens et al., naive users drifted toward more torque in Ingraham et al., and customization contributed only a quarter of the benefit in Poggensee and Collins. Whether an expert's judgment can be written down at all, on the evidence of one expert (De Witte et al., 2025).
Missing. A preregistered comparison of preferential optimization with expert or self-tuning on the same patients and time budget. Clinical populations beyond 2 people with paraplegia, a few with spinal cord injury, and a few amputees. Retinal implant studies with blind users. Any study of deep brain stimulation, cochlear implants, or functional electrical stimulation in rehabilitation, and any new peer-reviewed user study of preferential Bayesian optimization for hearing aids or cochlear implants from 2025 to September 2026. An independent evaluation of hearing-aid preference tuning. Measured expert tuning time and quality as the baseline for controller calibration; one abstract reports a comparison with manual tuning, and we could not read its measurements (Roveda et al., 2021). Preference tuning of quadrotors and quadrupeds judged by more than a few people: we found one quadruped study with a single user (Cosner et al., 2022) and one quadcopter study with three lab members (Menn et al., 2026b) (inference from the searches reported above).
Sources cited in Section 33.7 8
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Søgaard Jensen et al. (2019) Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- De Witte et al. (2025) How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation
- Roveda et al. (2021) Pairwise Preferences-Based Optimization of a Path-Based Velocity Planner in Robotic Sealing Tasks
- Cosner et al. (2022) Safety-Aware Preference-Based Learning for Safety-Critical Control
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
33.8 Exercises #
Liu et al. report metabolic reductions of 14.5% to 15.4% against wearing the exoskeleton with assistance off, and 8.6% to 13.0% against walking without it. Suppose wearing the exoskeleton without assistance costs 5% more energy than walking without it. Starting from a 15% reduction against assistance off, what reduction would you expect against no exoskeleton? What does this say about comparing Schäfer et al.'s 16.6% (against zero torque) with Ding et al.'s 17.4% (against no device), and with Liu et al.'s 14.5% to 15.4%?
Solution
Let be the energy of walking without the device. With the device worn but not assisting, . A 15% reduction against that gives , an 11% reduction against no device: carrying a device that does not assist costs energy, so the same assistance looks better against "assistance off" than against "no device".
Schäfer et al.'s 16.6% and Liu et al.'s 14.5% to 15.4% are both measured against the device worn but not assisting, so they can be set side by side (devices, parameters, and speeds still differ). Ding et al.'s 17.4% is measured against no device; restated against the device worn without assistance it would be larger, so the optimizer-tuned exosuit's advantage over the self-tuned hip exoskeleton is somewhat bigger than the raw numbers suggest, by an amount that depends on each device's penalty when worn without assistance. The 5% is an assumption for the exercise.
LineCoSpar's six participants agreed with the model on 4 validation preferences each, at 75%, 100%, 100%, 25%, 100%, and 100%. If a participant answered at random, what is the probability of agreeing on all 4? Of agreeing on at least 3? What does this say about how much a single participant's 100% tells you?
Solution
At random each answer agrees with probability 1/2, so all four agree with probability , and at least three agree with probability . A single participant's 100% on 4 checks is weak evidence on its own, roughly a one in sixteen event under pure guessing. Across six participants, four perfect scores are much less likely by chance, so the study as a whole is informative, but per-person validation with 4 checks cannot distinguish a person whose preferences the model captured from one who was lucky. This is one reason validation with a few internal checks is the weakest part of the evidence in Table 33.1.
Arens et al. measured just-noticeable differences of 13.1 N for lifting and 9.8 N for lowering assistance. Suppose the search range for lifting assistance is 0 to 80 N. If comparisons between settings closer than one JND carry no information, roughly how many distinguishable levels does the lifting range have? What does that suggest about how many pairwise comparisons the optimization of this one parameter can usefully use?
Solution
Roughly distinguishable levels. With about six levels, a search can locate the preferred level by something like bisection in about informative comparisons once noise is small, or a few times that with realistic noise; comparisons between candidates less than one JND apart add little. The perception threshold, not the optimizer, sets the useful resolution, which is why Arens et al. built the JND into the acquisition function and why the chapter suggests it as a stopping rule. The range of 0 to 80 N is an assumption for the exercise.
Further reading #
- Schäfer et al. (2026) is the clearest test of whether an algorithm is needed at all, and its discussion compares durations with metabolic optimization.
- Tucker et al. (2020a) and Li et al. (2021) show how the query was adapted to walking: one-dimensional lines for scale, ordinal labels and a region of interest for comfort.
- Søgaard Jensen et al. (2019) is the best-controlled evaluation in the chapter, a double-blind comparison against the manufacturer's own prescribed settings, with both positive and null outcomes.
- Schoinas et al. (2025) measures directly how far simulated users are from real ones.
- De Witte et al. (2025) shows, with one expert, why a person's choices may be easier to learn than their own description of what they want.
- Ingraham et al. (2022) is the reference point for how stable a wearer's own preference is.
References
- (2023). A Novel Framework to Facilitate User Preferred Tuning for a Robotic Knee Prosthesis. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §33.2
- (2008). Personalization of Hearing Aids through Bayesian Preference Elicitation. Trial registry, onderzoekmetmensen.nl. non-peer-reviewed Cited in §33.4
- (2025). Preference-based assistance optimization for lifting and lowering with a soft back exosuit. Science Advances. doi:10.1126/sciadv.adu2099. Cited in §33.1
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §33.1
- (2026). Context-Continuous Preference Learning for Exoskeleton Personalization. arXiv. preprint Cited in §33.1
- (2021). The Collaboration between Hearing Aid Users and Artificial Intelligence to Optimize Sound. Seminars in Hearing. doi:10.1055/s-0041-1735135. Cited in §33.4
- (2018). Efficient characterization of individual differences in compression ratio preference. The Journal of the Acoustical Society of America. doi:10.1121/1.5067390. Cited in §33.4
- (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Cited in §33.1 §33.6
- (2018). Batch Active Preference-Based Learning of Reward Functions. CoRL 2018. Cited in §33.6
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §33.6
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Cited in §33.6
- (2021). APReL: A Library for Active Preference-based Reward Learning Algorithms. arXiv. software Cited in §33.6
- (2022). Learning Reward Functions from Diverse Sources of Human Feedback: Optimally Integrating Demonstrations and Preferences. IJRR. Cited in §33.6
- (2025). Preference-Based Learning in Audio Applications: A Systematic Analysis. arXiv. preprint Cited in §33.4
- (2026). Adaptive Human-Robot Collaborative Painting Combining Preference-Based Optimization and Dynamic Motion Primitives. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3683596. Cited in §33.6
- (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. Cited in §33.6 §33.7
- (2024). Human-in-the-loop controller tuning using Preferential Bayesian Optimization. IFAC-PapersOnLine. doi:10.1016/j.ifacol.2024.08.306. Cited in §33.6
- (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. Cited in §33.6
- (2022). Learning Controller Gains on Bipedal Walking Robots via User Preferences. ICRA 2022. Cited in §33.6
- (2025). How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation. arXiv. preprint Cited in §33.6 §33.7
- (2026). User preference in the personalized control of an ankle prosthesis: a case study. Journal of NeuroEngineering and Rehabilitation. doi:10.1186/s12984-026-01931-w. Cited in §33.2
- (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. Cited in §33.1 §33.3
- (2025). An augmented preference-based Bayesian approach for optimizing neuromodulation stimulation parameters using meta learning. Journal of Neural Engineering. Cited in §33.5
- (2022). Human-in-the-loop optimization of visual prosthetic stimulation. Journal of Neural Engineering. Cited in §33.5
- (2024). Integrating Human Expertise in Continuous Spaces: A Novel Interactive Bayesian Optimization Framework with Preference Expected Improvement. 2024 21st International Conference on Ubiquitous Robots (UR). doi:10.1109/ur61395.2024.10597501. Cited in §33.6
- (2024). A Personalizable Controller for the Walking Assistive omNi-Directional Exo-Robot (WANDER). ICRA 2024. Cited in §33.1
- (2022). Cochlear Implant Compression Optimization for Musical Sound Quality in MED-EL Users. Ear & Hearing. Cited in §33.4
- (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Cited in §33.5
- (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. Cited in §33.4
- (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §33.3 §33.7
- (2023). Leveraging user preference in the design and evaluation of lower-limb exoskeletons and prostheses. Current Opinion in Biomedical Engineering. Cited in §33.1
- (2026). Multi-Objective Human-in-the-Loop Bayesian Optimization of a Lower-Limb Exoskeleton. arXiv. preprint Cited in §33.1
- (2025). Validation of Dynamic Bayesian Optimization for a Non-Stationary Human-in-the-Loop Optimization Problem. bioRxiv. preprint Cited in §33.1
- (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. Cited in §33.6
- (2023). User preference optimization for control of ankle exoskeletons using sample efficient active learning. Science Robotics. Cited in §33.1
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §33.1
- (2026d). Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization. arXiv. preprint Cited in §33.1 §33.3
- (2026). Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §33.1 §33.3
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §33.6 §33.7
- (2021). Learning Multimodal Rewards from Rankings. CoRL 2021. Cited in §33.6
- (2015). Perception-based Personalization of Hearing Aids using Gaussian Processes and Active Learning. IEEE/ACM Transactions on Audio, Speech, and Language Processing. Cited in §33.4
- (2019). Learning Reward Functions by Integrating Human Demonstrations and Preferences. RSS 2019. Cited in §33.6
- (2026). Simultaneous Forward and Inverse Human-in-the-Loop Optimization. CoRL 2026. Cited in §33.1
- (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §33.1
- (2023). GLISp-r: a preference-based optimization algorithm with convergence guarantees. Computational Optimization and Applications. Cited in §33.6
- (2025). Rapid Online Learning of Hip Exoskeleton Assistance Preferences. ICRA 2025. Cited in §33.1
- (2021). Pairwise Preferences-Based Optimization of a Path-Based Velocity Planner in Robotic Sealing Tasks. IEEE Robotics and Automation Letters. Cited in §33.6 §33.7
- (2017). Active Preference-Based Learning of Reward Functions. Robotics: Science and Systems XIII. Cited in §33.6
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §33.1 §33.3 §33.7
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §33.5 §33.7
- (2024). Coactive Preference-Guided Multi-Objective Bayesian Optimization: An Application to Policy Learning in Personalized Plasma Medicine. IEEE Control Systems Letters. Cited in §33.6
- (2022). Personalizing exoskeleton assistance while walking in the real world. Nature. Cited in §33.1
- (2024). On human-in-the-loop optimization of human–robot interaction. Nature. Cited in §33.1
- (2019). Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference. Trends in Hearing. doi:10.1177/2331216519847413. Cited in §33.4 §33.7
- (2017a). Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces. IJCAI 2017. Cited in §33.5
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §33.5
- (2024). An Individual Prosthesis Control Method with Human Subjective Choices. Biomimetics. doi:10.3390/biomimetics9020077. Cited in §33.2
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Cited in §33.2
- (2024). A Review of Machine Learning Approaches for the Personalization of Amplification in Hearing Aids. Sensors. doi:10.3390/s24051546. Cited in §33.4
- (2023). Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning. California Institute of Technology. doi:10.7907/j9hk-xa17. thesis Cited in §33.1
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §33.1
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §33.1
- (2021). Preference-Based Learning for User-Guided HZD Gait Generation on Bipedal Walking Robots. ICRA 2021. Cited in §33.6
- (2022). POLAR: Preference Optimization and Learning Algorithms for Robotics. arXiv. preprint Cited in §33.6
- (2022). Personalizing over-the-counter hearing aids using pairwise comparisons. Smart Health. doi:10.1016/j.smhl.2021.100231. Cited in §33.4
- (2017). Human-in-the-loop optimization of exoskeleton assistance during walking. Science. Cited in §33.1
- (2021). Optimization of Spinal Cord Stimulation Using Bayesian Preference Learning and Its Validation. IEEE Transactions on Neural Systems and Rehabilitation Engineering. doi:10.1109/tnsre.2021.3113636. Cited in §33.5