Tuning an Exoskeleton with a Person in the Loop
In Chapter 23 people chose the search space and then watched an optimizer work on measured yields. Here a person is part of the measurement. An ankle exoskeleton pushes the foot down at the end of each step, and if the push has the right size at the right moment, walking costs the wearer less energy. The right size and moment differ from person to person, so the device is tuned for each wearer, with the wearer walking in it and an optimizer choosing what to try next. This is optimization with a person inside the loop, and it is one of the best documented cases of it: the studies that established the approach date from 2017, and Chapter 33 reviews what they and their successors found.
This chapter takes the tuner's seat. It works one problem end to end: what is being optimized and what one evaluation costs, what the published studies did, a tuning session that you run yourself, and the two complications that make the case instructive, a person who adapts while being measured and the discovery that people tune such devices quite well by themselves.
One thing must be said plainly at the start. A reader cannot wear an exoskeleton through a web page, and no published study released a person's full metabolic landscape. So every interactive figure in this chapter runs on a simulated walker, a model whose numbers were set to match published measurements. The text says which values are measured and where they come from, and which are assumptions. What the simulation shows is what the published numbers imply when they are put together; it is not new evidence.
24.1 The problem #
24.1.1 The device and its parameters #
An exoskeleton is a wearable robot that applies a torque, a turning force, at a joint. The controller repeats a torque profile once per gait cycle, the interval from one heel strike to the next heel strike of the same foot. In the line of ankle exoskeleton experiments that began with Zhang et al. (2017), the profile is defined by four numbers: the peak torque, the time in the gait cycle at which the peak occurs, and the times the torque takes to rise and to fall (Slade et al., 2022). Each is confined to a safe range. The peak torque is normalized to body mass and allowed between 0 and 1 N·m per kilogram (Slade et al., 2022); in the experiments of Poggensee and Collins (2021) the peak, rise, and fall times were kept within 35% to 55%, 10% to 40%, and 5% to 20% of the gait cycle, as Kutulakos and Slade (2024) report in the preprint that reuses those data.
To keep the whole problem visible on a map, this chapter tunes two of the four: the peak torque, between 0 and 1 N·m per kilogram, and the peak time, between 35% and 55% of the gait cycle. A setting is a point in the unit square, with the peak torque and the peak time rescaled to . Two parameters is not a toy size for this field: Ding et al. (2018) tuned exactly two, the peak and offset timing of hip assistance.
24.1.2 What is optimized #
The classical objective is metabolic cost, the rate at which the body uses energy while walking. It is estimated by indirect calorimetry: the wearer breathes through a mask that measures the oxygen consumed and the carbon dioxide produced. A good setting lowers the metabolic rate below that of walking in the same device with the motors off. Using this measurement to drive an optimizer is called human-in-the-loop optimization.
The other objective is what the wearer prefers. It needs no mask, it includes comfort and the feeling of stability that a metabolic number leaves out, and it can be read much faster. It is also a different quantity: around the settings that participants chose by feel for a hip exoskeleton, the metabolic rate showed no clear minimum, and the authors assume that the participants were weighing more than effort, comfort for one (Schäfer et al., 2026). Both objectives appear below, the first as a measurement with noise, the second as a comparison the wearer feels.
24.1.3 What one evaluation costs #
The metabolic rate cannot be read off instantly. After a change of setting, the rate measured at the mouth moves toward its new level gradually, because the body's oxygen stores and transport delay the response. Selinger and Donelan (2014) showed that the response during walking is well described as a first-order system, one that closes a fixed fraction of the remaining gap per unit time, with a time constant of 42 ± 12 seconds (mean ± standard deviation across subjects), and that individual breaths scatter widely around it. Waiting for the rate to settle takes several minutes per setting. The laboratory protocol that followed Zhang et al. (2017) instead has the wearer walk each setting for two minutes while the breaths are recorded, and estimates from them the steady level the response is heading for; the two minutes are a compromise between the time per setting and the accuracy of the estimate (Slade et al., 2022).
The estimate is noisy. Kutulakos and Slade (2024) put its standard deviation at 4.6% of the metabolic rate, for a first-order fit to two minutes of data, and cite the experiments of Zhang et al. (2017) for it; the supplement of that study gives an average error of 4% for two-minute estimates. The figure below simulates one such bout.
Some things to try. At two minutes the estimate has a standard deviation of 4.6 points around the true value, by construction; the figure prints this value as "theory" next to the spread of the 60 bouts it happened to draw, which is 4.4 points. Shorten the bout to one minute and the standard deviation more than doubles, to 10.6 points, because the response has barely left its starting level and the fit must extrapolate. Lengthen it to six minutes and it falls to 1.9 points, at three times the cost. Then set the true change to −3% and look at the strip at two minutes: an improvement of that size is invisible in a single bout.
That last observation shapes the whole problem. The benefit of tuning for the individual, over a good generic setting, is a matter of a few points: in Poggensee and Collins (2021) a generic controller reduced the metabolic rate of trained users by 31%, and customized assistance by 39%. Near the optimum, neighboring settings differ by less than the noise of one estimate, so an optimizer has to average, through its model, over many evaluations (Exercise 24.1). At two minutes each, a one-hour session buys thirty. The metabolic optimization that Slade et al. (2022) ran as their laboratory reference took 128 minutes of walking.
24.1.4 The person changes #
The last ingredient is that the walker is not a fixed function. People learn to use an exoskeleton. In Poggensee and Collins (2021), naive users needed about 109 minutes of assisted walking to become expert; training contributed about half of the final 39% reduction and customization about one quarter; the generic controller that gave 31% after training gave only 10% before it; and the best peak torque kept growing slowly over the whole study, which the authors read as adaptation on a longer time scale. A tuning session of twenty minutes is therefore run on a person who will not exist an hour later. Section 24.4 returns to what that does to an optimizer.
Sources cited in Section 24.1 7
- Zhang et al. (2017) Human-in-the-loop optimization of exoskeleton assistance during walking
- Slade et al. (2022) Personalizing exoskeleton assistance while walking in the real world
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Selinger and Donelan (2014) Estimating instantaneous energetic cost during non-steady-state gait
24.2 What the studies did #
Three ways of tuning have been tried on people, and they differ in who judges a setting. Table 24.1 lists one or two studies of each kind with the numbers this chapter uses; Section 33.1.4 has the full set, with participants and validation, and Section 33.3 discusses how far the results can be compared.
| Who judges | Study | Device and parameters | Tuner | Time and result |
|---|---|---|---|---|
| metabolic estimate | Zhang et al. (2017) | one ankle, 4 parameters, 11 participants | evolution strategy (CMA-ES) | 64 min of walking for 9 of the 11; 24.2 ± 7.4% below zero torque |
| metabolic estimate | Ding et al. (2018) | hip exosuit, 2 timing parameters, 8 participants | Bayesian optimization | 40 min of optimization, converged after 21.4 ± 1.0 min; 17.4 ± 3.2% below walking without the device |
| metabolic estimate | Poggensee and Collins (2021) | both ankles, 4 parameters, naive users in three training groups | CMA-ES, as Kutulakos and Slade (2024) describe it, for the group with customized assistance | 39% below the device turned off, after about 109 min of training |
| wearable sensors | Slade et al. (2022) | ankles, 4 parameters | model that ranks settings from ankle motion, 30 s per setting | 32 min in the laboratory, where metabolic estimates took 128 min |
| the wearer's comparisons | Tucker et al. (2020a) | walking gaits, 6 parameters, 6 participants | preference learning along random lines | 30 gait trials and 6 validation trials each |
| the wearer's comparisons | Lee et al. (2023) | ankle, 4 parameters | evolutionary algorithm with a learned ranker | settings stable after 43 ± 7 comparisons |
| the wearer, by hand | Ingraham et al. (2022) | ankle, torque size and timing, 24 participants | self-tuning, blind to the values | converged in 105 seconds per trial |
| the wearer, by hand | Schäfer et al. (2026) | hip, 4 timing parameters, 11 participants | self-tuning with a thumbstick | 10.9 ± 0.9 min, 30.5 settings; 16.6 ± 1.1% below zero torque |
Two optimizers recur in the first rows. Bayesian optimization is the loop of Chapter 11: a Gaussian process models the metabolic rate over the settings, and an acquisition function picks the next setting to walk. Ding et al. (2018) chose it because it is suited to noisy signals and very few evaluations. The other is the covariance matrix adaptation evolution strategy, CMA-ES (Hansen and Ostermeier, 2001), introduced in Section 15.8.1. It keeps no model of the landscape. It draws a small generation of settings from a Gaussian distribution, has the wearer walk each one, moves the distribution's mean toward the better half, reshapes its covariance along the directions that helped, and repeats. Its estimate of the best setting is the mean. It forgets every generation after using it, which makes it wasteful when evaluations are scarce and, as we will see, forgiving when the person changes.
The preference studies replace the mask with the wearer's judgment. CoSpar and LineCoSpar asked wearers which of two gaits they preferred and let them suggest improvements (Tucker et al., 2020b; Tucker et al., 2020a). The authors of the first note that users had difficulty remembering more than two trials, a reason to compare each trial only with the one just before it. A 2026 preprint tuned six parameters of a hip exoskeleton from pairwise comparisons in sessions of 20.6 ± 4.6 minutes, validation included, for five participants (Liu et al., 2026d).
The self-tuning studies remove the optimizer too. The wearer holds a control, changes a parameter, feels the result, and stops when satisfied. In Schäfer et al. (2026) the participants spent 18.7 seconds per setting on average and changed only one parameter at a time in 97.5% of their adjustments.
Sources cited in Section 24.2 12
- Zhang et al. (2017) Human-in-the-loop optimization of exoskeleton assistance during walking
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
- Slade et al. (2022) Personalizing exoskeleton assistance while walking in the real world
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Lee et al. (2023) User preference optimization for control of ankle exoskeletons using sample efficient active learning
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Hansen and Ostermeier (2001) Completely Derandomized Self-Adaptation in Evolution Strategies
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Liu et al. (2026d) Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization
24.3 A session, re-enacted #
24.3.1 The simulated walker #
The simulated walker has a reduction : the fraction by which setting lowers the metabolic rate below zero-torque walking, after minutes of assisted walking. It is the product of an amplitude, a torque term, and a timing term,
where and are the walker's best torque and best peak time. The torque term is zero at zero torque, rises to one at as , and falls beyond it as , so too much torque can make walking cost more than no torque at all. The timing term is wide: a peak that is 4% of the gait cycle away from the best time costs an adapted walker about 3.7 points, less than the noise of one estimate.
Adaptation moves all three quantities. The walker's level of adaptation is , which reaches 95% at 108 minutes. The amplitude grows from for a novice to for an expert, the best peak time shifts by 2.4% to 4.8% of the gait cycle, and the best torque rises from roughly 0.45 to roughly 0.85 N·m per kilogram, half as fast as the rest. Each walker has their own and , drawn at random. A generic setting, the same for everyone, sits at a torque of 0.6 and a peak time of 45%.
Table 24.2 lists what the numbers were matched to.
| Quantity | Published value | In the simulation |
|---|---|---|
| reduction with a generic setting, before training | 10% (Poggensee and Collins, 2021) | 10.0% on average over all walkers the figure can draw, at |
| reduction with a generic setting, after training | 31% (Poggensee and Collins, 2021) | 31.2% on average over all adapted walkers the figure can draw (for the 40 walkers of Section 24.3.3, the mean is 32.1% and the median 32.9%) |
| reduction with customized assistance, after training | 39% (Poggensee and Collins, 2021) | 39% at every adapted walker's optimum |
| time to become an expert user | about 109 minutes; best peak torque grows more slowly (Poggensee and Collins, 2021) | 95% adapted at 108 minutes; best torque adapts half as fast |
| noise of a two-minute metabolic estimate | standard deviation 4.6% (Kutulakos and Slade, 2024) | Gaussian, standard deviation 0.046 |
| ranges of the two parameters | peak torque 0 to 1 N·m per kilogram (Slade et al., 2022); peak time 35% to 55% of the gait cycle (Kutulakos and Slade, 2024) | the same |
| time per self-tuned setting | 18.7 seconds (Schäfer et al., 2026) | the same |
| CMA-ES settings | per generation, initial step 30% of the range (Kutulakos and Slade, 2024) | 6 per generation for , step 0.3 |
| Bayesian optimization settings | Matérn kernel, confidence bound with exploration constant 2.6, or 0.93 for a less exploring variant (Kutulakos and Slade, 2024) | the same |
| a novice's best possible reduction | none found | 13.5% at (set so that the generic setting gives 10%) |
| what the wearer feels | none found | see below |
| time per felt comparison of two settings | none found | 37.4 seconds, two settings at 18.7 seconds |
The felt signal is the least certain part of the model. The wearer's felt effort is the metabolic cost plus a dislike of high torque, , which fades as the walker adapts (with the slower adaptation level of the torque). The dislike reflects a measured tendency, that naive users of an ankle exoskeleton preferred more torque as the experiment went on (Ingraham et al., 2022), but its size is invented. Each reading of the felt effort carries Gaussian noise of standard deviation 0.03, and a difference between two readings smaller than 0.025 is reported as "about the same". Those two numbers are assumptions; a perception floor of this kind is real, and for the stiffness of an ankle exoskeleton during walking the smallest noticeable change has been measured at about 42% (Maberry and Martin, 2026).
24.3.2 Four tuners #
Four tuners work on the walker, each on a fresh copy, for the same minutes of walking.
- You, by hand. You choose a setting, the walker tries it for 18.7 seconds, and you are told whether it feels easier, harder, or about the same as the previous setting. You never see a number.
- Bayesian optimization on two-minute metabolic estimates. A Gaussian process with a Matérn 5/2 kernel models the measured rate; after four spread-out starting settings, the next setting minimizes the lower confidence bound , the minimizing counterpart of the upper confidence bound of Section 12.4; the recommendation is the setting with the lowest posterior mean.
- CMA-ES on the same estimates, six settings per generation, starting at the center of the square. Its recommendation is its current mean.
- Preferential Bayesian optimization (PBO, Chapter 19) on what the walker feels. Each query is a pair of settings chosen by EUBO (Section 19.4); the walker tries both and says which feels easier; a comparison takes 37.4 seconds. The recommendation is the compared setting with the highest posterior mean utility.
In 24 minutes that is 77 settings by hand, 38 comparisons, or 12 metabolic estimates, which for CMA-ES is two generations.
Some things to try. Tune the default walker, who is new to the device, for five to ten minutes of walking time (about fifteen to thirty settings), and keep the setting you trust. Most readers find the same thing the simulated person in the next figure finds: large mistakes are easy to feel and to undo, and near the end almost every change feels "about the same". Then compare your row of the table with the others, and use Show trail of to see where each optimizer spent its evaluations. Switch to Already adapted and tune again: the landscape is larger and steeper, and the feel is more decisive. Finally try the 48-minute session on a novice and tune for all of it; the walker's optimum moves while you work, and the best torque at the end is not the best torque at the start.
For readers without a pointer, or without patience, the button Let a simulated person tune runs a simple self-tuner: it starts at the generic setting, changes one parameter at a time, keeps a change only if it feels easier, and shrinks its step after repeated failures. The figure below shows its session on the default walker.
Read the table in the figure. After 10.9 minutes the walker's best possible reduction is 20.2%, and the four tuners reach 20.0% (by hand), 18.3% (Bayesian optimization), 19.6% (CMA-ES, which has not finished a generation and still recommends its starting point), and 18.9% (preferences), against 18.9% for the generic setting that nobody tuned. The differences between the tuners, and between them and no tuning at all, are less than two points. With noise of 4.6 points per metabolic estimate, a real experiment would need dozens of evaluation bouts per setting to tell them apart (Exercise 24.1).
The last column is the uncomfortable one. Judged on the walker as they will be once adapted, the settings found in this session give 24% to 30%, while the untuned generic setting gives 31.2%. The session tuned the device for a novice who cannot yet use much torque, and the adapted walker wants far more: the faint cross in the map sits at a torque of 0.88, the settings of the four tuners between 0.35 and 0.5.
24.3.3 Across many walkers #
One walker is an anecdote. Running the same four tuners on 40 simulated walkers gives the following medians for walkers who are already adapted, so that the landscape stands still. After 12 minutes, Bayesian optimization reaches a reduction of 38.2%, PBO 37.9%, the simulated self-tuner 37.8%, and CMA-ES 35.9%; the best possible is 39.0% and the generic setting gives 32.9%. The self-tuner and the preference optimizer are already at 37.5% after six minutes, when Bayesian optimization has had three estimates and stands at 32.7%. CMA-ES reaches 37% after about 48 minutes and stays near it.
So on a walker who holds still, everything except the evolution strategy recovers most of the roughly six points between the generic setting and the optimum within a quarter of an hour, and the two tuners that use the wearer's feeling get there first, because they test six settings in the time one metabolic estimate takes. This agrees in kind with what the studies report: self-tuning in about eleven minutes (Schäfer et al., 2026), Bayesian optimization of two parameters in about twenty-one (Ding et al., 2018). It also depends on an assumption the simulation makes and the studies do not guarantee, that what the adapted wearer feels points at the metabolic optimum.
Sources cited in Section 24.3 7
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
- Slade et al. (2022) Personalizing exoskeleton assistance while walking in the real world
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
24.4 When the person adapts #
A Gaussian process posterior treats every observation as a reading of one fixed function (Chapter 8). A walker who is learning the device is a different function every few minutes. Measurements from the first ten minutes describe a person who has since changed, yet they stay in the data with full weight.
The evidence that this matters comes from several directions. The 109 minutes of Poggensee and Collins (2021) are far longer than a tuning session. When people meet a new exoskeleton behavior, the variability of their step frequency, ankle angle, and muscle activity first rises and then falls, on different time scales for different variables (Abram et al., 2022). Preferences move as well: naive users preferred higher torque as the experiment progressed (Ingraham et al., 2022). And in the preprint of Kutulakos and Slade (2024), which simulated a novice by blending one subject's fitted landscape into another's over 80 evaluations, Bayesian optimization, which had converged in about 60 evaluations on a fixed landscape, was slowed by the change, while CMA-ES reached the optimum at a similar rate with and without it.
The figure below repeats that experiment on the walkers of this chapter.
Three things can be read from it.
Adaptation is larger than tuning. With novices, judged minute by minute, all the lines rise together and, after the first ten minutes, stay within about two points of each other. At 24 minutes the best possible reduction is 25.9%, the tuners are at 23.2% to 23.6%, and the generic setting, which nobody tuned, is at 23.3%. At 60 minutes the tuners are at 31.4% to 32.5% and the generic setting at 31.3%. What moves the lines from 15% to 35% is the walker's own learning. This is the simulation's version of the finding it was calibrated to, that training contributed about half of the benefit and customization about a quarter (Poggensee and Collins, 2021).
A setting tuned early is stale later. Switch the judge to The walker once adapted. The settings recommended after 24 minutes would give the adapted walker 26.3% to 31.1%, below the 32.9% of the generic setting: the optimizers have faithfully found low-torque settings for a novice. The medians cross the generic setting only after 56 minutes for Bayesian optimization, 60 for CMA-ES, and 62 for the self-tuner, and within the hour it ran, PBO did not cross it. A short tuning session on a new user can leave them worse off, later, than no tuning at all.
How the tuner handles old data matters, and not always as expected. Turn on Show Bayesian optimization variants. Giving the Gaussian process the time of each measurement as a third input, so that old measurements count less for predictions about now, is the simplest model of drift (Section 46.5). It costs a little on adapted walkers and helps on novices: at 120 minutes the reduction on the adapted walker is 36.9% against 35.2% for the plain model. Exploring less, with the exploration constant lowered from 2.6 to 0.93, hurts here: that variant ends at 31.8%, below the generic setting, because it keeps returning to a region that was best for the novice. Kutulakos and Slade (2024) found the opposite on their landscapes: with a simulated novice, the less exploring variant reached the optimum in about 100 evaluations and the default variant more slowly. The authors suggest, without testing it, that its frequent evaluations near the estimated optimum let it follow the optimum as it moves; on fixed landscapes with 12 and 20 parameters the same variant settled on a worse setting. The two simulations differ in how the landscape moves, and neither is an experiment; the disagreement is a reason to test the exploration constant on the problem at hand, which those authors also advise, by simulation first and then in pilot sessions. CMA-ES and the self-tuner, which keep no long memory, track the change without any adjustment.
What should a practitioner do? The recommendations in Section 46.5 follow from this picture: let the person walk with the device before tuning begins, give early measurements less weight or model time explicitly, and tune again later. A preprint on a hip exoskeleton with 16 participants and three parameters, which optimized walking speed, reports that a Bayesian optimizer built for a changing response did better than the standard one in effectiveness, model accuracy, and personalization (Kim and Sergi, 2026). For preferences, the same problem has no tested solution: as of September 2026 we found no preferential optimization method with a model of a drifting utility (Section 29.10), and the preference optimizer in the figure is the slowest to recover.
Sources cited in Section 24.4 5
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Abram et al. (2022) General variability leads to specific adaptation toward optimal movement policies
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
- Kim and Sergi (2026) Validation of Dynamic Bayesian Optimization for Human-in-the-Loop Optimization of Exoskeleton Control at User-Driven Walking Speed
24.5 Against self-tuning #
The simulated self-tuner is a dozen lines of code with no model, and in the figures above it keeps pace with the optimizers. That mirrors the most informative comparison in the literature. In Schäfer et al. (2026), 11 people without prior exoskeleton experience tuned four hip timing parameters by thumbstick in 10.9 ± 0.9 minutes, and the resulting assistance lowered their metabolic rate by 16.6 ± 1.1% compared with zero torque. By the authors' own comparison, that is a quarter of the time of the optimization that Ding et al. (2018) ran for two parameters, and half of the time after which that optimization had converged. In Ingraham et al. (2022), 24 people tuned ankle torque and timing without seeing the values, converged in 105 seconds per trial, and repeated their own choice with a standard deviation of 1.7 N·m and 1.5% of the gait cycle. Section 33.3 discusses these studies and the caveats on comparing them with optimizer studies, which used other devices and baselines.
The simulation makes the reasons visible, and each is a property of the problem more than of any algorithm.
The wearer's sensor is fast. A felt comparison takes seconds; a metabolic estimate takes two minutes and is still noisy. In the same walking time the hand tuner sees six times as many settings.
The optimum is flat. Near the best setting, the reduction changes by less than the noise of a measurement and less than a person can feel. In Schäfer et al. (2026), shifting a timing by up to ±8% of the stride time did not change the metabolic reduction significantly. A coarse, fast search loses almost nothing to a precise one.
There are few parameters. With two to four parameters, changing one at a time works. Nothing here suggests it would with twenty.
The simulation also shows where self-tuning goes wrong, and these are the same places where preferential optimization goes wrong, since both listen to the same signal.
What is felt is not what is measured. The simulated novice dislikes torque, so tuning by feel, by hand or by PBO, ends at less torque than the metabolic optimum. For real wearers the gap between preference and metabolic cost is documented in both directions: the hip tuners of Schäfer et al. (2026) did not sit at a metabolic minimum, and prosthesis users preferred an ankle stiffness that made the motion of the prosthetic and the intact joint symmetric, with no significant relation to metabolic rate (Clites et al., 2021). Which of the two the device should serve is a decision about the objective, to be made before any tuning.
Below the perception floor, comparisons are noise. Once every change feels "about the same", further tuning by feel is a random walk. The floor can be measured, about 42% for exoskeleton stiffness during walking (Maberry and Martin, 2026), and it gives a principled point to stop asking (inference; the stopping rules of Section 46.6 do not yet include one).
People are not consistent with themselves. Two of three people with an amputation who explored the settings of a powered prosthesis chose different settings in different trials on the same day (Díaz et al., 2026).
What does this mean for the method? First, hand tuning is the baseline that a preferential optimizer has to beat, and Section 46.1 makes checking it the third question to ask before building one. No study has yet run the two on the same people with the same time budget and an endpoint fixed in advance; Section 47.4 describes that experiment, and the figures of this chapter are its simulation, with the simulation's caveat that the felt signal was written by us. Second, the simulation leaves out much of what matters to a wearer. It has no comfort beyond one penalty term, no safety, no fatigue, and no sense of agency, which in Schäfer et al. (2026) fell from 0.80 to 0.49 on a scale from 0 to 1 when assistance was switched on, even though the participants had chosen the assistance themselves. And its walkers are healthy adults, like nearly all participants of the studies it was calibrated to (Section 33.1.4).
In this problem the choice of optimizer is the smallest of the decisions. What an evaluation is (two noisy minutes of breathing, or a few seconds of feeling), when the person is tuned (before or after they have learned the device), and which objective the device should serve (measured effort or preference) each move the result by more than the difference between Bayesian optimization, an evolution strategy, and a careful hand.
The next chapter, Chapter 25, removes the instrument altogether: the only measurement is a person's choice between two versions of a photograph.
Sources cited in Section 24.5 6
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Clites et al. (2021) Understanding patient preference in prosthetic ankle stiffness
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Díaz et al. (2026) User preference in the personalized control of an ankle prosthesis: a case study
24.6 Exercises #
A two-minute metabolic estimate has a standard deviation of 4.6 points. Two settings in fact differ by 3 points. How many bouts per setting does it take to detect the difference by the usual standard of a two-sided test at the 5% level with 80% power, that is, a test that reports a difference in only 5% of experiments when there is none and in 80% of experiments when the difference is real? How many minutes of walking is that? Repeat for a difference of 8 points, the gap between generic and customized assistance in Poggensee and Collins (2021).
Solution
For two groups of independent estimates each, the difference of the means has standard deviation . Detecting a difference at the 5% level with 80% power needs , where 1.96 is the standard normal value exceeded in size with probability 5% and 0.84 the value exceeded with probability 20%, so . With and : , so 37 bouts per setting, 74 bouts, 148 minutes of walking. With : , so 6 bouts per setting, 24 minutes. A gap of 8 points can be confirmed in one session; the gaps of less than two points between tuners in Figure 24.3 cannot, and during those 148 minutes a new user would have adapted (Section 24.1.4).
A session lasts 24 minutes. Count what each of the four tuners of Section 24.3.2 gets to see. How many times does CMA-ES update its distribution? Why does that explain its slow start in Figure 24.4, and why does the same property protect it when the walker adapts?
Solution
By hand: settings. Preferences: comparisons. Metabolic estimates: , for Bayesian optimization and for CMA-ES alike. CMA-ES with six settings per generation completes two generations, so it updates its mean twice, the first time after 12 minutes. Until then its recommendation is its starting point. Bayesian optimization refits its model after every estimate. The property that makes CMA-ES slow, that it uses a generation once and then discards it, also means that measurements of the walker as a novice cannot mislead it an hour later: nothing older than one generation is in its state except through the mean and covariance they produced.
In the simulation, a felt reading has Gaussian noise of standard deviation 0.03, and a change is reported as "easier" when the previous reading exceeds the new one by more than 0.025. A change in fact lowers the felt effort by 0.02 (two points). With what probabilities is it reported as easier, about the same, and harder? What does the simple self-tuner of the text do in each case?
Solution
The difference of two readings is Gaussian with mean 0.02 and standard deviation . Easier: . Harder: . About the same: the remaining . The self-tuner keeps the change only in the first case, so it throws away a real two-point improvement more often than it accepts it, and each rejection costs two settings (the trial and the return). A two-point improvement is within reach of neither the wearer's perception nor a single metabolic estimate, which is the flat optimum seen from both sides.
Using for the amplitude and peak time and for the torque, compute the adaptation levels after a 24-minute session. A walker's best torque is 0.45 as a novice and 0.86 as an expert. Where is it after 24 minutes, and what does the answer imply for a setting tuned perfectly in that session?
Solution
and . The best torque is . A setting tuned perfectly at 24 minutes has torque 0.57, while the adapted walker's best is 0.86. On the adapted walker its torque term is , so even with perfect timing it gives about of the possible 39%. The session optimized for a person who was halfway through learning the device.
Sources cited in Section 24.6 1
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
Further reading #
- Zhang et al. (2017) is the study that established human-in-the-loop optimization of exoskeletons, with an evolution strategy on two-minute metabolic estimates; Ding et al. (2018) is the Bayesian optimization counterpart on a hip exosuit.
- Selinger and Donelan (2014) explain why a metabolic rate can be estimated before it settles, and measure the time constant of the response.
- Poggensee and Collins (2021) separate what training, adaptation, and customization each contribute, with naive users followed until they became expert.
- Kutulakos and Slade (2024), a preprint, fit landscapes to those data and simulate optimizers on them, including a novice who adapts; the idea of this chapter's simulation comes from it.
- Schäfer et al. (2026) and Ingraham et al. (2022) are the self-tuning studies: what people choose by feel, how fast, and how repeatably.
- Tucker et al. (2020b) and Tucker et al. (2020a) tune exoskeleton gaits from the wearer's comparisons; Chapter 33 reviews the whole literature, including prostheses and the comparison with manual tuning.
References
- (2022). General variability leads to specific adaptation toward optimal movement policies. Current Biology. Cited in §24.4
- (2021). Understanding patient preference in prosthetic ankle stiffness. Journal of NeuroEngineering and Rehabilitation. Cited in §24.5
- (2026). User preference in the personalized control of an ankle prosthesis: a case study. Journal of NeuroEngineering and Rehabilitation. doi:10.1186/s12984-026-01931-w. Cited in §24.5
- (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. Cited in §24.1 §24.2 §24.3 §24.5
- (2001). Completely Derandomized Self-Adaptation in Evolution Strategies. Evolutionary Computation. Cited in §24.2
- (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §24.2 §24.3 §24.4 §24.5
- (2026). Validation of Dynamic Bayesian Optimization for Human-in-the-Loop Optimization of Exoskeleton Control at User-Driven Walking Speed. bioRxiv. preprint Cited in §24.4
- (2024). Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance. bioRxiv. preprint Cited in §24.1 §24.2 §24.3 §24.4
- (2023). User preference optimization for control of ankle exoskeletons using sample efficient active learning. Science Robotics. Cited in §24.2
- (2026d). Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization. arXiv. preprint Cited in §24.2
- (2026). Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §24.3 §24.5
- (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §24.1 §24.2 §24.3 §24.4 §24.6
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §24.1 §24.2 §24.3 §24.5
- (2014). Estimating instantaneous energetic cost during non-steady-state gait. Journal of Applied Physiology. Cited in §24.1
- (2022). Personalizing exoskeleton assistance while walking in the real world. Nature. Cited in §24.1 §24.2 §24.3
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §24.2
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §24.2
- (2017). Human-in-the-loop optimization of exoskeleton assistance during walking. Science. Cited in §24.1 §24.2