arXiv:2501.08626cs.GTcs.HC2025-01被引 2

不求解逆问题,仅靠观察人类行为就找到最优控制策略。

A Learning Algorithm That Attains the Human Optimum in a Repeated Human-Machine Interaction Game

  • 基于博弈论设计学习算法,直接从人类动作推导最优解。
  • 在多轮人机交互中稳定收敛至预设成本函数的最小值。
  • 适用于需优化人体感知成本的智能辅助系统场景。

当人类与基于学习的控制系统交互时,通常目标是降低人类知晓但系统未知的成本函数。例如外骨骼可自适应调整助力以最小化人体运输代谢成本。传统方法需求解逆问题来推断人类成本,但此类问题常病态、难解或对数据敏感。本文提出一种仅通过观察人类行为即可定位成本最小点的游戏理论学习算法,避免了逆问题求解。我们在大量受试者实验中评估该算法性能,结果表明其在标量和多维情形下均能稳定收敛至指定人类成本函数的最小值。最后,我们展望了理论与实证方向的未来扩展。

原文摘要 · Abstract (English)

When humans interact with learning-based control systems, a common goal is to minimize a cost function known only to the human. For instance, an exoskeleton may adapt its assistance in an effort to minimize the human's metabolic cost-of-transport. Conventional approaches to synthesizing the learning algorithm solve an inverse problem to infer the human's cost. However, these problems can be ill-posed, hard to solve, or sensitive to problem data. Here we show a game-theoretic learning algorithm that works solely by observing human actions to find the cost minimum, avoiding the need to solve an inverse problem. We evaluate the performance of our algorithm in an extensive set of human subjects experiments, demonstrating consistent convergence to the minimum of a prescribed human cost function in scalar and multidimensional instantiations of the game. We conclude by outlining future directions for theoretical and empirical extensions of our results.

人机交互博弈学习优化控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。