arXiv:2602.04213cs.AI2026-02

让普通人也能轻松教AI学开车,通过互动调整策略和参数。

InterPReT: Interactive Policy Restructuring and Training Enable Effective Imitation Learning from Laypersons

  • 用户可交互输入指令与示范,系统自动重构策略并优化参数。
  • 34人实验表明,政策更稳健,且不降低易用性。
  • 适合无机器学习背景的普通用户训练可靠智能体。

模仿学习在许多任务中已取得成功,但现有方法多依赖技术专家提供的大规模示范,并需密切监控训练过程,这对普通用户而言难以实现。为降低教学门槛,我们提出交互式策略重构与训练(InterPReT),通过用户指令持续更新策略结构并优化参数,使终端用户可交互式提供指令与示范,实时监控代理性能,并审查其决策策略。一项包含34名用户的实验显示,在由非专业人士同时负责示范和决定停止时机的情况下,该方法生成的策略更具鲁棒性,且未损害系统可用性,优于通用模仿学习基线。这表明该方法更适合无机器学习背景的用户训练可靠的策略。

原文摘要 · Abstract (English)

Imitation learning has shown success in many tasks by learning from expert demonstrations. However, most existing work relies on large-scale demonstrations from technical professionals and close monitoring of the training process. These are challenging for a layperson when they want to teach the agent new skills. To lower the barrier of teaching AI agents, we propose Interactive Policy Restructuring and Training (InterPReT), which takes user instructions to continually update the policy structure and optimize its parameters to fit user demonstrations. This enables end-users to interactively give instructions and demonstrations, monitor the agent's performance, and review the agent's decision-making strategies. A user study (N=34) on teaching an AI agent to drive in a racing game confirms that our approach yields more robust policies without impairing system usability, compared to a generic imitation learning baseline, when a layperson is responsible for both giving demonstrations and determining when to stop. This shows that our method is more suitable for end-users without much technical background in machine learning to train a dependable policy

模仿学习交互式训练普通人可用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。