arXiv:2606.14862cs.RO2026-06

用语言精准控制机器人触觉行为,只需少量标注即可实现个性化操作。

TacStyle: Personalizing Tactile Robot Policies using Structured Behavior Representations

论文配图:TacStyle: Personalizing Tactile Robot Policies using Structured Behavior Representations
图 1 · 摘自论文原文
  • 通过学习用户偏好的结构化潜在表示,实现对行为风格的系统建模。
  • 在模拟和真实场景中,仅需少量标签即实现精确行为调整,优于传统方法。
  • 适合需要个性化交互的智能机器人应用,如家庭助手机器人。

能够协助人类的机器人系统应能根据个体用户偏好调整行为。例如,用户可能希望机械臂在折叠衣物或清洁家具时调整施加的力。自然语言为传达此类偏好提供了直观方式。近年来,语言条件化的机器人策略已使机器人能根据语言提示执行任务。然而,将该方法扩展到如何执行任务,需详细标注任务数据中轨迹的偏好或风格。这类标注不仅难以收集,且直接基于标签进行条件化可能无法对连续行为范围实现精细控制。例如,通过“比之前多施加一点压力”等抽象指令难以精确传达所需力度。因此,本文提出通过语言推理偏好行为,而非直接生成行为。我们首先学习一个结构化潜在表示,按对应轨迹差异组织用户偏好;然后利用基础模型解释该潜在空间,选择生成期望行为的值。在模拟与真实实验中,从直观结构化的潜在空间中选择行为,可更精确适应用户偏好,且所需偏好标签显著少于传统语言条件化策略。

原文摘要 · Abstract (English)

Robotic systems that assist humans should be capable of adapting their behaviors to individual user preferences. For instance, users may want a robot arm to adjust the amount of force it applies while folding their laundry or cleaning furniture. Natural language provides an intuitive way for humans to communicate such preferences. Recent progress in language-conditioned robot policies has shown that robots can successfully use language prompts to determine what task to perform. However, extending the same approach to realize how the task should be performed requires detailed labels describing the preferences or styles of trajectories in the task data. Not only is collecting such annotations challenging, but conditioning directly on these labels may also fail to provide fine-grained control over a continuous range of behaviors. For example, it can be difficult to convey the exact force that a robot must apply through abstract instructions like "apply a bit more pressure than before". Therefore, in this work, we propose using language to reason over preferred behaviors instead of directly generating them. We first learn a structured latent representation that organizes user preferences according to differences in the corresponding trajectories. Then, given a preference prompt, we use a foundation model to interpret this latent space and choose a value that produces the desired behavior. Through both simulation and real-world experiments, we show that selecting robot behaviors from an intuitively structured latent space enables more precise adaptation to user preferences while requiring significantly fewer preference labels than language-conditioned policies.

机器人控制个性化语言交互行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。