让智能体用自然语言解释动作,实现透明决策。
Joint Action Language Modelling for Transparent Policy Execution
- 将策略学习转化为语言生成任务,边说边做。
- 同时生成高质量语言和正确动作,效果更好。
- 适合关注可解释AI行为的研究者。
智能体的意图常被具身策略的黑箱特性掩盖。通过使用描述下一步动作的自然语言陈述进行沟通,可增强对智能体行为的透明度。我们旨在将透明行为直接融入学习过程,将策略学习问题转化为语言生成问题,并与传统的自回归建模结合。所提出的模型在语言-桌环境(Language-Table)中,能生成透明的自然语言语句,随后输出表示具体动作的离散化标记,以解决长时程任务。沿用前人工作,该模型以自回归方式学习生成由特殊离散标记表示的策略。我们特别关注动作预测与生成高质量透明语言之间的关系。发现当二者同时生成时,动作轨迹质量和语言表达质量均显著提升。
原文摘要 · Abstract (English)
An agent's intention often remains hidden behind the black-box nature of embodied policies. Communication using natural language statements that describe the next action can provide transparency towards the agent's behavior. We aim to insert transparent behavior directly into the learning process, by transforming the problem of policy learning into a language generation problem and combining it with traditional autoregressive modelling. The resulting model produces transparent natural language statements followed by tokens representing the specific actions to solve long-horizon tasks in the Language-Table environment. Following previous work, the model is able to learn to produce a policy represented by special discretized tokens in an autoregressive manner. We place special emphasis on investigating the relationship between predicting actions and producing high-quality language for a transparent agent. We find that in many cases both the quality of the action trajectory and the transparent statement increase when they are generated simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。