arXiv:2410.03959cs.CLcs.AI2024-10EMNLP被引 12

让智能体从不同视角生成和理解指代表达,提升多智能体协作沟通能力。

Grounding Language in Multi-Perspective Referential Communication

  • 设计多视角指代任务,要求智能体考虑彼此视觉差异进行交流。
  • 人类配对沟通成功率超模型,2970条表达与判断数据验证有效性。
  • 训练可改进的开放模型,通信成功率从58.9%升至69.3%,超越大厂模型。

我们提出一个在多智能体具身环境中进行指代表达生成与理解的新任务及数据集。在此任务中,两个共享场景的智能体需考虑彼此的视觉视角差异(可能不同于自身视角),以生成并理解场景中物体及其空间关系的指代表达。我们收集了2,970条人工撰写的指代表达,每条均配有对应的人类理解判断,并评估自动化模型作为说话者与听者与人类搭档配合的表现,发现模型在生成与理解方面的表现均落后于人类搭档组合。最后,我们通过证据驱动方式训练一个开源说话者模型,在与听者搭配时表现出通信成功,使成功率从58.9%提升至69.3%,甚至超过最强的专有模型。

原文摘要 · Abstract (English)

We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspective, which may be different from their own, to both produce and understand references to objects in a scene and the spatial relations between them. We collect a dataset of 2,970 human-written referring expressions, each paired with human comprehension judgments, and evaluate the performance of automated models as speakers and listeners paired with human partners, finding that model performance in both reference generation and comprehension lags behind that of pairs of human agents. Finally, we experiment training an open-weight speaker model with evidence of communicative success when paired with a listener, resulting in an improvement from 58.9 to 69.3% in communicative success and even outperforming the strongest proprietary model.

指代表达多智能体具身对话通信优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。