arXiv:2409.05721cs.CLcs.AI2024-09中稿 · publication at INL…被引 1

让对话中的指代表达更精准,兼顾视觉和上下文信息

Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding

  • 先生成候选表达,再根据对话上下文重排序
  • 重排序后表达的图文匹配准确率显著提升
  • 适合需要自然对话交互的智能助手场景

我们提出一种在视觉对话中生成指代表达(REG)的方法,旨在生成既具有区分性又符合语境的表达。该方法分为两阶段:第一阶段将REG建模为文本与图像条件下的下一个词预测任务,基于前文语言上下文和目标对象的视觉表示,自回归生成指代表达;第二阶段引入基于对话理解的重排序机制,通过生成-重排策略对候选表达进行重排序,依据其在特定对话语境下的区分能力。人工评估结果表明,所提两阶段方法能有效生成更具区分性的指代表达,经重排序后的表达在图文检索准确率上优于贪婪解码生成的结果。

原文摘要 · Abstract (English)

We propose an approach to referring expression generation (REG) in visually grounded dialogue that is meant to produce referring expressions (REs) that are both discriminative and discourse-appropriate. Our method constitutes a two-stage process. First, we model REG as a text- and image-conditioned next-token prediction task. REs are autoregressively generated based on their preceding linguistic context and a visual representation of the referent. Second, we propose the use of discourse-aware comprehension guiding as part of a generate-and-rerank strategy through which candidate REs generated with our REG model are reranked based on their discourse-dependent discriminatory power. Results from our human evaluation indicate that our proposed two-stage approach is effective in producing discriminative REs, with higher performance in terms of text-image retrieval accuracy for reranked REs compared to those generated using greedy decoding.

指代表达视觉对话重排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。