arXiv:2512.02791cs.CL2025-12

构建三阶段数据合成框架,提升对话指代理解的泛化能力

Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension

  • 分三阶段合成逼真且可控的对话指代表达数据
  • 在多个标准指标上显著优于现有方法
  • 适合需要强泛化性的视觉对话与指代理解研究者

基于对话的广义指代理解(GREC)要求模型在复杂视觉场景中定位表达内容,并在长对话上下文中解决指代消解问题。然而,现有系统在训练与评估域分布差异下表现不佳,主要因标注对话接地数据稀缺。本文提出一种三阶段数据合成方法,在真实性和可控性间取得平衡,生成可用于对话条件接地任务的大规模监督数据。在合成数据上微调模型后,各项标准评估指标均实现一致且显著提升。

原文摘要 · Abstract (English)

Dialogue-Based Generalized Referring Expression Comprehension (GREC) requires models to ground the expression and unlimited targets in complex visual scenes while resolving coreference across a long dialogue context. However, existing systems struggle under distribution shift between training and evaluation domains, a gap exacerbated by the scarcity of annotated dialogue grounding data. We address this challenge with a three-tier data-synthesis method that balances realism and controllability to produce scalable supervision for dialogue-conditioned grounding. Fine-tuning on the synthesized data yields consistent, substantial improvements over prior approaches across standard evaluation metrics.

指代理解对话系统数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。