arXiv:2604.27850cs.CL2026-04

让大模型通过物体描述推理,提升对话系统指代消解准确率

Reasoning over Object Descriptions Improves Coreference Resolution in Task-Based Dialogue Systems

论文配图:Reasoning over Object Descriptions Improves Coreference Resolution in Task-Based Dialogue Systems
图 1 · 摘自论文原文
  • 大模型在测试时基于物体元数据与对话历史进行逐步推理
  • 在SIMMC 2.1上实现跨域泛化,准确率优于传统监督模型
  • 少样本下对新场景和新物体仍有强适应性,适合复杂交互场景

任务型对话系统通过自然语言交互帮助用户完成特定目标。准确的指代消解至关重要,需识别对话中对物体的引用,但在视觉关联环境中因场景复杂、物体元数据多样而更具挑战。当前方法受限于领域泛化差及对标注数据的依赖,易过拟合数据集特异性特征。本文提出一种单模态测试时推理方法,使大语言模型(LLM)能基于详细物体元数据与对话历史推理以改进指代消解。在SIMMC 2.1数据集上的实证结果表明,LLM可生成有效的分步推理过程,精准关联对话上下文与场景中的物体。大量实验显示其能准确链接对话与物体;且在少样本设置下,对未见场景与新物体具备良好泛化能力,跨域评估中优于编码器式监督模型。这些发现凸显结构化元数据与精心提示工程对提升任务导向对话系统鲁棒性与泛化性的关键作用。

原文摘要 · Abstract (English)

Task-based dialogue systems assist users in achieving specific goals, such as executing actions or retrieving information, through natural language interactions. Accurate coreference resolution is essential, as it involves identifying object references within the dialogue - a task that becomes increasingly challenging in visually grounded environments characterized by complex scenes and diverse object metadata. However, coreference resolution in task-based dialogue remains limited by poor generalization across domains and heavy reliance on supervised models that often overfit to dataset-specific artifacts. In this work, we propose a unimodal test-time reasoning approach that enables large language models (LLMs) to reason over detailed object metadata and dialogue history to improve coreference resolution. Empirical results on the SIMMC 2.1 dataset demonstrate that LLMs can generate step-by-step reasoning processes that effectively align dialogue context with objects present in the scene. Extensive experiments highlight the models' ability to link conversations and objects accurately. Moreover, we show that test-time reasoning under few-shot settings generalizes effectively to unseen scenarios and novel objects, outperforming encoder-based supervised methods in cross-domain evaluations. These findings underscore the critical role of structured metadata and careful prompt engineering in enhancing the robustness and generalization of task-oriented dialogue systems.

指代消解大模型推理对话系统少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。