提出新框架,精准识别多模态长对话中的异常输入。
'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue
- 设计DIAEF框架,融合视觉语言模型与新评分机制。
- 在未见标签场景下,联合检测效果优于单一模态。
- 适合构建鲁棒的多轮对话系统,提升用户体验。
多模态上下文中的分布外(OOD)检测对于识别跨模态输入偏差至关重要,尤其在开放域对话系统或真实对话交互中。本文旨在通过高效检测多轮长对话及图像中的异常输入,改善用户交互体验。提出一种名为对话图像对齐增强框架(DIAEF)的新评分机制,结合视觉语言模型,用于检测两类关键场景:(1)对话与图像输入不匹配;(2)包含此前未见标签的输入对。实验结果表明,在多个基准测试中,联合使用图像与多轮对话的OOD检测,在未见标签场景下表现优于单独使用任一模态。在存在错配对时,所提评分能有效识别并具备强鲁棒性,有助于构建更具领域感知力和自适应能力的对话代理,并为未来研究提供基准。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection in multimodal contexts is essential for identifying deviations in combined inputs from different modalities, particularly in applications like open-domain dialogue systems or real-life dialogue interactions. This paper aims to improve the user experience that involves multi-round long dialogues by efficiently detecting OOD dialogues and images. We introduce a novel scoring framework named Dialogue Image Aligning and Enhancing Framework (DIAEF) that integrates the visual language models with the novel proposed scores that detect OOD in two key scenarios (1) mismatches between the dialogue and image input pair and (2) input pairs with previously unseen labels. Our experimental results, derived from various benchmarks, demonstrate that integrating image and multi-round dialogue OOD detection is more effective with previously unseen labels than using either modality independently. In the presence of mismatched pairs, our proposed score effectively identifies these mismatches and demonstrates strong robustness in long dialogues. This approach enhances domain-aware, adaptive conversational agents and establishes baselines for future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。