构建多轮多模态欺骗意图识别数据集,揭示大模型在复杂推理中的短板。
MISID: A Multimodal Multi-turn Dataset for Complex Intent Recognition in Strategic Deception Games
- 设计双层标注体系,支持长对话中多轮策略性欺骗的细粒度分析。
- 发现主流多模态大模型存在文本主导视觉幻觉、跨模态协同不足等问题。
- 提出FRACTAM框架,通过解耦-锚定-推理提升复杂因果链推理能力。
理解复杂多轮交互中的人类意图仍是人机交互与行为分析的核心挑战。现有意图识别数据集多聚焦单轮或简单对话,而真实场景常涉及参与者需在长时间内维持复杂欺骗叙事的高阶策略互动。为此,我们推出MISID——一个源自高风险社交策略游戏的多模态、多轮、多参与者的基准数据集。MISID采用细粒度双层多维标注方案,适用于长上下文话语分析与基于证据的因果追踪。对主流多模态大语言模型(MLLMs)在MISID上的系统评估揭示其在复杂场景中的关键缺陷:文本主导的视觉幻觉、跨模态协同能力下降、因果线索串联能力有限。为此,我们提出FRACTAM基线框架,采用“解耦-锚定-推理”范式,通过提取纯单模态事实表示减少文本偏见,使用两阶段检索实现长程事实锚定,并构建显式的跨模态证据链。大量实验表明,FRACTAM显著提升主流模型在复杂战略任务中的表现,增强隐含意图检测与推理能力,同时保持稳健的感知精度。数据集已公开于https://naislab.cn/datasets/MISID。
原文摘要 · Abstract (English)
Understanding human intent in complex multi-turn interactions remains a fundamental challenge in human-computer interaction and behavioral analysis. While existing intent recognition datasets focus mainly on single utterances or simple dialogues, real-world scenarios often involve sophisticated strategic interactions where participants must maintain complex deceptive narratives over extended periods. To address this gap, we introduce MISID, a comprehensive multimodal, multi-turn, and multi-participant benchmark for intent recognition. Sourced from high-stakes social strategy games, MISID features a fine-grained, two-tier multi-dimensional annotation scheme tailored for long-context discourse analysis and evidence-based causal tracking. Our systematic evaluation of state-of-the-art Multimodal Large Language Models (MLLMs) on MISID reveals critical deficiencies in complex scenarios, including text-prior visual hallucination, impaired cross-modal synergy, and limited capacity in chaining causal cues. Consequently, we propose FRACTAM as a baseline framework. Using a ``Decouple-Anchor-Reason'' paradigm, FRACTAM reduces text bias by extracting pure unimodal factual representations, employs two-stage retrieval for long-range factual anchoring, and constructs explicit cross-modal evidence chains. Extensive experiments demonstrate that FRACTAM enhances mainstream models' performance in complex strategic tasks, improving hidden intent detection and inference while maintaining robust perceptual accuracy. Our dataset is available at https://naislab.cn/datasets/MISID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。