arXiv:2603.18425cs.CL2026-03

提出多模态任务干扰基准,揭示模型切换任务时的性能下降规律

Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs

  • 构建多模态任务干扰评测基准,覆盖六类跨模态任务
  • 从纯文本转图像任务时性能骤降,反向影响小
  • 模态不匹配是主要干扰源,多维度不匹配会加剧问题

任务干扰指单次对话中任务切换导致的性能下降,此前研究仅限于纯文本场景,而多模态对话系统日益普及。本文提出一个用于评估多模态大模型任务干扰现象的基准,涵盖文本与视觉任务共六类,并系统性地在三个维度上变化历史-目标匹配度:模态不匹配、推理需求不匹配和答案格式不匹配。在开源与闭源模型上的实验表明,任务干扰具有高度方向性:从纯文本任务切换到图像目标时性能显著下降,反之则影响微弱。当多个维度的不匹配同时存在时,干扰进一步放大,其中模态差异影响最显著,其次为答案格式,推理需求变化的影响最小。

原文摘要 · Abstract (English)

Task interference, the performance degradation caused by task switches within a single conversation, has been studied exclusively in text-only settings despite the growing prevalence of multimodal dialogue systems. We introduce a benchmark for evaluating this phenomenon in multimodal LLMs, covering six tasks across text and vision with systematic variation of history-target along three axes: modality mismatch, reasoning mismatch, and answer format mismatch. Experiments on both open-weights and proprietary models reveal that task interference is highly directional: switching from text-only to image-based targets causes severe performance drops, while the reverse transition yields minimal degradation. Interference is further amplified when mismatches co-occur across multiple dimensions, and is driven most strongly by modality differences, followed by answer format, while reasoning requirement shifts cause minimal degradation.

多模态任务干扰模型评测指令对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。