arXiv:2604.14520cs.CV2026-04

提出动态调制的多模态融合框架,解决模型因结构僵化导致的性能下降问题。

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

论文配图:Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs
图 1 · 摘自论文原文
  • 通过动态切换输入路径,消除静态融合带来的位置偏差和对齐陷阱
  • 在多个基准上实现稳定提升,优于传统联合推理方式
  • 适合需要可靠多模态理解的复杂任务场景

全能型多模态大语言模型(Omni-MLLMs)旨在统一整合多种感官信息。然而,近期评估揭示了一个关键性能悖论:单模态基线模型经常优于联合多模态推理。我们追溯这一感知脆弱性源于当前模型普遍采用的静态融合拓扑,识别出两类结构性病理:序列输入中的位置偏差,以及交错格式中的对齐陷阱,这些缺陷会系统性扭曲注意力机制,与任务语义无关。为解决这种功能僵化,我们提出链式模态(Chain of Modality, CoM),一个代理式框架,将多模态融合从被动拼接转变为动态编排。CoM 自适应地调度输入拓扑,在并行、序列和交错路径间切换,以消除结构偏差。此外,CoM 将认知执行拆分为两条任务对齐路径:面向直接感知的“直觉决策”路径,以及面向分析审计的“推理决策”路径。该框架可在零训练或数据高效微调(SFT)设置下运行,在多个基准测试中均展现出鲁棒且一致的泛化能力。

原文摘要 · Abstract (English)

Omni-modal Large Language Models (Omni-MLLMs) promise a unified integration of diverse sensory streams. However, recent evaluations reveal a critical performance paradox: unimodal baselines frequently outperform joint multimodal inference. We trace this perceptual fragility to the static fusion topologies universally employed by current models, identifying two structural pathologies: positional bias in sequential inputs and alignment traps in interleaved formats, which systematically distort attention regardless of task semantics. To resolve this functional rigidity, we propose Chain of Modality (CoM), an agentic framework that transitions multimodal fusion from passive concatenation to dynamic orchestration. CoM adaptively orchestrates input topologies, switching among parallel, sequential, and interleaved pathways to neutralize structural biases. Furthermore, CoM bifurcates cognitive execution into two task-aligned pathways: a streamlined ``Direct-Decide'' path for direct perception and a structured ``Reason-Decide'' path for analytical auditing. Operating in either a training-free or a data-efficient SFT setting, CoM achieves robust and consistent generalization across diverse benchmarks.

多模态动态融合大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。