arXiv:2506.04755cs.CVcs.AI2025-06被引 11

只用9.3%数据,让多模态模型推理更高效

Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning

  • 通过识别关键认知样本,筛选出真正能激发推理的高质量数据
  • 仅用9.3%训练数据,性能超越全量数据集,节省超43%算力
  • 适合追求高效训练与低资源部署的多模态研究者

尽管多模态大语言模型(MLLMs)在复杂推理任务中已取得显著进展,但普遍认为需大量训练数据才能提升多模态推理能力,导致数据冗余与高昂计算成本。本文提出质疑:小规模高价值数据集能否媲美甚至超越全量数据?基于关键观察——仅少数样本(称为认知样本)能真正触发多模态推理,其余贡献甚微,我们提出新型数据选择范式「推理激活潜力(RAP)」。RAP通过两个互补估计器识别认知样本:1)因果差异估计器(CDE),基于潜在结果模型,通过对比多模态与纯文本输入输出,剔除过度依赖语言先验的样本;2)注意力置信度估计器(ACE),利用令牌级自注意力机制,剔除中间推理阶段中被无关但过度强调的令牌主导的样本。此外,引入难度感知替换模块(DRM),以更具挑战性的样本替换平凡实例,保障推理复杂性。六组数据集实验表明,使用仅9.3%训练数据的RAP方法,在性能上持续优于全量数据,计算成本降低超过43%。

原文摘要 · Abstract (English)

While multi-modal large language models (MLLMs) have made significant progress in complex reasoning tasks via reinforcement learning, it is commonly believed that extensive training data is necessary for improving multi-modal reasoning ability, inevitably leading to data redundancy and substantial computational costs. However, can smaller high-value datasets match or outperform full corpora for multi-modal reasoning in MLLMs? In this work, we challenge this assumption through a key observation: meaningful multi-modal reasoning is triggered by only a sparse subset of training samples, termed cognitive samples, whereas the majority contribute marginally. Building on this insight, we propose a novel data selection paradigm termed Reasoning Activation Potential (RAP)}, which identifies cognitive samples by estimating each sample's potential to stimulate genuine multi-modal reasoning by two complementary estimators: 1) Causal Discrepancy Estimator (CDE) based on the potential outcome model principle, eliminates samples that overly rely on language priors by comparing outputs between multi-modal and text-only inputs; 2) Attention Confidence Estimator (ACE), which exploits token-level self-attention to discard samples dominated by irrelevant but over-emphasized tokens in intermediate reasoning stages. Moreover, we introduce a Difficulty-aware Replacement Module (DRM) to substitute trivial instances with cognitively challenging ones, thereby ensuring complexity for robust multi-modal reasoning. Experiments on six datasets show that our RAP method consistently achieves superior performance using only 9.3% of the training data, while reducing computational costs by over 43%.

多模态数据筛选高效训练推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。