清理视觉干扰后,小模型通过分阶段训练也能逼近大模型表现。
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

- 用纯视觉探测剔除误导性题目,构建更公平的评估集
- 三阶段训练让3B模型性能接近30B大模型,无需更强教师
- 自蒸馏数据提升模型融合多模态能力,结果更可信
全模态语言模型旨在联合理解音频、视觉与语言信息,但当前基准测试结果常被视觉线索过度膨胀。本文研究现有全模态基准是否能区分视觉捷径与真正的音视频-语言融合能力,并在去视觉偏见评估下分析后训练行为。通过视觉仅探针审计九个基准,移除可仅凭视觉解答的问题,保留完整子集以确保比较稳定,得到包含8,551个保留题目的新评估集OmniClean(原16,968个题)。在OmniClean上,评估基于Qwen2.5-Omni-3B的三阶段后训练方案:混合双模态SFT、混合模态RLVR、自蒸馏数据上的SFT。均衡双模态SFT带来有限且不均收益,RLVR实现首次广泛提升,自蒸馏重塑基准表现。经过自蒸馏SFT后,3B模型整体性能达到并略微超过未使用更强全模态教师的Qwen3-Omni-30B-A3B-Instruct。结果表明,控制视觉泄露后,全模态进展更易解读,小模型亦可通过分阶段后训练与自蒸馏监督获显著提升。
原文摘要 · Abstract (English)
Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni-modal benchmarks separate visual shortcuts from genuine audio-visual-language evidence integration, and how post-training behaves under a visually debiased evaluation setting. We audit nine omni-modal benchmarks with visual-only probing, remove visually solvable queries, and retain full subsets when filtering is undefined or would make comparisons unstable. This yields OmniClean, a cleaned evaluation view with 8,551 retained queries from 16,968 audited queries. On OmniClean, we evaluate OmniBoost, a three-stage post-training recipe based on Qwen2.5-Omni-3B: mixed bi-modal SFT, mixed-modality RLVR, and SFT on self-distilled data. Balanced bi-modal SFT gives limited and uneven gains, RLVR provides the first broad improvement, and self-distillation reshapes the benchmark profile. After SFT on self-distilled data, the 3B model reaches performance comparable to, and in aggregate slightly above, Qwen3-Omni-30B-A3B-Instruct without using a stronger omni-modal teacher. These results show that omni-modal progress is easier to interpret when evaluation controls visual leakage, and that small omni-modal models can benefit from staged post-training with self-distilled omni-query supervision. Project page: https://cheliu-computation.github.io/omni/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。