arXiv:2602.23730cs.AI2026-02

让多模态大模型在东南亚场景下实现感知与逻辑的协同,揭示二者权衡机制。

Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off

  • 分步训练:先分离再融合感知与推理能力,提升区域适配性。
  • 用小成本生成-判断-修正流程,将文本推理迁移到多模态任务中。
  • 发现推理越强,低层感知越不稳定,存在时序漂移和视觉过解读现象。

近期多模态大语言模型追求全感知能力,但如何将稳健的感官基础与复杂推理结合仍是挑战,尤其在欠代表地区。本文介绍面向东南亚(SEA)的10B参数多语言全感知模型MERaLiON2-Omni(Alpha)的研究预览。我们提出一种渐进式训练流程,显式解耦并重新整合“系统1”(感知)与“系统2”(推理)能力。首先,通过正交模态适配,将本地音视频特征(如新加坡英语混杂、本土文化地标)与多语言大模型对齐,建立稳健的感知基座。其次,为低成本注入认知能力,提出生成-判断-修正流水线,利用超大模型过滤幻觉并以共识机制解决冲突,合成高质量银数据,使文本链式思维推理迁移至多模态场景。在新提出的SEA-Omni基准套件上的综合评估揭示效率-稳定性悖论:推理作为非线性放大器显著提升抽象任务表现(数学与指令遵循),却引入低层感知不稳定性。具体表现为长上下文音频中的时间漂移,以及逻辑过度解释导致像素级现实被忽略。本文详述架构设计、高效训练方案及感知与推理间权衡的诊断分析。

原文摘要 · Abstract (English)

Recent advancements in Multimodal Large Language Models (MLLMs) pursue omni-perception capabilities, yet integrating robust sensory grounding with complex reasoning remains a challenge, particularly for underrepresented regions. In this report, we introduce the research preview of MERaLiON2-Omni (Alpha), a 10B-parameter multilingual omni-perception tailored for Southeast Asia (SEA). We present a progressive training pipeline that explicitly decouples and then integrates "System 1" (Perception) and "System 2" (Reasoning) capabilities. First, we establish a robust Perception Backbone by aligning region-specific audio-visual cues (e.g., Singlish code-switching, local cultural landmarks) with a multilingual LLM through orthogonal modality adaptation. Second, to inject cognitive capabilities without large-scale supervision, we propose a cost-effective Generate-Judge-Refine pipeline. By utilizing a Super-LLM to filter hallucinations and resolve conflicts via a consensus mechanism, we synthesize high-quality silver data that transfers textual Chain-of-Thought reasoning to multimodal scenarios. Comprehensive evaluation on our newly introduced SEA-Omni Benchmark Suite reveals an Efficiency-Stability Paradox: while reasoning acts as a non-linear amplifier for abstract tasks (boosting mathematical and instruction-following performance significantly), it introduces instability in low-level sensory processing. Specifically, we identify Temporal Drift in long-context audio, where extended reasoning desynchronizes the model from acoustic timestamps, and Visual Over-interpretation, where logic overrides pixel-level reality. This report details the architecture, the data-efficient training recipe, and a diagnostic analysis of the trade-offs between robust perception and structured reasoning.

多模态推理感知东南亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。