让多模态模型精准选择社交视频中的关键证据,提升推理准确率。
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

- 基于认知结构分解输入,动态路由不同模态专家处理任务。
- 在社交视频问答中达59.38%准确率,较最强基线提升26.77个百分点。
- 适合需要跨模态协同、时间对齐与社会语境理解的研究者。
多模态模型可处理视频、音频和文本,但统一接入并不意味着能正确使用证据。这一差距在社交视频问答中尤为明显,答案可能依赖手势、语气、时间线索或言行不一致。我们提出CogniRoute,一种基于认知结构引导的专家混合框架。该方法通过训练时的纯认知结构对每个样本进行跨模态关系、推理需求和时间范围的分解,并在监督微调中对齐全局路由信号。我们进一步引入路由感知强化学习,联合优化生成结果与专家分配,奖励包括答案正确性、模态一致性推理及认知时间定位。为支持训练与评估,我们构建了OmniSocialBench,一个包含11.8万条结构化训练样本的诊断性社交视频问答数据集,附有可追溯的推理路径、结构标签、时间证据片段和人工验证的测试集。CogniRoute在OmniSocialBench上达到59.38%平均准确率,相比最强专有基线提升15.33个百分点,相比最强开源多模态基线提升26.77个百分点,尤其在需音视频协调、冲突解决和时间定位的社会推理任务上提升显著。
原文摘要 · Abstract (English)
Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence. This gap is especially pronounced in social video question answering, where the answer may hinge on a gesture, vocal tone, temporal cue, or mismatch between what is said and what is visually expressed. We introduce CogniRoute, a schema-guided Mixture-of-Experts framework for social omni reasoning. CogniRoute uses a training-only cognitive schema that factorizes each example by cross-modal relation, reasoning demand, and temporal scope, and aligns global routing signatures with this structure during supervised fine-tuning. We further introduce route-aware reinforcement learning, which jointly optimizes token generation and expert allocation using rewards for answer correctness, modality-consistent reasoning, and cognitive temporal grounding. To support training and evaluation, we construct OmniSocialBench, a diagnostic social video QA resource with 118K structured training examples, grounded reasoning traces, schema labels, temporal evidence spans, and a manually verified evaluation split. CogniRoute achieves 59.38\% average accuracy on OmniSocialBench, improving over the strongest proprietary baseline by 15.33 percentage points and the strongest open-source omni baseline by 26.77 points, with the largest gains on questions requiring audio-visual coordination, conflict resolution, and temporally grounded social inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。