用视觉信息增强音频推理,提升模型判断准确率
VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

- 融合音视频多模态特征,补充音频中的线索
- 通过投票与一致性检查,稳定推理结果
- 适合需要高精度音频理解的智能系统研发
音频推理需对时序动态、声学混叠信号进行多步、基于证据的推理,超出传统语音识别或图像描述等感知任务。我们提出VISA,作为Interspeech 2026音频推理挑战(代理赛道)的参赛系统,依据MMAR评分标准评估正确性与推理质量。在‘将大音频语言模型作为工具’范式下,VISA通过辅助多模态证据强化大音频语言模型,避免复杂调度。系统包含三个模块:用于互补音视频线索的多模态特征提取、基于一致性检查的模型投票推理,以及细粒度类别感知路由以解决分歧并选择符合评分标准的推理链。在官方代理赛道排行榜中,VISA总分位列第二,获得66.23%的MMAR评分;同时取得77.40%的准确率,是单模型与代理赛道所有系统中的最高值。
原文摘要 · Abstract (English)
Audio reasoning requires multi-step, evidence-grounded inference over temporally dynamic and acoustically mixed signals, exceeding conventional perception tasks such as ASR or captioning. We present VISA, our submission to the Interspeech 2026 Audio Reasoning Challenge (Agent Track), evaluated via the MMAR Rubrics for correctness and reasoning quality. Under a "LALM as a Tool" paradigm, VISA strengthens large audio language models with auxiliary multi-modal evidence while avoiding heavy orchestration. The system integrates three components: multi-modal feature extraction for complementary audio and acoustic-visual clues, model-voting inference with consistency checking for stable predictions, and fine-grained category-aware routing to resolve disagreements and select rubric-aligned reasoning chains. On the official Agent Track leaderboard, VISA ranks 2nd overall with a 66.23% Rubrics score. It also achieves 77.40% Accuracy, the highest among all systems listed across both the Single Model and Agent tracks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。