用多源证据融合提升语音问答推理可信度
Multi-Source Evidence Fusion for Audio Question Answering
- 双大模型生成独立观察,文本模型交叉验证25个声学工具输出
- 每步推理均附带可靠度标签,推理链更完整可验证
- 在语音推理挑战赛中排名第一,推理质量远超对手
大型音频语言模型(LALMs)能回答关于语音、音乐和环境声音的问题,但其内部推理过程高度不透明,难以验证。本文介绍TalTech在Interspeech 2026语音推理挑战赛代理赛道的解决方案,系统评估聚焦于推理过程质量,包括事实准确性、逻辑严谨性和推理链完整性。我们的多源集成流水线使用两个LALM生成独立观测,一个独立的文本推理模型则将这些结果与25个分属不同可靠度层级的声学工具输出进行交叉验证。通过将每个推理步骤锚定在明确且标注可靠度的证据上,系统生成密集且可验证的推理链。该系统在挑战赛中排名第一,其推理质量指标显著优于所有竞争系统。
原文摘要 · Abstract (English)
Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We describe TalTech's solution to the Agent Track of the Interspeech 2026 Audio Reasoning Challenge, in which systems are evaluated on reasoning process quality, specifically the factual accuracy, logical soundness, and completeness of their reasoning chains. Our multi-source ensemble pipeline uses two LALMs that generate independent observations, while a separate text-only reasoning model cross-checks these against outputs from 25 acoustic tools organized into reliability tiers. By grounding every inference step in explicit, reliability-tagged evidence, the system produces dense, verifiable reasoning chains. Our system ranked first in the challenge, outperforming all competing systems by a wide margin in challenge's reasoning quality metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。