arXiv:2512.09066cs.SDcs.AI2025-12Transactions of th…

构建可信赖的音频问答答案评估框架,支持开放回答与争议检测。

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

  • 三阶段标注流程融合人工判断与人机协同校正
  • 在3699组问答上实现0.91的评判相关性与0.85泛化能力
  • 能识别人类分歧高的问题,适合模型评测与数据清洗

大音频语言模型(LALMs)能力的可靠评估对推动技术进步至关重要。随着基准测试日益包含复杂推理与主观任务,模型需生成开放回答。本文提出开放回答正确性评估(ORCA)——一种轻量且可靠的基于模型的评估方法,用于判断答案正确性并建模人类分歧。通过结合人工判断、结构化反馈与人机校正的三阶段标注流程,我们在三个音频理解与推理基准上收集了9,663条标注,覆盖15个LALMs,Krippendorff's alpha达0.82。实验表明,采用课程学习的ORCA模型在已见基准上与平均人类评分的斯皮尔曼相关性为0.91,在未见基准上仍保持0.85,优于包括Gemini 2.5 Flash在内的多个大语言模型判官基线。此外,ORCA预测的方差与人类分歧高度相关,可有效识别有问题的评测条目。

原文摘要 · Abstract (English)

Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increasingly necessitate open-ended responses from LALMs. We present Open-ended Response Correctness Assessment (ORCA) -- a reliable and lightweight model-based approach for answer correctness and disagreement modeling. We employ a three-stage annotation pipeline combining human judgment, structured feedback, and human-AI correction, yielding 9,663 annotations across 3,699 question-answer pairs from 15 LALMs on three audio understanding and reasoning benchmarks (achieving a Krippendorff's alpha of 0.82). Our experiments employing curriculum learning show that ORCA models achieve a Spearman correlation of 0.91 with average human correctness ratings on seen benchmarks and generalize to unseen benchmarks with a score of 0.85, outperforming several LLM judge baselines including Gemini 2.5 Flash. Furthermore, we demonstrate that ORCA's predicted variance correlates strongly with human disagreement, allowing it to effectively identify problematic benchmark items.

音频理解模型评估开放回答人类分歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。