arXiv:2508.07295cs.CL2025-08AAAI被引 6

构建多语言多模态事实性评估基准,助力提升语音文本模型可靠性

CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation

  • 设计跨语言跨模态平行语音文本问答数据集,覆盖8种语言
  • 现有大模型在多语言语音事实性任务上表现仍不理想
  • 仅用5个示例即可实现英文LLM能力向多语言语音问答迁移

随着大语言模型在多语言环境中的普及,确保其生成内容无幻觉、保持事实准确性变得至关重要。然而,现有用于评估多模态大模型(MLLMs)可靠性的基准主要集中在英文文本或视觉模态,缺乏对多语言输入尤其是语音模态的系统评估。为此,我们提出首个跨语言跨模态事实性评估基准CCFQA,包含8种语言的平行语音-文本事实性问答对,可系统评估MLLM在跨语言与跨模态场景下的事实性能力。实验表明,当前主流MLLM在该基准上仍面临显著挑战。我们进一步提出一种少样本迁移学习策略,仅需5个示例即可将英文大模型的问答能力有效迁移到多语言语音问答任务中,性能接近GPT-4o-mini-Audio。我们已开源代码与数据集(https://github.com/yxduir/ccfqa),以推动具备更强语音理解能力的可信多模态大模型发展。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual modalities with a primary emphasis on English, which creates a gap in evaluation when processing multilingual input, especially in speech. To bridge this gap, we propose a novel Cross-lingual and Cross-modal Factuality benchmark (CCFQA). Specifically, the CCFQA benchmark contains parallel speech-text factual questions across 8 languages, designed to systematically evaluate MLLMs' cross-lingual and cross-modal factuality capabilities. Our experimental results demonstrate that current MLLMs still face substantial challenges on the CCFQA benchmark. Furthermore, we propose a few-shot transfer learning strategy that effectively transfers the Question Answering (QA) capabilities of LLMs in English to multilingual Spoken Question Answering (SQA) tasks, achieving competitive performance with GPT-4o-mini-Audio using just 5-shot training. We release CCFQA as a foundational research resource to promote the development of MLLMs with more robust and reliable speech understanding capabilities. Our code and dataset are available at https://github.com/yxduir/ccfqa.

多语言语音问答事实性评估少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。