构建首个跨语言跨模态幻觉检测基准,评估大模型在多场景下的可靠性。
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
- 设计联合跨语言与跨模态的幻觉检测任务,覆盖双维度挑战。
- 测试主流开源与闭源模型均表现不佳,说明现有技术仍不成熟。
- 适合研究大模型鲁棒性、多模态应用落地的研究者使用。
探究大语言模型(LLMs)在跨语言和跨模态场景中的幻觉问题,对推动其在真实世界应用的大规模部署具有重要意义。然而,现有研究仅聚焦单一场景(跨语言或跨模态),缺乏对二者联合场景下幻觉的系统探索。为此,我们提出首个联合跨语言与跨模态幻觉检测基准CCHall,同时涵盖跨语言与跨模态幻觉场景,可用于评估大模型在多维情境下的能力。我们对主流开源与闭源大模型进行了全面评估,结果表明当前模型在CCHall上仍面临显著挑战。我们希望CCHall能成为评估大模型在联合跨语言跨模态场景中表现的重要资源。
原文摘要 · Abstract (English)
Investigating hallucination issues in large language models (LLMs) within cross-lingual and cross-modal scenarios can greatly advance the large-scale deployment in real-world applications. Nevertheless, the current studies are limited to a single scenario, either cross-lingual or cross-modal, leaving a gap in the exploration of hallucinations in the joint cross-lingual and cross-modal scenarios. Motivated by this, we introduce a novel joint Cross-lingual and Cross-modal Hallucinations benchmark (CCHall) to fill this gap. Specifically, CCHall simultaneously incorporates both cross-lingual and cross-modal hallucination scenarios, which can be used to assess the cross-lingual and cross-modal capabilities of LLMs. Furthermore, we conduct a comprehensive evaluation on CCHall, exploring both mainstream open-source and closed-source LLMs. The experimental results highlight that current LLMs still struggle with CCHall. We hope CCHall can serve as a valuable resource to assess LLMs in joint cross-lingual and cross-modal scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。