arXiv:2509.06065cs.CL2025-09

首个菲律宾语版真实度评测,揭示大模型多语言可靠性差距

KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

  • 将TruthfulQA翻译为菲律宾语,构建多语言评测基准
  • 新模型在菲律宾语中表现优于旧模型,但整体仍低于英文水平
  • 部分问题类型在跨语言迁移中更易出错,提示需更全面评估

大型语言模型在各类任务中表现优异,但幻觉问题限制了其可靠应用。现有如TruthfulQA等评测主要面向英语,缺乏对低资源语言的评估。为此,我们提出KatotohananQA,即TruthfulQA的菲律宾语版本。采用二选一框架,评估了七款免费级商用模型。结果表明,英语与菲律宾语之间的真实度表现存在显著差距;较新的OpenAI模型(GPT-5和GPT-5 mini)展现出较强的多语言鲁棒性。同时,不同问题特征下表现差异明显,某些问题类型、类别和主题在跨语言迁移中更不稳健,凸显了开展更广泛多语言评估的重要性,以确保大模型使用的公平性与可靠性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are primarily available in English, leaving a gap in evaluating LLMs in low-resource languages. To address this, we present KatotohananQA, a Filipino translation of the TruthfulQA benchmark. Seven free-tier proprietary models were assessed using a binary-choice framework. Findings show a significant performance gap between English and Filipino truthfulness, with newer OpenAI models (GPT-5 and GPT-5 mini) demonstrating strong multilingual robustness. Results also reveal disparities across question characteristics, suggesting that some question types, categories, and topics are less robust to multilingual transfer which highlight the need for broader multilingual evaluation to ensure fairness and reliability in LLM usage.

多语言评估真实性评测低资源语言LLM幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。