arXiv:2507.22752cs.CL2025-07Transactions of th…被引 4

构建多模态区域问答数据集,评估大模型在本地知识上的回答能力

CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset

  • 基于三国民众编写的中英文双语问题,融合文本与视觉理解任务
  • 顶尖开源模型在纯文本题上准确率超40%,视觉题低于30%
  • 发现传统字符串匹配法因命名实体多而表现意外出色

我们提出 CUS-QA,一个用于评估开放型区域问答的基准数据集,涵盖文本和视觉两种模态。数据集由捷克、斯洛伐克和乌克兰的母语者手工创建,基于 Wikipedia 内容,并提供英文翻译。问题包括纯文本类和需视觉理解的类型。我们通过提示工程评估前沿大语言模型(LLMs),并引入人工判断答案正确性。基于人工评估,分析现有自动评价指标的可靠性。基线结果显示,即使最佳开源模型在文本类问题上准确率也仅超过40%,视觉类问题则低于30%。基于 LLM 的评价指标与人工判断高度相关,而传统字符串重叠指标表现意外良好,可能源于答案中大量出现命名实体。

原文摘要 · Abstract (English)

We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines using state-of-the-art large language models (LLMs). Our dataset consists of manually curated questions and answers grounded in Wikipedia, created by native speakers from Czechia, Slovakia, and Ukraine, with accompanying English translations. It includes both purely textual questions and those requiring visual understanding. We evaluate state-of-the-art LLMs through prompting and add human judgments of answer correctness. Using these human evaluations, we analyze the reliability of existing automatic evaluation metrics. Our baseline results show that even the best open-weight LLMs achieve only over 40% accuracy on textual questions and below 30% on visual questions. LLM-based evaluation metrics show strong correlation with human judgment, while traditional string-overlap metrics perform surprisingly well due to the prevalence of named entities in answers.

问答系统多模态区域知识大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。