arXiv:2509.11303cs.CL2025-09被引 1

构建韩语物理常识数据集,融入韩式文化元素提升模型理解力

Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context

  • 从海量网页筛选出1.1万道韩语常识题,经多轮AI与人工精修得441组高质量数据
  • 83.22%最佳模型准确率仍远低于人类水平,文化相关题型错误率超40%
  • 含19.7%本土文化元素如泡菜、韩服、泡菜冰箱,推动更具文化敏感性的研究

现有物理常识推理数据集如PIQA主要基于英语,缺乏文化多样性。本文提出Ko-PIQA,一个融合韩国文化背景的韩语物理常识推理数据集。从301万条网络爬取问题出发,通过三阶段过滤流程(使用三个语言模型)筛选出11,553道符合PIQA风格的问题。经GPT-4o优化及人工验证后,最终获得441个高质量问答对。其核心特点是文化嵌入性:19.7%的问题包含传统韩国食物(如泡菜)、服饰(韩服)及专用电器(泡菜冰箱)等文化特异性元素,需超越直译的文化理解能力。在该数据集上评估七种语言模型,最佳模型准确率达83.22%,最差仅59.86%,表明模型在文化相关场景下表现显著不足,凸显文化多样性数据集的重要性。Ko-PIQA可作为韩语模型基准,并为更包容的常识推理研究奠定基础。数据集与代码将公开。

原文摘要 · Abstract (English)

Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning dataset that incorporates cultural context. Starting from 3.01 million web-crawled questions, we employed a multi-stage filtering approach using three language models to identify 11,553 PIQA-style questions. Through GPT-4o refinement and human validation, we obtained 441 high-quality question-answer pairs. A key feature of Ko-PIQA is its cultural grounding: 19.7% of questions contain culturally specific elements like traditional Korean foods (kimchi), clothing (hanbok), and specialized appliances (kimchi refrigerators) that require culturally-aware reasoning beyond direct translation. We evaluate seven language models on Ko-PIQA, with the best model achieving 83.22% accuracy while the weakest reaches only 59.86%, demonstrating significant room for improvement. Models particularly struggle with culturally specific scenarios, highlighting the importance of culturally diverse datasets. Ko-PIQA serves as both a benchmark for Korean language models and a foundation for more inclusive commonsense reasoning research. The dataset and code will be publicly available.

常识推理韩语数据文化感知多语言NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。