构建韩语物理常识数据集,融入韩式文化元素提升模型理解力
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
- 从海量网页筛选出1.1万道韩语常识题,经多轮AI与人工精修得441组高质量数据
- 83.22%最佳模型准确率仍远低于人类水平,文化相关题型错误率超40%
- 含19.7%本土文化元素如泡菜、韩服、泡菜冰箱,推动更具文化敏感性的研究
现有物理常识推理数据集如PIQA主要基于英语,缺乏文化多样性。本文提出Ko-PIQA,一个融合韩国文化背景的韩语物理常识推理数据集。从301万条网络爬取问题出发,通过三阶段过滤流程(使用三个语言模型)筛选出11,553道符合PIQA风格的问题。经GPT-4o优化及人工验证后,最终获得441个高质量问答对。其核心特点是文化嵌入性:19.7%的问题包含传统韩国食物(如泡菜)、服饰(韩服)及专用电器(泡菜冰箱)等文化特异性元素,需超越直译的文化理解能力。在该数据集上评估七种语言模型,最佳模型准确率达83.22%,最差仅59.86%,表明模型在文化相关场景下表现显著不足,凸显文化多样性数据集的重要性。Ko-PIQA可作为韩语模型基准,并为更包容的常识推理研究奠定基础。数据集与代码将公开。
原文摘要 · Abstract (English)
Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning dataset that incorporates cultural context. Starting from 3.01 million web-crawled questions, we employed a multi-stage filtering approach using three language models to identify 11,553 PIQA-style questions. Through GPT-4o refinement and human validation, we obtained 441 high-quality question-answer pairs. A key feature of Ko-PIQA is its cultural grounding: 19.7% of questions contain culturally specific elements like traditional Korean foods (kimchi), clothing (hanbok), and specialized appliances (kimchi refrigerators) that require culturally-aware reasoning beyond direct translation. We evaluate seven language models on Ko-PIQA, with the best model achieving 83.22% accuracy while the weakest reaches only 59.86%, demonstrating significant room for improvement. Models particularly struggle with culturally specific scenarios, highlighting the importance of culturally diverse datasets. Ko-PIQA serves as both a benchmark for Korean language models and a foundation for more inclusive commonsense reasoning research. The dataset and code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。