从罗马尼亚版《谁想成为百万富翁》视频中构建文化丰富语料库,揭示模型在本土文化题上表现显著下降
A Culturally-Rich Romanian NLP Dataset from "Who Wants to Be a Millionaire?" Videos
- 通过OCR+人工校验提取游戏节目问答对,标注领域、文化属性和难度
- 罗马尼亚本土文化题准确率仅50-75%,国际题达80-95%
- 适合关注多语言模型文化偏见与教育应用的研究者
大型语言模型在不同语言和文化背景下的表现差异显著。本研究基于罗马尼亚电视游戏节目《谁想成为百万富翁?》(Vrei să fii Milionar?)的视频数据,构建了一个新型、富含文化信息的多语言语料库。采用光学字符识别(OCR)、自动化文本提取与人工验证相结合的方法,收集问题-答案对,并添加领域(如生物、历史)、文化相关性(罗马尼亚特有或国际性)及难度等元数据。在该语料库上对主流大模型(包括罗马尼亚语适配模型)进行基准测试,发现模型在国际性问题上的准确率普遍达到80%-95%,而在罗马尼亚本土文化问题上仅为50%-75%。进一步通过将罗马尼亚语问题翻译为英语以及使用法语类似数据集进行跨语言测试,探究了这一性能差异。结果表明,文化背景与数据来源显著影响模型表现,为构建更具文化敏感性的多语言自然语言处理系统提供了实践指导,尤其适用于教育领域。该数据集已公开于Hugging Face。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate varying performance across languages and cultural contexts. This study introduces a novel, culturally-rich, multilingual dataset derived from video recordings of the Romanian game show "Who Wants to Be a Millionaire?" (Vrei să fii Milionar?). We employed an innovative process combining optical character recognition (OCR), automated text extraction, and manual verification to collect question-answer pairs, enriching them with metadata including question domain (e.g., biology, history), cultural relevance (Romanian-specific vs. international), and difficulty. Benchmarking state-of-the-art LLMs, including Romanian-adapted models, on this dataset revealed significant performance disparities: models consistently achieve higher accuracy (80-95%) on international questions compared to Romanian-specific cultural questions (50-75%). We further investigate these differences through experiments involving machine translation of Romanian questions into English and cross-lingual tests using a comparable dataset in French. Our findings underscore the impact of cultural context and data source on LLM performance and offer practical insights for building robust, culturally-aware multilingual NLP systems, especially in educational domains. The dataset is publicly available at Hugging Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。