arXiv:2510.27672cs.CL2025-10EMNLP被引 1

用用户与模型协作方式,挖掘大模型缺失的文化知识。

Culture Cartography: Mapping the Landscape of Cultural Knowledge

  • 模型主动提出低置信度问题,引导用户填补文化空白
  • 新数据使Llama-3.1-8B在文化基准上提升19.2%准确率
  • 适合做文化知识补全或跨文化AI训练的研究者

为服务全球用户,大语言模型需具备文化特定知识,但这些知识往往未在预训练中习得。如何发现对本族群体重要却未被模型掌握的知识?现有方法多为单向:研究者设计问题由用户被动回答(传统标注),或用户主动产出数据供研究者构建基准(知识提取)。我们提出混合倡议方法CultureCartography:由模型生成其置信度低的问题,显式暴露自身知识边界与盲区,再由人类用户直接编辑填补空白并引导话题方向。我们实现该方法为工具CultureExplorer。相比传统人类回答模型提问的基线,CultureExplorer更有效挖掘出DeepSeek R1和GPT-4o等主流模型仍缺失的知识,即使开启网络搜索亦然。以该数据微调Llama-3.1-8B,在相关文化评测中最高提升19.2%准确率。

原文摘要 · Abstract (English)

To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find such knowledge that is (1) salient to in-group users, but (2) unknown to LLMs? The most common solutions are single-initiative: either researchers define challenging questions that users passively answer (traditional annotation), or users actively produce data that researchers structure as benchmarks (knowledge extraction). The process would benefit from mixed-initiative collaboration, where users guide the process to meaningfully reflect their cultures, and LLMs steer the process towards more challenging questions that meet the researcher's goals. We propose a mixed-initiative methodology called CultureCartography. Here, an LLM initializes annotation with questions for which it has low-confidence answers, making explicit both its prior knowledge and the gaps therein. This allows a human respondent to fill these gaps and steer the model towards salient topics through direct edits. We implement this methodology as a tool called CultureExplorer. Compared to a baseline where humans answer LLM-proposed questions, we find that CultureExplorer more effectively produces knowledge that leading models like DeepSeek R1 and GPT-4o are missing, even with web search. Fine-tuning on this data boosts the accuracy of Llama-3.1-8B by up to 19.2% on related culture benchmarks.

文化知识混合协作模型补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。