通过几何错位自动挖掘文化种子,提升大模型的文化理解能力。
C-Mining: Unsupervised Discovery of Seeds for Cultural Data Synthesis via Geometric Misalignment

- 利用跨语言嵌入空间的几何错位识别文化特异性区域
- 无需人工或LLM标注,自动提取高保真文化点,成本降低150倍以上
- 适合需要大规模高质量文化数据合成的研究者和开发者
大语言模型的文化对齐日益依赖合成数据生成。其中最关键的初始步骤是种子的选取,但现有方法缺乏可量化的筛选标准。当前方法依赖不可扩展的人工筛选或存在偏见的LLM提取,将文化特异性视为抽象概念而非可度量信号。本文提出C-Mining,一种无监督框架,将文化种子发现从主观选择转化为可计算的数据挖掘问题。该方法利用预训练嵌入空间中文化概念的跨语言几何错位作为可量化的发现信号,系统识别具有显著语言排他性和几何隔离性的区域,并主动过滤噪声,从而无需人类或LLM监督,从原始多语言语料中自动提取高保真文化点(CPs),准备成本降低超过150倍。进一步利用挖掘的知识引导多样化指令微调数据集的生成。大量实验表明,该以种子为中心的方法显著提升文化理解和推理能力,在CulturalBench-Hard上提升6.03分,超越现有最优基线,为高质量文化数据合成提供了可扩展、可量化的解决方案。
原文摘要 · Abstract (English)
Achieving cultural alignment in Large Language Models (LLMs) increasingly depends on synthetic data generation. For such synthesis, the most vital initial step is seed curation; however, current methods lack quantifiable standards for selecting these seeds. Existing approaches rely on unscalable manual curation or bias-prone LLM extraction, treating cultural specificity as an abstract concept rather than a measurable signal. In this paper, we address this "quantification gap" by proposing C-Mining, an unsupervised framework that transforms the discovery of cultural seeds from a subjective selection process into a computable data mining formulation. Our approach exploits a novel geometric insight, leveraging the cross-lingual misalignment of cultural concepts within pre-trained embedding spaces as a quantifiable discovery signal. By systematically identifying these regions characterized by pronounced linguistic exclusivity and geometric isolation, while actively filtering out noise, C-Mining automatically extracts high-fidelity Culture Points (CPs) from raw multilingual corpora without reliance on human or LLM supervision, reducing preparation costs by more than 150-fold. We further leverage the mined knowledge to steer the synthesis of diverse instruction-tuning datasets. Extensive experiments demonstrate that this seed-centric approach significantly enhances cultural understanding and reasoning capabilities, achieving a +6.03 point improvement on CulturalBench-Hard and surpassing state-of-the-art baselines, providing a scalable, quantifiable solution for high-quality cultural data synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。