arXiv:2602.01161cs.CL2026-02被引 1

研究发现,语料的词汇特性比语义复杂度更影响大模型文化表现。

Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models

  • 通过主成分分析提取语料的词汇、句法和语义结构特征
  • 词汇多样性成分在跨模型测试中表现最稳定,提升文化理解力
  • 过度强调语义连贯或风格多样反而可能降低文化适应性

大语言模型在全球部署中面临文化错位问题,但用于文化适配的微调语料的语言特性尚不明确。本文从语料中心视角出发,探究不同语言(阿拉伯语、中文、日语)语料的哪些语言属性与文化表现相关,是否能在训练前预测性能,以及这些效应在不同模型间的差异。对三类语言语料计算轻量级语言学、语义和结构指标,并分别进行主成分分析,确保成分反映同语言内部差异而非跨语言差异。结果发现,主成分对应可解释的轴线:语义连贯性、表面词汇与句法多样性、词汇或结构丰富性,其构成随语言而异。在三个主流模型族(LLaMA、Mistral、DeepSeek)上进行微调并评估文化知识、价值观与规范基准表现。结果显示,主成分与下游性能相关,但关联性强依赖于模型。通过控制子集干预证明,以词汇为导向的成分(PC3)最具鲁棒性,在多个模型与基准间表现一致;而强调语义或多样性极端(PC1-PC2)往往效果中性甚至有害。

原文摘要 · Abstract (English)

The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adaptation remain poorly understood. We adopt a dataset-centric view of cultural alignment and ask which linguistic properties of fine-tuning data are associated with cultural performance, whether these properties are predictive prior to training, and how these effects vary across models. We compute lightweight linguistic, semantic, and structural metrics for Arabic, Chinese, and Japanese datasets and apply principal component analysis separately within each language. This design ensures that the resulting components capture variation among datasets written in the same language rather than differences between languages. The resulting components correspond to broadly interpretable axes related to semantic coherence, surface-level lexical and syntactic diversity, and lexical or structural richness, though their composition varies across languages. We fine-tune three major LLM families (LLaMA, Mistral, DeepSeek) and evaluate them on benchmarks of cultural knowledge, values, and norms. While PCA components correlate with downstream performance, these associations are strongly model-dependent. Through controlled subset interventions, we show that lexical-oriented components (PC3) are the most robust, yielding more consistent performance across models and benchmarks, whereas emphasizing semantic or diversity extremes (PC1-PC2) is often neutral or harmful.

文化对齐语料分析大模型语言结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。