探明大模型在学术写作中是否编码英中文化差异,发现隐藏层可精准识别国籍特征。
Nationality encoding in language model hidden states: Probing culturally differentiated representations in persona-conditioned academic text
- 用逻辑回归探测隐藏状态,发现第18层对国籍分类准确率达96.8%
- 英式写作多用被动语态和评价性词汇,中式写作倾向名词化和国际表达
- 适合语言教学与跨文化写作研究者参考
大型语言模型在英语学术写作(EAP)中的应用日益广泛,但其生成文本是否蕴含文化差异仍不明确。本研究检验了Gemma-3-4b-it在英国与中文学术人设下生成研究论文引言时,隐藏状态中是否编码国籍区分信息。基于45个提示模板与6种人设构成的2×3实验设计,生成270篇文本。在全部35层隐藏状态上训练逻辑回归探测器,并设置随机标签基线、表面文本基准、跨模型测试及句级基线作为对照。使用Stanza NLP工具包标注探测到的关键词位置,分析其结构、词汇与立场特征。结果显示,国籍探测器在第18层达到0.968的交叉验证准确率,且在保留集上实现完全分类。国籍编码呈现非单调层间变化:结构影响集中于中上层网络,词汇领域效应更早显现。高信号词位上,英式模式表现出更多后置修饰、缓和、强化、被动语态及评价性/过程导向词汇;中式模式则更多前置修饰、名词谓语及社会文化或国际化词汇。然而,句级分析未发现表面文本存在显著国籍差异。研究拓展了探测方法在社会语言学属性上的应用,对EAP与语言教学具实践意义。
原文摘要 · Abstract (English)
Large language models are increasingly used as writing tools and pedagogical resources in English for Academic Purposes, but it remains unclear whether they encode culturally differentiated representations when generating academic text. This study tests whether Gemma-3-4b-it encodes nationality-discriminative information in hidden states when generating research article introductions conditioned by British and Chinese academic personas. A corpus of 270 texts was generated from 45 prompt templates crossed with six persona conditions in a 2 x 3 design. Logistic regression probes were trained on hidden-state activations across all 35 layers, with shuffled-label baselines, a surface-text skyline classifier, cross-family tests, and sentence-level baselines used as controls. Probe-selected token positions were annotated for structural, lexical, and stance features using the Stanza NLP pipeline. The nationality probe reached 0.968 cross-validated accuracy at Layer 18, with perfect held-out classification. Nationality encoding followed a non-monotonic trajectory across layers, with structural effects strongest in the middle to upper network and lexical-domain effects peaking earlier. At high-signal token positions, British-associated patterns showed more postmodification, hedging, boosting, passive voice, and evaluative or process-oriented vocabulary, while Chinese-associated patterns showed more premodification, nominal predicates, and sociocultural or internationalisation vocabulary. However, sentence-level analysis found no significant nationality differences in the full generated surface text. The findings extend probing methodology to a sociolinguistic attribute and have practical implications for EAP and language pedagogy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。