arXiv:2604.17718cs.CLcs.SI2026-04

大模型能隐含理解文化并调整说话方式,但效果有限。

Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation

论文配图:Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
图 1 · 摘自论文原文
  • 通过情境线索测试模型是否能自动适应文化语用
  • 仅约19.6%的语用变化能在无明示时复现
  • 对权威态度的适应最强,对群体个体框架最弱

现有基准多测试大模型对文化知识的直接问答能力。本文研究另一问题:当文化仅通过情境暗示时,模型是否会调整表达方式?我们在五种语言(英语、德语、印地语、尼泊尔语、乌尔都语)中评估了60个与文化相关的对话场景,三种条件:中性提示(Prompt A)、显式文化指令(Prompt B)、隐含情境提示(Prompt C)。评分涵盖12项语用特征,包括对权威的尊重、个体与集体框架、不确定性管理等。定义语用情境敏感度(PCS)为从提示A→B的语用变化中,在提示A→C下重现的比例。在四个部署的LLM上,平均PCS为0.196(标准差=0.113),表明模型仅能恢复约五分之一的显式指令下的语用变化。对权威相关线索的迁移最强(PCS=0.299),对个体-群体框架最弱(PCS=0.120)。不确定性行为表现不一:所有五种语言中,修饰密度的显式差距为负值,说明对齐训练反而抑制了目标行为。由于印地语和乌尔都语共享核心语法但对应不同文化群体,我们以此为自然对照;配对分析未发现基线差异(t=0.96, p=0.339, dz=0.06),表明模型主要响应语言结构而非语言所承载的文化关联。结论:多语言文化语用是显式与隐式使用的问题,而不仅是事实知识问题。

原文摘要 · Abstract (English)

Many benchmarks show that large language models can answer direct questions about culture. We study a different question: do they also change how they speak when culture is only implied by the situation? We evaluate 60 culturally grounded conversational scenarios across five languages in three conditions: a neutral baseline (Prompt A), an explicit cultural instruction (Prompt B), and implicit situational cueing (Prompt C). We score responses on 12 pragmatic features covering deference to authority, individual-versus-group framing, and uncertainty management. We define Pragmatic Context Sensitivity (PCS) as the fraction of the Prompt A->B shift that reappears under Prompt A->C. Across four deployed LLMs and five languages (English, German, Hindi, Nepali, Urdu), the primary stable-only PCS mean is 0.196 (SD = 0.113), indicating that the models recover only about one-fifth of the pragmatic shift they can produce when instructed explicitly. Transfer is strongest for authority-related cues (0.299) and weakest for individual-versus-group framing (0.120). Uncertainty-related behaviour is mixed: hedging density exhibits negative explicit gaps in all five languages, suggesting that alignment training actively suppresses the target behaviour. Because Hindi and Urdu share core grammar yet index distinct cultural communities, we use them as a natural control; a paired analysis finds no reliable baseline difference (t = 0.96, p = 0.339, dz = 0.06), suggesting that models respond primarily to linguistic structure rather than to the cultural associations a language carries. We argue that multilingual cultural pragmatics is an explicit-versus-implicit deployment problem, not only a factual knowledge problem.

大模型文化理解语用适应多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。