研究发现,用本地语言提示会降低大模型对印度口头传统的忠实度。
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

- 用句子嵌入和余弦相似度量化三种印度叙事传统在大模型中的表现差异
- 跨传统相似度高达0.52-0.66,显示文化同质化趋势明显
- 本地语言提示反而降低忠实度,最高下降27个百分点
大型语言模型主要基于英语互联网文本训练,可能过度代表某些文化叙事,导致非西方故事传统被压缩为单一同质化原型。本研究开展一项初步计算实验,考察印度三大截然不同的区域口头与文学传统:拉贾斯坦邦的帕布吉史诗、古典泰米尔桑加姆诗歌和孟加拉民间故事。我们为每种传统收集了真实参考语料(分别为11、21和10段),并使用两种LLM(Claude Sonnet和Gemini)针对每种传统生成54个样本,涵盖通用、文化特定和本土语言三类提示。通过Sentence-BERT嵌入与余弦相似度,测量参考漂移(输出与其自身传统的真实文本的贴近程度,相对于其他两个传统)和跨传统趋同性。结果表明,尽管输出仍更接近自身传统,但跨传统相似度(0.52–0.66)远高于各传统间真实距离所预测值,表明存在部分同质化。令人意外的是,使用本地语言(印地语、泰米尔语或孟加拉语)提示,反而比英文提示显著降低对真实传统的忠实度,降幅最高达27个百分点。我们讨论这一现象与以往多语言提示研究的矛盾,认为其反映了在激发普遍文化多样性与模拟少数、记录较少的口头传统之间的差异。本研究作为博士课题的一部分,可视为近期大规模人工标注研究的轻量级、可扩展补充。
原文摘要 · Abstract (English)
Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language. Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition's authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition's reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions' genuine distance would predict, indicating partial homogenisation. Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition. We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。