arXiv:2512.00274cs.CL2025-12

测试大模型对印度方言情绪表达的理解能力,发现直接处理本地语言效果极差。

Lost without translation -- Can transformer (language models) understand mood states?

  • 用直接嵌入法处理印地语等11种本地语言,结果几乎无法区分情绪状态。
  • 翻译成英文或中文后嵌入,性能显著提升,最高得分达0.67(人类翻译+中文模型)。
  • 现有专用印地语模型表现不佳,说明需构建真正理解本地语言的模型。

背景:大型语言模型在精神健康领域有潜力,但主要基于英语。它们对其他语言中情绪表达的理解能力尚不明确,因不同语言有各自的痛苦表述方式。目的:量化语言模型对四种情绪状态(抑郁、正常、欣快躁狂、忧郁躁狂)在11种印地语系语言中的独特表达(痛苦习语)的忠实表征能力。方法:收集247个独特短语,测试七种条件:(a) 使用多语言和印地语专用模型直接嵌入原生或罗马化文本;(b) 嵌入英译或中译文本。性能以调整兰德指数、标准化互信息、同质性与完整性组成的综合得分衡量。结果:直接嵌入印地语语言表现极差(综合得分=0.002)。所有翻译方法均有显著提升,其中使用Gemini翻译的英文嵌入(gemini-001模型)得分为0.60,人工翻译英文嵌入得分为0.61。令人意外的是,人工翻译英文后再转为中文,用中文模型嵌入,得分最高(0.67)。专用印地语模型(IndicBERT、Sarvam-M)表现较差。结论:当前模型无法有效理解印地语系语言中的情绪状态,构成在印度进行诊断或治疗应用的根本障碍。高质量翻译可暂时弥补缺陷,但依赖专有模型或复杂翻译流程不可持续。模型必须首先具备对多样化本地语言的理解能力,才能在心理健康领域实现全球应用。

原文摘要 · Abstract (English)

Background: Large Language Models show promise in psychiatry but are English-centric. Their ability to understand mood states in other languages is unclear, as different languages have their own idioms of distress. Aim: To quantify the ability of language models to faithfully represent phrases (idioms of distress) of four distinct mood states (depression, euthymia, euphoric mania, dysphoric mania) expressed in Indian languages. Methods: We collected 247 unique phrases for the four mood states across 11 Indic languages. We tested seven experimental conditions, comparing k-means clustering performance on: (a) direct embeddings of native and Romanised scripts (using multilingual and Indic-specific models) and (b) embeddings of phrases translated to English and Chinese. Performance was measured using a composite score based on Adjusted Rand Index, Normalised Mutual Information, Homogeneity and Completeness. Results: Direct embedding of Indic languages failed to cluster mood states (Composite Score = 0.002). All translation-based approaches showed significant improvement. High performance was achieved using Gemini-translated English (Composite=0.60) and human-translated English (Composite=0.61) embedded with gemini-001. Surprisingly, human-translated English, further translated into Chinese and embedded with a Chinese model, performed best (Composite = 0.67). Specialised Indic models (IndicBERT and Sarvam-M) performed poorly. Conclusion: Current models cannot meaningfully represent mood states directly from Indic languages, posing a fundamental barrier to their psychiatric application for diagnostic or therapeutic purposes in India. While high-quality translation bridges this gap, reliance on proprietary models or complex translation pipelines is unsustainable. Models must first be built to understand diverse local languages to be effective in global mental health.

精神健康多语言情绪识别语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。