多语言模型过度自信,用户易盲目信任,存在全球安全风险。
Humans overrely on overconfident language models, across languages
- 分析模型生成的不确定性表达,发现其在五种语言中均表现过度自信。
- 跨语言测试显示,用户对不确定表达的依赖率差异显著,日语使用者更易忽视谨慎表述。
- 强调需结合文化语言背景评估模型安全性,避免统一标准带来的风险。
随着大语言模型在全球范围部署,其跨语言的不确定性表达校准至关重要。已有研究指出,英语中的大型语言模型存在语言过度自信问题,导致用户过度依赖其生成内容。然而,不同语言对认知标记(如‘我认为’)的使用和解读存在显著差异。本文研究了五种语言中多语言语言(误)校准、过度自信及过度依赖的风险,评估模型在全局环境下的安全性。研究发现,跨语言过度依赖风险较高:模型在多种语言中均频繁生成强化词,即使在错误回答中也是如此。同时,模型对语言间差异敏感——例如,在日语中生成最多不确定标记,而在德语和中文中生成最多确定性标记。进一步的人类依赖实验表明,跨语言依赖行为存在显著差异:日本参与者比英语使用者更倾向于忽略不确定表达的缓和功能,从而更可能依赖含此类表达的生成结果。综合来看,跨语言环境下对过度自信模型生成内容的依赖风险极高。研究揭示了多语言语言校准的挑战,并强调必须进行文化和语言情境化的模型安全评估。
原文摘要 · Abstract (English)
As large language models (LLMs) are deployed globally, it is crucial that their responses are calibrated across languages to accurately convey uncertainty and limitations. Prior work shows that LLMs are linguistically overconfident in English, leading users to overrely on confident generations. However, the usage and interpretation of epistemic markers (e.g., 'I think it's') differs sharply across languages. Here, we study the risks of multilingual linguistic (mis)calibration, overconfidence, and overreliance across five languages to evaluate LLM safety in a global context. Our work finds that overreliance risks are high across languages. We first analyze the distribution of LLM-generated epistemic markers and observe that LLMs are overconfident across languages, frequently generating strengtheners even as part of incorrect responses. Model generations are, however, sensitive to documented cross-linguistic variation in usage: for example, models generate the most markers of uncertainty in Japanese and the most markers of certainty in German and Mandarin. Next, we measure human reliance rates across languages, finding that reliance behaviors differ cross-linguistically: for example, participants are significantly more likely to discount expressions of uncertainty in Japanese than in English (i.e., ignore their 'hedging' function and rely on generations that contain them). Taken together, these results indicate a high risk of reliance on overconfident model generations across languages. Our findings highlight the challenges of multilingual linguistic calibration and stress the importance of culturally and linguistically contextualized model safety evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。