arXiv:2410.13237cs.CLcs.AI2024-10NAACL被引 8

提出量化语言混淆的熵值指标,揭示大模型在多语言生成中的脆弱性。

Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis

  • 引入语言混淆熵,基于语言类型学和词汇变异量化混淆程度
  • 发现不同大模型存在可复现的语言混淆模式,与安全漏洞相关
  • 为多语言对齐与防御攻击提供语言相似性先验思路

语言混淆指大语言模型生成的内容既非目标语言,也非语境恰当的语言,表现为不可预测的异常行为。我们假设该现象存在语言规律,并揭示了大模型间语言混淆的模式。提出一种新度量——语言混淆熵,基于语言类型学与词汇变异的语言分布来直接量化混淆程度。与语言混淆基准(Marchisio et al., 2024)的全面对比验证了该指标的有效性,揭示了大模型间语言混淆的共性模式。进一步将语言混淆与大模型安全关联,发现其在多语言嵌入反演攻击中呈现特定规律。分析表明,语言类型学提供了理论基础,有助于利用语言相似性作为先验知识,优化大模型对齐与安全防护。

原文摘要 · Abstract (English)

Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate language. This phenomenon presents a critical challenge in text generation by LLMs, often appearing as erratic and unpredictable behavior. We hypothesize that there are linguistic regularities to this inherent vulnerability in LLMs and shed light on patterns of language confusion across LLMs. We introduce a novel metric, Language Confusion Entropy, designed to directly measure and quantify this confusion, based on language distributions informed by linguistic typology and lexical variation. Comprehensive comparisons with the Language Confusion Benchmark (Marchisio et al., 2024) confirm the effectiveness of our metric, revealing patterns of language confusion across LLMs. We further link language confusion to LLM security, and find patterns in the case of multilingual embedding inversion attacks. Our analysis demonstrates that linguistic typology offers theoretically grounded interpretation, and valuable insights into leveraging language similarities as a prior for LLM alignment and security.

大模型安全语言混淆语言类型学量化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。