测试翻译对复合危害的跨语言传递效果,发现印地语等语言中攻击成功率显著上升。
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
- 构建多语言基准CompositeHarm,融合结构化攻击与真实场景危害数据
- 印地语等印地语系语言中攻击成功率大幅上升,达原英语水平的2.3倍
- 采用轻量化推理策略,兼顾效率与跨语言评估保真度
当前大语言模型的安全评估仍以英语为主。翻译常被用作探测多语言行为的捷径,但难以全面反映有害意图或结构在语言间的演变。部分危害在翻译后几乎保持不变,而另一些则发生扭曲或消失。为此,我们提出CompositeHarm——一个基于翻译的基准,用于考察安全对齐在语法与语义双重变化下的表现。该基准融合AttaQ(结构化对抗攻击)与MMSafetyBench(情境性真实危害)两个英文数据集,并扩展至英语、印地语、阿萨姆语、马拉地语、卡纳达语和古吉拉特语共六种语言。使用三个大模型进行测试发现,在对抗性语法下,印地语等印地语系语言中的攻击成功率显著上升,最高达英语基线的2.3倍;而情境性危害转移则相对温和。为提升可扩展性与能效,研究采用受边缘智能启发的轻量级推理策略,减少冗余评估步骤,同时保持跨语言一致性。该设计使大规模多语言安全测试在计算上可行且环境友好。结果表明,翻译基准是必要起点,但不足以构建真正适应语言差异的安全系统。
原文摘要 · Abstract (English)
Most safety evaluations of large language models (LLMs) remain anchored in English. Translation is often used as a shortcut to probe multilingual behavior, but it rarely captures the full picture, especially when harmful intent or structure morphs across languages. Some types of harm survive translation almost intact, while others distort or disappear. To study this effect, we introduce CompositeHarm, a translation-based benchmark designed to examine how safety alignment holds up as both syntax and semantics shift. It combines two complementary English datasets, AttaQ, which targets structured adversarial attacks, and MMSafetyBench, which covers contextual, real-world harms, and extends them into six languages: English, Hindi, Assamese, Marathi, Kannada, and Gujarati. Using three large models, we find that attack success rates rise sharply in Indic languages, especially under adversarial syntax, while contextual harms transfer more moderately. To ensure scalability and energy efficiency, our study adopts lightweight inference strategies inspired by edge-AI design principles, reducing redundant evaluation passes while preserving cross-lingual fidelity. This design makes large-scale multilingual safety testing both computationally feasible and environmentally conscious. Overall, our results show that translated benchmarks are a necessary first step, but not a sufficient one, toward building grounded, resource-aware, language-adaptive safety systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。