研究多语言大模型如何遵守指令优先级,发现语言差异会显著影响指令执行可靠性。
Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

- 构建跨语言指令优先级评测基准XIHBench,覆盖六种语言与多种冲突场景。
- 语言间冲突的指令遵守率高于同语言冲突,存在'语言边界效应'。
- 模型偏好的语言中低优先级指令更难被覆盖,带来安全风险,适合多语言部署研究者参考。
指令优先级(IH)要求模型按来源优先级排序指令,确保高优先级指令覆盖低优先级。尽管对安全可控部署至关重要,现有评估几乎仅限于英文,难以判断多语言环境下IH合规性是否稳定。我们提出XIHBench,一个涵盖六种语言、四个领域、三种IH设置的多语言IH评测基准,包含同语言与跨语言冲突。实验发现两个一致模式:第一,IH合规性存在明显的语言依赖不对称性——某语言在高优先级时提升合规性,却可能在低优先级时造成干扰;第二,跨语言冲突下的合规率高于同语言冲突,称为‘语言边界效应’。此外,语言专精使模型偏好语言中的低优先级指令更难被覆盖,引发多语言环境下的可靠性和安全风险。
原文摘要 · Abstract (English)
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。