arXiv:2602.13867cs.CL2026-02中稿 · AAAI被引 2

为南半球语言构建高效文化适配的安全对齐方法

Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages

  • 提出参数高效安全调优,适配低资源与混合语境
  • 发现英语安全机制在非英语场景失效,文化危害仍存
  • 倡导本地社区参与,推动公平AI落地

大语言模型正被部署于全球南方地区,日常使用涉及低资源语言、代码混用及文化特异性规范。然而,当前的安全管道、评估基准与对齐方法仍主要针对英语和少数高资源语言,隐含假设安全与事实性可跨语言迁移,但证据表明其并不成立。我们综合最新研究发现:(i) 安全防护在低资源和代码混用输入上显著弱化;(ii) 即使标准毒性评分合格,文化有害行为仍可能持续存在;(iii) 仅基于英语的知识修正与安全补丁难以迁移到低资源语言。为此,我们为全球南方的研究者与学生提出务实路径:参数高效的安全引导、基于文化的评估与偏好数据构建,以及赋能本地社区参与的协作流程。目标是将多语言安全作为代表性不足区域公平AI的核心要求,而非附加项。

原文摘要 · Abstract (English)

Large language models (LLMs) are being deployed across the Global South, where everyday use involves low-resource languages, code-mixing, and culturally specific norms. Yet safety pipelines, benchmarks, and alignment still largely target English and a handful of high-resource languages, implicitly assuming safety and factuality ''transfer'' across languages. Evidence increasingly shows they do not. We synthesize recent findings indicating that (i) safety guardrails weaken sharply on low-resource and code-mixed inputs, (ii) culturally harmful behavior can persist even when standard toxicity scores look acceptable, and (iii) English-only knowledge edits and safety patches often fail to carry over to low-resource languages. In response, we outline a practical agenda for researchers and students in the Global South: parameter-efficient safety steering, culturally grounded evaluation and preference data, and participatory workflows that empower local communities to define and mitigate harm. Our aim is to make multilingual safety a core requirement-not an add-on-for equitable AI in underrepresented regions.

多语言安全文化适配低资源语言公平AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。