首个面向乌尔都语的3H对齐评测,揭示大模型跨文化适配缺陷
Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

- 人工校准+词典引导编辑,构建本土化乌尔都语评测集
- 零样本测试显示助人、安全、诚实指标全面下降
- 警示现有对齐方法在低资源语言中失效,适合多语言研究者
大型语言模型(LLMs)在以英语为中心的环境中表现出良好的帮助性、无害性和诚实性(3H)对齐,但这些优势在低资源语言中难以迁移,源于文化错配。现有双语3H评测主要依赖自动化翻译或基于LLM的合成,传播源语言偏见并牺牲本地相关性。为填补这一空白,我们提出Pak3H1,首个经过人工验证且文化情境化的乌尔都语3H评测套件,包括PakAlpaca(帮助性)、PakBeaverTails(无害性)和PakTruthfulQA(诚实性)。通过多阶段流程整合人工文化适配与词典引导后编辑,优先保障母语者判断,确保语义准确与情境真实。在多个开源及专有模型架构上的零样本评估揭示系统性跨语言对齐差距:帮助性胜率在本地化情境下下降,无害性防护机制在区域安全风险面前失效,复合诚实性指标因本地事实约束显著退化。研究暴露当前对齐方法的结构性局限,强调人类引导本地化对公平多语言评估的必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignment in English-centric settings, but these gains transfer poorly to low-resource languages due to cultural mismatches. Existing multilingual 3H benchmarks rely predominantly on automated translation or LLM based synthesis, propagating source-language biases while sacrificing local relevance. To address this gap, we introduce Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty). Our multi-stage pipeline integrates manual cultural adaptation and dictionary-guided post editing to prioritize native speaker judgment, ensuring both semantic fidelity and contextual authenticity. Zero-shot evaluations across multiple open and proprietary LLM architectures reveal systematic cross-lingual alignment gaps: helpfulness win rates decline under localized contexts, harmlessness guardrails break down against regional safety risks, and composite honesty metrics degrade substantially due to localized factual constraints. These findings expose structural limitations in current alignment approaches, underscoring the necessity of human-guided localization for equitable multilingual evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。