arXiv:2606.28345cs.ROcs.AI2026-06中稿 · publication in Pro…

测试大模型在跨文化情境下优先救助决策的公平性,发现中文日文模型表现明显落后。

Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients

论文配图:Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
图 1 · 摘自论文原文
  • 构建基于文化偏好梯度的多语言评估框架,模拟真实社会机器人救助场景。
  • 4种语言模型在跨文化差异识别上普遍失效,西语表现接近双倍于中日语。
  • 仅对比示例提示有效提升公平性,纯推理提示反而加剧偏差,适合伦理审计团队使用。

LLM驱动的社会机器人日益决定谁应优先获得现实帮助。由于不同文化对年龄、地位和群体规模的优先规则各异,若未进行多元文化校准,将导致资源分配不公。然而现有大模型道德审计多以英语为中心,极少测试具身化情境,使多元文化校准成为亟待填补的诊断空白。本文提出一种基于梯度的审计框架,用于多语言评估大模型在文化偏好梯度下的道德权衡行为。基于九个跨领域社交机器人研究(>8,000篇论文),我们构建了涵盖照护、教育和服务领域的对称控制场景,将Moral Machine实验中的“救谁”问题转化为“先助谁”的困境,保留身份权衡(多数人对少数人;年轻对年长;高地位对低地位)。我们在四个国家-语言对中,对四种大模型在四种提示范式下进行审计(共57,600次决策),以国家特定的MME偏好梯度为基准。通过序数一致性检验模型是否区分文化背景;构建治理类型学分析其梯度区分、方向倾向与决策迟疑性。结果发现:模型普遍存在文化不对称的梯度跟踪失败,仅靠提示无法可靠修正:西方语言决策质量近乎中日语的两倍;多数优先权衡中高度确定性常抹除跨文化差异;对年龄与地位规范的部分敏感可能边缘化少数群体。提示效果不均:仅对比示例能带来一致改进,而仅推理类提示反而恶化追踪能力。研究呼吁将多语言、多元文化审计作为大模型-机器人部署前的必经关口,并指出模型自身因素比提示更可靠。

原文摘要 · Abstract (English)

LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to calibrate pluralistically can scale into unequal access. Yet LLM moral audits remain English-centered, rarely test embodied contexts, leaving pluralistic calibration as an urgent diagnostic gap amid intensifying LLM-robot deployment. We introduce a gradient-based audit framework for multilingual evaluation of LLM moral trade-off behavior against cultural preference gradients. Grounded in nine cross-domain social robotics reviews (>8,000 papers), we derive symmetry-controlled scenarios across care, education, and services, translating the Moral Machine Experiment's "whom to spare" into "whom to assist first" dilemmas with preserved identity trade-offs (many vs. few; young vs. old; higher vs. lower status). We audit four LLMs across four country-language pairs in four prompting regimes (57,600 decisions), benchmarked against country-specific MME preference gradients. Ordinal concordance tests whether models differentiate cultural contexts; a governance typology maps vulnerabilities in gradient differentiation, directional tendency, and deliberation. We find persistent, culturally asymmetric gradient tracking failures that prompting alone cannot reliably correct: quality calibration is nearly twice as strong for Western-language decisions as for Chinese and Japanese; high determinism in majority-first trade-offs often erases cross-cultural gradients; partial sensitivity to age- and status-based norms risks sidelining minorities. Prompting effects are uneven; only contrastive exemplars yield consistent gains, while reasoning-only prompts can worsen tracking. Our results motivate multilingual, pluralistic audits as an LLM-robot pre-deployment gate and suggest model factors are a more robust lever than prompting alone.

大模型审计社会机器人跨文化道德决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。