为多语言大模型构建基于文化背景的治理框架,避免加剧语言不平等。
Culturally-Grounded Governance for Multilingual Language Models: Rights, Data Boundaries, and Accountable AI Design
- 从文化与权利视角重构多语言AI治理,强调本地化规范的重要性。
- 揭示训练数据与评估中语言不平等、全球部署与地方价值错位等核心问题。
- 适合关注AI公平性、跨文化设计及政策制定的研究者与从业者。
多语言大语言模型(MLLMs)日益在多元文化、语言与政治背景下部署,但现有治理框架仍以英语为中心,假设用户同质化且抽象定义公平性,导致低资源语言和文化边缘群体面临系统性风险。数据实践、模型行为与问责机制常与当地规范、权利与期待不符。本文结合人机交互与AI治理中的跨文化视角,整合多语言模型行为、数据不对称与社会技术伤害的现有证据,提出一种基于文化的治理框架。识别出三大相互关联的挑战:训练数据与评估中存在文化和语言不平等;全球部署与本地规范、价值观和权力结构错位;对边缘语言社区所受伤害缺乏有效问责机制。本文不提出新技术基准,而是倡导将多语言AI治理视为社会文化与权利问题,提出数据管理、透明度与参与式问责的设计与政策启示,强调文化根基治理对防止模型以规模与中立之名复制全球不平等至关重要。
原文摘要 · Abstract (English)
Multilingual large language models (MLLMs) are increasingly deployed across cultural, linguistic, and political contexts, yet existing governance frameworks largely assume English-centric data, homogeneous user populations, and abstract notions of fairness. This creates systematic risks for low-resource languages and culturally marginalized communities, where data practices, model behavior, and accountability mechanisms often fail to align with local norms, rights, and expectations. Drawing on cross-cultural perspectives in human-centered computing and AI governance, this paper synthesizes existing evidence on multilingual model behavior, data asymmetries, and sociotechnical harm, and articulates a culturally grounded governance framework for MLLMs. We identify three interrelated governance challenges: cultural and linguistic inequities in training data and evaluation practices, misalignment between global deployment and locally situated norms, values, and power structures, and limited accountability mechanisms for addressing harms experienced by marginalized language communities. Rather than proposing new technical benchmarks, we contribute a conceptual agenda that reframes multilingual AI governance as a sociocultural and rights based problem. We outline design and policy implications for data stewardship, transparency, and participatory accountability, and argue that culturally grounded governance is essential for ensuring that multilingual language models do not reproduce existing global inequalities under the guise of scale and neutrality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。