arXiv:2410.12880cs.CLcs.AI2024-10NAACL被引 7

构建文化敏感评估与训练数据集,提升小模型跨文化表现

Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models

  • 设计文化伤害测试集,覆盖多文化场景下的敏感问题
  • 用多元标注者反馈构建对齐文化偏好的微调数据集
  • 显著降低模型生成文化冒犯内容的概率,适合伦理向AI研究者

随着大语言模型在全球应用日益广泛,文化敏感性愈发重要,需确保不同背景用户感到被尊重与理解。当模型未能契合特定文化规范时,可能引发误述或违背文化价值观的问题。本文针对小参数模型因缺乏充分训练数据而难以捕捉全球文化细微差别的挑战,提出两项关键贡献:(1) 构建文化伤害测试集,通过多文化情境评估模型输出的敏感性;(2) 构建文化对齐偏好数据集,基于多元标注者的反馈进行微调,以恢复模型的文化敏感性。实验表明,引入文化对齐反馈可显著改善模型行为,大幅降低生成文化不敏感或有害内容的可能性。本工作为实现更包容、尊重的AI系统铺平道路,助力大语言模型安全、伦理地穿越多元文化格局。

原文摘要 · Abstract (English)

As LLMs are increasingly deployed in global applications, the importance of cultural sensitivity becomes paramount, ensuring that users from diverse backgrounds feel respected and understood. Cultural harm can arise when these models fail to align with specific cultural norms, resulting in misrepresentations or violations of cultural values. This work addresses the challenges of ensuring cultural sensitivity in LLMs, especially in small-parameter models that often lack the extensive training data needed to capture global cultural nuances. We present two key contributions: (1) A cultural harm test dataset, created to assess model outputs across different cultural contexts through scenarios that expose potential cultural insensitivities, and (2) A culturally aligned preference dataset, aimed at restoring cultural sensitivity through fine-tuning based on feedback from diverse annotators. These datasets facilitate the evaluation and enhancement of LLMs, ensuring their ethical and safe deployment across different cultural landscapes. Our results show that integrating culturally aligned feedback leads to a marked improvement in model behavior, significantly reducing the likelihood of generating culturally insensitive or harmful content. Ultimately, this work paves the way for more inclusive and respectful AI systems, fostering a future where LLMs can safely and ethically navigate the complexities of diverse cultural landscapes.

文化敏感模型评估微调数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。