arXiv:2601.07973cs.CYcs.AI2026-01被引 2

构建跨文化规范框架,自动检测人机对话中的行为违规。

Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations

  • 提出多维度规范分类体系,区分人际与人机互动规范。
  • 实测主流模型在不同国家/情境下违规率差异显著。
  • 适合关注AI安全、跨文化交互的开发者与研究者。

生成式AI需在跨文化场景中兼具实用性与安全性。关键挑战在于理解模型对社会文化规范的遵守情况。尽管该问题在自然语言处理领域受关注,现有工作在规范理解与评估上仍缺乏细致性与全面性。本文提出一种规范分类体系,明确其语境(如区分人类间规范与人机互动规范)、领域(如适用领域)及执行机制(如约束方式)。我们展示了如何将该分类体系转化为可自动评估模型在真实、开放场景中规范遵守情况的流程。探索性分析表明,当前最先进的模型普遍存在规范违反现象,且违规率随模型、互动语境及国家而异。此外,提示意图与情境设定也显著影响违规率。该分类体系与演示评估流程,为现实场景下文化规范遵守的精细化、情境敏感性评估提供了支持。

原文摘要 · Abstract (English)

Generative AI models ought to be useful and safe across cross-cultural contexts. One critical step toward this goal is understanding how AI models adhere to sociocultural norms. While this challenge has gained attention in NLP, existing work lacks both nuance and coverage in understanding and evaluating models' norm adherence. We address these gaps by introducing a taxonomy of norms that clarifies their contexts (e.g., distinguishing between human-human norms that models should recognize and human-AI interactional norms that apply to the human-AI interaction itself), specifications (e.g., relevant domains), and mechanisms (e.g., modes of enforcement). We demonstrate how our taxonomy can be operationalized to automatically evaluate models' norm adherence in naturalistic, open-ended settings. Our exploratory analyses suggest that state-of-the-art models frequently violate norms, though violation rates vary by model, interactional context, and country. We further show that violation rates also vary by prompt intent and situational framing. Our taxonomy and demonstrative evaluation pipeline enable nuanced, context-sensitive evaluation of cultural norm adherence in realistic settings.

AI安全跨文化规范检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。