arXiv:2607.05405cs.CYcs.AI2026-07

评测大模型在健康对话中对隐性文化线索的适应能力

CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

论文配图:CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries
图 1 · 摘自论文原文
  • 将文化视为规范遵循连续体,构建60个跨六文化的健康对话角色
  • 五款主流模型仅20%-30%响应符合文化情境,阿富汗场景低至8.8%
  • 模型更倾向遵循内置偏见而非响应文化线索,尤其在隐性风格上表现稍好

为实现公平且无刻板印象的用户交互,AI模型需具备文化胜任力——即推断并适应用户隐性传递的文化价值观,而非依赖静态人口属性。本文提出CCBENCH框架,将文化视为规范遵循的连续状态而非二元归属。以健康为案例,构建CCBENCH-Health,包含60个基于理论的人物设定,涵盖六种文化,每人均参与18轮真实对话,共生成3,120次互动。评估使用52个来自真实用户论坛的医疗问题。基准测试显示,即使最优模型也仅在20%-30%情况下给出文化恰当回应。当显式提示关注对话历史中的文化线索(CoT)时,性能平均提升3-5%。研究发现,模型在人物回避文化规范时表现更好,揭示其倾向于遵循固有偏见而非响应文化线索的持续不对称性。该现象在阿富汗语境尤为显著(平均仅8.8%),文化线索极少带来恰当建议。此外,模型有时更易适应隐性的文化对话风格,而非显性的文化实践,但表现因文化而异。

原文摘要 · Abstract (English)

To interact with users fairly and without stereotyping, AI models must display cultural competency, i.e., the ability to infer and adapt to a user's implicitly signaled cultural values, rather than relying on static demographic traits. We introduce CCBENCH, a framework for evaluating cultural competency in large language models (LLMs), treating culture as a continuum of norm adherence states rather than as a binary state of cultural belongingness. As a case study on health, we create CCBENCH-Health, which includes 60 theoretically grounded personas exhibiting varied norm-adherence states across six cultures, each engaging in 18 realistic dialogues. Each persona is evaluated on 52 authentic healthcare questions drawn from real user forums, yielding 3,120 unique interactions. Benchmarking five leading models reveals that even the best achieve culturally appropriate responses only 20-30% of the time. When explicitly prompted to focus on culturally relevant cues from the conversational history (CoT), performance improves modestly by 3-5% on average. We find that models perform best when personas avoid cultural norms rather than follow them, revealing a persistent asymmetry, suggesting a preference in the models to align with built-in biases than adapt to cultural cues. This is especially observed in the Afghan context (Avg: 8.8%), where cultural cues rarely yield appropriate health advice. Finally, we find that models sometimes adapt more readily to implicit, cultural conversational styles than to explicitly stated cultural practices, though this varies across cultures.

大模型评测文化胜任力健康AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。