arXiv:2603.01952cs.AI2026-03ACL

用动态社会模拟测试大模型在多文化场景下的行为恰当性。

LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations

  • 将大模型作为代理嵌入模拟城市,评估其任务完成与文化规范遵守。
  • 发现不同文化下模型表现差异显著,且任务效率与文化敏感度存在权衡。
  • 提出可量化规范违规与评价者不确定性的自动评估框架,适合跨文化研究者使用。

大型语言模型(LLMs)越来越多地被用作自主代理,但现有评估主要关注任务成功率,而忽视文化适当性或评估者可靠性。我们提出了LiveCultureBench,一个包含多文化、动态交互的基准测试平台,将LLM作为代理置于模拟城镇中,同时评估其任务完成情况与社会文化规范的遵守程度。该仿真构建了一个小型城市的地点图谱,包含具有多样化人口统计与文化背景的合成居民。每轮实验为一名居民分配日常目标,其余居民提供社会背景。基于LLM的验证器生成结构化判断,涵盖规范违反与任务进展,并聚合为衡量任务-规范权衡及验证器不确定性的指标。通过在多种模型与文化背景下应用LiveCultureBench,我们研究了(i)LLM代理的跨文化鲁棒性,(ii)其在有效性与规范敏感性之间的平衡,以及(iii)LLM作为评判者在自动化评估中的可靠性边界,何时需人类干预。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability. We introduce LiveCultureBench, a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms. The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles. Each episode assigns one resident a daily goal while others provide social context. An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty. Using LiveCultureBench across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed.

大模型评估多文化社会模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。