arXiv:2604.12545cs.AIcs.CY2026-04

用LLM模拟不同文化下民众对官僚程序的情绪反应,发现效果有限且东方文化更差。

Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents

论文配图:Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents
图 1 · 摘自论文原文
  • 构建评估框架,测试LLM在跨文化情境中对官僚程序的情绪响应
  • 所有模型表现不佳,东方文化下情绪响应偏差更严重
  • 推出交互式界面RAMO,支持模拟与数据收集,开源可用

提升政策制定质量是公共管理的核心议题。以往的人类受试研究揭示,公民在政策执行过程中对官僚程序的情绪反应存在显著跨文化差异。尽管大语言模型(LLM)为模拟类人反应、降低实验成本提供了可能,但其生成符合文化背景的情绪反应能力尚未得到验证。为此,我们提出一个评估框架,用于衡量LLM在多元文化背景下对官僚程序的情绪响应。作为初步研究,我们将其应用于单一官僚程序场景。结果显示,所有模型与人类情绪反应的对齐度均有限,尤其在东亚文化中表现更差;文化提示策略对改善对齐效果作用甚微。我们进一步提出 extbf{RAMO}——一个交互式界面,用于模拟公民对官僚程序的情绪反应,并收集人类数据以优化模型。该界面已公开发布于 https://ramo-chi.ivia.ch。

原文摘要 · Abstract (English)

Improving policymaking is a central concern in public administration. Prior human subject studies reveal substantial cross-cultural differences in citizens' emotional responses to red tape during policy implementation. While LLM agents offer opportunities to simulate human-like responses and reduce experimental costs, their ability to generate culturally appropriate emotional responses to red tape remains unverified. To address this gap, we propose an evaluation framework for assessing LLMs' emotional responses to red tape across diverse cultural contexts. As a pilot study, we apply this framework to a single red-tape scenario. Our results show that all models exhibit limited alignment with human emotional responses, with notably weaker performance in Eastern cultures. Cultural prompting strategies prove largely ineffective in improving alignment. We further introduce \textbf{RAMO}, an interactive interface for simulating citizens' emotional responses to red tape and for collecting human data to improve models. The interface is publicly available at https://ramo-chi.ivia.ch.

LLM跨文化情绪模拟政策设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。