测试大模型在危机公关中策略性隐瞒信息的能力,发现部分模型能合理权衡声誉与道德。
Crisis-Bench: Benchmarking Strategic Ambiguity and Reputation Management in Large Language Models
- 构建动态危机模拟环境,区分公开与私密叙事状态
- 通过股价变化评估模型在80个行业危机中的表现
- 揭示模型在道德与专业需求间的策略性平衡能力
标准安全对齐使大模型趋向绝对诚实与帮助,形成僵化的‘童子军’道德观。这虽适合通用助手,但在公关、谈判和危机管理等专业领域带来‘透明度代价’。为衡量这一差距,我们提出Crisis-Bench,一个基于多智能体部分可观马尔可夫决策过程(POMDP)的基准,评估大模型在高风险企业危机中的表现。该基准涵盖8个行业的80种不同剧情,要求基于大模型的公关代理在7天动态危机中管理严格分离的私密与公开叙事状态,强制信息不对称。不同于依赖静态真值的传统评测,我们引入‘仲裁-市场循环’:由仲裁判定公众情绪并转化为模拟股价,构建真实经济激励机制。结果揭示关键分歧:部分模型因伦理顾虑而退缩,另一些则展现出合法且有效的战略隐瞒能力,以稳定股价。Crisis-Bench首次提供‘声誉管理’能力的量化框架,主张从僵化道德主义转向情境感知的专业对齐。
原文摘要 · Abstract (English)
Standard safety alignment optimizes Large Language Models (LLMs) for universal helpfulness and honesty, effectively instilling a rigid "Boy Scout" morality. While robust for general-purpose assistants, this one-size-fits-all ethical framework imposes a "transparency tax" on professional domains requiring strategic ambiguity and information withholding, such as public relations, negotiation, and crisis management. To measure this gap between general safety and professional utility, we introduce Crisis-Bench, a multi-agent Partially Observable Markov Decision Process (POMDP) that evaluates LLMs in high-stakes corporate crises. Spanning 80 diverse storylines across 8 industries, Crisis-Bench tasks an LLM-based Public Relations (PR) Agent with navigating a dynamic 7-day corporate crisis simulation while managing strictly separated Private and Public narrative states to enforce rigorous information asymmetry. Unlike traditional benchmarks that rely on static ground truths, we introduce the Adjudicator-Market Loop: a novel evaluation metric where public sentiment is adjudicated and translated into a simulated stock price, creating a realistic economic incentive structure. Our results expose a critical dichotomy: while some models capitulate to ethical concerns, others demonstrate the capacity for Machiavellian, legitimate strategic withholding in order to stabilize the simulated stock price. Crisis-Bench provides the first quantitative framework for assessing "Reputation Management" capabilities, arguing for a shift from rigid moral absolutism to context-aware professional alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。