评测大模型模拟真实Reddit讨论的逼真度,发现现有模型仍差距明显。
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions

- 基于4292个真实Reddit线程构建评测基准,从多个维度对比生成与真实对话
- 当前五种模型生成内容在重复性、叙事结构等方面与真实数据存在显著差异
- 适合关注大模型社会行为模拟真实性的研究者和开发者使用
大语言模型代理正被广泛用于模拟现实世界互动,但其生成行为是否保留了真实人类互动的内容模式与交互动态尚不明确。现有评估体系零散,难以系统比较或衡量进展。本文以Reddit讨论为具体切入点,构建了包含4,292个真实帖子的MiroBench基准,通过统计检验在四个核心维度——重复性与语义一致性、叙事内容、毒性与攻击性、结构复杂性——对比生成与真实讨论。跨五个领域、五种模型的实验表明,当前模拟器在分布上仍显著偏离真实数据,轻量级提示优化仅带来有限提升。MiroBench为衡量、诊断与改进基于LLM的社会仿真真实性提供了可操作的基准。
原文摘要 · Abstract (English)
LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented, which makes it difficult to compare systems or measure progress. In this paper, we focus on Reddit discussions as a concrete first step toward evaluating real-world social simulation. Reddit threads provide public, topic-grounded, multi-party interactions where people share experiences, debate, seek advice, express emotion, and collectively respond to products, events, and social issues. These discussions offer an observable window into broader social behavior, making them a useful setting for testing whether LLM agents can reproduce not only fluent text, but also the distributional patterns and interaction dynamics of real online communities. We introduce MiroBench, a benchmark for Reddit discussion simulation built from 4,292 real Reddit threads. MiroBench uses statistical tests to compare generated and real discussions across four major aspects: repetition and semantic uniformity, narrative content, toxicity and aggression, and structural complexity. Experiments across five domains and five models show that current simulators remain distributionally mismatched with real Reddit threads, while a lightweight prompt-based improvement procedure provides only limited gains. MiroBench offers a concrete benchmark for measuring, diagnosing, and improving realism in LLM-based social simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。