评估大模型生成的西语新闻评论是否真实,发现多数模型表现不佳。
Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

- 用5个大模型生成西班牙新闻评论,对比真实数据
- 原生模型严重低估仇恨言论,且情感倾向偏差明显
- 微调后效果不均,通义千问3最平衡,Mistral7B过量生成仇恨内容
大模型驱动的社交代理被广泛用于模拟线上社交行为,但其真实性难以验证。现有研究多依赖通用基准,较少关注新闻评论这类短时反应性话语。本文以Hatemedia数据集为基础,将5,631条新闻与58,555条真实用户评论配对,并在相同实验设置下用五种大模型生成对应合成数据集。从仇恨言论、情感倾向和语义一致性三个维度比较真实与合成评论。结果表明,未经微调的模型作为真实评论的代理效果差:显著低估仇恨言论,引入模型特异性情感偏见,且分布上远离人类回复。微调虽提升部分性能,但效果不均衡;通义千问3(Qwen3)表现最均衡,而Mistral7B虽在情感与语义对齐上最优,却过度高估仇恨言论。即使生成的回复看似合理,也不一定符合公共话语的整体分布特征。
原文摘要 · Abstract (English)
LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has largely relied on general-purpose benchmarks, while less attention has been paid to short, reactive discourse such as audience replies to online news. In this paper, we evaluate whether LLM-generated reactions to Spanish online news reproduce measurable properties of real audience discourse. Using the Hatemedia dataset, we pair 5,631 news items with 58,555 real audience reactions, and generate a matched synthetic dataset using five LLMs under a shared experimental setting. We compare real and synthetic reactions across three dimensions: hate speech, sentiment, and semantic alignment, considering both off-the-shelf and fine-tuned generation. Results show that off-the-shelf models are poor proxies for real audience reactions: they strongly underproduce hate speech, introduce model-specific sentiment biases, and remain distributionally distant from human replies. Fine-tuning improves fidelity unevenly. Qwen3 provides the most balanced approximation, while Mistral7B achieves the strongest sentiment and semantic alignment but overshoots hate prevalence. Plausible synthetic replies do not necessarily reproduce the distributional properties of public discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。