arXiv:2512.01896cs.CL2025-12中稿 · CMC-Computers, Mat…被引 3

构建首个公开舆情报告生成评测基准,推动大模型在危机响应中的应用

OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation

  • 定义自动化舆情报告生成任务,设计事件中心数据集
  • 涵盖463个危机事件,含新闻、社交媒体与参考摘要
  • 提出模拟专家评估的智能体框架,与人工判断高度一致

在线公共舆情报告通过整合新闻与社交媒体内容,为政府和企业及时应对危机提供支持。尽管大语言模型已使自动化报告生成成为可能,但该领域缺乏系统性研究,尤其缺少明确的任务定义与评测基准。为此,我们定义了自动化在线公共舆情报告生成(OPOR-GEN)任务,构建了以事件为中心的OPOR-BENCH数据集,覆盖463个危机事件,包含对应的新闻文章、社交媒体帖子及参考摘要。为评估报告质量,我们提出OPOR-EVAL——一种基于智能体的评估框架,通过在上下文中分析生成报告来模拟人类专家评价。前沿模型实验表明,该框架与人工判断具有高相关性。本研究提供的任务定义、基准数据集与评估体系,为该关键领域的后续研究奠定了坚实基础。

原文摘要 · Abstract (English)

Online Public Opinion Reports consolidate news and social media for timely crisis management by governments and enterprises. While large language models have made automated report generation technically feasible, systematic research in this specific area remains notably absent, particularly lacking formal task definitions and corresponding benchmarks. To bridge this gap, we define the Automated Online Public Opinion Report Generation (OPOR-GEN) task and construct OPOR-BENCH, an event-centric dataset covering 463 crisis events with their corresponding news articles, social media posts, and a reference summary. To evaluate report quality, we propose OPOR-EVAL, a novel agent-based framework that simulates human expert evaluation by analyzing generated reports in context. Experiments with frontier models demonstrate that our framework achieves high correlation with human judgments. Our comprehensive task definition, benchmark dataset, and evaluation framework provide a solid foundation for future research in this critical domain.

大模型评测舆情分析自动报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。