评测智能体在新闻写作中的真实工作流表现
Benchmarking Agentic Newswriting via Journalistic Workflows
- 设计新闻写作任务链,模拟记者主动搜证与迭代写作
- 6000个真实新闻案例中,智能体查准率高但叙事整合差
- 适合研究具身智能体在信息密集型任务中的能力边界
近期工业界推出的自主数字代理(如Manus AI和Gemini研究模式)展现出通过自主决策与任务分解完成结构化任务的潜力,但其在真实世界信息密集型工作流中的表现仍不明确。本文以新闻写作为例,研究此类系统在需要迭代规划、上下文推理及主动发现缺失背景的复杂流程中的表现。我们提出NEWSAGENT基准,用于评估代理如何搜索原始素材、筛选相关信息,并通过核心新闻功能迭代修订稿件。给定写作指令和部分一手材料,代理需识别叙事视角、基于关键词提问、检索历史背景并生成完整新闻稿。不同于常规摘要或检索任务,关键背景信息并不直接提供,必须主动发现,贴近真实报道约束。NEWSAGENT包含6000个经人工验证的真实新闻样本。我们评估了使用常见代理框架的开源与闭源大模型,结果显示代理在事实检索方面表现良好,但在规划与叙事整合上存在明显短板。我们认为NEWSAGENT为评估代理在网页数据操作方面的实际生产力提供了真实测试环境。
原文摘要 · Abstract (English)
Recent advances in autonomous digital agents from industry (e.g., Manus AI and Gemini's research mode) highlight their potential for structured tasks through autonomous decision-making and task decomposition, but it remains unclear how well such systems support real-world information-intensive workflows. We study this question in journalism, where newswriting requires iterative planning, contextual reasoning, and active discovery of missing background to produce a coherent article. We introduce NEWSAGENT, a benchmark for evaluating how agents search raw materials, select relevant information, and iteratively revise drafts through core journalistic functions. Given a writing instruction and partial firsthand materials, agents must identify narrative perspectives, issue keyword-based queries, retrieve historical context, and generate complete news articles. Unlike typical summarization or retrieval tasks, essential context is not directly available and must be actively discovered, reflecting real-world reporting constraints. NEWSAGENT consists of 6k human-verified examples derived from real news. We evaluate open- and closed-sourced LLMs with commonly-used agentic frameworks on NEWSAGENT, which shows that agents are capable of retrieving relevant facts but struggling with planning and narrative integration. We believe that NEWSAGENT serves a realistic testbed for iterating and evaluating agent capabilities in terms of web data manipulation to real-world productivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。