评测大模型研究报告的逻辑可靠性,让内容可信可查。
ReportLogic: Evaluating Logical Quality in Deep Research Reports
- 从结构、论述到证据链,三层评估报告逻辑性
- 构建人工标注数据集,训练可扩展的逻辑判别模型
- 发现现有大模型易被表面文字误导,需加强推理审查
用户越来越多地依赖大语言模型(LLMs)进行深度研究,将多源信息整合为结构化报告以支持理解和决策。在此背景下,报告的实际可靠性取决于逻辑质量:其论点和论证是否明确有据,能否作为后续应用的可信基础,而不仅是表面流畅或信息丰富。然而,现有评估框架普遍忽视这一需求。为此,我们提出ReportLogic,一个基于读者视角可审计性的报告级逻辑质量基准。该基准采用分层分类体系,评估读者能否:(1) 追踪具有统一分析主线的在题报告结构(宏观逻辑),(2) 理清必要的上下文推进(表述逻辑),(3) 通过显式论点-支持关系验证结论(结构逻辑)。基于此体系,我们构建了人工标注的评分指导数据集,并训练了一个开源的LogicJudge模型以实现规模化评估。进一步通过对抗攻击测试法官鲁棒性,发现现成的LLM裁判常受冗长等表面特征影响,且推理模式可能掩盖支持关系断裂。总体结果为构建更稳健的逻辑评估器及提升大模型生成报告的逻辑可靠性提供了可操作建议。
原文摘要 · Abstract (English)
Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action. In this context, the practical reliability of such reports hinges on logical quality: whether the report's claims and arguments are explicitly supported and can be trusted as a basis for downstream use, rather than merely appearing fluent or informative. However, current evaluation frameworks largely overlook this requirement. To bridge this gap, we introduce ReportLogic, a benchmark that quantifies report-level logical quality through a reader-centric lens of auditability. Specifically, ReportLogic adopts a hierarchical taxonomy that evaluates whether readers can (1) trace an on-topic report structure with a unified analytical arc (Macro-Logic), (2) understand the progression with necessary context (Expositional-Logic), and (3) verify conclusions via explicit claim--support (Structural-Logic). Based on this taxonomy, we construct a human-annotated rubric-guided dataset and train an open-source LogicJudge for scalable evaluation. We further evaluate judge robustness via adversarial attacks, showing that off-the-shelf LLM judges are frequently influenced by superficial cues (e.g., verbosity), and reasoning modes can mask broken support relations. Overall, our results provide actionable guidance for building more robust logic evaluators and improving the logical reliability of LLM-generated reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。