测试大模型能否准确根据证据强度生成科研摘要,发现其判断常过于保守。
CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models

- 设计可审计的诊断框架CalBrief,评估模型对证据强度与范围的把握能力。
- 4类标签比2类标签导致63%的过度保守,主要因标签空间扩大所致。
- 标签细化虽难匹配,但事后合并可提升表现,提示需分开评估判断与组织能力。
大型语言模型(LLMs)被越来越多地用作研究助手,但尚不清楚它们是否能根据支持证据的强度和范围来校准研究结论。本文研究证据校准型科学简报:给定一组相关论文,系统应生成包含证据强度、范围边界及缺失证据警示的综合结论。我们构建了一个经验证的试点基准,包含16个异构科学证据包和96个经人工验证的结论,并采用CalBrief(一种可审计的角色/缺口/强度框架)作为诊断工具,定位简报失效环节。在公平模式评估下,结构化组织提升了角色与缺口推理能力,但显式强度校准策略系统性地过于保守,低于多数基线与直接调用基线。通过控制实验,对比GPT-4o、Claude Sonnet、Gemini Flash三种闭源模型,分离出三种潜在原因:约63%的保守差距源于将标签空间从二元{中等, 弱}扩展至四元{中等, 弱, 不确定, 证据不足}(所有模型p < 0.001);仅1%归因于缺口/范围信号注入(不显著);剩余36%来自流程策略本身。此外,四分类预测可事后合并为二分类,效果可匹敌甚至超越直接二分类提示,表明额外标签包含严格匹配所隐藏的信息。标签级强度判断与可审计证据组织是当前存在张力的两种能力,应在评估中分别对待。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence. We study evidence-calibrated scientific briefing: given a bounded package of related papers, a system should generate package-level takeaways with evidence strength, scope boundaries, and missing-evidence caveats. We contribute a verified pilot benchmark of 16 heterogeneous scientific evidence packages and 96 human-verified takeaways, and we use CalBrief, an auditable role/gap/strength framework, as a diagnostic probe to locate where briefing breaks down. Under a fair-schema evaluation, structured organization improves role and gap reasoning, but an explicit strength-calibration policy is systematically over-conservative and falls below majority and direct-LLM baselines. To explain why, we run a controlled diagnostic across three closed-model backbones (GPT-4o, Claude Sonnet, Gemini Flash) that separates three potential causes of conservatism. Approximately 63% of the conservatism gap is attributable to expanding the label space from binary {moderate, weak} to four-way {moderate, weak, uncertain, insufficient_evidence} (p < 0.001 across all backbones); only 1% is attributable to gap/scope signal injection (not significant); the remaining 36% arises from the pipeline policy itself. We also find that 4-way predictions can be post-hoc collapsed back to binary and then match or exceed direct binary prompting, so the extra labels carry information that strict matching hides. Label-level strength judgment and auditable evidence organization are distinct abilities currently in tension, and should be evaluated separately for LLM research assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。