arXiv:2509.04499cs.CLcs.AI2025-09被引 12

评测大模型研究系统在引用与证据上的可靠性,发现其常过度自信且引文不实。

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

  • 构建八维度审计框架,从答案、来源到引文全程分析
  • 超40%的陈述无自身引文支持,部分系统引文准确率仅40%
  • 适合关注AI可信度、内容真实性与引证质量的研究者

生成式搜索引擎和深度研究大模型代理承诺提供可信赖、基于来源的综合信息,但用户常遭遇过度自信、弱溯源和混乱的引用行为。我们提出DeepTRACE,一个融合社会技术视角的审计框架,将社区识别的失败案例转化为八个可量化的维度,覆盖回答文本、来源与引文。该框架采用逐语句分析(分解与置信度评分),构建引文与事实支持矩阵,实现对系统推理与证据归属的端到端审计。通过自动化提取管道评估多个公开模型(如GPT-4.5/5、You.com、Perplexity、Copilot/Bing、Gemini),并使用经验证与人类标注员一致的LLM裁判,我们评估了网络搜索引擎与深度研究配置的表现。结果表明,生成式搜索引擎和深度研究代理在争议性问题上频繁输出片面且高度自信的回答,且大量陈述缺乏自身列出来源的支持。尽管深度研究配置降低了过度自信并提升了引文完整性,但在争议问题上仍严重片面,且仍存在高比例未被支持的陈述,引文准确率在40%至80%之间波动。

原文摘要 · Abstract (English)

Generative search engines and deep research LLM agents promise trustworthy, source-grounded synthesis, yet users regularly encounter overconfidence, weak sourcing, and confusing citation practices. We introduce DeepTRACE, a novel sociotechnically grounded audit framework that turns prior community-identified failure cases into eight measurable dimensions spanning answer text, sources, and citations. DeepTRACE uses statement-level analysis (decomposition, confidence scoring) and builds citation and factual-support matrices to audit how systems reason with and attribute evidence end-to-end. Using automated extraction pipelines for popular public models (e.g., GPT-4.5/5, You.com, Perplexity, Copilot/Bing, Gemini) and an LLM-judge with validated agreement to human raters, we evaluate both web-search engines and deep-research configurations. Our findings show that generative search engines and deep research agents frequently produce one-sided, highly confident responses on debate queries and include large fractions of statements unsupported by their own listed sources. Deep-research configurations reduce overconfidence and can attain high citation thoroughness, but they remain highly one-sided on debate queries and still exhibit large fractions of unsupported statements, with citation accuracy ranging from 40--80% across systems.

AI审计引文可靠性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。