arXiv:2605.10125cs.AIcs.HC2026-05被引 1

AI工具助研究高效起步,但精准验证仍需人工把关。

Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research

  • 结合人机指标构建评估框架,覆盖可用性与生成透明度
  • 问答工具摘要准确但精确信息提取不可靠,解释性差
  • 文献工具适合探索性搜索,系统综述不适用

人工智能(AI)工具正被引入科研流程,用于提升文档分析、问答(Q&A)和文献检索等任务的效率。然而,其输出难以验证,生成过程缺乏透明度,且易出错。现有基准测试未能充分捕捉可用性、可解释性和工作流整合等以人为本的标准。为此,本文提出并应用一个融合人机指标的评估框架,用于评测面向研究的AI Q&A与文献综述工具。结果显示,问答工具能提供有价值的概览和总体准确的摘要,但在精确信息提取上可靠性不足;可解释性AI(xAI)准确性尤其低,突出显示的来源段落常与答案不符,迫使研究者承担验证责任。文献综述工具支持探索性搜索,但复现性差,对所选来源和数据库透明度低,源质量不一致,不适合系统综述。比较两类工具发现:尽管AI能在研究早期阶段和浅层任务中提升效率,其输出仍需人工验证。研究强调解释性功能对提升透明度、验证效率的重要性,并呼吁重视人类中心评估以保障实际应用价值。

原文摘要 · Abstract (English)

Artificial intelligence (AI) tools are being incorporated into scientific research workflows with the potential to enhance efficiency in tasks such as document analysis, question answering (Q&A), and literature search. However, system outputs are often difficult to verify, lack transparency in their generation and remain prone to errors. Suitable benchmarks are needed to document and evaluate arising issues. Nevertheless, existing benchmarking approaches are not adequately capturing human-centered criteria such as usability, interpretability, and integration into research workflows. To address this gap, the present work proposes and applies a benchmarking framework combining human-centered and computer-centered metrics to evaluate AI-based Q&A and literature review tools for research use. The findings suggest that Q&A tools can offer valuable overviews and generally accurate summaries; however, they are not always reliable for precise information extraction. Explainable AI (xAI) accuracy was particularly low, meaning highlighted source passages frequently failed to correspond to generated answers. This shifted the burden of validation back onto the researcher. Literature review tools supported exploratory searches but showed low reproducibility, limited transparency regarding chosen sources and databases, and inconsistent source quality, making them unsuitable for systematic reviews. A comparison of these tool groups reveals a similar pattern: while AI tools can enhance efficiency in the early stages of the research workflow and shallow tasks, their outputs still require human verification. The findings underscore the importance of explainability features to enhance transparency, verification efficiency and careful integration of AI tools into researchers' workflows. Further, human-centered evaluation remains an important concern to ensure practical applicability.

AI评估可解释性科研工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。