AI在文献综述中提取数据时,主要问题不是胡编乱造,而是理解差异。
Hallucination vs interpretation: rethinking accuracy and precision in AI-assisted data extraction for knowledge synthesis
- 用大模型自动提取文献信息,对比人类与AI的响应一致性。
- AI错误率仅1.51%,远低于人类的4.37%,多数差异源于解释不同。
- 适合需要高效处理大量文献的研究者,尤其关注解释性任务的场景。
知识综合(文献综述)对健康专业教育至关重要,但数据提取耗时费力。人工智能辅助提取虽提升效率,却引发准确性担忧。本文构建基于大语言模型的提取平台,对比了AI与人类在187篇文献、17个问题上的表现。通过评分一致性和主题相似性评估,发现AI在明确陈述的问题(如标题、目的)上与人类高度一致,但在需主观解读或文本未提及的问题(如基尔帕特里克结果、研究动机)上表现较差。人类-人类一致性并不优于人类-AI,且同样受问题类型影响。769/3179(24.2%)的分歧主要源于解释差异(18.3%),而AI错误仅占1.51%,人类不准确率则达4.37%。结果表明,AI变异性更多来自可解释性而非幻觉。重复运行可识别出解释复杂或模糊的任务,优化后续人工审查流程。AI可在知识综合中作为透明可信的伙伴,但仍需保留关键的人类洞察。
原文摘要 · Abstract (English)
Knowledge syntheses (literature reviews) are essential to health professions education (HPE), consolidating findings to advance theory and practice. However, they are labor-intensive, especially during data extraction. Artificial Intelligence (AI)-assisted extraction promises efficiency but raises concerns about accuracy, making it critical to distinguish AI 'hallucinations' (fabricated content) from legitimate interpretive differences. We developed an extraction platform using large language models (LLMs) to automate data extraction and compared AI to human responses across 187 publications and 17 extraction questions from a published scoping review. AI-human, human-human, and AI-AI consistencies were measured using interrater reliability (categorical) and thematic similarity ratings (open-ended). Errors were identified by comparing extracted responses to source publications. AI was highly consistent with humans for concrete, explicitly stated questions (e.g., title, aims) and lower for questions requiring subjective interpretation or absent in text (e.g., Kirkpatrick's outcomes, study rationale). Human-human consistency was not higher than AI-human and showed the same question-dependent variability. Discordant AI-human responses (769/3179 = 24.2%) were mostly due to interpretive differences (18.3%); AI inaccuracies were rare (1.51%), while humans were nearly three times more likely to state inaccuracies (4.37%). Findings suggest AI variability depends more on interpretability than hallucination. Repeating AI extraction can identify interpretive complexity or ambiguity, refining processes before human review. AI can be a transparent, trustworthy partner in knowledge synthesis, though caution is needed to preserve critical human insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。