AI写文献综述易偏颇,选材和立场受模型影响,需专家把关。
Writing literature reviews with AI: principles, hurdles and some lessons learned
- 用相同论文集生成不同风格综述,显示选文影响结论方向
- 人机选文仅20%重合,主流观点占主导,批判性不足
- 依赖AI易忽略盲点,适合有领域知识者审阅辅助生成内容
我们定性比较了不同程度使用AI辅助生成的文献综述。同一LLM基于280篇论文,因输入选择不同,产出从主流中立到批判性、后殖民视角的迥异结果,且均非预设意图。尽管输出初看流畅有深度,细读发现存在认知盲区、偏见与深度缺失。六版对比揭示多重陷阱:(1)无知偏见——未获取的文献无法察觉;(2)对齐与数字谄媚——商业模型迎合用户暗示,强化偏见;(3)主流化倾向——因统计特性偏好主流内容,人类与模型选文仅20%重合;(4)创造性重构能力弱,表述模糊;(5)缺乏批判视角,源于远距离阅读与政治正确。多数问题可通过提示缓解,但前提为使用者具备足够领域知识以识别缺陷。悖论在于:要写出高质量的AI辅助综述,需先掌握文献,而这正是AI本应减少的工作。总体而言,AI可提升综述广度与质量,但省时效果有限,盲目交由AI处理是灾难。论文最后提出撰写与评估此类综述的建议。
原文摘要 · Abstract (English)
We qualitatively compared literature reviews produced with varying degrees of AI assistance. The same LLM, given the same corpus of 280 papers but different selections, produced dramatically different reviews, from mainstream and politically neutral to critical and post-colonial, though neither orientation was intended. LLM outputs always appear at first glance to be well written, well informed and thought out, but closer reading reveals gaps, biases and lack of depth. Our comparison of six versions shows a series of pitfalls and suggests precautions necessary when using AI assistance to make a literature review. Main issues are: (1) The bias of ignorance (you do not know what you do not get) in the selection of relevant papers. (2) Alignment and digital sycophancy: commercial AI models slavishly take you further in the direction they understand you give them, reinforcing biases. (3) Mainstreaming: because of their statistical nature, LLM productions tend to favor mainstream perspectives and content; in our case there was only 20% overlap between paper selections by humans and the LLM. (4) Limited capacity for creative restructuring, with vague and ambiguous statements. (5) Lack of critical perspective, coming from distant reading and political correctness. Most pitfalls can be addressed by prompting, but only if the user knows the domain well enough to detect them. There is a paradox: producing a good AI-assisted review requires expertise that comes from reading the literature, which is precisely what AI was meant to reduce. Overall, AI can improve the span and quality of the review, but the gain of time is not as massive as one would expect, and a press-button strategy leaving AI to do the work is a recipe for disaster. We conclude with recommendations for those who write, or assess, such LLM-augmented reviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。