arXiv:2605.16000cs.SIcs.AI2026-05

用AI辅助人工审核论文引用,提升学术编辑的文献质量把关效率。

CitePrism: Human-in-the-Loop AI for Citation Auditing and Editorial Integrity

论文配图:CitePrism: Human-in-the-Loop AI for Citation Auditing and Editorial Integrity
图 1 · 摘自论文原文
  • 结合大模型推理与语义相似度,自动分析引用上下文和元数据。
  • 在104篇道路工程文献中,识别出所有不相关引用,误报需人工复核。
  • 适合期刊编辑部进行引用质量筛查,尤其关注伦理与准确性。

期刊编辑和审稿人需确保论文引用相关、准确、及时且合乎伦理,但当前引用审核仍以人工为主,流程分散且难以扩展。引用上下文、元数据质量、自引模式及参考文献完整性均影响引用是否合理支持论点。我们提出CitePrism,一种透明的混合决策支持框架,融合大模型上下文推理、嵌入式语义相似度、元数据验证、完整性标记及人工介入审核。CitePrism提取引用邻域,丰富参考文献元数据,计算融合相关性得分,提示元数据与自引问题,并支持可配置阈值筛选。在一项包含104个参考文献的道路工程案例研究中,与人工标注的相关性判断一致性达到Cohen's kappa = 0.429。在阈值tau = 17时,系统成功标记了所有人工判定为无关的引用,同时产生部分需分析师复核的误报。结果表明CitePrism可辅助保守型编辑筛查与引用质量分级,但尚未证明普遍适用性。该系统定位为试点阶段的决策支持工具,非自主违规检测或自动化编辑系统。需在更多论文、领域、标注者及部署场景下开展广义验证后方可投入实际使用。

原文摘要 · Abstract (English)

Editors and reviewers are expected to ensure that manuscripts cite relevant, accurate, current, and ethically appropriate literature, yet manuscript-level citation auditing remains largely manual, fragmented, and difficult to scale. Citation context, metadata quality, self-citation patterns, and bibliographic integrity all affect whether a reference appropriately supports a local claim. We present CitePrism, a transparent hybrid decision-support framework for editorial citation auditing that combines LLM-assisted contextual reasoning, embedding-based semantic similarity, metadata verification, integrity-oriented flags, and human-in-the-loop analyst review. CitePrism extracts citation neighborhoods, enriches reference metadata, computes fused relevance scores, surfaces metadata and self-citation review prompts, and supports configurable threshold-based triage. In a preliminary validation on a single case-study manuscript with 104 references from pavement engineering, agreement with human binary relevance labels reached Cohen's kappa = 0.429. At operating threshold tau = 17, CitePrism flagged all human-labeled irrelevant citations, while also producing false positives requiring analyst review. These results suggest that CitePrism may support conservative editorial screening and citation-quality triage, but they do not establish general editorial performance. CitePrism is intended as pilot-stage decision support, not as an autonomous misconduct detector or automated editorial decision system. Broader validation across manuscripts, domains, annotators, baselines, and deployment settings is required before operational use.

引用审计AI辅助学术诚信编辑工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。