arXiv:2409.05367cs.CL2024-09ACL综述被引 5

让论文评审变得可解释,通过结构化推理追踪专家思考过程

STRICTA: Structured Reasoning in Critical Text Assessment for Peer Review and Beyond

  • 将评审过程拆解为因果关联的步骤图,实现透明化评估
  • 收集40位专家对20余篇生物医学论文的4000+推理步骤
  • 验证大模型能否模仿专家推理,助力人机协同评审

关键文本评估是事实核查、同行评审和作文评分等专家活动的核心。然而,现有方法将其视为黑箱,限制了可解释性与人机协作。为此,我们提出结构性推理在关键文本评估中的应用(STRICTA),一种将文本评估建模为显式、分步推理过程的新框架。STRICTA基于因果理论(Pearl, 1995),将评估分解为相互连接的推理节点构成的图结构,并依据专家交互数据填充该图,用于研究评估流程并促进人机协作。我们正式定义了STRICTA,并应用于生物医学论文评估研究,构建了一个包含4000余条推理步骤的数据集,来自约40位生物医学专家对20多篇论文的评估。利用该数据集,我们实证研究了专家在关键文本评估中的推理模式,并探究大语言模型是否能在这些工作流中模仿并支持专家。由此产生的工具与数据集为研究文本评估中的人机协同推理铺平了道路,适用于同行评审及其他领域。

原文摘要 · Abstract (English)

Critical text assessment is at the core of many expert activities, such as fact-checking, peer review, and essay grading. Yet, existing work treats critical text assessment as a black box problem, limiting interpretability and human-AI collaboration. To close this gap, we introduce Structured Reasoning In Critical Text Assessment (STRICTA), a novel specification framework to model text assessment as an explicit, step-wise reasoning process. STRICTA breaks down the assessment into a graph of interconnected reasoning steps drawing on causality theory (Pearl, 1995). This graph is populated based on expert interaction data and used to study the assessment process and facilitate human-AI collaboration. We formally define STRICTA and apply it in a study on biomedical paper assessment, resulting in a dataset of over 4000 reasoning steps from roughly 40 biomedical experts on more than 20 papers. We use this dataset to empirically study expert reasoning in critical text assessment, and investigate if LLMs are able to imitate and support experts within these workflows. The resulting tools and datasets pave the way for studying collaborative expert-AI reasoning in text assessment, in peer review and beyond.

文本评估人机协作推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。