arXiv:2505.17832cs.CL2025-05中稿 · the 3rd TRR 318 Co…

构建首个生物领域人类式解释数据集,助力AI可解释性研究

Emerging categories in scientific explanations

  • 从生物技术文献中提取真实人类解释语句,构建高质量数据集
  • 提出6类与3类分类体系,3类标注一致性达0.667(Krippendorf Alpha)
  • 适合从事AI可解释性、科学文本理解的研究者使用

清晰有效的解释对人类理解与知识传播至关重要。近年来,科学解释研究的范畴已从社会科学扩展至机器学习与人工智能领域。机器学习决策的解释需具备影响力且贴近人类表达,但现有缺乏聚焦于人类生成、类人化解释的大规模数据集。本文通过从生物技术与生物物理领域的多种文献源(如PubMed的PMC开放获取子集)中提取说明性句子,构建了一个公开可用的数据集。该数据集包含两种分类体系(6类与3类),并基于数据归纳出多类别标注标准。通过评估标注者一致性,3类标注体系获得0.667的Krippendorf Alpha值,验证了分类框架的可靠性。

原文摘要 · Abstract (English)

Clear and effective explanations are essential for human understanding and knowledge dissemination. The scope of scientific research aiming to understand the essence of explanations has recently expanded from the social sciences to machine learning and artificial intelligence. Explanations for machine learning decisions must be impactful and human-like, and there is a lack of large-scale datasets focusing on human-like and human-generated explanations. This work aims to provide such a dataset by: extracting sentences that indicate explanations from scientific literature among various sources in the biotechnology and biophysics topic domains (e.g. PubMed's PMC Open Access subset); providing a multi-class notation derived inductively from the data; evaluating annotator consensus on the emerging categories. The sentences are organized in an openly-available dataset, with two different classifications (6-class and 3-class category annotation), and the 3-class notation achieves a 0.667 Krippendorf Alpha value.

可解释性科学文本数据集自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。