arXiv:2601.09753cs.CYcs.AI2026-01

为AI科研评估工具建立可问责的批判性实践规范

Critically Engaged Pragmatism: Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools

  • 提出批判性务实主义框架,强调评估工具需针对具体用途验证可靠性
  • 要求开发者透明披露设计、训练与评测细节,以识别误差与偏见来源
  • 适合关注AI科研可信度、评估工具伦理的研究者与政策制定者

AI科学评估工具旨在衡量研究可信度。然而,如同传统影响因子,其使用常脱离语境且被误用。为此,本文提出批判性务实主义作为科学规范,敦促科学共同体审视评估工具的目的及其目的特定可靠性。为推动该规范,工具开发者应透明、完整地报告设计、训练与基准测试细节,以支持对目的特定可靠性、各类误差风险及偏见的评估。随着新形式的错误、偏见和投机行为的出现,最佳实践报告标准也应持续更新。在此框架下,AI科学评估工具并非客观的可信度仲裁者,而是批判性话语实践的对象,其最终支撑科学共同体的可信度基础。

原文摘要 · Abstract (English)

AI science evaluation tools aim to assess research credibility. As with traditional metrics such as impact factors, their edicts can be decontextualised and repurposed in problematic ways. To address this, I propose Critically-Engaged Pragmatism as a scientific norm enjoining scientific communities to scrutinise the purposes and purpose-specific reliability of AI science evaluation tools. To foster Critically Engaged Pragmatism, creators of AI science evaluation tools should transparently and fully report design, training, and benchmarking details to facilitate assessments of purpose-specific reliability, liability to different types of error, and bias. What count as best practices for the transparent reporting of AI science evaluation tools should be updated as new forms of error, bias, and gamesmanship are discovered. Under this framework, AI science evaluation tools are not objective arbiters of scientific credibility. Rather, they are the object of critical discursive practices that ultimately ground the credibility of scientific communities.

AI评估科学伦理批判性思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。