arXiv:2410.04981cs.CL2024-10EMNLP被引 4

用数据驱动方法自动识别科研写作中的严谨性标准,提升论文可信度评估。

On the Rigour of Scientific Writing: Criteria, Analysis, and Insights

  • 基于关键词提取与定义生成,构建可迁移的严谨性评估框架
  • 实证发现:明确表述能增强严谨感,模糊措辞会削弱可信度
  • 适用于机器学习、自然语言处理等多领域,助力论文质量分析

严谨性对科学研究至关重要,关乎结果的可复现性与有效性。然而,现有研究极少从计算角度建模严谨性,也缺乏对其在实际中能否有效衡量论文严谨性的分析。本文提出一种自下而上的数据驱动框架,用于自动识别并定义科研写作中的严谨性标准,包括关键词抽取、定义生成与关键标准识别。该框架具备领域无关性,可适配不同学科的严谨性评估需求。我们在机器学习与自然语言处理领域的两大顶会(ICLR 和 ACL)数据集上开展全面实验,验证了框架的有效性。进一步的语言学分析表明:明确表达确定性有助于提升严谨性感知,而使用建议性语气或概率不确定性则会降低这种感知。

原文摘要 · Abstract (English)

Rigour is crucial for scientific research as it ensures the reproducibility and validity of results and findings. Despite its importance, little work exists on modelling rigour computationally, and there is a lack of analysis on whether these criteria can effectively signal or measure the rigour of scientific papers in practice. In this paper, we introduce a bottom-up, data-driven framework to automatically identify and define rigour criteria and assess their relevance in scientific writing. Our framework includes rigour keyword extraction, detailed rigour definition generation, and salient criteria identification. Furthermore, our framework is domain-agnostic and can be tailored to the evaluation of scientific rigour for different areas, accommodating the distinct salient criteria across fields. We conducted comprehensive experiments based on datasets collected from two high impact venues for Machine Learning and NLP (i.e., ICLR and ACL) to demonstrate the effectiveness of our framework in modelling rigour. In addition, we analyse linguistic patterns of rigour, revealing that framing certainty is crucial for enhancing the perception of scientific rigour, while suggestion certainty and probability uncertainty diminish it.

科研写作严谨性评估自然语言分析数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。