arXiv:2503.11376cs.CLcs.AI2025-03中稿 · Publication in the…被引 4

用规则匹配检测学术文本中的不确定性表达,效果优于大模型。

Annotating Scientific Uncertainty: A comprehensive model using linguistic patterns and comparison with existing approaches

  • 结合模式匹配与句法分析,识别文本中的不确定性表述。
  • 准确率达0.808,在科学不确定性检测中优于大模型。
  • 适合需要可解释性与领域适配性的研究场景。

UnScientify 是一个用于检测学术全文中科学不确定性的系统。该系统采用弱监督技术,识别科学文本中口头表达的不确定性及其作者引用。其核心方法基于多阶段流水线,整合了片段模式匹配、复杂句分析和作者引用核查。该方法简化了标注流程,覆盖多种不确定性表达形式,支持信息检索、文本挖掘和科学文档处理等应用。评估结果显示,尽管现代大语言模型表现强劲,但使用传统技术的UnScientify在科学不确定性检测任务中表现更优,准确率达到0.808。这一发现凸显了规则基础与模式匹配策略在资源效率、可解释性和领域适应性关键场景中的持续优势。

原文摘要 · Abstract (English)

UnScientify, a system designed to detect scientific uncertainty in scholarly full text. The system utilizes a weakly supervised technique to identify verbally expressed uncertainty in scientific texts and their authorial references. The core methodology of UnScientify is based on a multi-faceted pipeline that integrates span pattern matching, complex sentence analysis and author reference checking. This approach streamlines the labeling and annotation processes essential for identifying scientific uncertainty, covering a variety of uncertainty expression types to support diverse applications including information retrieval, text mining and scientific document processing. The evaluation results highlight the trade-offs between modern large language models (LLMs) and the UnScientify system. UnScientify, which employs more traditional techniques, achieved superior performance in the scientific uncertainty detection task, attaining an accuracy score of 0.808. This finding underscores the continued relevance and efficiency of UnScientify's simple rule-based and pattern matching strategy for this specific application. The results demonstrate that in scenarios where resource efficiency, interpretability, and domain-specific adaptability are critical, traditional methods can still offer significant advantages.

不确定性检测规则匹配可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。