构建工具链,让机器自动识别科学文本中的失实信息。
Machine Understanding of Scientific Language
- 设计新方法检测需核查的科学主张,支持零样本事实验证
- 提出对抗性主张生成与跨源域适应技术,提升小数据下性能
- 适合关注科学传播、虚假信息检测的研究者与政策制定者
科学信息表达了人类对自然的理解,主要以论文、新闻及社交媒体讨论等形式传播。然而,并非所有科学文本都忠实反映真实科学。近年来,网络上科学文本数量激增,自动识别其可信度已成为社会重要议题。本文致力于构建数据集、方法与工具,实现科学语言的机器理解,以大规模分析科学传播。研究涵盖自动事实核查、有限数据学习和科学文本处理三大方向,提出多项创新:可核查主张识别、对抗性主张生成、多源域适应、众包标签学习、引用价值检测、零样本科学事实核查、夸大性科学主张检测,以及科学传播中信息变化程度建模。关键成果表明,这些方法能有效利用少量科学文本,识别误导性陈述并揭示科学传播机制的新洞见。
原文摘要 · Abstract (English)
Scientific information expresses human understanding of nature. This knowledge is largely disseminated in different forms of text, including scientific papers, news articles, and discourse among people on social media. While important for accelerating our pursuit of knowledge, not all scientific text is faithful to the underlying science. As the volume of this text has burgeoned online in recent years, it has become a problem of societal importance to be able to identify the faithfulness of a given piece of scientific text automatically. This thesis is concerned with the cultivation of datasets, methods, and tools for machine understanding of scientific language, in order to analyze and understand science communication at scale. To arrive at this, I present several contributions in three areas of natural language processing and machine learning: automatic fact checking, learning with limited data, and scientific text processing. These contributions include new methods and resources for identifying check-worthy claims, adversarial claim generation, multi-source domain adaptation, learning from crowd-sourced labels, cite-worthiness detection, zero-shot scientific fact checking, detecting exaggerated scientific claims, and modeling degrees of information change in science communication. Critically, I demonstrate how the research outputs of this thesis are useful for effectively learning from limited amounts of scientific text in order to identify misinformative scientific statements and generate new insights into the science communication process
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。