提出无需真实新颖性标签的新颖性评估基准,验证指标有效性。
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
- 基于三个公理设计可控数据池操作测试新颖性评分
- 发现人类评审新颖性评分与质量存在混淆,现有指标多依赖表面冗余
- 建议融合嵌入与大模型方法,适合研究新颖性评估的学者
科学论文新颖性的严格评估对人类科学家而言仍具挑战。随着人工智能科学家兴起,该任务自动化和可靠性愈发重要,避免资源浪费在已有研究上。由于难以量化真实新颖性,现有新颖性指标常依赖噪声大、有偏倚的信号,如引用次数或同行评审分数。本文提出一个基准,可在无需明确新颖性标签的情况下比较新颖性指标。该基准通过三组受控数据池操作检验评分是否合理:当数据池覆盖更多论文内容时评分下降,当池子相关性降低时评分上升,当池子时间更晚时评分下降。我们发现,即便是最直接的人类信号——ICLR评审的新颖性评分——也与质量纠缠不清。在十种系统(从嵌入度量到AI科学家新颖性检测)中,表面冗余已基本解决,但概念冗余仍未克服;嵌入式与大模型方法各有优势,未来指标应结合二者以提升评估效果。我们公开了基准与评估代码,推动该领域研究。
原文摘要 · Abstract (English)
The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interest in AI scientists, it is becoming more and more important that this task be automatable and reliable, lest attention and compute be wasted on ideas that have already been explored. Due to the challenge of quantifying ground-truth novelty, however, existing novelty metrics generally validate against noisy, confounded signals such as citation counts or peer review scores. We introduce a benchmark that compares novelty metrics without requiring explicit novelty labels. It tests whether scores move correctly under controlled pool manipulations, organized under three axioms requiring that scores fall as the pool covers more of a paper's content, rise as the pool loses relevance, and fall as the pool moves later in time. Indeed, we show that even for the most direct human signal, ICLR reviewer novelty scores, the axis of novelty is entangled with quality. Across ten systems, from embedding metrics to AI scientist novelty checks, we find that surface redundancy is largely solved but conceptual redundancy is not, and that embedding metrics and LLM-based metrics both have their places; future metrics should combine such approaches for better novelty evaluation. We release our benchmark and evaluation code to enable this research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。