arXiv:2603.29045cs.CV2026-03被引 3

用动态反例逼迫模型真正理解机制,而非仅通过测试

Let the Abyss Stare Back Adaptive Falsification for Autonomous Scientific Discovery

  • 设计可自适应生成反例的评估框架,主动挑战候选方案
  • 在可控实验中发现能跨环境迁移的新型损失函数FNG-CE
  • 适合追求真实科学发现而非单纯性能提升的研究者

自主科学发现正进入更危险的阶段:一旦评估标准固化,强大的搜索过程可能学会通过考试却未掌握任务本应揭示的机制。本文提出‘让深渊回望’——将评估从被动验证转为动态反例驱动的主动挑战。我们构建DASES框架,其中创新者、深渊反例生成器与机制因果提取器在固定科学契约下协同进化,生成可执行的科学成果与合乎科学的反例环境。在一个仅有一个可编辑位置的控制型损失发现任务中,DASES否决了静态验证通过的方案,识别出首个通过可接受反例检验的候选,发现了可跨合成发现环境迁移的FNG-CE损失函数,在包含ImageNet在内的标准基准上,其表现持续优于CE和CE+L2。

原文摘要 · Abstract (English)

Autonomous scientific discovery is entering a more dangerous regime: once the evaluator is frozen, a sufficiently strong search process can learn to win the exam without learning the mechanism the task was meant to reveal. This is the idea behind our title. To let the abyss stare back is to make evaluation actively push against the candidate through adaptive falsification, rather than passively certify it through static validation. We introduce DASES, a falsification-driven framework in which an Innovator, an Abyss Falsifier, and a Mechanistic Causal Extractor co-evolve executable scientific artifacts and scientifically admissible counterexample environments under a fixed scientific contract. In a controlled loss-discovery problem with a single editable locus, DASES rejects artifacts that static validation would have accepted, identifies the first candidate that survives the admissible falsification frontier, and discovers FNG-CE, a loss that transfers beyond the synthetic discovery environment and consistently outperforms CE and CE+L2 under controlled comparisons across standard benchmarks, including ImageNet.

科学发现反例生成自适应评估损失函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。