arXiv:2504.02486cs.AI2025-04

提出为真实实验数据加水印,提升科研数据可信度与模型稳定性。

We Need Improved Data Curation and Attribution in AI for Scientific Discovery

  • 建议对真实实验数据加水印,增强数据溯源能力。
  • 约四分之三的开源实验数据使用率低,需提升可发现性。
  • 仅对部分真实数据加水印,即可显著提升模型鲁棒性。

随着人类生成数据与合成数据的交互日益复杂,科学发现面临数据完整性与模型稳定性挑战。本文对比分析了合成数据与真实实验数据在科学研究中的作用。研究发现,开放平台上近四分之三的实验数据采用率较低,存在通过自动化手段提升其可发现性和可用性的空间。同时,辨别合成数据与真实实验数据的难度持续上升。为此,我们建议在现有合成数据检测工作基础上,加强真实实验数据的水印技术应用,以强化数据可追溯性与完整性。估算表明,即便仅对年产量不足一半的真实数据进行水印处理,也能有效维持模型稳健性,并推动合成数据与人工数据的均衡融合。

原文摘要 · Abstract (English)

As the interplay between human-generated and synthetic data evolves, new challenges arise in scientific discovery concerning the integrity of the data and the stability of the models. In this work, we examine the role of synthetic data as opposed to that of real experimental data for scientific research. Our analyses indicate that nearly three-quarters of experimental datasets available on open-access platforms have relatively low adoption rates, opening new opportunities to enhance their discoverability and usability by automated methods. Additionally, we observe an increasing difficulty in distinguishing synthetic from real experimental data. We propose supplementing ongoing efforts in automating synthetic data detection by increasing the focus on watermarking real experimental data, thereby strengthening data traceability and integrity. Our estimates suggest that watermarking even less than half of the real world data generated annually could help sustain model robustness, while promoting a balanced integration of synthetic and human-generated content.

数据溯源合成数据科研可信度水印技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。