arXiv:2602.15712cs.CVcs.AI2026-02

提出先定标准后赋语义的图像结构发现方法,提升跨领域可复现性。

Criteria-first, semantics-later: reproducible structure discovery in image-based sciences

  • 先根据显式标准提取无语义结构,再映射到具体领域标签
  • 在跨传感器、跨机构场景下保持结构稳定,避免标签漂移问题
  • 适合长期监测与数字孪生,支持可复现科学与AI应用

在自然与生命科学中,图像已成为主要测量手段,但主流分析仍采用语义先行范式——通过预测或强制特定领域标签来恢复结构。该范式在开放科学发现、跨传感器/站点可比性及长期监测等关键场景中系统性失效,因领域本体与标签集会随文化、机构和生态变化而漂移。本文提出一种演绎式反转:标准优先、语义后置。构建统一框架,将标准定义的、无语义的结构提取与下游语义映射分离,为图像科学提供通用可复现分析架构。可复现科学要求第一层分析基于标准驱动、无语义的结构发现,生成由明确最优性准则定义的稳定分区、结构场或层级。语义不被抛弃,而是作为从已发现结构产物到领域本体或词汇的显式映射,支持多重解释与显式跨映射,无需重写上游提取。该思想基于控制论、观察即区分及信息论中信息与意义的分离,跨领域证据表明:当标签无法扩展时,标准优先组件必然出现。最后,讨论了超越分类准确率的验证方式,以及将结构产物视为可复现、适配AI的数字对象用于长期监测与数字孪生的前景。

原文摘要 · Abstract (English)

Across the natural and life sciences, images have become a primary measurement modality, yet the dominant analytic paradigm remains semantics-first. Structure is recovered by predicting or enforcing domain-specific labels. This paradigm fails systematically under the conditions that make image-based science most valuable, including open-ended scientific discovery, cross-sensor and cross-site comparability, and long-term monitoring in which domain ontologies and associated label sets drift culturally, institutionally, and ecologically. A deductive inversion is proposed in the form of criteria-first and semantics-later. A unified framework for criteria-first structure discovery is introduced. It separates criterion-defined, semantics-free structure extraction from downstream semantic mapping into domain ontologies or vocabularies and provides a domain-general scaffold for reproducible analysis across image-based sciences. Reproducible science requires that the first analytic layer perform criterion-driven, semantics-free structure discovery, yielding stable partitions, structural fields, or hierarchies defined by explicit optimality criteria rather than local domain ontologies. Semantics is not discarded; it is relocated downstream as an explicit mapping from the discovered structural product to a domain ontology or vocabulary, enabling plural interpretations and explicit crosswalks without rewriting upstream extraction. Grounded in cybernetics, observation-as-distinction, and information theory's separation of information from meaning, the argument is supported by cross-domain evidence showing that criteria-first components recur whenever labels do not scale. Finally, consequences are outlined for validation beyond class accuracy and for treating structural products as FAIR, AI-ready digital objects for long-term monitoring and digital twins.

图像分析可复现科学结构发现数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。