arXiv:2510.05744cs.CLastro-ph.IM2025-10中稿 · Ontology Matching …

用多源数据和大模型实现天文观测设施名称标准化。

Adaptive and Multi-Source Entity Matching for Name Standardization of Astronomical Observation Facilities

  • 融合八类语义资源,用NLP技术计算实体匹配分
  • 通过大模型验证映射合理性,确保数据可发现
  • 适合天文数据集成与跨平台知识库构建者

本研究致力于构建天文观测设施的多源映射方法。为比较两个实体,采用可调参数的评分机制与自然语言处理技术(词袋、序列、表面匹配)对来自八个语义资源(包括Wikidata和天文学专有资源)提取的实体进行匹配。利用标签、定义、描述、外部标识符及波段、发射日期、资助机构等特定属性。最终借助大语言模型(LLM)判断映射建议是否合理并提供解释,保障所验证同义对的合理性与FAIR性。生成的映射包含多源同义集,每个实体仅对应一个标准化标签,将用于我们的名称解析API,并集成至国际虚拟天文台联盟(IVOA)词汇表与OntoPortal-Astro平台。

原文摘要 · Abstract (English)

This ongoing work focuses on the development of a methodology for generating a multi-source mapping of astronomical observation facilities. To compare two entities, we compute scores with adaptable criteria and Natural Language Processing (NLP) techniques (Bag-of-Words approaches, sequential approaches, and surface approaches) to map entities extracted from eight semantic artifacts, including Wikidata and astronomy-oriented resources. We utilize every property available, such as labels, definitions, descriptions, external identifiers, and more domain-specific properties, such as the observation wavebands, spacecraft launch dates, funding agencies, etc. Finally, we use a Large Language Model (LLM) to accept or reject a mapping suggestion and provide a justification, ensuring the plausibility and FAIRness of the validated synonym pairs. The resulting mapping is composed of multi-source synonym sets providing only one standardized label per entity. Those mappings will be used to feed our Name Resolver API and will be integrated into the International Virtual Observatory Alliance (IVOA) Vocabularies and the OntoPortal-Astro platform.

名称标准化多源融合天文数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。