arXiv:2605.12226cs.IR2026-05

用众包提升本体匹配验证质量,解决大模型带来的标注负担。

Crowd-OM: Crowdsourcing for Ontology Matching Validation

论文配图:Crowd-OM: Crowdsourcing for Ontology Matching Validation
图 1 · 摘自论文原文
  • 设计三类领域专用机制保障众包质量
  • 支持不同标注者与场景,有效减少误差
  • 适合需要高精度验证的本体集成项目

大型语言模型(LLMs)在本体匹配(OM)中展现出强大能力,能发现更多匹配候选,但传统依赖领域专家的验证方式已不堪重负。尽管众包可扩大验证参与范围,但标注者差异带来的偏见与错误会降低质量。本文探索众包用于本体匹配验证,并提出新型质量保障系统 Crowd-OM。设计了三种领域特定机制:差异信任度、一致性预填充和时间依赖意见,以确保众包结果可靠性。Crowd-OM 可无缝集成至现有 OM 系统,实现人机协同验证。评估表明其在应对多样标注者与不同场景时表现有效。文章还讨论了两个实际应用案例及当前局限性,为未来优化提供方向。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) pose new challenges for ontology matching (OM). While OM systems built on LLMs have shown remarkable capabilities in discovering more matching candidates, traditional OM validation that relies on domain experts has become overwhelming. Although crowdsourcing can expose OM validation to large online communities, bias and errors from diverse annotators can reduce the quality of crowdsourcing for OM validation. In this study, we explore the use of crowdsourcing for OM validation and introduce a novel quality assurance crowdsourcing system called Crowd-OM. We propose three domain-specific mechanisms (differential trustworthiness, coherence pre-filling, and time-dependent opinion) to ensure the quality of crowdsourcing for OM validation. Crowd-OM can be integrated with existing OM systems to enable human-in-the-loop validation. The evaluation shows its effectiveness in handling various annotators and different annotation scenarios. We discuss two real-world use cases and current limitations for improvement.

本体匹配众包验证大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。