arXiv:2606.27342cs.DBcs.AI2026-06被引 1

剖析低资源实体匹配中领域感知分布对齐的效果与限制

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

论文配图:Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching
图 1 · 摘自论文原文
  • 分析BEACON框架在不同数据条件下的表现机制
  • 发现分布对齐在少样本场景下显著提升匹配准确率
  • 适合关注少样本实体匹配与领域自适应的研究者

实体匹配(EM)是数据集成中的核心步骤,用于判断来自不同来源的记录是否指向同一真实实体。近期研究引入了领域信息和低资源学习技术,以增强EM系统在真实场景中的适应性。尽管这些方法表现出色,但其在不同数据约束和监督水平下的实际表现仍不清晰。本文深入研究了一种先进的低资源、领域感知实体匹配方法——BEACON,通过一系列针对性实验,评估其在不同算法选择和数据可用性条件下的性能变化,进一步揭示了分布对齐的作用机制及BEACON框架的行为特征。

原文摘要 · Abstract (English)

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine whether they refer to the same real-world entity. Recent work has incorporated domain information and low-resource learning techniques to better adapt EM systems to realistic settings. While these approaches have demonstrated strong performance, it remains unclear how they behave under varying data constraints and levels of supervision in practice. In this paper, we investigate a state-of-the-art method for low-resource, domain-aware EM--BEACON--and study how its performance is affected by different algorithmic choices and data availability conditions. We conduct a series of targeted experiments to evaluate these variations, providing deeper insight into the role of distribution alignment and the behavior of the BEACON framework.

实体匹配低资源学习分布对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。