arXiv:2604.04939cs.AI2026-04

提出新方法判断多源数据是否描述同一实体,无需转换特征值即可比较。

Proximity Measure of Information Object Features for Solving the Problem of Their Identification in Information Systems

  • 用概率和可能性分别衡量定量与定性特征的相似性
  • 无需特征值归一化或转换,直接比较不同来源数据
  • 适用于多源异构信息系统的实体识别,尤其适合不确定数据

本文提出一种新的定量-定性接近度度量方法,用于判断从多个独立来源进入统一信息资源的信息对象特征是否属于同一物理对象(观测对象)。该方法考虑了因测量误差导致的特征值差异,对定量特征采用概率度量,对定性特征使用可能性度量。通过验证其满足度量应有的公理体系,证明了方法的合理性。与已有方法不同,本方法无需对特征值进行变换以保证可比性。论文还提出了基于多样化特征组合的多种信息对象接近度计算变体,适用于复杂信息系统中的实体匹配任务。

原文摘要 · Abstract (English)

The paper considers a new quantitative-qualitative proximity measure for the features of information objects, where data enters a common information resource from several sources independently. The goal is to determine the possibility of their relation to the same physical object (observation object). The proposed measure accounts for the possibility of differences in individual feature values - both quantitative and qualitative - caused by existing determination errors. To analyze the proximity of quantitative feature values, the author employs a probabilistic measure; for qualitative features, a measure of possibility is used. The paper demonstrates the feasibility of the proposed measure by checking its compliance with the axioms required of any measure. Unlike many known measures, the proposed approach does not require feature value transformation to ensure comparability. The work also proposes several variants of measures to determine the proximity of information objects (IO) based on a group of diverse features.

信息融合实体识别相似度度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。