比较两种距离度量在偏斜特征密度下的分类效果
Supervised Pattern Recognition Involving Skewed Feature Densities
- 用k近邻法对比欧氏距离与相似性指数的分类性能
- 偏斜特征密度下,相似性指数表现更优
- 适合研究特征分布对分类影响的学者
模式识别是众多科学与技术活动的基础任务,但也面临诸多挑战,如特征选择与变换。本文通过k近邻监督分类方法,比较了欧氏距离与基于重合相似性指数的差异性度量,在一维与二维对称密度经不同变换后的特征上的分类性能。针对两组具有或无重叠的密度分布,评估了两类方法在区分相邻组交界点时的准确率。结果显示,对于右偏特征密度的数据集,基于重合相似性指数的差异性度量具有更强的分类潜力;同时发现,数据元素间的对比锐度可独立于监督分类性能。结果揭示了特征分布形态对分类器设计的重要影响。
原文摘要 · Abstract (English)
Pattern recognition constitutes a particularly important task underlying a great deal of scientific and technologica activities. At the same time, pattern recognition involves several challenges, including the choice of features to represent the data elements, as well as possible respective transformations. In the present work, the classification potential of the Euclidean distance and a dissimilarity index based on the coincidence similarity index are compared by using the k-neighbors supervised classification method respectively to features resulting from several types of transformations of one- and two-dimensional symmetric densities. Given two groups characterized by respective densities without or with overlap, different types of respective transformations are obtained and employed to quantitatively evaluate the performance of k-neighbors methodologies based on the Euclidean distance an coincidence similarity index. More specifically, the accuracy of classifying the intersection point between the densities of two adjacent groups is taken into account for the comparison. Several interesting results are described and discussed, including the enhanced potential of the dissimilarity index for classifying datasets with right skewed feature densities, as well as the identification that the sharpness of the comparison between data elements can be independent of the respective supervised classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。