提出新评估方法,让无监督词发现的词典质量更可信
Revisiting Lexicon Evaluation in Unsupervised Word Discovery
- 改进编辑距离,按聚类大小加权评估内部一致性
- 新增逆向指标,衡量真实词在聚类间分布是否合理
- 在合成与真实数据上验证,结果更贴近真实词典分布
从发现的词单元构建词典是零资源语音处理的核心目标。但现有评估指标——归一化编辑距离——存在固有偏差:对大聚类过度敏感,且忽略真实词在聚类间的分布情况。基于聚类理论,我们提出两个新指标:一是加权聚类大小的修正版内部一致性度量;二是反向度量,用于评估真实词在各聚类间的分布均匀性。在合成与真实词典上的实验表明,这两个指标联合使用时:(1) 与词典与真实分布的相似性相关性更高;(2) 对评价偏差更具鲁棒性。
原文摘要 · Abstract (English)
Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality? A common metric, normalized edit distance, averages the phoneme edit distances between discovered units in each cluster. We show that this metric has an inherent bias toward the quality of large clusters, inhibiting fair evaluation. Moreover, it ignores how well true classes are distributed across clusters. Based on established theory in clustering literature, we propose two metrics that address these shortcomings: a modified metric that weighs cluster size when assessing within-cluster consistency, and an inverse metric that assesses how true words are spread across clusters. Through experiments on synthetic and real-world lexicons, we demonstrate that combined, these metrics are: (1) more closely correlated with how similar a lexicon is to the ground-truth distribution, and (2) more robust to biases that skew lexicon evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。