新指标SemR-p兼顾语义相关性与排序位置,更贴近人类对关键词重要性的判断。
Meaning in Order, Order in Meaning: Semantic R-precision for Keyphrase Evaluation

- 将语义相似度融入排名感知的R-Precision框架,关注早期出现的相关关键词。
- 在多个模型和数据集上验证,能更好区分不同生成质量的关键词列表。
- 适合评估关键词生成任务中用户中心的语义相关性与排序合理性。
自动关键词生成的质量评估仍具挑战性。传统指标或依赖精确词面匹配,或考虑语义相似度但忽略预测排序,均与人类对信息量和相关性的判断不符。本文提出语义R-Precision(SemR-p),将语义相似度融入排名感知的R-Precision框架。该指标从人类视角出发,受信息检索指标启发,奖励早期出现的语义相关关键词。通过广泛分析其语义敏感性、排序感知能力及判别力,结果表明SemR-p为关键词预测提供了互补评估视角,有助于更准确反映以用户为中心的相关性概念,同时补充传统的词面与语义匹配指标。
原文摘要 · Abstract (English)
Evaluating the quality of automatically generated keyphrases remains a complex challenge. Traditional metrics either rely on exact lexical matching or consider semantic similarity while ignoring prediction ranking, both of which misalign with how humans judge informativeness and relevance. We introduce Semantic R-Precision (SemR-p), a novel evaluation metric that integrates semantic similarity into the rank-aware R-Precision framework. Designed from a human-centric perspective and inspired by Information Retrieval metrics, SemR-p rewards semantically relevant keyphrases that appear early in the output list. We conducted extensive analyses to assess its semantic sensitivity, ranking awareness, and discriminative power across models and datasets. The results suggest that SemR-p offers a complementary lens for evaluating keyphrase predictions, helping to better reflect user-centred notions of relevance alongside traditional lexical and semantic matching metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。