让文档相似度计算聚焦具体方面,效果远超传统全局方法。
The Critical Role of Aspects in Measuring Document Similarity
- 基于指定方面计算相似度,提升可解释性
- 与人类判断一致率高出约80%(对比全局方法)
- 适合需要精准语义比较的研究者和应用
我们提出ASPECTSIM,一种简单且可解释的框架,要求在显式指定的方面下计算文档相似度,不同于传统的整体性方法。在新构建的26,000个方面-文档对基准上实验发现,使用直接GPT-4o提示实现的ASPECTSIM,相比无显式方面的整体相似度,人类与机器判断一致性显著提升约80%。这些结果强调了在测量文档相似度时显式考虑方面的必要性,并呼吁修订现有标准做法。随后,我们利用16个小型开源大模型和9个嵌入模型进行了大规模元评估,旨在使ASPECTSIM可访问且可复现。直接提示大模型生成ASPECTSIM评分效果不佳(人类-机器一致性仅20-30%),但两阶段优化使其一致性提升约140%。尽管如此,其表现仍远低于GPT-4o模型,表明小型开源大模型在捕捉方面条件相似性方面仍有差距。
原文摘要 · Abstract (English)
We introduce ASPECTSIM, a simple and interpretable framework that requires conditioning document similarity on an explicitly specified aspect, which is different from the traditional holistic approach in measuring document similarity. Experimenting with a newly constructed benchmark of 26K aspect-document pairs, we found that ASPECTSIM, when implemented with direct GPT-4o prompting, achieves substantially higher human-machine agreement ($\approx$80% higher) than the same for holistic similarity without explicit aspects. These findings underscore the importance of explicitly accounting for aspects when measuring document similarity and highlight the need to revise standard practice. Next, we conducted a large-scale meta-evaluation using 16 smaller open-source LLMs and 9 embedding models with a focus on making ASPECTSIM accessible and reproducible. While directly prompting LLMs to produce ASPECTSIM scores turned out be ineffective (20-30% human-machine agreement), a simple two-stage refinement improved their agreement by $\approx$140%. Nevertheless, agreement remains well below that of GPT-4o-based models, indicating that smaller open-source LLMs still lag behind large proprietary models in capturing aspect-conditioned similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。