arXiv:2509.21310cs.AI2025-09被引 5

SAGE benchmark揭示嵌入模型在语义理解中的真实短板。

SAGE: A Realistic Benchmark for Semantic Understanding

  • 构建五维对抗性评测框架,覆盖人类偏好、噪声鲁棒性等维度
  • 发现顶尖模型在人类偏好上得分0.682,但信息敏感度仅0.794
  • 揭示模型性能存在显著权衡,适合评估真实场景部署能力

随着大语言模型在传统基准上表现优异,亟需更严苛的评估框架来深入探测语义理解能力。我们提出SAGE(语义对齐与泛化评估),一个针对嵌入模型和相似性度量的严谨基准,涵盖五大维度:人类偏好对齐、变换鲁棒性、信息敏感性、聚类性能与检索鲁棒性。不同于聚焦单一能力的现有基准,SAGE通过对抗性条件、噪声变换及细微的人类判断任务,在30多个数据集上进行评估。对9种嵌入模型和经典度量的全面测试显示显著性能差距:例如,OpenAI text-embedding-3-large在人类偏好对齐上达0.682,优于最佳经典度量(0.591);但在信息敏感性任务中,Jaccard相似度达0.905,远超顶级嵌入模型的0.794。同时发现,text-embedding-3-small聚类性能最高(0.483),但鲁棒性最低(0.011)。SAGE揭示了当前语义理解能力的关键局限,为实际部署提供更真实的模型鲁棒性评估。

原文摘要 · Abstract (English)

As large language models (LLMs) achieve strong performance on traditional benchmarks, there is an urgent need for more challenging evaluation frameworks that probe deeper aspects of semantic understanding. We introduce SAGE (Semantic Alignment & Generalization Evaluation), a rigorous benchmark designed to assess both embedding models and similarity metrics across five categories: Human Preference Alignment, Transformation Robustness, Information Sensitivity, Clustering Performance, and Retrieval Robustness. Unlike existing benchmarks that focus on isolated capabilities, SAGE evaluates semantic understanding through adversarial conditions, noisy transformations, and nuanced human judgment tasks across 30+ datasets. Our comprehensive evaluation of 9 embedding models and classical metrics reveals significant performance gaps, with no single approach excelling across all dimensions. For instance, while state-of-the-art embedding models like OpenAI's text-embedding-3-large dominate in aligning with human preferences (0.682 vs. 0.591 for the best classical metric), they are significantly outperformed by classical metrics on information sensitivity tasks, where Jaccard Similarity achieves a score of 0.905 compared to the top embedding score of 0.794. SAGE further uncovers critical trade-offs: OpenAI's text-embedding-3-small achieves the highest clustering performance (0.483) but demonstrates extreme brittleness with the lowest robustness score (0.011). SAGE exposes critical limitations in current semantic understanding capabilities and provides a more realistic assessment of model robustness for real-world deployment.

语义理解嵌入模型评测基准鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。