arXiv:2605.28190cs.CL2026-05被引 1

提出动态评估框架HTEB,多维度测试文本嵌入模型的鲁棒性。

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

论文配图:The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness
图 1 · 摘自论文原文
  • 用LLM随机变换输入,从词汇、长度、语言三轴动态测试模型鲁棒性。
  • 16个模型在32个数据集上表现分化,规模提升不弥补鲁棒性差距。
  • 英语数据更易受干扰,凸显当前基准的局限性,适合模型评估者参考。

现有嵌入基准(如MTEB)为每个模型提供单一分数,隐含将鲁棒性视为静态标量属性。我们认为嵌入鲁棒性是多维的,因模型对不同变化类型响应各异,需动态评估以揭示静态基准掩盖的失败。本文提出更难的文本嵌入基准(HTEB),通过大语言模型在评估时随机变换输入,沿词汇/风格、长度和语言三个实际可解释轴动态挑战模型鲁棒性。在覆盖42种语言的32个数据集上评估16个开源嵌入模型,基于4,800次英文子样本的人工评分验证,发现三大规律:(1) 模型在各轴表现出特定且部分解耦的鲁棒性特征;(2) 三种模型家族中,规模增大虽提升绝对得分,但未缩小原始与变换评估间的差距,且仅在语言轴上改善明显;(3) 英语数据集比多语言数据集更易受HTEB变换影响。结果表明,HTEB能识别模型在部署相关轴上的优劣势,挑战现有基准,主张采用多维、动态的鲁棒性评估方式。

原文摘要 · Abstract (English)

Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding robustness is multidimensional, since models respond differently to different types of variation, and requires dynamic evaluation to expose failures hidden by static benchmarks. We introduce the Harder Text Embedding Benchmark (HTEB), a dynamic evaluation framework that challenges model robustness along three practically interpretable axes (Lexical/Stylistic, Length and Language) by stochastically transforming inputs at evaluation time with an LLM. Evaluating 16 open-weight embedding models on 32 datasets covering 42 languages under transformations validated by 4,800 human ratings on an English subsample, we find three patterns: (1) Models exhibit specific, partly decoupled robustness profiles across axes. (2) Across three model families, scale increases absolute scores but does not close the gap between original and transformed evaluations. Here, scaling tends to improve specifically the Language axis. (3) English datasets are more sensitive to HTEB transformations than multilingual datasets. This demonstrates that HTEB identifies strengths and weaknesses of models along deployment-relevant axes, challenging current embedding benchmarks and arguing for multidimensional, dynamic robustness evaluation.

嵌入模型鲁棒性评估多维度动态测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。