arXiv:2409.12060cs.CLcs.AI2024-09被引 9

构建多维度评测基准,精准评估文本同义检测模型真实能力

PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models

  • 设计多维度评测框架,覆盖10个数据集的细粒度测试
  • 发现模型在不同场景下存在性能权衡,单一数据集无法全面评估
  • 支持按使用场景调整严格度,适配大模型应用需求

文本同义检测任务长期面临挑战。现有研究对同义的理解过于简化,难以反映其复杂性。为此,我们构建了PARAPHRASUS——一个用于多维度评估、基准测试与模型选择的同义检测评测基准。该基准涵盖3项挑战、超过10个数据集(含8个重构和2个新标注),可揭示模型在细粒度评估下的性能权衡。同时,支持针对不同应用场景进行提示校准,使大模型适应特定严格程度。我们已将该基准及评测工具库开源至https://github.com/impresso/paraphrasus。

原文摘要 · Abstract (English)

The task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase dataset can leave uncertainty about their true semantic understanding. To alleviate this, we create PARAPHRASUS, a benchmark designed for multi-dimensional assessment, benchmarking and selection of paraphrase detection models. We find that paraphrase detection models under our fine-grained evaluation lens exhibit trade-offs that cannot be captured through a single classification dataset. Furthermore, PARAPHRASUS allows prompt calibration for different use cases, tailoring LLM models to specific strictness levels. PARAPHRASUS includes 3 challenges spanning over 10 datasets, including 8 repurposed and 2 newly annotated; we release it along with a benchmarking library at https://github.com/impresso/paraphrasus

同义检测评测基准大模型多维度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。