arXiv:2411.01322cs.LGstat.ML2024-11被引 10

提出评估嵌入技术的标准化框架,覆盖三种典型应用场景。

FEET: A Framework for Evaluating Embedding Techniques

  • 构建三类使用场景:固定嵌入、少样本嵌入、全微调嵌入
  • 在情感分析与医疗领域开展案例验证,全面评估模型表现
  • 为表示学习研究提供可复现的基准协议,适合模型开发者参考

本研究提出FEET,一种用于指导基础模型开发与评测的标准化协议。尽管已有众多基准数据集用于评估这些模型,我们通过三个不同场景构建结构化评估流程,以全面理解其实际性能。定义了三大核心应用场景:固定嵌入、少样本嵌入与全微调嵌入。每个场景均通过两个案例研究加以说明:一项在情感分析任务中,另一项在医疗领域,展示该评估如何深入检验基础模型在科研应用中的有效性。建议将此协议作为未来推进表示学习模型研究的标准。

原文摘要 · Abstract (English)

In this study, we introduce FEET, a standardized protocol designed to guide the development and benchmarking of foundation models. While numerous benchmark datasets exist for evaluating these models, we propose a structured evaluation protocol across three distinct scenarios to gain a comprehensive understanding of their practical performance. We define three primary use cases: frozen embeddings, few-shot embeddings, and fully fine-tuned embeddings. Each scenario is detailed and illustrated through two case studies: one in sentiment analysis and another in the medical domain, demonstrating how these evaluations provide a thorough assessment of foundation models' effectiveness in research applications. We recommend this protocol as a standard for future research aimed at advancing representation learning models.

嵌入评估基础模型评测协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。