arXiv:2508.15837cs.CLcs.AI2025-08被引 3

用对比分析验证大模型能否跨数据集迁移,省去重复训练

Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading

  • 比较STSB、Mohler与SPRAG三数据集的语义相似性
  • 发现现有模型在新数据上仍可保持高精度表现
  • 适合想少训练多复用的NLP研究者参考

构建特定数据集的模型需反复微调,成本高昂。本研究探究先进模型在已知数据集上训练后,能否有效迁移到未知文本数据集。选取两个成熟基准数据集STSB和Mohler,以及新提出的SPRAG作为未探索领域,通过稳健的相似性度量与统计方法进行细致对比分析。目标是深入理解先进模型在新场景下的适用性与可迁移性。研究结果有望重塑自然语言处理领域,使现有模型得以广泛复用于不同数据集,减少对资源密集型定制训练的需求,从而加速NLP发展并提升模型部署效率。

原文摘要 · Abstract (English)

Developing dataset-specific models involves iterative fine-tuning and optimization, incurring significant costs over time. This study investigates the transferability of state-of-the-art (SOTA) models trained on established datasets to an unexplored text dataset. The key question is whether the knowledge embedded within SOTA models from existing datasets can be harnessed to achieve high-performance results on a new domain. In pursuit of this inquiry, two well-established benchmarks, the STSB and Mohler datasets, are selected, while the recently introduced SPRAG dataset serves as the unexplored domain. By employing robust similarity metrics and statistical techniques, a meticulous comparative analysis of these datasets is conducted. The primary goal of this work is to yield comprehensive insights into the potential applicability and adaptability of SOTA models. The outcomes of this research have the potential to reshape the landscape of natural language processing (NLP) by unlocking the ability to leverage existing models for diverse datasets. This may lead to a reduction in the demand for resource-intensive, dataset-specific training, thereby accelerating advancements in NLP and paving the way for more efficient model deployment.

模型迁移语义相似性NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。