arXiv:2505.24455cs.CL2025-05EMNLP被引 1
小而专的语料也能训练出高质量模型表示。
Domain Pre-training Impact on Representations
- 用小规模专业语料预训练,效果不输通用语料
- 混合语料效果取决于任务与专业语料分布相似度
- 适合关注领域适配的模型研究者
本实证研究分析了预训练语料对学习到的Transformer表示质量的影响。我们聚焦于仅通过预训练产生的表示质量。实验表明,使用小规模、专业的语料进行预训练也能获得有效的表示;而将通用语料与专业语料结合的效果,取决于目标任务与专业语料之间的分布相似性。
原文摘要 · Abstract (English)
This empirical study analyzes the effects of the pre-training corpus on the quality of learned transformer representations. We focus on the representation quality induced solely through pre-training. Our experiments show that pre-training on a small, specialized corpus can yield effective representations, and that the success of combining a generic and a specialized corpus depends on the distributional similarity between the target task and the specialized corpus.
预训练表示学习领域适应
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。