测试捷克语嵌入模型发现,语义相似性好不等于翻译评估表现优。
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
- 用变换数据集和语义相似度任务评估嵌入模型能力
- 内在评估优的模型在翻译任务中表现不一致,甚至不如平滑嵌入
- 提示需更关注下游任务设计,而非单纯优化语义探针
本文通过内在与外在评估范式对比捷克语专用与多语言句子嵌入模型。内在评估使用Costra(复杂句式变换数据集)及多个语义文本相似度(STS)基准,检验嵌入对语义相似性、时间特征和风格变化的捕捉能力。外在评估中,以COMET为基础的指标微调各嵌入模型进行机器翻译评估。实验揭示出有趣断层:内在语义相似性表现优异的模型,在下游翻译评估任务中并未持续领先;相反,嵌入空间看似过度平滑的模型经微调后反而取得出色结果。这表明语义属性探测与下游任务性能间关系复杂,凸显了构建‘可操作化语义’或深化下游任务数据集(如翻译评估)研究的必要性。
原文摘要 · Abstract (English)
In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra, a complex sentence transformation dataset, and several Semantic Textual Similarity (STS) benchmarks to assess the ability of the embeddings to capture linguistic phenomena such as semantic similarity, temporal aspects, and stylistic variations. In the extrinsic evaluation, we fine-tune each embedding model using COMET-based metrics for machine translation evaluation. Our experiments reveal an interesting disconnect: models that excel in intrinsic semantic similarity tests do not consistently yield superior performance on downstream translation evaluation tasks. Conversely, models with seemingly over-smoothed embedding spaces can, through fine-tuning, achieve excellent results. These findings highlight the complex relationship between semantic property probes and downstream task, emphasizing the need for more research into 'operationalizable semantics' in sentence embeddings, or more in-depth downstream tasks datasets (here translation evaluation)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。