综述语义文本相似度最新进展,涵盖模型、方法与应用挑战
Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey
- 系统梳理六类技术:基于Transformer、对比学习、领域适配等
- 指出FarSSiBERT、AspectCSE等模型达到新基准,医疗金融领域有定制方案
- 适合关注NLP语义理解、模型优化的研究者与工程师参考
自2021年以来,语义文本相似度(STS)研究迅速发展,得益于Transformer架构、对比学习和领域特定技术的进步。本文综述了六大关键方向:基于Transformer的模型、对比学习、领域聚焦解决方案、多模态方法、图结构方法以及知识增强技术。近期模型如FarSSiBERT和DeBERTa-v3表现出卓越性能,而AspectCSE等对比方法建立了新基准。针对医疗文本的CXR-BERT与金融领域的Financial-STS展示了领域适配的有效性。此外,多模态、图结构及知识融合模型进一步提升了语义理解能力。通过整合分析这些进展,本综述为当前方法、实际应用及未解决问题提供了洞见,旨在帮助研究者与实践者应对快速演进的STS领域,揭示新兴趋势与未来机遇。
原文摘要 · Abstract (English)
Semantic Textual Similarity (STS) research has expanded rapidly since 2021, driven by advances in transformer architectures, contrastive learning, and domain-specific techniques. This survey reviews progress across six key areas: transformer-based models, contrastive learning, domain-focused solutions, multi-modal methods, graph-based approaches, and knowledge-enhanced techniques. Recent transformer models such as FarSSiBERT and DeBERTa-v3 have achieved remarkable accuracy, while contrastive methods like AspectCSE have established new benchmarks. Domain-adapted models, including CXR-BERT for medical texts and Financial-STS for finance, demonstrate how STS can be effectively customized for specialized fields. Moreover, multi-modal, graph-based, and knowledge-integrated models further enhance semantic understanding and representation. By organizing and analyzing these developments, the survey provides valuable insights into current methods, practical applications, and remaining challenges. It aims to guide researchers and practitioners alike in navigating rapid advancements, highlighting emerging trends and future opportunities in the evolving field of STS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。