专为电信领域优化的向量模型,提升术语理解与检索精度。
T-VEC: A Telecom-Specific Vectorization Model with Enhanced Semantic Understanding via Deep Triplet Loss Fine-Tuning
- 基于gte-Qwen2-1.5B微调,使用三元组损失增强语义表征
- 在IETF RFC与厂商手册上测试,超越MPNet等主流模型
- 适合需要精准理解电信术语的工程与研发人员
电信行业的专业术语和复杂概念给通用自然语言处理模型带来持续挑战。通用嵌入模型常难以准确表达电信领域的语义,限制了其在信息检索和下游任务中的应用。我们提出T-VEC(Telecom Vectorization Model),一个基于gte-Qwen2-1.5B-instruct骨干网络、通过三元组损失目标微调的领域适配嵌入模型。微调在T-Embed数据集上进行,该数据集为高质量、大规模的电信概念、标准与运维场景数据集,虽含部分机密内容无法完全公开,但已开源75%以支持领域表示学习研究。在包含1500个查询-段落对的自定义基准测试中,T-VEC在IETF RFC与厂商手册数据上表现优于MPNet、BGE、Jina和E5,展现出更强的领域关联性与语义精确度。嵌入可视化进一步显示电信相关概念聚类紧密。我们发布T-VEC及其分词器,以支持电信领域内语义忠实的NLP应用。
原文摘要 · Abstract (English)
The specialized vocabulary and nuanced concepts of the telecommunications industry pose persistent challenges for standard Natural Language Processing (NLP) models. Generic embedding models often struggle to represent telecom-specific semantics, limiting their utility in retrieval and downstream tasks. We present T-VEC (Telecom Vectorization Model), a domain-adapted embedding model fine-tuned from the gte-Qwen2-1.5B-instruct backbone using a triplet loss objective. Fine-tuning was performed on T-Embed, a high-quality, large-scale dataset covering diverse telecom concepts, standards, and operational scenarios. Although T-Embed contains some proprietary material and cannot be fully released, we open source 75% of the dataset to support continued research in domain-specific representation learning. On a custom benchmark comprising 1500 query-passage pairs from IETF RFCs and vendor manuals, T-VEC surpasses MPNet, BGE, Jina and E5, demonstrating superior domain grounding and semantic precision in telecom-specific retrieval. Embedding visualizations further showcase tight clustering of telecom-relevant concepts. We release T-VEC and its tokenizer to support semantically faithful NLP applications within the telecom domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。