arXiv:2505.14906cs.CLcs.SY2025-05被引 4

用大模型提取6G通信知识中的结构化实体,效率提升9倍

Understanding 6G through Language Models: A Case Study on LLM-aided Structured Entity Extraction in Telecom Domain

  • 基于令牌高效表示和分层并行解码,提升实体识别精度
  • 在6GTech数据集上准确率超基线方法,处理速度提升5至9倍
  • 适合通信智能、AI原生网络等方向的研究者参考

知识理解是未来6G网络推进网络智能化与人工智能原生架构的基础。在此范式下,信息抽取在将分散的通信知识转化为结构化格式方面起关键作用,助力各类AI模型更好理解网络术语。本文提出一种基于语言模型的信息抽取技术——电信结构化实体抽取(TeleSEE),采用令牌高效的表示方法预测实体类型与属性键,减少输出令牌数量并提升预测准确率。同时,TeleSEE引入分层并行解码机制,在标准编码器-解码器架构中融合额外提示与解码策略,优化实体抽取任务。为更精准评估该技术在通信领域的性能,研究构建了名为6GTech的数据集,包含2390个句子和23747个词,来源于100多篇6G相关技术文献。实验表明,所提TeleSEE方法在准确率上优于其他基线方法,并实现5至9倍的样本处理速度提升。

原文摘要 · Abstract (English)

Knowledge understanding is a foundational part of envisioned 6G networks to advance network intelligence and AI-native network architectures. In this paradigm, information extraction plays a pivotal role in transforming fragmented telecom knowledge into well-structured formats, empowering diverse AI models to better understand network terminologies. This work proposes a novel language model-based information extraction technique, aiming to extract structured entities from the telecom context. The proposed telecom structured entity extraction (TeleSEE) technique applies a token-efficient representation method to predict entity types and attribute keys, aiming to save the number of output tokens and improve prediction accuracy. Meanwhile, TeleSEE involves a hierarchical parallel decoding method, improving the standard encoder-decoder architecture by integrating additional prompting and decoding strategies into entity extraction tasks. In addition, to better evaluate the performance of the proposed technique in the telecom domain, we further designed a dataset named 6GTech, including 2390 sentences and 23747 words from more than 100 6G-related technical publications. Finally, the experiment shows that the proposed TeleSEE method achieves higher accuracy than other baseline techniques, and also presents 5 to 9 times higher sample processing speed.

6G大模型知识抽取通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。