arXiv:2511.13131cs.AIcs.CV2025-11

为电信领域打造多模态大模型评测基准与适配模型

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

  • 构建电信场景下的多模态任务评测集,涵盖文本与图像
  • 微调后模型在实际任务中性能显著提升
  • 揭示现有多模态模型短板,指导后续研究

大型语言模型(LLMs)在自动化复杂推理与决策任务方面展现出强大能力。在电信领域,它们有望实现网络优化、故障自动排查、客户支持增强及合规性保障。然而,其部署面临领域特有挑战,需专门适配。为克服这些障碍并加速LLMs在电信领域的应用,我们提出MM-Telco——一个针对电信领域的综合性多模态基准与模型套件。该基准包含多种任务(文本与图像驱动),覆盖网络运维、网络管理、文档质量提升、相关文本与图像检索等真实应用场景。此外,我们对多种LLMs与视觉语言模型(VLMs)进行了基线实验。在我们的数据集上微调的模型表现出显著性能提升。实验还揭示了当前先进多模态模型的薄弱环节,为未来研究提供方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they hold the potential to transform network optimization, automate troubleshooting, enhance customer support, and ensure regulatory compliance. However, their deployment in telecom is hindered by domain-specific challenges that demand specialized adaptation. To overcome these challenges and to accelerate the adaptation of LLMs for telecom, we propose MM-Telco, a comprehensive suite of multimodal benchmarks and models tailored for the telecom domain. The benchmark introduces various tasks (both text based and image based) that address various practical real-life use cases such as network operations, network management, improving documentation quality, and retrieval of relevant text and images. Further, we perform baseline experiments with various LLMs and VLMs. The models fine-tuned on our dataset exhibit a significant boost in performance. Our experiments also help analyze the weak areas in the working of current state-of-art multimodal LLMs, thus guiding towards further development and research.

多模态电信AI大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。