arXiv:2509.21949cs.NIcs.CL2025-09中稿 · the IEEE GLOBECOM …被引 1

评测两款开源大模型在通信技术问答中的表现,发现性能有差异且需领域适配。

Evaluating Open-Source Large Language Models for Technical Telecom Question Answering

  • 构建105个通信领域问答对,用语义和判别模型评估效果。
  • Gemma在语义准确性和正确率上更优,DeepSeek在词汇一致性略胜。
  • 揭示模型幻觉与判断不一致问题,强调需专用模型支持工程应用。

大型语言模型(LLMs)在多个领域展现出强大能力,但在电信等技术领域的表现仍待深入探索。本文评估了两款开源LLM——Gemma 3 27B和DeepSeek R1 32B——在基于高级无线通信材料的事实性与推理类问题上的表现。我们构建了一个包含105个问答对的基准测试集,并采用词汇匹配、语义相似度及LLM作为裁判者评分等多种方式评估性能。同时通过源归属分析与得分方差,考察模型的一致性、判断可靠性及幻觉现象。结果显示,Gemma在语义保真度和裁判模型评定的正确率方面表现更佳,而DeepSeek则略有更高的词汇一致性。进一步发现当前模型在电信应用中仍存局限,亟需领域适配以支撑可信赖的AI工程助手。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities across various fields. However, their performance in technical domains such as telecommunications remains underexplored. This paper evaluates two open-source LLMs, Gemma 3 27B and DeepSeek R1 32B, on factual and reasoning-based questions derived from advanced wireless communications material. We construct a benchmark of 105 question-answer pairs and assess performance using lexical metrics, semantic similarity, and LLM-as-a-judge scoring. We also analyze consistency, judgment reliability, and hallucination through source attribution and score variance. Results show that Gemma excels in semantic fidelity and LLM-rated correctness, while DeepSeek demonstrates slightly higher lexical consistency. Additional findings highlight current limitations in telecom applications and the need for domain-adapted models to support trustworthy Artificial Intelligence (AI) assistants in engineering.

大模型评测通信技术开源模型问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。