构建6G时代语义通信评估基准,测试AI模型的决策推理能力
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
- 基于30类6G标准任务设计多轮推理题,覆盖复杂网络场景
- 筛选出3722道高置信度难题,涵盖不确定性和最坏情况优化
- 适配大模型评测,助力开放科研与可复现性研究
本文提出6G-Bench,一个面向人工智能原生6G网络中语义通信与网络级推理的开源评估基准。该基准从3GPP、IETF、ETSI、ITU-T及O-RAN联盟的6G与AI代理标准化活动中提取30项决策任务(T1–T30),并按标准化方向划分为五类能力范畴。从113,475个场景生成10,000道高难度多选题,采用任务条件提示确保在不确定性下进行多步定量推理及多轮次最坏后悔最小化。经自动过滤与专家人工验证,保留3,722道高置信度题目作为评测集,完整题库公开以支持6G专用模型训练与微调。使用6G-Bench评估了22种基础模型,涵盖密集型与专家混合架构、短/长上下文设计(最长达1M token)以及开源自研与专有系统。各模型确定性单次准确率(pass@1)范围为0.22至0.82,凸显语义推理能力差异显著;领先模型在意图与策略推理上准确率达0.87–0.89,而在推理密集任务上选择性鲁棒性分析显示pass@5值介于0.20至0.91之间。为支持开放科学与可复现性,6G-Bench数据集已发布于GitHub:https://github.com/maferrag/6G-Bench
原文摘要 · Abstract (English)
This paper introduces 6G-Bench, an open benchmark for evaluating semantic communication and network-level reasoning in AI-native 6G networks. 6G-Bench defines a taxonomy of 30 decision-making tasks (T1--T30) extracted from ongoing 6G and AI-agent standardization activities in 3GPP, IETF, ETSI, ITU-T, and the O-RAN Alliance, and organizes them into five standardization-aligned capability categories. Starting from 113,475 scenarios, we generate a balanced pool of 10,000 very-hard multiple-choice questions using task-conditioned prompts that enforce multi-step quantitative reasoning under uncertainty and worst-case regret minimization over multi-turn horizons. After automated filtering and expert human validation, 3,722 questions are retained as a high-confidence evaluation set, while the full pool is released to support training and fine-tuning of 6G-specialized models. Using 6G-Bench, we evaluate 22 foundation models spanning dense and mixture-of-experts architectures, short- and long-context designs (up to 1M tokens), and both open-weight and proprietary systems. Across models, deterministic single-shot accuracy (pass@1) spans a wide range from 0.22 to 0.82, highlighting substantial variation in semantic reasoning capability. Leading models achieve intent and policy reasoning accuracy in the range 0.87--0.89, while selective robustness analysis on reasoning-intensive tasks shows pass@5 values ranging from 0.20 to 0.91. To support open science and reproducibility, we release the 6G-Bench dataset on GitHub: https://github.com/maferrag/6G-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。