统一评估标准,解决跨模态语义通信的度量难题
Rethinking Communication Metrics: How Should We Measure Meaning?

- 从评估视角重构语义通信指标体系
- 梳理六类指标并对比其适用场景与局限
- 适合关注语义通信质量评估的研究者
语义通信将通信目标从符号准确重建转向意义保留、任务完成和高效信息交换。然而,其评估仍分散于通信、自然语言处理、计算机视觉和机器学习领域,缺乏跨模态、任务和信道条件的统一评价指标。本文从评估中心视角,系统梳理文本与图像语义通信系统的各项关键性能指标(KPI),按通信目标、源模态、接收端输出、参考依赖性、评估层级及信道/资源约束进行分类。综述了基于重建、任务导向、无参考、表示级、感知和信道感知等六类指标,比较其在跨模态中的角色、优势与不足。分析了未解的语义-指标挑战对监控、质量保障、资源优化、故障诊断与标准化的影响。核心开放问题包括:缺乏通用语义成功标准与标准化语义真值、语义漂移、无参考评估受限、机器学习指标与通信约束融合弱、关系级与多模态指标缺失。最后提出未来研究方向:构建标准化、可解释、自适应、任务感知、通信感知的评估框架。
原文摘要 · Abstract (English)
Semantic communication shifts the objective of communication systems from accurate symbol reconstruction toward meaning preservation, task accomplishment, and efficient information exchange. However, its evaluation remains fragmented across telecommunications, natural language processing, computer vision, and machine learning, and no single metric can characterize semantic quality across modalities, tasks, and channel conditions. This article surveys key performance indicators (KPIs) for text- and image-based semantic communication systems from a unified, evaluation-centered perspective. Unlike prior surveys primarily organized around architectures, applications, or transmission strategies, this work focuses on how semantic success should be defined and measured. Existing KPIs are classified according to communication goal, source modality, receiver output, reference availability, evaluation level, and channel or resource constraints. The survey reviews reconstruction-based, task-oriented, reference-free, representation-level, perceptual, and channel-aware metrics, and presents a cross-modality comparison of their roles, strengths, and limitations. It further analyzes how unresolved semantic-KPI challenges affect monitoring, quality assurance, resource optimization, fault diagnosis, and standardization. Key open problems include the absence of universal semantic success criteria and standardized semantic ground truth, semantic drift, limited reference-free evaluation, weak integration of machine-learning metrics with communication constraints, and the lack of relation-level and multimodal KPIs. Finally, future research directions are outlined toward standardized, interpretable, adaptive, task-aware, and communication-aware evaluation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。