arXiv:2603.04413cs.CLcs.AI2026-03被引 1

用符号学与解释学构建新指标,评估大模型摘要的语义准确性。

Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries

  • 融合符号学与解释学,提出基于内容分析的定性评估方法。
  • 在5个数据集上发现大模型摘要语义准确率低于人类,尤其依赖上下文时。
  • 适合关注生成文本深层意义、追求可解释评估的研究者。

人类语言的意义具有关系性、情境依赖性和涌现性,源于动态符号系统而非固定词义映射。现有计算方法难以捕捉这种复杂性。本文提出跨学科框架,结合符号学、解释学与质性研究方法,分析大模型中语言符号如何转化为向量表示,并揭示统计近似与人类理解之间的差距。为此,我们引入归纳概念评分(ICR)指标,基于归纳内容分析与反思主题分析,评估大模型输出的语义准确性和意义对齐度,超越传统词汇相似性度量。在五个数据集(样本量N = 50至800)上,对比大模型与人类生成的主题摘要,结果显示大模型虽具高语言相似性,但在语义准确性上表现不佳,尤其在捕捉上下文相关意义方面;性能随数据量增大而提升,但模型间差异显著,可能反映概念重复频率与连贯性的不同。结论强调,在评估大模型输出意义时,应采用系统化的质性解读范式。

原文摘要 · Abstract (English)

Meaning in human language is relational, context dependent, and emergent, arising from dynamic systems of signs rather than fixed word-concept mappings. In computational settings, this semiotic and interpretive complexity complicates the generation and evaluation of meaning. This article proposes an interdisciplinary framework for studying meaning in large language model (LLM) generated language by integrating semiotics and hermeneutics with qualitative research methods. We review prior scholarship on meaning and machines, examining how linguistic signs are transformed into vectorized representations in static and contextualized embedding models, and identify gaps between statistical approximation and human interpretive meaning. We then introduce the Inductive Conceptual Rating (ICR) metric, a qualitative evaluation approach grounded in inductive content analysis and reflexive thematic analysis, designed to assess semantic accuracy and meaning alignment in LLM-outputs beyond lexical similarity metrics. We apply ICR in an empirical comparison of LLM generated and human generated thematic summaries across five datasets (N = 50 to 800). While LLMs achieve high linguistic similarity, they underperform on semantic accuracy, particularly in capturing contextually grounded meanings. Performance improves with larger datasets but remains variable across models, potentially reflecting differences in the frequency and coherence of recurring concepts and meanings. We conclude by arguing for evaluation frameworks that leverage systematic qualitative interpretation practices when assessing meaning in LLM-generated outputs from reference texts.

语义评估大模型评价符号学质性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。