arXiv:2606.11222cs.CLcs.IT2026-06

用几何方法量化文本语义信息,揭示了三种核心维度的权衡关系。

A Geometric Profile of Semantic Information in Text: Frame-Conditional Uniqueness and a Trade-Off Triangle for Scalar Summaries

  • 基于句向量结构构建语义轮廓,定义新颖性、广度和整合度三维度
  • 证明单一数值摘要无法同时满足稳定性、尺度鲁棒性和跨模型可比性
  • 提出两个实用指标,均在23类合成数据和5部经典小说上表现优异

文本承载多少语义信息?香农理论关注符号不确定性,忽略意义;而类似BERTScore的成对度量仅比较两文本。本文构建几何框架,从句子嵌入结构衡量语义内容。首先,在固定嵌入与基线前提下,六个自然公理唯一确定一个标量度量(至比例),即帧条件唯一性定理;但该标量过于粗糙,需更丰富表示。其次,提出三坐标语义轮廓:新颖性(偏离通用话语的位移)、广度(不同思想多样性)和整合度(思想间连通性),并定义离散最小单元(语义量子),其分辨率由聚类阈值τ决定。第三,证明一个不可能定理:不存在任何标量能同时满足分析稳定性(改写/拼接不变)、序数鲁棒性(跨文本尺度一致)和跨表示可比性。展示两个实用标量$S_{\mathrm{minmax}}$和$S_{\mathrm{rank}}$,分别占据该权衡三角的两个顶点。在23类合成数据、5部Project Gutenberg小说及3种嵌入模型上的验证确认该权衡。推荐的秩归一化配置在28项序数检验中通过25项(21项经Benjamini-Hochberg校正后仍通过),优于七种基线,包括词频熵与基于BERTScore的新颖性信号。另通过变分结果将广度坐标关联到确定点过程的对数行列式(斯皮尔曼ρ=0.985,覆盖507章),为广度提供优化理论基础。

原文摘要 · Abstract (English)

How much meaning does a text carry? Shannon's theory measures uncertainty over symbols and is intentionally indifferent to meaning, while pairwise metrics such as BERTScore compare two texts rather than characterizing one. We develop a geometric framework that measures semantic content from the structure of a text's sentence embeddings. The framework has three parts. First, within a fixed embedding and baseline, six natural axioms uniquely determine a scalar measure up to scale, a frame-conditional uniqueness theorem. The resulting scalar is empirically too coarse, motivating a richer representation. Second, we propose a three-coordinate semantic profile capturing novelty (displacement from generic discourse), breadth (diversity of distinct ideas), and integration (connectedness among them), together with a discrete minimal unit (the semantic quantum) whose resolution is fixed by a clustering threshold $τ$. Third, we prove a no-go theorem: no scalar summary of the profile can simultaneously satisfy analytic stability under paraphrase and concatenation, ordinal robustness across text scales, and cross-representation comparability. We exhibit two practical scalars, $S_{\mathrm{minmax}}$ and $S_{\mathrm{rank}}$, each occupying a distinct corner of this trade-off triangle. Validation across 23 synthetic categories, 5 Project Gutenberg novels, and 3 embedding models confirms the trade-off. The recommended rank-normalized configuration passes 25 of 28 ordinal checks as point estimates (21 of 28 after Benjamini-Hochberg correction), outperforming seven baselines including unigram entropy and a BERTScore-based novelty signal. A separate variational result connects the breadth coordinate to the log-determinant of a determinantal point process (Spearman $ρ= 0.985$ over 507 Gutenberg chapters), giving an optimization-theoretic foundation for breadth.

语义量化文本表征几何建模评价指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。