arXiv:2604.17659cs.CLcs.AI2026-04

提升每字语义密度,能显著提高大模型回答准确率

Semantic Density Effect (SDE): Maximizing Information Per Token Improves LLM Accuracy

  • 通过去除低信息量词,提升每令牌的语义密度来优化提示
  • 语义密度超0.80的提示平均准确率高出8.4个百分点
  • 无需增加计算量,适合追求精准输出的研究与应用

我们提出语义密度效应(SDE):在所有主流大模型中,单位令牌携带更高语义信息的提示,始终产生更准确、更聚焦且幻觉更少的输出。SDE定义为语义负载词数与总提示词数之比,已调整冗余度和具体性。不同于以往通过添加词(思维链)、重复提示或调整顺序优化的方法,SDE通过移除或替换低信息词,在不损失语义信号的前提下提升表现。在五款前沿模型和七个基准测试上评估,语义密度高于0.80的超密集提示,相比稀疏提示平均提升8.4个百分点,且不增加令牌数或延迟。结合指令位置效应(IPE),最高可提升11.7个百分点。

原文摘要 · Abstract (English)

We introduce the Semantic Density Effect (SDE): the empirical finding that prompts carrying higher semantic information per token consistently produce more accurate, focused, and less hallucinated outputs across all major LLM families. SDE is defined as the ratio of semantically loaded tokens to total prompt tokens, adjusted for redundancy and concreteness. Unlike prior prompt optimization techniques that add tokens (Chain of Thought), duplicate the prompt (Prompt Repetition), or reorder components (Instruction Placement Effect), SDE improves performance by removing or replacing low-information tokens while preserving or sharpening the semantic signal. Evaluated across five frontier models and seven benchmarks, ultra-dense prompts (SDE > 0.80) outperform diluted counterparts by an average of +8.4 percentage points with 0 additional tokens and 0 latency overhead. Combined with Instruction Placement Effect (IPE), the gain reaches +11.7 percentage points

提示优化语义密度大模型准确率提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。