arXiv:2511.15210cs.CLcs.AI2025-11Conference of the …被引 1

揭示文本内在维度:科学写作比小说更简单

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

  • 通过语言特征与稀疏自编码器分析文本内在维度
  • 科学文本内在维度约8,小说达10.5,差异显著
  • 正式语气降维,情感叙事升维,影响模型表征复杂度

内在维度(ID)是现代大模型分析的重要工具,用于研究训练动态、缩放行为和数据集结构,但其文本决定因素仍不明确。本文首次通过交叉编码器分析、语言特征与稀疏自编码器(SAEs),系统探究了文本属性对ID的影响。发现:第一,ID与熵类指标互补,在控制长度后无相关性,ID捕捉的是几何复杂度而非预测质量;第二,文本类型呈现稳定分层:科学论述的ID约为8,百科内容为9,创意/观点类写作高达10.5,表明当前大模型认为科学文本“表征简单”而小说需更多自由度;第三,利用SAEs识别出因果特征:正式语气、报告模板、统计数据降低ID,个性化、情绪与叙事提升ID,操控实验验证其因果性。研究为正确使用和解读内在维度提供了实证依据。

原文摘要 · Abstract (English)

Intrinsic dimension (ID) is an important tool in modern LLM analysis, informing studies of training dynamics, scaling behavior, and dataset structure, yet its textual determinants remain underexplored. We provide the first comprehensive study grounding ID in interpretable text properties through cross-encoder analysis, linguistic features, and sparse autoencoders (SAEs). In this work, we establish three key findings. First, ID is complementary to entropy-based metrics: after controlling for length, the two are uncorrelated, with ID capturing geometric complexity orthogonal to prediction quality. Second, ID exhibits robust genre stratification: scientific prose shows low ID (~8), encyclopedic content medium ID (~9), and creative/opinion writing high ID (~10.5) across all models tested. This reveals that contemporary LLMs find scientific text "representationally simple" while fiction requires additional degrees of freedom. Third, using SAEs, we identify causal features: scientific signals (formal tone, report templates, statistics) reduce ID; humanized signals (personalization, emotion, narrative) increase it. Steering experiments confirm these effects are causal. Thus, for contemporary models, scientific writing appears comparatively "easy", whereas fiction, opinion, and affect add representational degrees of freedom. Our multi-faceted analysis provides practical guidance for the proper use of ID and the sound interpretation of ID-based results.

大模型分析内在维度文本复杂度表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。