让大模型用不人类可读的紧凑文本交流,效果仍很好。
Large Language Models Do Not Always Need Readable Language

- 用非标准压缩文本(BabelTele)替代自然语言,模型仍能理解。
- 文本压缩到原长27.9%时,语义保留率达99.5%。
- 适合追求低上下文开销的模型间通信场景。
大型语言模型(LLMs)通常使用人类可读的自然语言进行提示和交互,即使目标读者是其他模型。本文探究了是否可用紧凑、非标准的文本形式编码语义信息,牺牲人类可读性但保持模型可恢复。这类面向模型的文本表示称为BabelTele,不是固定协议,而是一种对模型生成与理解能力的实证探测。通过可读性诊断、模型似然度测量、人工问卷和下游任务评估,发现BabelTele可大幅偏离常规自然语言,同时为指令微调的LLM保留核心语义。作为一种任务无关的表征范式,其信息密度高,在文本量缩减至原长27.9%时仍保持99.5%的语义保真度。进一步评估其在跨模型迁移、智能体记忆和多智能体通信中的鲁棒性,结果表明其可降低上下文开销并维持可靠下游性能,但效果依赖于压缩器-阅读器配对与任务设定。研究显示,人类可读性、自然语言典型性与模型侧语义可恢复性可部分解耦,为未来大模型系统中模型原生表征的探索开辟路径。
原文摘要 · Abstract (English)
Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whether semantic information can be encoded in compact, non-standard textual forms that sacrifice human readability while remaining recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMs' capacity to generate and interpret such representations. Through readability diagnostics, model likelihood measures, human questionnaires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserving core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information density, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communication. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typicality, and model-side semantic recoverability can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。