arXiv:2601.05411cs.CL2026-01

用语言模型估算公文信息熵,可视化提升可读性

Glitter: Visualizing Lexical Surprisal for Readability in Administrative Texts

  • 结合多模型估计文本信息熵,实现可读性量化
  • 通过可视化发现公文中的高困惑度段落
  • 适合政策起草者和文书优化者使用

本研究探讨如何利用文本信息熵衡量可读性。提出一种可视化框架,通过多个语言模型近似计算文本的信息熵,并以图形方式呈现结果。目标是评估并改进行政或官僚文本的可读性和清晰度。该工具已作为开源软件发布于 https://github.com/ufal/Glitter。

原文摘要 · Abstract (English)

This work investigates how measuring information entropy of text can be used to estimate its readability. We propose a visualization framework that can be used to approximate information entropy of text using multiple language models and visualize the result. The end goal is to use this method to estimate and improve readability and clarity of administrative or bureaucratic texts. Our toolset is available as a libre software on https://github.com/ufal/Glitter.

可读性信息熵公文优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。