用点过程建模文本生成中的不确定性传播,提升长序列输出一致性。
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation

- 用多变量霍克斯过程建模生成节点的激活时序与影响关系。
- 在GDELT数据集上,压缩提示记忆下晚期语义对齐提升23%。
- 适合需要长序列推理与稳定输出的智能体系统开发者。
智能体式文本生成系统按顺序写作,每一步输出可能成为后续步骤的上下文,导致不确定性具有路径依赖性:早期模糊会持续影响后期结果。本文提出HawkesLLM框架,将时间影响建模与文本生成分离。将生成过程视为由文本生成智能体组成的网络,用多变量霍克斯过程建模节点激活时序及哪些先前输出应影响当前提示。语言模型基于该时序模型选出的紧凑记忆生成新内容。在预留的全球事件、语言与语气数据库(GDELT)新闻级联案例中进行评估。诊断方法追踪与局部保留参考的语义对齐度,并区分局部漂移与全局漂移。在紧凑提示记忆预算下,HawkesLLM显著提升了后期语义对齐效果。
原文摘要 · Abstract (English)
Agentic text-simulation systems write in sequence, with each item becoming possible context for later steps. That makes uncertainty path-dependent: an early ambiguity can affect later outputs. This paper studies this problem with HawkesLLM, a framework that separates temporal influence modeling from text generation. We represent the cascade as a network whose nodes are text-generating agents. A multivariate Hawkes process models how these nodes activate over time and which earlier node outputs should influence later prompts. A language model then writes each new event from the compact memory selected by this temporal model. We evaluate the framework on a held-out Global Database of Events, Language, and Tone (GDELT) news-cascade case study. The diagnostics track semantic alignment with local held-out references and separate local drift from global drift. In this setting, HawkesLLM improves late-stage semantic alignment under a compact prompt-memory budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。