arXiv:2506.11095cs.CLcs.AI2025-06ACL被引 3

用拓扑分析文本结构,预测读者好奇心。

Persistent Homology of Topic Networks for the Prediction of Reader Curiosity

  • 通过主题建模与持久同调分析文本语义网络的动态拓扑变化
  • 在49人实验中,拓扑特征解释了73%的读者好奇度方差
  • 适合研究文本互动性、阅读体验或情感计算的研究者

读者好奇心——寻求信息的动力——对文本参与度至关重要,但在自然语言处理中仍相对未被充分探索。基于洛温斯坦的信息缺口理论,我们提出一个框架,通过量化文本语义结构中的语义信息缺口来建模读者好奇心。该方法结合受BERTopic启发的主题建模与持久同调技术,分析由文本片段构建的动态语义网络的拓扑特性(连通分量、环、空洞),并将这些特征作为信息缺口的代理指标。为实证评估该流程,我们收集了49名参与者在阅读苏·柯林斯《饥饿游戏》小说时的好奇度评分。随后,利用该流程提取的拓扑特征作为自变量,预测这些评分,实验表明其显著优于基线模型(解释偏差分别为73%与30%),验证了该方法的有效性。该流程为分析文本结构及其与读者参与度的关系提供了新的计算方法。

原文摘要 · Abstract (English)

Reader curiosity, the drive to seek information, is crucial for textual engagement, yet remains relatively underexplored in NLP. Building on Loewenstein's Information Gap Theory, we introduce a framework that models reader curiosity by quantifying semantic information gaps within a text's semantic structure. Our approach leverages BERTopic-inspired topic modeling and persistent homology to analyze the evolving topology (connected components, cycles, voids) of a dynamic semantic network derived from text segments, treating these features as proxies for information gaps. To empirically evaluate this pipeline, we collect reader curiosity ratings from participants (n = 49) as they read S. Collins's ''The Hunger Games'' novel. We then use the topological features from our pipeline as independent variables to predict these ratings, and experimentally show that they significantly improve curiosity prediction compared to a baseline model (73% vs. 30% explained deviance), validating our approach. This pipeline offers a new computational method for analyzing text structure and its relation to reader engagement.

文本分析读者行为拓扑学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。