arXiv:2504.07459cs.CL2025-04被引 3

用语言学特征提升叙事文本因果图生成质量

Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts

  • 结合语言学特征与RoBERTa模型识别因果关系
  • 在100篇故事中优于GPT-4o和Claude 3.5
  • 适合需要可解释因果链的文本分析场景

我们提出一种从叙事文本生成因果图的新框架,连接高层级因果与具体事件间的关系。方法首先利用大语言模型(LLM)摘要提取以主体为中心的节点;引入包含七个语言学特征的“专家指数”,融入情境-任务-行为-后果(STAC)分类模型;该混合系统结合RoBERTa嵌入与专家指数,在因果链接识别上表现优于纯LLM方法。最后通过五轮结构化提示过程精炼并构建连通因果图。在100篇叙事章节和短篇故事上的实验表明,该方法在因果图质量上持续优于GPT-4o和Claude 3.5,同时保持可读性。开源工具提供了可解释、高效的叙事因果链捕捉方案。

原文摘要 · Abstract (English)

We propose a novel framework for generating causal graphs from narrative texts, bridging high-level causality and detailed event-specific relationships. Our method first extracts concise, agent-centered vertices using large language model (LLM)-based summarization. We introduce an "Expert Index," comprising seven linguistically informed features, integrated into a Situation-Task-Action-Consequence (STAC) classification model. This hybrid system, combining RoBERTa embeddings with the Expert Index, achieves superior precision in causal link identification compared to pure LLM-based approaches. Finally, a structured five-iteration prompting process refines and constructs connected causal graphs. Experiments on 100 narrative chapters and short stories demonstrate that our approach consistently outperforms GPT-4o and Claude 3.5 in causal graph quality, while maintaining readability. The open-source tool provides an interpretable, efficient solution for capturing nuanced causal chains in narratives.

因果图语言学特征叙事理解可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。