arXiv:2602.14274cs.LGcs.AI2026-02

用文本数据做因果推断,让无结构信息也能指导商业决策。

Integrating Unstructured Text into Causal Inference: Empirical Evidence from Real Data

  • 用Transformer模型从文本中提取因果信息,替代传统结构化数据。
  • 在群体、组别和个体层面,文本推断结果与结构化数据一致。
  • 适合缺乏表格数据但有大量文本的场景,如舆情分析、用户反馈研究。

因果推断是支持商业决策的关键工具,传统上依赖结构化数据。但在许多真实场景中,此类数据可能不完整或不可用。本文提出一个框架,利用基于Transformer的语言模型,从非结构化文本中进行因果推断。我们在人群、群体和个体三个层级上,将基于文本的因果估计结果与结构化数据的结果进行了对比。研究发现两种方法得出的结果具有一致性,验证了非结构化文本在因果推断任务中的潜力。该方法拓展了因果推断在仅有文本数据可用场景下的应用范围,使在结构化表格数据稀缺时仍能实现数据驱动的商业决策。

原文摘要 · Abstract (English)

Causal inference, a critical tool for informing business decisions, traditionally relies heavily on structured data. However, in many real-world scenarios, such data can be incomplete or unavailable. This paper presents a framework that leverages transformer-based language models to perform causal inference using unstructured text. We demonstrate the effectiveness of our framework by comparing causal estimates derived from unstructured text against those obtained from structured data across population, group, and individual levels. Our findings show consistent results between the two approaches, validating the potential of unstructured text in causal inference tasks. Our approach extends the applicability of causal inference methods to scenarios where only textual data is available, enabling data-driven business decision-making when structured tabular data is scarce.

因果推断自然语言文本分析商业决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。