从文本中高效提取数量事实,通过解析描述+弱监督提升准确率。
Towards Efficient Quantity Retrieval from Text:An Approach via Description Parsing and Weak Supervision
- 将文本描述转为结构化(描述, 数量)对,便于检索
- 在财报数据上将准确率从30.98%提升至64.66%
- 适合需要从文档中挖数量信息的金融/政务场景
定量事实由企业和政府持续生成,支持数据驱动决策。尽管常见事实已结构化,许多长尾定量事实仍埋藏于非结构化文档中,难以获取。我们提出数量检索任务:给定定量事实的描述,系统返回相关数值及支持证据。理解上下文中的数量语义是该任务的关键。我们提出基于描述解析的框架,将文本转化为结构化的(描述,数量)对以实现高效检索。为提升学习效果,我们利用数量共现关系构建大规模伪标签数据集,采用弱监督方法。我们在大型财务年报语料库和新标注的数量描述数据集上评估了该方法,显著提升顶1检索准确率,从30.98%提高到64.66%。
原文摘要 · Abstract (English)
Quantitative facts are continually generated by companies and governments, supporting data-driven decision-making. While common facts are structured, many long-tail quantitative facts remain buried in unstructured documents, making them difficult to access. We propose the task of Quantity Retrieval: given a description of a quantitative fact, the system returns the relevant value and supporting evidence. Understanding quantity semantics in context is essential for this task. We introduce a framework based on description parsing that converts text into structured (description, quantity) pairs for effective retrieval. To improve learning, we construct a large paraphrase dataset using weak supervision based on quantity co-occurrence. We evaluate our approach on a large corpus of financial annual reports and a newly annotated quantity description dataset. Our method significantly improves top-1 retrieval accuracy from 30.98 percent to 64.66 percent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。