通过保留表格结构提升科学文献问答准确率
Extracting Information from Scientific Literature via Visual Table Question Answering Models
- 保留表格结构的识别方法优于传统OCR和通用视觉问答模型
- 在7个预设问题上,结构保留方法准确率显著更高
- 适合需要精准提取论文数据的研究者与系统开发者
本研究探索三种处理科学论文中表格数据的方法,以提升抽取式问答性能并开发系统性文献综述工具。评估方法包括:(1) 光学字符识别(OCR)提取文档信息,(2) 预训练文档视觉问答模型,(3) 表格检测与结构识别,将表格内容与文本合并以回答抽取式问题。在初步实验中,对10份含表格的电磁场(RF-EMF)相关论文样本,添加7组预定义的抽取式问答对进行测试。结果表明,保持表格结构的方法表现更优,尤其在内容呈现与组织方面。准确识别文档中的特定符号与标记成为提升性能的关键因素。研究结论指出,保持表格结构完整性对于提升科学文献中抽取式问答的准确性与可靠性至关重要。
原文摘要 · Abstract (English)
This study explores three approaches to processing table data in scientific papers to enhance extractive question answering and develop a software tool for the systematic review process. The methods evaluated include: (1) Optical Character Recognition (OCR) for extracting information from documents, (2) Pre-trained models for document visual question answering, and (3) Table detection and structure recognition to extract and merge key information from tables with textual content to answer extractive questions. In exploratory experiments, we augmented ten sample test documents containing tables and relevant content against RF- EMF-related scientific papers with seven predefined extractive question-answer pairs. The results indicate that approaches preserving table structure outperform the others, particularly in representing and organizing table content. Accurately recognizing specific notations and symbols within the documents emerged as a critical factor for improved results. Our study concludes that preserving the structural integrity of tables is essential for enhancing the accuracy and reliability of extractive question answering in scientific documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。