测试了报告转AI可用格式时的信息损失,发现图表信息严重丢失。
Meet Your New Client: Writing Reports for AI -- Benchmarking Information Loss in Market Research Deliverables
- 将报告转为Markdown后让大模型问答,评估信息保留度
- 文本提取可靠,但图表等复杂内容丢失严重
- 适合做市场研究的AI兼容性优化,或知识管理团队
随着组织采用检索增强生成(RAG)系统管理知识,传统市场研究报告面临新挑战。这些报告不仅面向人类阅读,还需被AI系统解析以回答问题。本研究评估了报告在输入RAG系统过程中出现的信息损失。通过端到端基准测试,比较了将PDF和PPTX文档转换为Markdown后,大语言模型(LLM)回答事实性问题的表现。结果表明,虽然文本内容可被准确提取,但图表、图形等复杂对象中的信息存在显著丢失。这提示需要开发专为AI设计的新型报告格式,以确保研究洞察不因格式转换而流失。
原文摘要 · Abstract (English)
As organizations adopt retrieval-augmented generation (RAG) for their knowledge management systems (KMS), traditional market research deliverables face new functional demands. While PDF reports and slides have long served human readers, they are now also "read" by AI systems to answer user questions. To future-proof reports being delivered today, this study evaluates information loss during their ingestion into RAG systems. It compares how well PDF and PowerPoint (PPTX) documents converted to Markdown can be used by an LLM to answer factual questions in an end-to-end benchmark. Findings show that while text is reliably extracted, significant information is lost from complex objects like charts and diagrams. This suggests a need for specialized, AI-native deliverables to ensure research insights are not lost in translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。