arXiv:2605.13110cs.MAcs.AI2026-05

用多智能体自动完成风投尽调,精准提取官方财报数据。

A Multi-Agent Orchestration Framework for Venture Capital Due Diligence

  • 通过程序化管道逆向解析希腊工商注册系统接口,抓取动态财务文件。
  • 采用布局感知OCR解析财报,避免生成虚假财务数据。
  • 全流程公开可复现,适合风投机构与量化研究者使用。

我们提出一个完全自动化的多智能体框架,用于风险投资中的企业尽职调查与市场分析。系统基于事件驱动的编排架构,结合大语言模型(LLMs)与实时网页检索,将非结构化数据转化为结构化投资情报。核心技术贡献在于一个程序化提取管道,逆向工程希腊工商注册系统(Γ.Ε.ΜΗ.)的前后端通信,查询动态接口获取官方财务文件,并通过布局感知OCR提取器进行解析。系统设有结构化回退机制,明确标记数据缺失,而非生成未经验证的数值,直接应对金融场景中的幻觉问题。所有工作流成果均公开,支持复现。

原文摘要 · Abstract (English)

We present a fully automated multi-agent framework for corporate due diligence and market analysis in venture capital. The system runs on an event-driven orchestration architecture, combining Large Language Models (LLMs) with real-time web retrieval to synthesize unstructured data into structured investment intelligence. A central technical contribution is a programmatic extraction pipeline that reverse-engineers the frontend-to-backend communication of the Greek Business Registry ($Γ$.E.MH.), querying dynamic endpoints to retrieve official financial filings that are then parsed using a layout-aware OCR extractor. A structural fallback mechanism explicitly flags data absence rather than generating unverified figures, directly targeting hallucination in financial contexts. All workflow artifacts are publicly available to support replication.

风投多智能体信息提取自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。