arXiv:2604.09237cs.CL2026-04ACL

让AI自动从文档中提取结构化数据,回答复杂研究问题。

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

论文配图:ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
图 1 · 摘自论文原文
  • 用大模型交互式发现数据模式,自动生成可查询的数据库
  • 在法律与生物信息学领域验证,支持真实世界分析任务
  • 开源工具+网页界面,专家可快速上手自己的数据

许多学科面对大型文献集合中的自然语言研究问题,其答案通常需要结构化证据,传统方法依赖人工设计标注模式并逐条标注语料库,过程缓慢且易出错。我们提出ScheMatiQ,通过调用基础大模型,将问题与文档集合输入,自动生成数据模式和可支撑分析的结构化数据库,并提供网页界面供用户交互调整与修正。与领域专家协作验证表明,ScheMatiQ生成的结果可有效支持法律与计算生物学中的实际分析任务。我们已将ScheMatiQ开源,包含公开网页界面、源代码及演示视频,网址为www.ScheMatiQ-ai.com,欢迎各领域专家使用其处理自身数据。

原文摘要 · Abstract (English)

Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the corpus, a slow and error-prone process. We introduce ScheMatiQ, which leverages calls to a backbone LLM to take a question and a corpus to produce a schema and a grounded database, with a web interface that lets steer and revise the extraction. In collaboration with domain experts, we show that ScheMatiQ yields outputs that support real-world analysis in law and computational biology. We release ScheMatiQ as open source with a public web interface, and invite experts across disciplines to use it with their own data. All resources, including the website, source code, and demonstration video, are available at: www.ScheMatiQ-ai.com

知识提取大模型应用结构化数据交互式工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。