用时空坐标统一结构化非结构化数据,提升知识提取准确率。
STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System
- 基于时空维度构建多维数据仓库,结合大模型进行上下文感知抽取。
- 在公共卫生数据上,实体提取F1提升4.37%(GPT-4o-mini)。
- 支持交互式分析,适合需要时空推理的跨领域应用。
从非结构化数据中提取结构化知识仍面临实际挑战:实体与事件抽取流程脆弱,知识图谱构建需高成本本体工程,跨域泛化能力普遍不达生产标准。相比之下,空间与时间提供普适的上下文锚点,能自然对齐异构信息,并促进检索与推理等下游任务。我们提出 extbf{STIndex},一个端到端系统,将非结构化内容转化为多维时空数据仓库。用户可自定义领域分析维度及其层级结构,由大语言模型完成上下文感知的抽取与定位。 extbf{STIndex}整合文档级记忆、地理编码校正与质量验证机制,提供交互式可视化仪表板,支持聚类、突现检测与实体网络分析。在公开公共卫生基准测试中, extbf{STIndex}相较基线模型,时空实体提取F1分别提升4.37%(GPT-4o-mini)和3.60%(Qwen3-8B)。实时演示与开源代码见https://stindex.ai4wa.com/dashboard。
原文摘要 · Abstract (English)
Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowledge graph construction requires costly ontology engineering, and cross-domain generalization is rarely production-ready. In contrast, space and time provide universal contextual anchors that naturally align heterogeneous information and benefit downstream tasks such as retrieval and reasoning. We introduce \textbf{STIndex}, an end-to-end system that structures unstructured content into a multidimensional spatiotemporal data warehouse. Users define domain-specific analysis dimensions with configurable hierarchies, while large language models perform context-aware extraction and grounding. \textbf{STIndex} integrates document-level memory, geocoding correction, and quality validation, and offers an interactive analytics dashboard for visualization, clustering, burst detection, and entity network analysis. In evaluation on a public health benchmark, \textbf{STIndex} improves spatiotemporal entity extraction F1 by 4.37\% (GPT-4o-mini) and 3.60\% (Qwen3-8B). A live demonstration and open-source code are available at https://stindex.ai4wa.com/dashboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。