FLOWER自动构建实体关系模型,提升数据理解与洞察效率。
FLOWER: Flow-Oriented Entity-Relationship Tool
- 基于动态采样与数据分析,自动识别显式隐式依赖
- 比抽样方法快2.15倍,约束学习效率提升2.6倍
- 支持23种语言,适配多场景数据库分析
跨数据源的关系探索对实体识别至关重要。由于数据库存储大量合成与真实数据,正确处理所有对象是关键任务。然而,实体关系模型的构建常依赖人工判断。本文提出FLOWER:首个端到端解决方案,可即时消除处理、创建和可视化显式/隐式依赖的繁琐与资源消耗问题,支持主流SQL方言。启动后,FLOWER自动检测内置约束,并通过动态采样与鲁棒数据分析技术构建准确且必要的模型。该方法能提升实体关系建模与数据叙事能力,帮助用户更深入理解数据基础并发现新洞见。在STATS基准测试中,其分布表示性能优于分层抽样2.4倍,约束学习提升2.6倍,加速比达2.15倍;数据叙事方面,准确率提高1.19倍,上下文减少1.86倍,优于LLM。工具支持23种语言,兼容CPU与GPU,表明其在真实数据场景下具备更高质量、可扩展性与适用性。
原文摘要 · Abstract (English)
Exploring relationships across data sources is a crucial optimization for entities recognition. Since databases can store big amount of information with synthetic and organic data, serving all quantity of objects correctly is an important task to deal with. However, the decision of how to construct entity relationship model is associated with human factor. In this paper, we present flow-oriented entity-relationship tool. This is first and unique end-to-end solution that eliminates routine and resource-intensive problems of processing, creating and visualizing both of explicit and implicit dependencies for prominent SQL dialects on-the-fly. Once launched, FLOWER automatically detects built-in constraints and starting to create own correct and necessary one using dynamic sampling and robust data analysis techniques. This approach applies to improve entity-relationship model and data storytelling to better understand the foundation of data and get unseen insights from DB sources using SQL or natural language. Evaluated on state-of-the-art STATS benchmark, experiments show that FLOWER is superior to reservoir sampling by 2.4x for distribution representation and 2.6x for constraint learning with 2.15x acceleration. For data storytelling, our tool archives 1.19x for accuracy enhance with 1.86x context decrease compare to LLM. Presented tool is also support 23 languages and compatible with both of CPU and GPU. Those results show that FLOWER can manage with real-world data a way better to ensure with quality, scalability and applicability for different use-cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。