解决企业数据资产检索错漏问题,提升精准度与使用指导能力。
A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery

- 构建双层知识库与图引导检索,结合场景标注提升召回精度。
- 检索准确率从19.1%升至96.6%,知识覆盖率达77%,延迟低于5.4秒。
- 适合需要高精度数据发现与使用指导的企业级数据分析团队。
企业数据分析代理面临两大结构性缺陷:通用RAG检索错误率高(Hit@10=19.1%),且无法提供使用知识以避免指标误读——根源来自语义鸿沟、实体歧义、模式漂移和资产使用信息缺失等四个因素。我们提出两层解决方案,部署于小红书商业广告数据仓库(5,300+ Hive 表,14个领域)。采用三层双用途知识库(179篇文档,八部分注释模板),支持检索与生成,配合闭环更新流水线实现日级新鲜度(一次确认,30秒热重载)。图引导检索器(GGR)利用2,859节点知识图谱作为候选门控,结合意图路由,实现71.6倍的令牌压缩。场景感知排序器(SAR)引入19类实体识别与显式场景标注;仅负向知识即带来25个百分点的Hit@10提升。在两个100题基准测试中,Hit@10从19.1%提升至96.6%(+77.5个百分点),知识覆盖率从56%增至77%,端到端延迟为4.84–5.33秒。
原文摘要 · Abstract (English)
Enterprise data analytics agents face two structural failures: generic RAG retrieves the wrong asset (Hit@10=19.1%) and delivers no usage knowledge to prevent metric misinterpretation---stemming from four root causes (C1--C4) ranging from semantic gap and entity ambiguity to schema drift and asset-usage gap. We present a two-layer solution deployed in the commercial advertising data warehouse at Xiaohongshu (5,300+ Hive tables, 14 domains). A three-tier dual-purpose knowledge base (179 documents, eight-section annotation template) serves both retrieval and generation, with a closed-loop refresh pipeline maintaining day-level freshness (one yes/no approval, 30s hot-reload). The Graph-Guided Retriever (GGR) uses a 2,859-node knowledge graph as a candidate gate with intent routing to deliver 71.6x token reduction. The Scene-Aware Ranker (SAR) applies 19-class entity recognition and explicit scenario annotations; negative knowledge alone contributes 25 percentage points of Hit@10 gain. On two 100-question benchmarks, Hit@10 rises from 19.1% to 96.6% (+77.5pp) and knowledge coverage from 56% to 77%, at 4.84--5.33s end-to-end latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。