让大模型更懂数据库字段,通过互动补充语义信息。
ISEE: Interactive Semantic Enrichment for Database Fields

- 引入交互式语义增强系统,结合用户知识补全字段描述
- 实验证明可显著降低认知负担,提升描述质量与下游任务表现
- 适合需要理解复杂数据字段的研究者和数据工程师
基于大语言模型的智能体在数据理解、探索和检索等任务中日益普及,但其性能高度依赖数据语义的清晰度与完整性。实践中,许多字段描述模糊或不完整,因关键上下文(如自定义字段含义)多源自用户领域知识,极少公开记录。这一差距限制了智能体在实体链接等下游任务中的表现。为此,我们提出全新的交互式语义增强系统(ISEE)。给定字段描述,ISEE通过评分机制评估其质量,主动收集领域知识,并与用户协作丰富语义。通过用户研究、自动化用户模拟、定量评估与案例分析,我们证明ISEE能显著降低认知负荷,提升描述质量,并增强下游任务性能。
原文摘要 · Abstract (English)
LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their performance heavily depends on the clarity and completeness of data semantics. In practice, many field descriptions remain ambiguous or incomplete, as much of the essential context (e.g., the meaning of a customized field) originates from users' domain knowledge and is rarely documented publicly. This gap restricts the agents' task performance in downstream tasks, such as entity-linking. To bridge this gap, we introduce a novel and comprehensive Interactive SEmantic Enrichment system (ISEE). Given a data field description, ISEE measures its quality through a scoring system, gathers domain knowledge, and collaboratively enriches the semantics with users. Through a user study, automated user simulation, quantitative evaluation, and case study, we demonstrate that ISEE significantly reduces cognitive load, improves description quality, and enhances downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。