arXiv:2608.10330cs.AI2026-08

用分层组合性提升智能助手理解模糊指代的能力

Hierarchical Compositionality for An Assistive AI Agent

论文配图:Hierarchical Compositionality for An Assistive AI Agent
图 1 · 摘自论文原文
  • 基于人类验证的语义特征构建分层属性体系,实现对象表征
  • 在用户交互历史中自动识别概念层级,准确率显著优于现有方法
  • 适合需要个性化理解的辅助系统开发,尤其擅长处理语义歧义

AI助手在各类应用中日益重要,当前主流依赖大语言模型等深度网络架构,但其资源消耗大、决策不透明,在新情境下常做出任意判断。本文提出一种基于早期人工智能核心原则的新型架构设计,聚焦于解决人类对话中指代模糊的问题。人类通过利用领域上下文和对方偏好等组合性知识来化解歧义。受此启发,我们构建了一个嵌入分层组合性的架构:以人类验证的语义特征为基础表示领域对象,结合有限用户交互历史中自动识别出的概念与属性层级。助手通过推理该组合层次结构、领域动态规则,以及语义兼容性、会话显著性和用户主题偏好模型,必要时请求人类澄清。实验表明,该方法在多种场景下持续优于数据驱动基线,展现出对特定用户画像的良好适应能力。

原文摘要 · Abstract (English)

AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents based on core principles that can be traced back to the early pioneers of AI but are not fully utilized in modern AI methods. We do so in this paper in the context of the core problem of AI agents addressing ambiguity in the objects being referred to by the human participants. Humans address such ambiguity by heuristically leveraging compositional knowledge of domain context and the preferences of the other human participants. Drawing inspiration from this observation, we describe an architecture that embeds the principle of hierarchical compositionality and uses simple heuristics to achieve the desired disambiguation. Specifically, domain objects are represented in terms of primitive attributes drawn from human-validated semantic feature norms, and a hierarchical combination of attributes and concepts automatically identified from a limited observed history of interactions of an assistive agent with specific users. The assistive agent then achieves the desired disambiguation by reasoning with knowledge of this compositional hierarchy; axioms governing domain dynamics; and models of semantic compatibility, session salience, and user-specific thematic preference, requesting human clarification when necessary. Experiments show that our approach consistently outperforms state of the art data-driven baselines, supporting adaptation to specific user profiles.

智能助手分层组合语义消歧用户建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。