根据任务上下文给物体可用性排序,提升智能体决策准确性
Leverage Task Context for Object Affordance Ranking
- 用任务上下文建模物体优先级,融合图像与文本信息
- 在50k图像、66万+物体上验证,显著优于现有模型
- 适合需要精准物体选择的机器人、AI助手场景
智能体根据物体的可用性完成不同任务,但如何依据任务上下文选择合适物体仍缺乏研究。现有方法将同一类可用性中的物体视为等价,忽略了其优先级随任务上下文变化的特点,限制了复杂环境下的准确决策。为此,本文提出基于任务上下文的物体可用性排序方法:给定复杂场景图像和任务上下文的文本描述,揭示任务-物体关系并确定检测到物体的优先级顺序。我们设计了包含任务关系挖掘模块和图组更新模块的上下文嵌入分组排序框架,实现任务上下文的深度融合与全局相对关系传播。由于缺乏此类数据,我们构建了首个大规模面向任务的可用性排序数据集,涵盖25个常见任务、超过50,000张图像和661,000+个物体。实验表明,该方法在显著优于当前先进模型,在显著性排序和多模态目标检测任务中表现优异。源代码与数据集将公开发布。
原文摘要 · Abstract (English)
Intelligent agents accomplish different tasks by utilizing various objects based on their affordance, but how to select appropriate objects according to task context is not well-explored. Current studies treat objects within the affordance category as equivalent, ignoring that object affordances vary in priority with different task contexts, hindering accurate decision-making in complex environments. To enable agents to develop a deeper understanding of the objects required to perform tasks, we propose to leverage task context for object affordance ranking, i.e., given image of a complex scene and the textual description of the affordance and task context, revealing task-object relationships and clarifying the priority rank of detected objects. To this end, we propose a novel Context-embed Group Ranking Framework with task relation mining module and graph group update module to deeply integrate task context and perform global relative relationship transmission. Due to the lack of such data, we construct the first large-scale task-oriented affordance ranking dataset with 25 common tasks, over 50k images and more than 661k objects. Experimental results demonstrate the feasibility of the task context based affordance learning paradigm and the superiority of our model over state-of-the-art models in the fields of saliency ranking and multimodal object detection. The source code and dataset will be made available to the public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。