arXiv:2506.14823cs.CVcs.AI2025-06

用符号推理让机器读懂动物图像和自然语言问题

ViLLa: A Neuro-Symbolic approach for Animal Monitoring

  • 视觉+语言+逻辑三模块协同,分离感知与推理
  • 可准确回答数量、位置等结构化问题
  • 适合需要透明可解释性的生态监测场景

自然环境中监测动物种群需同时理解视觉数据与人类语言查询。本文提出ViLLa(视觉-语言-逻辑方法),一种神经符号框架,实现可解释的动物监测。该框架包含三个核心组件:用于识别图像中动物及其空间位置的视觉检测模块,用于理解自然语言查询的语言解析器,以及基于逻辑推理的符号推理层。给定一张图像和问题如“场景中有多少只狗?”或“野牛在哪里?”,系统将视觉检测结果转化为符号事实,并利用预定义规则推断出关于数量、存在性与位置的准确答案。与端到端黑箱模型不同,ViLLa将感知、理解与推理分离开来,具备模块化与透明性。在多种动物图像任务上评估表明,该系统能有效连接视觉内容与结构化的人类可读查询。

原文摘要 · Abstract (English)

Monitoring animal populations in natural environments requires systems that can interpret both visual data and human language queries. This work introduces ViLLa (Vision-Language-Logic Approach), a neuro-symbolic framework designed for interpretable animal monitoring. ViLLa integrates three core components: a visual detection module for identifying animals and their spatial locations in images, a language parser for understanding natural language queries, and a symbolic reasoning layer that applies logic-based inference to answer those queries. Given an image and a question such as "How many dogs are in the scene?" or "Where is the buffalo?", the system grounds visual detections into symbolic facts and uses predefined rules to compute accurate answers related to count, presence, and location. Unlike end-to-end black-box models, ViLLa separates perception, understanding, and reasoning, offering modularity and transparency. The system was evaluated on a range of animal imagery tasks and demonstrates the ability to bridge visual content with structured, human-interpretable queries.

动物监测神经符号可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。