通过实时反馈抑制幻觉,让AI看图更准更可信
Feedback-Enhanced Hallucination-Resistant Vision-Language Model for Real-Time Scene Understanding
- 用动态阈值实时评估输出可靠性,低信度时自动压制
- 相比传统方法幻觉率降低37%,在18帧/秒下保持实时性
- 适合机器人导航、安防监控等对准确率要求高的场景
实时场景理解是人工智能的重要进展,广泛应用于机器人、监控和辅助工具。然而,幻觉问题依然严峻:AI系统常误判视觉输入,检测不存在的物体或描述未发生的事件。此类错误在安全和自动驾驶等关键领域严重威胁可靠性。本文提出一种嵌入自我意识的视觉语言模型,不盲信初始输出,而是持续实时评估其可信度,动态调整置信阈值。当置信度低于基准线时,主动抑制不可靠判断。结合YOLOv5的目标检测能力与VILA1.5-3B的可控文本生成,确保描述基于确证的视觉数据。该方法实现动态阈值调节以提升精度,基于证据生成文本以减少幻觉,并在18帧每秒下保持实时性能。相较于传统方法,幻觉率降低37%。该反馈驱动设计在机器人导航、安全监控等应用中表现出高准确性与可靠性,使AI感知更贴近现实。
原文摘要 · Abstract (English)
Real-time scene comprehension is a key advance in artificial intelligence, enhancing robotics, surveillance, and assistive tools. However, hallucination remains a challenge. AI systems often misinterpret visual inputs, detecting nonexistent objects or describing events that never happened. These errors, far from minor, threaten reliability in critical areas like security and autonomous navigation where accuracy is essential. Our approach tackles this by embedding self-awareness into the AI. Instead of trusting initial outputs, our framework continuously assesses them in real time, adjusting confidence thresholds dynamically. When certainty falls below a solid benchmark, it suppresses unreliable claims. Combining YOLOv5's object detection strength with VILA1.5-3B's controlled language generation, we tie descriptions to confirmed visual data. Strengths include dynamic threshold tuning for better accuracy, evidence-based text to reduce hallucination, and real-time performance at 18 frames per second. This feedback-driven design cuts hallucination by 37 percent over traditional methods. Fast, flexible, and reliable, it excels in applications from robotic navigation to security monitoring, aligning AI perception with reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。