arXiv:2511.20795cs.CVcs.AI2025-11

轻量化复现知识增强视觉问答模型,提升边缘设备推理能力

Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models

  • 用极少参数复现KRISP,适配手机等边缘设备
  • 在DAQUAR数据集上达原模型75%性能,避免AI幻觉
  • 揭示原模型设计缺陷,适合资源受限场景研究

Facebook AI Research提出的KRISP将外部结构化知识融入视觉语言推理流程。尽管有效,但原模型依赖工业级训练、计算成本高且紧耦合大型主干网络。本文从新角度重审KRISP,提出轻量化复现版本,参数显著减少。虽性能约为原模型的75%,复现过程揭示多项设计缺陷、实际应用陷阱及隐含问题。通过系统消融实验,包括合成VQA数据的验证和对DAQUAR数据集的评估,探讨知识增强型VQA架构在资源受限下的可扩展性与有效性。所提模型采用低参数配置并受限于外部知识图谱领域,确保输出仅限该领域内,有效防止AI幻觉;极简参数使其可在智能手机、AR-VR等边缘设备上运行,提升离线视觉推理能力。

原文摘要 · Abstract (English)

Facebook AI Research introduced KRISP [4], which integrates structured external knowledge into pipelines for vision-language reasoning. Despite its effectiveness, the original model has been developed for industrial-scale training, is computationally demanding, and is tightly connected to a large backbone. In this work, we reexamine KRISP from a different angle and offer a lightweight reproduction with significantly fewer parameters. Even though our replicated model performs about 75 % of the original, the replication process uncovers a number of design flaws, real-world pitfalls, and implicit problems that were not fully covered in the original paper. We offer insights into the scalability and efficacy of knowledge-enhanced VQA architectures under resource constraints through systematic ablation studies, which include a proof-of-concept on synthetic VQA data and evaluation on the DAQUAR dataset. Our model, configured with a low parameter setup and constrained by the external Knowledge graph domain, prevents AI hallucinations and generates outputs solely within that domain. Minimal parameters allow us to function on edge devices like smartphones and AR-VR, further improving offline visual reasoning.

视觉问答知识增强轻量化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。