arXiv:2502.09913cs.AIcs.HC2025-02

用大模型实现网页端零样本源定位,无需人工即可快速找隐患

AutoS$^2$earch: Unlocking the Reasoning Potential of Large Models for Web-based Source Search

  • 基于链式思维提示,让大模型模拟人类推理决策
  • 性能接近人机协作,响应速度更快且无需众包
  • 适合工业安全、风险管控等实时性要求高的场景

基于网络的管理系统广泛应用于风险控制与工业安全领域。然而,如何有效集成源定位功能以帮助决策者快速发现并处理隐患(如气体泄漏)仍是一大挑战。以往方法依赖网络众包或AI算法,存在人力成本高、响应慢等问题。为此,本文提出AutoS²earch框架,利用大模型在网页应用中实现零样本源定位。该框架在简化的可视化界面中运行,通过链式思维提示模拟人类推理,多模态大语言模型将视觉观测转化为语言描述,使大模型对四个方向选择进行语言推理。大量实验表明,AutoS²earch性能接近人类-人工智能协同搜索,同时摆脱了对众包劳动力的依赖。本研究为其他工业场景中自主系统的设计提供了重要启示。

原文摘要 · Abstract (English)

Web-based management systems have been widely used in risk control and industrial safety. However, effectively integrating source search capabilities into these systems, to enable decision-makers to locate and address the hazard (e.g., gas leak detection) remains a challenge. While prior efforts have explored using web crowdsourcing and AI algorithms for source search decision support, these approaches suffer from overheads in recruiting human participants and slow response times in time-sensitive situations. To address this, we introduce AutoS$^2$earch, a novel framework leveraging large models for zero-shot source search in web applications. AutoS$^2$earch operates on a simplified visual environment projected through a web-based display, utilizing a chain-of-thought prompt designed to emulate human reasoning. The multi-modal large language model (MLLMs) dynamically converts visual observations into language descriptions, enabling the LLM to perform linguistic reasoning on four directional choices. Extensive experiments demonstrate that AutoS$^2$earch achieves performance nearly equivalent to human-AI collaborative source search while eliminating dependency on crowdsourced labor. Our work offers valuable insights in using web engineering to design such autonomous systems in other industrial applications.

大模型源定位工业安全零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。