用多智能体协作生成安全场景合成数据,解决真实数据难获取问题。
AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data
- 双智能体协同:评估者与编辑者通过大模型实现语义一致性迭代优化
- 生成符合安全规范的真实感图像,提升合成数据的可用性与真实性
- 适合安全监控、建筑施工等高风险领域数据增强需求
危险情境数据稀缺严重制约了安全关键应用(如建筑安全)中AI系统的训练,因伦理与物流限制难以采集真实数据。为此,我们提出一种端到端多智能体框架AgentSGEN,通过两个智能体在循环中的协作生成合成数据:评估者智能体作为基于LLM的判官,确保场景的语义一致性和安全约束;编辑者智能体则根据反馈生成并优化场景。依托大模型的推理与常识能力,该设计可生成符合真实规格的安全场景图像,弥补现有方法在语义深度上的不足。实验表明,该方法能有效生成兼具安全性与视觉语义质量的合成图像,为多媒体安全应用中的数据短缺问题提供可行解决方案。
原文摘要 · Abstract (English)
The scarcity of data depicting dangerous situations presents a major obstacle to training AI systems for safety-critical applications, such as construction safety, where ethical and logistical barriers hinder real-world data collection. This creates an urgent need for an end-to-end framework to generate synthetic data that can bridge this gap. While existing methods can produce synthetic scenes, they often lack the semantic depth required for scene simulations, limiting their effectiveness. To address this, we propose a novel multi-agent framework that employs an iterative, in-the-loop collaboration between two agents: an Evaluator Agent, acting as an LLM-based judge to enforce semantic consistency and safety-specific constraints, and an Editor Agent, which generates and refines scenes based on this guidance. Powered by LLM's capabilities to reasoning and common-sense knowledge, this collaborative design produces synthetic images tailored to safety-critical scenarios. Our experiments suggest this design can generate useful scenes based on realistic specifications that address the shortcomings of prior approaches, balancing safety requirements with visual semantics. This iterative process holds promise for delivering robust, aesthetically sound simulations, offering a potential solution to the data scarcity challenge in multimedia safety applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。