arXiv:2510.10633cs.AI2025-10

多智能体协作生成图文一致图像,提升细节与多样性。

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

  • 分领域智能体协同,通过强化学习优化文本与图像生成。
  • 生成内容词数增1614%,ROUGE-1下降69.7%,语义更对齐。
  • 适合研究跨模态生成、图像风格化与智能体协作的学者。

多模态文本到图像生成仍受限于在不同视觉领域中保持语义一致性与专业级细节。我们提出一种多智能体强化学习框架,协调专注于建筑、人物肖像和景观等领域的专用智能体,在文本增强模块与图像生成模块中分别引入多模态集成组件。各智能体采用近端策略优化(PPO)训练,基于综合奖励函数平衡语义相似性、语言视觉质量与内容多样性。通过对比学习、双向注意力及文本与图像间的迭代反馈实现跨模态对齐。在六组实验中,系统显著丰富生成内容(词数增加1614%),同时将ROUGE-1得分降低69.7%。在融合方法中,基于Transformer的策略取得最高综合评分(0.521),尽管偶有稳定性问题;多模态集成表现出中等一致性(相关系数0.444至0.481),反映出跨模态语义锚定的持续挑战。这些结果凸显了协作式、专业化架构在推进可靠多模态生成系统方面的潜力。

原文摘要 · Abstract (English)

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose a multi-agent reinforcement learning framework that coordinates domain-specialized agents (e.g., focused on architecture, portraiture, and landscape imagery) within two coupled subsystems: a text enhancement module and an image generation module, each augmented with multimodal integration components. Agents are trained using Proximal Policy Optimization (PPO) under a composite reward function that balances semantic similarity, linguistic visual quality, and content diversity. Cross-modal alignment is enforced through contrastive learning, bidirectional attention, and iterative feedback between text and image. Across six experimental settings, our system significantly enriches generated content (word count increased by 1614%) while reducing ROUGE-1 scores by 69.7%. Among fusion methods, Transformer-based strategies achieve the highest composite score (0.521), despite occasional stability issues. Multimodal ensembles yield moderate consistency (ranging from 0.444 to 0.481), reflecting the persistent challenges of cross-modal semantic grounding. These findings underscore the promise of collaborative, specialization-driven architectures for advancing reliable multimodal generative systems.

文本生成图像多智能体跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。