用场景提示动态控制大模型实时行为,无需重训。
ARdena: Scenario-driven control of real-time LLM agents
- 通过结构化提示实现运行时行为调控
- 不同场景定义可引发显著不同的交互行为
- 适合需要灵活响应的实时交互系统
大语言模型(LLMs)已推动对话代理能力提升,但在实时交互环境中可靠控制其行为仍具挑战。现有方法多依赖难以适应变化需求的微调或对齐过程。本文提出分层场景驱动的LLM控制框架,通过结合持久上下文与场景特定约束,在不修改底层模型的前提下实现交互过程中的行为调整。该框架被实现为ARdena——一个支持语音交互、视觉感知、工具使用和化身响应生成的实时多模态具身代理。评估显示,仅通过场景定义即可产生显著不同的交互行为,同时保持稳定实时运行,验证了场景驱动提示在控制LLM代理中的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge. Existing approaches often rely on model fine-tuning or alignment procedures that are difficult to adapt to changing interaction requirements. This paper introduces layered scenario-driven LLM control, a framework that enables runtime behavior control through structured prompting. By combining persistent context with scenario-specific constraints, the approach allows agent behavior to be modified during interaction without changing the underlying model. The framework is implemented in ARDena, a real-time multimodal embodied agent that integrates speech interaction, visual perception, tool use, and avatar-based response generation. The proposed approach is evaluated with respect to control effectiveness, response latency, and operational stability. The results demonstrate that scenario definitions alone can produce substantially different interaction behaviors while maintaining stable real-time operation, highlighting the effectiveness of scenario-driven prompting for controlling LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。