arXiv:2603.20633cs.AI2026-03被引 54

Seed1.8让AI能多轮交互、用工具、执行任务,像人一样处理现实问题。

Seed1.8 Model Card: Towards Generalized Real-World Agency

  • 统一接口支持搜索、代码生成与图形界面操作
  • 多轮交互下保持强语言与视觉理解能力
  • 适合开发真实场景下的智能代理系统

我们提出Seed1.8,一个面向通用现实世界智能体的基础模型:从单轮预测扩展至多轮交互、工具使用与多步执行。该模型在保持强大大语言模型和视觉语言性能的同时,支持统一的智能体接口——包括搜索、代码生成与执行、图形界面交互。部署方面,提供低延迟、低成本推理能力,支持可配置的思考模式及针对图像与视频的优化视觉编码。我们在标准基准与应用对齐的工作流上进行了评估,涵盖基础能力、多模态理解与智能体行为。Seed1.8已发布,以推动交互式真实场景应用的研究与开发。

原文摘要 · Abstract (English)

We present Seed1.8, a foundation model aimed at generalized real-world agency: going beyond single-turn prediction to multi-turn interaction, tool use, and multi-step execution. Seed1.8 keeps strong LLM and vision-language performance while supporting a unified agentic interface-search, code generation and execution, and GUI interaction. For deployment, it offers latency- and cost-aware inference, including configurable thinking modes and optimized visual encoding for images and video. We report evaluations on standard benchmarks and application-aligned workflows spanning foundational skills, multimodal understanding, and agentic behavior. Seed1.8 is released to support further research and development on interactive, real-world use cases.

智能体多模态交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。