arXiv:2605.21411cs.CV2026-05中稿 · CVPR

让道路视频生成的描述可调节语气,提升沟通实效性。

RoadTones: Tone Controllable Text Generation from Road Event Videos

论文配图:RoadTones: Tone Controllable Text Generation from Road Event Videos
图 1 · 摘自论文原文
  • 构建多语气标注数据集,支持语气可控的视频描述生成
  • 提出可生成中间推理过程的可控模型,提升生成可解释性
  • 设计联合评估体系,兼顾事实准确与语气契合度

现有视频-语言模型能生成道路事件的事实性描述,但缺乏对表达语气、紧迫感或风格的控制,限制了在通信关键场景中的应用。为此,我们构建了一个涵盖数据、模型与评估的完整方案,用于实现语气可控的道路视频描述生成。通过人工验证的数据生成流程,扩展了道路视频语料库,引入多样化的语气标注和多语气描述,形成 RoadTones-51K 数据集。我们提出 RoadTones-VL-CoT 模型,支持生成语气条件下的思维链中间草稿,增强可解释性。同时,设计 RoadTones-Eval 评估体系,联合衡量事实一致性与语气贴合度。用户研究结果验证了生成内容的质量、语气控制能力及事实准确性。这些贡献共同奠定了情境敏感的语气可控视频描述生成基础。

原文摘要 · Abstract (English)

Existing video-language models can generate factual descriptions of road events but lack control over how these events are expressed: their tone, urgency, or style. This limits deployment in communication-critical settings where the effectiveness of a message depends on both content and presentation, not just factual accuracy. To mitigate this, we introduce a comprehensive dataset-model-evaluation suite for tone-controllable road video captioning. Our human-validated data generation pipeline expands road-video corpora with diverse tonal annotations and multi-tone captions, yielding the RoadTones-51K dataset. We propose RoadTones-VL-CoT, a controllable video-to-text model that also generates tone-conditioned Chain-of-Thought intermediate drafts for interpretability. We also introduce RoadTones-Eval, a new evaluation suite that jointly measures factual consistency and tone adherence. In addition, we conducted a user study whose results validate caption quality, tone control, and factual consistency. Together, these contributions lay the foundation for context-sensitive tone-controllable video captioning.

视频生成语气控制多模态评估体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。