EmoS构建高保真多模态情感理解数据集,支持细粒度情绪追踪。
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

- 融合静态片段与动态语流,提升数据生态效度与信号清晰度
- 双层人工标注确保连续情绪演变的可靠标签,准确率显著优于基线
- 适合训练情感识别与共情模型,尤其关注真实场景下的情绪分析
在高压力、老龄化社会背景下,具备同理心支持能力的大规模情感模型需求日益迫切。然而现有基准难以同时兼顾生态效度、信号清晰度与可靠的细粒度标注。我们提出EmoS,一个高保真双语基准,通过严格筛选的静态片段与动态语流独白子集相结合,克服现有数据集的生态效度不足与噪声问题。依托严谨的双层人工标注流程,EmoS提供可信的真实情绪演化标注。实证结果表明,在EmoS上微调多模态大语言模型(MLLMs)相比零样本基线有显著性能提升,为未来情感识别与共情模型的训练与评估奠定基础。数据集与代码已公开:https://github.com/NLP2CT/EmoS。
原文摘要 · Abstract (English)
In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more critical than ever. However, existing benchmarks fail to simultaneously achieve ecological validity, signal clarity, and reliable fine-grained labeling. We introduce EmoS, a high-fidelity bilingual benchmark designed to resolve the limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset. Supported by a rigorous dual-layer human annotation pipeline, EmoS provides trusted ground truth that captures continuous emotional evolution. Empirical results show that fine-tuning MLLMs (multimodal large language models) on EmoS yields significant gains over zero-shot baselines, laying the foundation for the training and evaluation of future emotion recognition models and empathy models. The dataset and code are publicly available at https://github.com/NLP2CT/EmoS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。