语音特征能提升大模型对速配中吸引力的预测效果。
Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating

- 融合文本与语音特征,提升配对吸引力排序准确率。
- 语音补充使配对排名准确率显著提高,但个体相关性提升不显著。
- 语音优势在语音预测更准的参与者中尤为明显,具条件性。
大型语言模型(LLMs)可基于对话文本预测人际吸引力,但语音信息能否带来额外增益仍不明确。本文利用日语速配对话数据,将仅依赖文本的LLM预测与有监督语音预测结果结合,评估参与者对其伴侣喜爱程度的预测表现。结果显示,语音信息可有效补充文本模型,在所有测试条件下均显著提升配对排名准确率;然而,个体层面的皮尔逊相关系数(Pearson $r$)提升在不同轮次和评价方向中存在差异,校正后无一显著。进一步分析发现,这些相关性提升集中于语音预测表现更优的参与者。表明语音虽在整体上增强预测能力,但其互补性具有条件性,关键在于何时何地有效。
原文摘要 · Abstract (English)
Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants' reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。