arXiv:2601.00940cs.CV2026-01

提出首个真实场景液体分割数据集与模型,解决透明/反光液体难识别问题。

Learning to Segment Liquids in Real-world Images

  • 设计跨注意力机制,让边界分支增强主分割分支的预测能力
  • 在5000张真实图像上达到领先性能,14类液体分割准确率显著提升
  • 适合机器人视觉、自动驾驶等需要精准感知液体的场景

水、酒、药物等液体无处不在,但其分割研究仍有限,制约了机器人安全避障与交互能力。液体外观与形态多样,且常透明或反光,会映射背景与环境。为此,我们构建了包含5000张真实图像、14个类别的液体数据集LQDS,并提出新型液体检测模型LQDM,通过专用边界分支与主分割分支间的跨注意力机制提升掩码预测精度。大量实验表明,LQDM在LQDS测试集上优于现有方法,为液体语义分割建立了强基线。我们相信LQDS和LQDM将推动该领域研究,并促进机器人等实际应用。数据集与代码已公开于https://lonaslee.github.io/LQDM/。

原文摘要 · Abstract (English)

Liquids like water, wine and medicine are everywhere. However, limited attention has been given to the task of segmenting liquids, hindering the ability of robots to safely avoid and interact with them. The segmentation of liquids is difficult because liquids come in diverse appearances and shapes; moreover, they can be both transparent or reflective, taking on arbitrary objects and scenes from their background and surroundings. To take on this challenge, we construct a liquid dataset, LQDS, consisting of 5000 real-world images annotated into 14 distinct classes, and design a novel liquid detection model, LQDM, which leverages cross-attention between a dedicated boundary branch and the main segmentation branch to enhance mask predictions. Extensive experiments demonstrate the effectiveness of LQDM on the testing set of LQDS, outperforming state-of-the-art methods to establish a strong baseline for the semantic segmentation of liquids. We believe that LQDS and LQDM will facilitate future research in liquid segmentation and enable practical applications in robotics. Our dataset and code is released at https://lonaslee.github.io/LQDM/.

液体分割真实图像机器人视觉跨注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。