Embodied AI Multi-Agent System Laboratory Automation Vision-Language-Action Task Planning
摘要

针对湿实验室自动化中协议非结构化及视觉环境复杂等挑战,本文提出 BioProVLA-Agent。该系统基于视觉 - 语言 - 动作模型,通过定制 LLM 将实验协议解析为可验证子任务,结合视觉状态验证与具身执行形成闭环工作流。为解决透明器皿等视觉干扰,开发了 AugSmolVLA 在线增强策略。实验表明,该系统在多种生物操作任务中显著提升了执行稳定性与鲁棒性,为低成本、以协议为中心的具身 AI 提供了新路径。

AI 推荐理由

论文核心在于将非结构化协议解析为可验证的子任务,并执行多步闭环工作流,属于典型的任务规划。

研究机构
Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education, East China University of Science and Technology, Shanghai, CN Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, CN Department of Laboratory Medicine, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, CN School of Information Science and Technology, Shihai University, Shijiezi, CN
论文信息
作者 Zhaohui Du, Zhe Wang, Hongmei Fei, Xiwen Cao, Ting Xiao et al.
发布日期 2026-05-08
arXiv ID 2605.07306
相关性评分 9/10 (高度相关)