NEWS / 摘要

MiMo-VL Technical Report

来源摘要

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on…

摘要由机器生成(来自来源站点),可能存在偏差; 本站不转载全文,请以原文为准。