NEWS / 摘要

Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

来源摘要

Reinforcement learning (RL) has emerged as a crucial approach for enhancing the capabilities of large language models. However, in Mixture-of-Experts (MoE) models, the routing mechanism often introduces instability, even leading to catastrophic RL training collapse. We analyze th…

摘要由机器生成(来自来源站点),可能存在偏差; 本站不转载全文,请以原文为准。