NEWS / 摘要
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
来源摘要
Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose perfo…
摘要由机器生成(来自来源站点),可能存在偏差; 本站不转载全文,请以原文为准。