NEWS / 摘要

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

来源摘要

Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose perfo…

摘要由机器生成(来自来源站点),可能存在偏差; 本站不转载全文,请以原文为准。