NEWS / 摘要

MiMo-Audio: Audio Language Models are Few-Shot Learners

来源摘要

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown that scaling next-token prediction pretr…

摘要由机器生成(来自来源站点),可能存在偏差; 本站不转载全文,请以原文为准。