Perplexity 公布了一项新研究:通过提示引导的自蒸馏,对一个 Computer 模型进行后训练,让它从自身错误中学习。
在一次线上 A/B 测试中,较晚训练的 checkpoint 相比早期 checkpoint,将工具调用失败率降低了 21.2%。
![]()
这项研究的思路是让模型从自身错误中学习,而不是依赖外部标注。后训练过程中,提示引导的自蒸馏成为关键手段。
线上 A/B 测试的结果直接体现在工具调用失败率上:较晚训练的 checkpoint 相比早期 checkpoint 降低了 21.2%。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.