AI写代码越来越快,审代码的AI却一直没个统一考卷。GitHub这次把尺子摆出来了。
GitHub发布开放基准ReviewBench,专门用来评估AI代码审查智能体。评测集不是随手凑的:187个开源仓库、219个PR、覆盖19种语言。
![]()
为什么是这219个PR
![]()
关键在于分布参考。ReviewBench的题目设计参照了1.039亿次pull request的真实分布,让评测集尽量贴近实际代码审查场景,而不是挑一批理想化的样本。
对做代码审查智能体的团队来说,这意味着多了一个可对齐的公开坐标。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.