GPT-6 Astra 在 ARC-AGI 3 基准测试中拿下 98.6%,而上一代 GPT-5.6 Sol 只有 7.8%。这个跨度大到原作者连发多个问号,直呼难以置信。
不只是推理测试,其他榜单同样夸张:
![]()
- FrontierMath Tier 4(v2):97.6%
- ExploitBench:100%
- SRE-Bench(四次尝试):99.2%
从 7.8% 到 98.6%,这已经不是迭代,是换了个物种。问题来了:是测试本身被摸透了,还是模型真的质变了?评论区吵翻了。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.