τ^τ-Bench把智能体构建本身设为任务
Sierra与普林斯顿大学联合发布τ^τ-Bench基准,测试开发者智能体能否独立交付可用的客服智能体。开发者智能体会拿到真实业务记录、掌握需求的客户、生产API、待继承代码库,以及服务成本与模型限制。
![]()
最终交付的客服智能体要用保留的模拟用户部署评分,覆盖4个领域共53个任务。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.