![]()
报告人:汪军教授
主持人:戴琼海院士、陶建华教授(网络直播平台)
报告时间:
2026年6月4日(周四)19:30-21:00
报告地点:
腾讯会议(ID:907-232-706 Password:260604)
直播链接:
https://cc207e9.livec.shangzhibo.tv/watch/11848999
主办单位:北京信息科学与技术国家研究中心
报告人简介
Jun Wang is Professor at the Computer Science department, University College London. Prof. Jun Wang is a leading expert in AI, Machine Learning, and Multiagent Systems, with over 200 publications. His research has earned eight Best Paper awards, including SIGIR Test of Time and Honourable Mentions, and has led to widely adopted algorithms used by Ray and CERN for particle discovery. He won the first global real-time bidding contest (2013) and NeurIPS 2020 black-box optimisation challenge, with solutions now deployed in industry. His patents with BT enhance personalisation in recommender systems by dynamically adjusting training data. As co-founder and Chief Scientist of UCL spinout MediaGamma (2013–2020), he led the development of AI-driven audience decision tools, helping the company secure £5.8M in funding before its acquisition in 2020.
报告摘要
We study continual experiential learning in Large Language Model (LLM) agents that integrate episodic memory with reinforcement learning. The central mechanism is reflection, where an agent uses past experiences to guide future decisions without modifying model parameters. Building on ideas from case-based reasoning and the Memento framework (Memento-1 and, Memento-2), we model learning as a memory-driven process in which agents store trajectories, cases, and reusable skills and retrieve them to improve decision making in new situations. To formalise this idea, we introduce the Stateful Reflective Decision Process (SRDP), where an agent maintains evolving memory and performs two operations: write, storing outcomes of interactions (policy evaluation), and read, retrieving relevant experiences to guide actions (policy improvement). We show how this read–write reflective learning can be integrated with reinforcement learning through retrieval-augmented policy iteration and prove that, as memory grows and increasingly covers the state space, the resulting policy converges to the optimal solution. This framework provides a principled foundation for memory-based LLM agents capable of continual adaptation during deployment. We will present our recent practical Memento-Skills agent system that is naturally integrated with existing industry scale LLM application.
鸣谢:中国人工智能学会和清华大学-福州数据技术联合研究院
系列交叉论坛
扫码观看直播
了解更多论坛信息
精彩内容,设个星标
让我们一起探索更多AI前沿资讯!
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.