斯坦福大学一项新研究提出名为Prefix Sliding的方法,在生成阶段丢弃推理轨迹中间的token,只保留包含指令与工具的前缀以及末尾数千个token。这样处理之后,内存占用被限制在一个固定上限内。
无需额外训练即可提速
![]()
该方法不需要对现有模型进行训练,就能让模型速度提升3倍,同时性能与全注意力机制持平。研究还支持超过100,000 token的强化学习rollout,为长上下文推理提供了新的实现路径。
相关论文已发布在arxiv.org,编号为2608.26070。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.