V-RAE 用冻结视觉基础模型表征替代传统 VAE 潜空间
谢赛宁转发介绍 V-RAE。该模型直接采用冻结视觉基础模型 DINOv3、SigLIP2、EUPE、V-JEPA 2.1 的表征作为视频生成潜空间,而不是传统 VAE 潜空间。
![]()
匹配设置下性能与收敛速度
在匹配设置下,V-RAE 在 Kinetics-600 上达到 2.13 rFVD,在 UCF101 上达到 117.86 gFVD。收敛速度比 VAE 潜空间快 6 倍。
引入 tFVD 评估时序平滑度
V-RAE 还引入 tFVD 来评估时序平滑度。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.