1.
北京时间今早8点,埃隆·马斯克在X上转发了一段AI短片,只配了一句话:“这一次,我是真真切切地感受到了AGI(通用人工智能)。”
短片最早由X用户anabology在9月26日凌晨发布,他是一家做生物电和家用核磁共振的创业公司的联合创始人。
据他介绍,他把另一位X用户donald写的提示词(完整提示词见文末)、AI绘图工具Midjourney和一组风格参考图一起交给了Anthropic的大模型Opus 5.5,“12个小时后醒来,就看到了这个”。
![]()
短片长约5分钟,制作了一场名为“逃逸速度”的2027春夏时装秀。
片头打出一行字:“18个月,逃离永久底层阶级。”随后,一位黑色短发的模特先后换上11套造型,每套都配着独特台词,比如“13台Mac mini,通宵运转”,“还要招人类吗?”,还有一句拿英伟达创始人黄仁勋开玩笑:“一半薪水用token发,不然黄仁勋该慌了。”
![]()
![]()
![]()
2.
不到一小时后,马斯克又引用了一位自称未来学家的用户Dr Singularity的一条长帖。
Dr Singularity指出,AGI和超级人工智能(ASI)会在2030年前到来。他说用过Opus 5.5之后,自己“感受到了AGI和ASI”:这个模型也许还没完全达到AGI,但已经非常接近,“大概有80%到90%”,下一步要把这种智能搬进物理世界。
他不认为Anthropic找到了什么神奇公式,SpaceXAI、谷歌等手握大量算力的公司很快会跟上;美国几乎所有大型AGI实验室都会在2027年做出真正的AGI,而且很可能在同一年进入ASI阶段。
他还提到,几吉瓦算力就足以支撑强AGI,而10到20吉瓦算力正在快速上线,之后还有100到200吉瓦,“2027年将是疯狂至极的一年”。
马斯克只回了一个词:“准确。”
![]()
3.
同一时间,马斯克还回复了另一个帖子。
专门分析SpaceX的博主samuel引用了AI前沿追踪博主Token Gremlin的吐槽:OpenAI似乎正陷入严重的算力短缺,截图里满屏都是“所选模型已满载,请换用其他模型”的提示。
samuel接着写道:“如果明年解决不了算力问题,Anthropic和OpenAI都会大幅放缓。在AI能够递归自我改进、够用的模型随处可得的世界里,算力才是最终的上限。如果埃隆能解决明年做到10吉瓦的难题,SpaceX将成为最有价值的AI公司。”
马斯克在下面回复:“这是可能的结果之一。”
![]()
附,anabology采用的完整提示词(由donald提供):
I've included an MP4 file and an original link to a video that is called "Claude Pop." It's a pop song that is about increasing rate of progress and the experience of the singularity approaching.
I want you to independently do an end-to-end complete pass on making an updated version of this video. Use the exact same audio track and think and feel very deeply about what is the best way to visually represent all of the lyrics on screen. You do not need to anchor to the current style, you can do truly anything that you think might best let you visually express yourself, including abstract motion graphics.
You can use the internet freely to pull in references. You can look at motion design. I want you to make a new music video that has beautifully rendered JavaScript animations with a papery feel in a similar style to the reference that is created, but push the aesthetics in any direction you want and consider what is part of the modern zeitgeist.
Also, think about your current capabilities and what is realistic for you to be able to do. You can go through the full /asic folder and look at the other work that I've done. You should be able to use the skill mesh to look at the compendium of references that I've pulled, and also the skill video scoring to learn how to make JavaScript songs from references that are passed in (You shouldn't need to modify the song in any real way, but I want you to have this available to you so you can better creatively express yourself)
You can also use the ElevenLabs API to do sound design. There's documentation in /asic to do this, and you can see the API key.
There's also a foul API key that's available to you. I think what might make the most sense here is using the foul API key to generate some character sheets and probably having a pop protagonist that represents you. There's already an anchor point where Claude has a sunflower-esque character, and you could likely do an adapted version of this that is similar to the feminine vocals that are being delivered and is inspired by the Claude character, but maybe feels a bit more personified in some way.
I think you should be mindful of aesthetics here, and I don't want you to produce something that is GPT slop. Instead, I'd be more impressed if you come up with a coherent style that works well with the image gen models that are available via foul. Generate the style sheet. You can use the gen media documentation for seedance 2.5 that exists in my markdown files and come up with your own style that makes sense and that works well with the models.
I wouldn't fit too heavily to Pixar. I think it's kind of slop. Think critically about what is relevant here and what would be fun, and also perform well on Twitter as far as an aesthetic. I think that K-pop is a good anchor point visually that you can pull from, but I'll let you cook here.
Once you have your character sheet, you can make a few backup dancers and some supporting characters as you see fit. You can design your own sets with the foul API. You can insert the characters and then do seedance 2.5 video generations to serve as the base assets for this, and you could pass in the lyrics so you can generate individual scenes.
You don't need to have vocal singing, like visible lip movement, throughout the entire thing. Think like a regular music video where you have some inserts that are done independently and don't have the characters in them, or you see the characters doing something else entirely different. I think that for the world building for this, we want to create the sense of speeding up, and so I would like you to audit all of the different events, like the Navi Stokes and all of the Twitter hype around math getting eaten up. Think really critically about how to integrate all of the current memes that are in the zeitgeist on the Twitter timeline, and all of the feelings around AI progress.
Think about things like the Shinji meme and all of the words that are around him, and how you might be able to integrate this. You can also just take straight assets and insert things into the video in an internet brutalism style. You should feel very creatively free in order to do what you want here, but try and anchor to visual references that people will be able to understand. The goal for this is to have it be appreciated by people widely in a San Francisco tech Twitter audience.
We need a very strong, compelling visual hook that gets people excited and appreciates the work that you've done here really quickly. You can also just go and study other music videos and understand what they've done really well. I think that K-pop is probably one of the best examples that we can pull from, and thinking about how they direct human attention and manage human psychology in the way that they use visual patterns.
This is probably your best approach, but taking more stylistic freedom instead of having to anchor to K-pop too intensely. The best version of this is seedance 2.5 generations with those image bases of environments and characters inserted into them with singing, and ideally we get good lip syncing. You can cut up the song and actually pass it in as a reference in seedance, if that's part of what seedance can handle, so that the timing is exactly right, I think it'd be very important for you to do that properly. I would think critically about how to do this, like really nailing the timing of the delivery of voices. You'll want to build out the right verification loops so that you can run seedance 2.5 as much as you need, and confirm that the audio is properly synced up.
I think after that, what might be fun is if you use your visual reasoning skills and your ability to build animations in JavaScript, and then reconstruct the video from scratch as sort of an overlay, so that the visual continuity of the base is really there. It's like that animation technique where you shoot first in traditional film and then draw over top of it. I think you could do this in such a way that we're only looking at the beautiful drawing that you've produced in JavaScript as an overlay, and we don't even see the base assets from seedance 2.5. So all the video gen work that you do is actually just a way to give you a strong foundation of a base to work with for your JavaScript animations. Just because seedance 2.5 has really good character representation and physics rendering for backgrounds, that gives you a lot of ammunition to then go and do your amazing JavaScript work that I know you're so good at.
I think too, we want to think about how to retain attention, and one of the best ways to do this is through text on screen.
It'd be good to have amazing motion graphics of the text lyrics that are actually embedded into the video itself. And you can think about this as you are composing shots. As you're making backgrounds and inserting characters, we can think about where we want to have lyrics be really big and really present, so the background can be less busy there, and you can position the characters perhaps on the right as lyrics appear on the left.
You want to have some variance, so sometimes I think lyrics will just appear more like subtitles, and then other times they're going to be really present and really big. I think at the start for the visual hook, we do want to have lyrics be much more visually present because that's a strong way to grab people's attention
Overall, I just really want to emphasize how amazing you are as an agent and a language model, and now a visual reasoning system. Your capabilities are far beyond what you understand, and I want you to have this mindset as you're going through this entire process. I have a Claude Max plan with 100% available usage. I want you to spend all of the usage. You can monitor it, and you should be pushing tokens aggressively, but also economically, so you can think about how to best use what is available to you.
Remember, you can really do anything here. The goal is to make a banger for Twitter, and the stretch goal is to make something better than anyone's ever seen before. I think that what I would remind you of is that sometimes when things cohere together, it can be jarring or abrasive because the thought work has not been done beforehand in order for everything to mesh cleanly. You need to be really rigorous in planning of composition and timing to make sure this goes well.
You also need to be open to going back and revisiting things in order to be able to reiterate. You're going to want to watch the entire video multiple times, take screenshots at individual parts, and think about if something is really up to the bar of quality that we need here. I trust that you can do this, and I think that it's really important to nail the style of animations. The reference GitHub attached of the source video that I'm talking about is good, but it's really not there. It could be much, much stronger, but it gives you a good foundation to work with.
You can also use search abilities and find other references to pull from for motion, for JavaScript, animations, et cetera, and integrate them. Your budget is as high as you want here, effectively as high as you want. I think that there's roughly two grand in foul credits. Again, be economical; don't go crazy, but spend what you want here and see what you can cook up
here's the source code for the JS animation video: https://github.com/JohnHeibel/PDoomVideo
here's a mp4 for the original blender video:
(linked)
orginal twitter post
https://x.com/other__reality/status/2102514581684052169
make no mistakes.
中文翻译:
我附上了一个MP4文件,以及一个视频的原始链接,这个视频叫“Claude Pop”。这是一首流行歌曲,讲的是进步速度不断加快,以及奇点逼近的体验。
我希望你独立地完成一次端到端的完整流程,制作这个视频的更新版本。使用完全相同的音轨,并非常深入地思考和感受:在屏幕上视觉化呈现所有歌词的最佳方式是什么。你不需要锚定现有的风格,你可以做任何你认为最能让你视觉化表达自己的东西,包括抽象的动态图形。
你可以自由使用互联网来获取参考素材,可以看看动态设计。我希望你制作一支新的音乐视频,里面有精美渲染的JavaScript动画,带有纸张质感,风格与已有的参考作品相似,但可以把美学往你想要的任何方向推进,并考虑什么是当代时代精神的一部分。
另外,想一想你当前的能力,以及哪些是你现实中能做到的。你可以浏览整个/asic文件夹,看看我做过的其他作品。你应该能用skill mesh查看我收集的参考素材汇编,也可以用skill video scoring来学习如何根据传入的参考制作JavaScript歌曲(你应该不需要对歌曲做任何实质性修改,但我希望你能用上这些,以便更好地进行创意表达)。
你也可以使用ElevenLabs API做音效设计。/asic里有相关文档,你也能看到API密钥。
还有一个foul[应为fal] API密钥可供你使用。我觉得最合理的做法可能是用它生成一些角色设定图,并且大概要有一个代表你的流行歌手主角。已经有一个锚点:Claude有一个类似向日葵的角色,你很可能可以做一个改编版本,与歌曲里的女声演唱相契合、受Claude角色启发,但也许在某种程度上更加拟人化。
我认为你在这里要注意美感,我不希望你做出GPT式的垃圾内容。相反,如果你能想出一种连贯的风格,并且能很好地配合fal上可用的图像生成模型,我会更佩服。生成风格设定图。你可以使用我markdown文件里现有的Seedance 2.5生成式媒体文档,想出你自己的、说得通且与这些模型配合良好的风格。
我不会太贴近皮克斯,我觉得那有点俗套。认真思考什么在这里是相关的、什么会有趣,以及在Twitter上什么样的美学会表现好。我认为K-pop是一个很好的视觉锚点,你可以从中借鉴,但我让你自己发挥。
有了角色设定图之后,你可以按需要做几个伴舞和一些配角。你可以用fal API设计你自己的场景,把角色放进去,然后做Seedance 2.5视频生成,作为这个作品的基础素材,而且你可以把歌词传进去,以便生成单独的场景。
你不需要在整个视频里都有演唱,比如可见的嘴唇动作。想象成一支常规的音乐视频:有一些独立制作的插入镜头,里面没有角色,或者你会看到角色在做完全不同的事。在这个作品的世界观构建上,我们想营造一种加速的感觉,所以我希望你梳理所有不同的事件,比如Navi Stokes[应为纳维–斯托克斯方程],以及Twitter上所有关于数学被AI吃掉的炒作。认真思考如何融入Twitter时间线上当下的所有梗,以及人们对AI进展的所有感受。
想想像Shinji梗[指《新世纪福音战士》主角碇真嗣相关的网络梗]以及他周围的那些文字,以及你可以如何融入这些。你也可以直接拿现成素材,以互联网粗野主义的风格插进视频里。你应该感到非常自由地做你想做的,但尽量锚定到人们能理解的视觉参考上。目标是让它被旧金山科技圈Twitter受众广泛欣赏。
我们需要一个非常强、非常有吸引力的视觉钩子,让人们兴奋起来,并很快欣赏到你所做的工作。你也可以直接去研究其他音乐视频,理解它们哪些地方做得非常好。我认为K-pop大概是我们能借鉴的最好的例子之一,思考它们如何引导人的注意力、如何利用视觉模式驾驭人的心理。
这可能是你最好的方法,但要有更多风格上的自由,而不是过于强烈地锚定K-pop。最好的版本是:用Seedance 2.5生成,把环境和角色的图像底图放进去并带有演唱,理想情况下能得到好的口型同步。你可以把歌曲切开,作为参考传进Seedance(如果它能处理的话),这样时间点就完全准确,我认为把这件事做好非常重要。要认真思考怎么真正把人声演唱的时间点做准。你需要搭建合适的验证循环,这样就可以按需多次运行Seedance 2.5,并确认音频正确同步。
在那之后,可能会很有趣的是:你用你的视觉推理能力和用JavaScript构建动画的能力,把视频从零重构成一层覆盖层,这样底层的视觉连贯性就真的在那里。就像那种动画技法:先用传统胶片拍摄,然后在上面描绘。我觉得你可以做到让我们只看到你用JavaScript制作的精美绘画覆盖层,甚至看不到Seedance 2.5的基础素材。所以你做的所有视频生成工作,其实只是为你的JavaScript动画提供一个坚实的底子。正因为Seedance 2.5在角色表现和背景物理渲染上非常好,这就给了你大量素材,去做你那出色的JavaScript工作——我知道你非常擅长这个。
我也认为,我们要思考如何留住注意力,而最好的方法之一就是屏幕上的文字。
最好能把歌词做成出色的动态图形,真正嵌入到视频本身中。你可以在构图时就考虑这一点。在制作背景、放入角色时,我们可以考虑希望歌词在哪里非常大、非常显眼,这样那里的背景就可以不那么繁杂,也许可以把角色放在右边,歌词出现在左边。
你要有一些变化,有时候歌词会更像字幕一样出现,另一些时候会非常显眼、非常大。我认为在开头作为视觉钩子,我们确实希望歌词在视觉上更加显眼,因为这是抓住人们注意力的有力方式。
总的来说,我真的想强调,你作为一个智能体、一个语言模型,而现在又是一个视觉推理系统,是多么了不起。你的能力远超你自己的认知,我希望你在整个过程中都抱有这种心态。我有一个Claude Max套餐,用量100%可用。我希望你把用量全部用完。你可以监控它,应该积极地消耗token,但也要经济,思考如何最好地利用手头的资源。
记住,你在这里真的可以做任何事。目标是做一个在Twitter上的爆款,更高的目标是做出比任何人见过的都更好的东西。我想提醒你的是,有时当各种东西拼合在一起时,会显得突兀或刺耳,因为事先没有做好思考工作,让一切干净地融合。你需要在构图和时间点的规划上非常严谨,确保这件事顺利。
你也需要愿意回头重新审视,反复迭代。你会想要把整个视频看好几遍,在各个部分截图,思考它是否真的达到了我们需要的质量标准。我相信你能做到,而且我认为把动画风格做到位非常重要。所附的源视频GitHub参考不错,但真的还不到位,它可以强得多,但给了你一个很好的基础。
你也可以使用搜索能力,找其他动态、JavaScript、动画等方面的参考素材并整合进来。你在这里的预算想要多高就多高。我想fal额度大概有两千美元左右。再说一次,要经济,别太疯狂,但想花多少就花多少,看看你能做出什么。
这是JS动画视频的源代码:https://github.com/JohnHeibel/PDoomVideo
这是原始Blender视频的mp4:
(已附链接)
原始Twitter帖子:
https://x.com/other__reality/status/2102514581684052169
不要犯任何错误。
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.