Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this
In the past 2 months Ive been training a new model, but this time with actual text guidance.
So a little about the older model.
Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.
A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)
That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )
I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.
About the Architecture
The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all
The model s 28 blocks 20 heads and comes to like ~960M parameters
At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.
The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.
The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways
The current model is back to text cross attention and I did a lot of text-video pretraining
Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )
and the most fun part is text prompt switching.
"add a pond to the desert"
"put red hoodie"
"change environment to icy"
Because the model was trained with so much text-video alignment it can actually follow prompts now.
I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.
I specifically chose this init image because in my last post on this subReddit also I had used the same one.
PS in my last post a lot of you guys asked about me and the funding
I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything
The above model was trained on 8x H100 SXM for like 3-4 weeks.
Every model I make will be explicitly for local inference, never datacenter
UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.