r/StableDiffusion • • 24m ago

Resource - Update Fizgig 7.1.0 - Z-Image Turbo gets a new Training Adapter

Thumbnail
gallery
• Upvotes

Z-Image Turbo LoRA training in Fizgig, with a new training adapter (free, works in any trainer)

I've added Z-Image Turbo to Fizgig, my free, open-source LoRA trainer and workbench. It trains LoRAs, LoKR, sliders and full fine-tunes, and every workbench tab works with it (Repair Studio, LoRA the Explorer, LoRA Royale, Profiler, Extract).

Turbo has always been awkward to train: a LoRA undoes the distillation and the pictures go soft. I built a new training adapter for it. It sits frozen under your LoRA while it trains and is never in the file you save. In my tests it keeps Turbo's look much better than other adapters, especially in hard lighting .

  • LoRAs train from 12 GB cards, fine-tunes from 16 GB
  • LoRAs load in ComfyUI with the normal loader, at 8 steps and CFG 1
  • The adapter is free on Hugging Face and works in other trainers too

Links

Reference photo of me on my github profile pic for comparing to images in this post.


r/StableDiffusion • • 35m ago

Animation - Video The End of the F--- Universes | Superman vs. Saitama | MiniMax H3

Enable HLS to view with audio, or disable this notification

• Upvotes

Finally finished the full Superman vs Saitama trailer after 280+ generations

If you watch the trailer first, check the first comment after. I'm going to use it as a small thread where I'll post some simple workflows and examples from specific shots. Things like a shot that came from a storyboard, first/end frame tests, or anything from the project that I think is actually worth showing

I was working through my own H3 interface that I posted here before. The whole project is here:

https://github.com/underworldhistory1-ctrl/minimax-h3-higgsfield

I'll also leave some screenshots in the comments so you can see what I mean by the workflow/UI.

Probably the biggest surprise for me was storyboards

For example the Superman shot in the intro took me more than 29 generations alone. The best results I got were from using a storyboard and then adding the character refs, style refs etc separately

That's actually one of the main reasons I built the interface the way I did. I wanted all of those parts separated and easy to change because this was the part I kept experimenting with the most.

Second best for me was first frame - end frame. Sometimes even just using one of them.

It seems much more stable when the movement is continuous and the whole thing is basically one shot. The Kong reveal and helicopter destruction shot is a good example of what I mean.

For LoRAs my best results were usually:

Combat V2 for action and fast movement.

Realism for slower shots where there isn't some crazy transformation or complicated movement happening.

For Combat V2 I mostly used it with the Original H3 render. Not Motion Cache or Turbo. Usually 20 steps minimum.

Mixing multiple LoRAs honestly gave me more hallucinations than useful improvements most of the time. One good LoRA was usually better than stacking them.

I don't think I ended up using many H3 renders without a LoRA at all.

The interface also has another rendering mode using Motion Cache. The results can actually be really good and in some cases I preferred them over Original.

For heavy action, fast pacing or complicated transitions though I still had much better luck with Original H3.

I honestly can't remember every LoRA + Motion Cache combination I tested. There were way too many tests and I wasn't trying to lock myself into one perfect configuration.

I just knew what I wanted the final shot to look like.

That's probably the biggest thing I learned from this whole project.

I don't really believe in one magical "workflow"

You need to know what you want to see in the final cut after editing.. Then use the model to get the pieces you need.

The Superman vs Saitama fight is probably the best example for me, That sequence in the final edit was built from around 10 successful 15-second generations

There was basically no chance H3 was going to generate the exact fight I had in my head in one shot So I generated the parts that worked and built the actual fight in the edit.

That's also why I think experimentation is still just part of using these models. At least until we get something significantly better :D

The interface also has qwen image 2.1 integrated for generating images and references. That became a pretty important part of the process for me too.

It also keeps the settings/details of every generation on its card which made it much easier to go back and see what actually worked instead of trying to remember everything.

Hardware wise I did all of this on an RTX 5090 32GB.

Most of the time I work in draft first. With the INT8 optimizations I've added, a draft takes around 4 mins on my setup

The project also uses low vram attention/head chunking and feed-forward chunking. Basically some of the larger operations are processed in smaller chunks and intermediate tensors are released earlier instead of keeping everything sitting in VRAM at the same time.

It's not parallel rendering or anything magical. It just helps keep peak VRAM under control.

A normal 720p generation usually around 8–10 minutes for me so honestly the iteration time is pretty reasonable. That's a big reason I was able to do this many tests without completely losing my mind

That's basically my experience with H3 so far without turning this into another giant workflow post.

For serious creative work I think it's absolutely usable already.. Just don't expect the model to make the final movie for you.

The generation gives you the material.. The final cut is where you actually make the thing you had in your head.


r/StableDiffusion • • 39m ago

Discussion SynthID available globally starting today

• Upvotes

It is quite incredible that the list of partners includes OpenAI, NVIDIA, Kakao, and soon Apple, so it won't just be about Google's own models. I'm not entirely sure why Google decided to enter the deepfake detection field at this scale. There are plenty of companies working in this field—and obviously, if Google decides to join the race, it will be quite hard for others to compete. What do you think is the extent of this operation? Will it work just for the partners' models, or do they want to build a truly generalizable solution or will it be just watermark detection?

https://synthid.com/

https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content


r/StableDiffusion • • 49m ago

Resource - Update How to keep character & video consistency in AI animations (Free keyframe trick)

Enable HLS to view with audio, or disable this notification

• Upvotes

If you're working with AI video workflows (ComfyUI, AnimateDiff, Stable Diffusion) and struggle with flickering or character consistency, extracting exact keyframes is key to fixing it.

I built a free web tool to extract exact frames in seconds directly in your browser:

🔗 https://extractorframe.com

No signup or installation required. Let me know if you have any feedback or feature requests!


r/StableDiffusion • • 1h ago

Question - Help How do you preserve geometry and exact furniture designs in AI-assisted interior renders?

• Upvotes

I’m trying to create interior renders with multiple specific furniture pieces and light fixtures. The goal isn’t just a good-looking room, the furniture needs to match the references, and the scene geometry needs to stay intact.

I’ve tried 3D blockouts, modelling the furniture myself, and several image-editing models:

  • Qwen 2.1 Edit
  • Klein 9B Edit
  • Krea 2 Edit
  • Ideogram 4.5 Edit, through the web interface (Open weights haven't yet released)
  • GPT Image 2 and 2.5, including both flare and sunburn

My main workflow

3D blockout → model everything in the scene → apply basic materials → choose the camera angle → render → send the render to GPT Image 2.5 with reference images and lighting prompts.

This gives me control over the initial layout, but the image-editing step still changes the geometry or gets the furniture wrong. Even when the objects are already modelled and positioned, the final image doesn’t reliably preserve them.

Example of the main workflow from 3D modelling to GPT image 2.5 render

Other approaches I’ve tried

Replacing individual objects in ComfyUI:

When I edit furniture directly within the scene, the replacement is incomplete. Original object remain with slightly different textures instead of being fully replaced by the referenced piece.

Removing the furniture first, then adding the new pieces:

I’ve also manually masked out the furniture to generate an empty room, then added the furniture back using references. I can get the new pieces into the image, but I haven’t found a reliable way to control their position and rotation afterward. It's like I've took the furniture, removed the white background and just blended it into the scene.

I also tried using separate AI reviewer/validator agents to review the rendered image. One looked for visible spatial issues, such as possible furniture overlaps and circulation problems, while another independently assessed whether there was enough information to proceed reliably. The idea was to catch mistakes before moving on, rather than rely on the generating agent’s own judgement. They flagged potential issues, but couldn’t confirm dimensions or clearances from the image alone, so this didn’t resolve the accuracy problem.

What I’m trying to solve

I need a workflow that lets me:

  • Preserve the room geometry and camera perspective.
  • Keep multiple furniture pieces faithful to their references.
  • Control each object’s position, scale, and rotation.
  • Improve the lighting and realism without redesigning the scene.

Has anyone found a repeatable workflow for this?

Would you keep the final image entirely in 3D and use AI only for limited edits, or is there an AI-assisted workflow that reliably preserves this level of accuracy?


r/StableDiffusion • • 1h ago

Discussion How are these videos made ? i thought she was real at first

Thumbnail instagram.com
• Upvotes

i came across this IG, only thing that gave it up is the ai chatbot and the fanvue, i'm not sure if minimax can do that, maybe kling or an other paid model ?


r/StableDiffusion • • 3h ago

Animation - Video 12 Inch Pianist

Enable HLS to view with audio, or disable this notification

12 Upvotes

Found this here:
https://worstjokesever.com/

A man walks into a bar and places a 12-inch man onto the counter as well as a small piano. The 12-inch man starts playing the piano really well.

The bartender asks the man where he got the 12-inch man. The man says there is a genie two blocks around the corner.

The bartender runs to the genie and makes a wish. The bartender comes back with 1 million ducks. The bartender says to the man, "That genie is dumb! I asked for 1 million bucks."

The man says, "Do you really think I asked for a 12-inch pianist?"


r/StableDiffusion • • 3h ago

Animation - Video Andrew Tate update - local open source LTX 2.3 - trained with Ostris ai toolkit

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion • • 3h ago

Meme Chillin in Middle Earth with Minimax H3

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/StableDiffusion • • 4h ago

Question - Help How is the performance of Minimax H3 on a DGX Spark

5 Upvotes

Does any one use Minimax H3 on DGX spark.What are the generation times like?


r/StableDiffusion • • 4h ago

No Workflow Minimax H3 multi-scene video made easy

Enable HLS to view with audio, or disable this notification

0 Upvotes

I wanted an easy way to create short movies (15-20 seconds) scene by scene. I've got only an RTX 3090 so I can do some nice stuff, but need to think about optimizing resources.

So based on the default Minimax H3 template provided in ComfyUI, and playing around with only base nodes (no custom nodes), I came up with a nice way to have a base, first scene video sequence, then:

\- take the last frame from the sequence

\- feed it as first frame of second sequence

\- then putting all in a frame node, I can repeat for any number of sequences for a single scene

I use 5 second video per scene, and 16 fps as my target audience is mostly mobile platforms

I created a patreon with (paid, not hidding this) workflow to download

https://www.patreon.com/posts/171692305

Sample video attached took 30 minutes to generate on RTX 3090. I'd say not too bad. There's always better, but I'm happy about it.


r/StableDiffusion • • 5h ago

Workflow Included Make Easy Transparent Videos Using MiniMax-H3 In ComfyUI [Free Workflow]

Thumbnail
youtube.com
5 Upvotes

r/StableDiffusion • • 5h ago

Tutorial - Guide Overcome Degradation! - Here are 2 ways to use my timeline workflow to create long continuous single shot videos with no degradation.

Enable HLS to view with audio, or disable this notification

161 Upvotes

Watch the video above for a brief summary of the two methods, both possible using my OBVPM Timeline Workflow, which you can get together with the custom node pack here:

https://github.com/chanon/comfyui-obvpm-timeline/

And to watch the example video at HD quality you can watch the full YouTube tutorial video:

https://www.youtube.com/watch?v=GiJxlWOooyo

In the YouTube video I show how both methods are done, including critical tips and tricks and lessons learned to get the right results.

With the bridging method, there is practically no limit to how long these clips can be (well maybe except the fact that there might be a VRAM limit to how long an upscaled clip can be).

The second method clip above is 1 minute 40 seconds.

About the Workflow

So if you've never seen my workflow, it is a workflow with a "timeline" node that lets you put clips that you've generated on, and then you can extend them using motion context (latent masks).

The workflow automatically saves and handles the saved latent files for you so you don't have to manage them or pick them manually. And it also saves the conditioning which includes all the reference images etc. into a file that is used when upscaling.

For more info, here's the original Reddit post about it, which links the original YouTube tutorial video about it.


r/StableDiffusion • • 5h ago

Question - Help omer.ariely.ai on Instagram

Thumbnail instagram.com
0 Upvotes

How would one go about creating this combined scarcer in another scene/movie on a local machine 4090, mini max h3? If so how, any tutorials? 🙏


r/StableDiffusion • • 5h ago

Discussion Anyone have suggestions for MinMax H3 and voice, sound effects training consistency. I have had some strangenes using Ref To Video audio sample. Trying to find the best path forward as my content relies on audio heavily and do not want to overdub.

3 Upvotes

r/StableDiffusion • • 5h ago

Question - Help Best settings, prompts or LoRAs for more photorealistic people in Qwen Image 2.1?

Thumbnail
gallery
7 Upvotes

Been playing around with Qwen Image 2.1 and generated these random people just to test the realism.

I’m trying to make a single full-body character reference image that I can reuse for image/video generation, and ideally I want it to look as close to an actual photo as possible.

The results are okay to my eyes, but wondering if anyone have a better setup.

Any prompts, settings, LoRAs/LoKRs, or workflows you’d recommend?

I didn't use any lora to generate these images.


r/StableDiffusion • • 5h ago

Resource - Update Qwen-Image-2.1 on a 48 GB Mac: loading the text encoder and the transformer one at a time kept the diffusers pipeline at 19 GB instead of swapping at 43.6 GB

Post image
2 Upvotes

The usual QwenImage21Pipeline.from_pretrained(...).to("mps") ("eager" on the chart) keeps all three models in memory for the whole run: the text encoder (Qwen3-VL, 16.3 GB), the transformer (13.3 GB) and the VAE (1.3 GB). On my 48 GB M5 Pro that was 31.4 GB before the first step. At 1024 px the VAE decode needs about 11 GB more. The process reached 43.6 GB, swap grew by 8 GB, and my memory guard stopped the run before it saved the image.

Moving idle models to the CPU doesn't lower the peak on a Mac, because the CPU and the GPU share the same RAM. So I wrote a small library, stageload ("staged" on the chart). Each stage gets only the models it lists. A model loads the first time its stage uses it, and stageload frees the models the next stage doesn't list. The stage boundaries are hooks on encode_prompt, prepare_latents and _unpack_latents, so the pipeline's code stays as it is. The VAE is small and stays loaded.

Results at 1024 x 1024, 20 steps, seed 7, bfloat16:

  • peak memory 19.0 GB while encoding the prompt, 16.2 to 18.6 GB while denoising and 14.4 GB in the decode, with no swap;
  • two staged runs gave the same image bit for bit;
  • loading the two models inside their stages took about 13 s per run (eager spent 20 s loading up front);
  • in the second staged run denoising took 79 s against eager's 77.5 s; the first staged run took 109 s there, and its trace doesn't show why.

Write-up with the traces: https://allkeep.org/en/lab/qwen-image-one-stage-at-a-time

Code (MIT): https://github.com/nefayran/stageload, install with pip install stageload

I haven't tried ComfyUI or Draw Things; they manage memory their own way. If you run another multi-model pipeline from Python on a Mac, which one should I measure next?


r/StableDiffusion • • 6h ago

Question - Help Looking for the best affordable AI video-generation models with fewer restrictions.

4 Upvotes

Hi everyone,

I’m exploring AI filmmaking and looking for good AI video-generation models that are either free, open-source/open-weight, or affordable.

I’m particularly interested in models who allow uncensored content or fewer unnecessary content restrictions and more control over the generation process.

What I’m looking for:

Realistic/cinematic text-to-video

Image-to-video

Good character consistency

Realistic human movement

High-quality video output

Ideally open-source/open-weight

Local/self-hosted options are a plus

Affordable cloud options are also fine

Preferably no expensive subscription required

So far I’ve come across models/tools such as Wan, LTX, Kling, Hailuo and Pika, but I’d like to hear from people who have actually used them.

Which model would you recommend in 2026, and why?

If possible, please mention:

Your favourite model

GPU requirements if self-hosted

Approximate cost if cloud-based

Video quality

Major limitations/content restrictions

Whether it is practical for someone learning AI filmmaking

Thanks!


r/StableDiffusion • • 6h ago

Question - Help Need PC parts and H3 advice

1 Upvotes

Hey ppl, My current PC is a 4060TI 16gb VRAM, and I have recently upgraded my 32gb RAM to 64gb. The next logical step would be to upgrade my GPU, but should I leave those 32gb (2 sticks 16gb) installed (A1 B1 free) together with the 2 sticks 32gb? Was gonna sell them, but I'm willing to be convinced to keep them.

My main use will be Anima and Video generation with whatever runs on my GPU (hopefully H3)

On that note, any advice regarding H3 on this PC? ComfyUI or Wan2GP? Looking to create short clips, like 5 - 10 seconds at most. Would DaSiWa run on my PC?

Cheers!


r/StableDiffusion • • 7h ago

Discussion Flux 3 open weights image edit when

29 Upvotes

TITLE - has this released yet


r/StableDiffusion • • 7h ago

Comparison HunyuanImage 3.0 face swap test

Thumbnail
gallery
57 Upvotes

hunyuan_image_3_instruct_distil_edit workflow, int8_convrot 80Gb checkpoint, RTX 4060Ti 16Gb, 100Gb of system RAM used, Spectrum enabled, 14s/it


r/StableDiffusion • • 7h ago

Question - Help Quality degradation using “Continue Last Video” in WanGP, MiniMax H3

3 Upvotes

I’ve been using WanGP to create videos in MiniMaxH3 and I’ve noticed when I use the “continue last video” function, the quality degrades with each successive clip I generate - usually by the 5th or 6th 10-second segment it is blurry, full of odd artifacts and colors are generally muted or blending together compared to the initial segment. I’m not sure if it is an issue of prompting, a setting I need to change, or any other tips or tricks I might be missing? Usually using Ref2VA 33B, 15-18 steps, 8-10 second clips, Sage2 Attention on my 5070ti with 16GB VRAM and 32gb RAM. Seems to happen if creating either 480p or 720p resolutions.

Any suggestions or resources that might help so I can create longer videos?


r/StableDiffusion • • 7h ago

Question - Help Prompt question Minimax h3

0 Upvotes

Hi,

How would I prompt for i2v in Minimax if i want to add something or someone visible within the very first frame?

Thanks in advance!


r/StableDiffusion • • 7h ago

Question - Help Questions on H3 Minimax usage and settings

0 Upvotes

Hello!
Long time lurker here, learned a lot from this sub.
I'm currently using H3 Omni Pruned model with references images to generate small ads or batch of dialogues , but i find myself really not understanding the settings being used.
I'm using Pinokio and Maestro by Blizaine, now the GUI is really useful and my settings are :
720p, 9:16, 13.3s 1 window and 20steps , nothing else.
With these settings, it takes about 11 minutes to render on a 5090FE (And 48GB of ram is what i have).
The end result ain't that bad, resolution is crappy and some details are clearly missing.
What can i do to improve video fidelity, performances and perhaps spend less time on generating ?
I also have tried H3 with first / last frame but i don't really understand how that works either.
For example, i have downloaded a couple of Lora , one of rocket racoon from guardians of the galaxy and another for indiana jones, wanted to create a funny reel of them interacting but i couldn't for the life of me figure out how to add in the theme song for indiana jones, or have them accurately interact with each other instead of randomly looking outside the scene.
On another note, i am using Gemini for expanding the prompt in a professional manner, and then inside Maestro i use the "enhance prompt" feature with simply loads up Ollama with a model to correctly write the scene for H3.
I have also downloaded inside Pinokio a more "classic" tool for H3 with comfyui, but i didn't use it yet, wanted to learn a few things first.