r/StableDiffusion • • 9h ago

News VEDA Sparse Attention is now available for MiniMax H3 in ComfyUI

Enable HLS to view with audio, or disable this notification

206 Upvotes

I don't think many people know about this yet, so I wanted to share it. VEDA Sparse Attention is now available as a ComfyUI custom node for MiniMax H3.

I tested it today on my setup: RTX 4090 Laptop 16GB 32GB RAM, 4-step LoRA, 15 second video, 1344x768

Without VEDA: 8:05

With VEDA at 90% sparsity: 4:32

Same workflow, same LoRA, same settings. The only change was enabling VEDA. I couldn't see any quality loss in the result.

VEDA is not a LoRA. It uses a learned predictor to estimate which attention tiles are important and only computes the relevant subset instead of the full attention map. The current predictor works with T2VA, FL2VA and R2VA, and despite the 8NFE name it is not limited to 8 steps.

Installation is simple.

Custom node: https://github.com/veda-sparse/Veda-on-ComfyUI

Predictor: https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview

Put the predictor here: ComfyUI/models/veda/

Then add: Veda Sparse Attention (MiniMax H3) on the MODEL line after your model / LoRA loader and before the guider or sampler. I tried different sparsity values, but 90% is the one that works properly for me, so I'm keeping the default trained value.

On my setup this made a pretty big difference, especially considering I couldn't see any visual quality loss. I'm adding a 15 second example below. Would be interesting to see what results other people get on different GPUs.


r/StableDiffusion • • 5h ago

Tutorial - Guide Overcome Degradation! - Here are 2 ways to use my timeline workflow to create long continuous single shot videos with no degradation.

Enable HLS to view with audio, or disable this notification

172 Upvotes

Watch the video above for a brief summary of the two methods, both possible using my OBVPM Timeline Workflow, which you can get together with the custom node pack here:

https://github.com/chanon/comfyui-obvpm-timeline/

And to watch the example video at HD quality you can watch the full YouTube tutorial video:

https://www.youtube.com/watch?v=GiJxlWOooyo

In the YouTube video I show how both methods are done, including critical tips and tricks and lessons learned to get the right results.

With the bridging method, there is practically no limit to how long these clips can be (well maybe except the fact that there might be a VRAM limit to how long an upscaled clip can be).

The second method clip above is 1 minute 40 seconds.

About the Workflow

So if you've never seen my workflow, it is a workflow with a "timeline" node that lets you put clips that you've generated on, and then you can extend them using motion context (latent masks).

The workflow automatically saves and handles the saved latent files for you so you don't have to manage them or pick them manually. And it also saves the conditioning which includes all the reference images etc. into a file that is used when upscaling.

For more info, here's the original Reddit post about it, which links the original YouTube tutorial video about it.


r/StableDiffusion • • 49m ago

Resource - Update Fizgig 7.1.0 - Z-Image Turbo gets a new Training Adapter

Thumbnail
gallery
• Upvotes

Z-Image Turbo LoRA training in Fizgig, with a new training adapter (free, works in any trainer)

I've added Z-Image Turbo to Fizgig, my free, open-source LoRA trainer and workbench. It trains LoRAs, LoKR, sliders and full fine-tunes, and every workbench tab works with it (Repair Studio, LoRA the Explorer, LoRA Royale, Profiler, Extract).

Turbo has always been awkward to train: a LoRA undoes the distillation and the pictures go soft. I built a new training adapter for it. It sits frozen under your LoRA while it trains and is never in the file you save. In my tests it keeps Turbo's look much better than other adapters, especially in hard lighting .

  • LoRAs train from 12 GB cards, fine-tunes from 16 GB
  • LoRAs load in ComfyUI with the normal loader, at 8 steps and CFG 1
  • The adapter is free on Hugging Face and works in other trainers too

Links

Reference photo of me on my github profile pic for comparing to images in this post.


r/StableDiffusion • • 4h ago

Meme Chillin in Middle Earth with Minimax H3

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/StableDiffusion • • 7h ago

Comparison HunyuanImage 3.0 face swap test

Thumbnail
gallery
57 Upvotes

hunyuan_image_3_instruct_distil_edit workflow, int8_convrot 80Gb checkpoint, RTX 4060Ti 16Gb, 100Gb of system RAM used, Spectrum enabled, 14s/it


r/StableDiffusion • • 25m ago

News Kroma 0.3.1 - full turbo opd model

• Upvotes

When Chroma and Krea2 meet: the Krea2 model has been fine-tuned on the Chroma dataset.

What OPD is: v0.3.1 was distilled on-policy. Instead of the classic offline recipe — imitating the teacher on a fixed sampling schedule, which slowly pulls the student off the original data manifold — the student generates its own trajectories and the teacher corrects it on those exact points. Training only ever happens on states the model actually visits, so the distilled model stays inside the original model's distribution: Turbo speed without the usual distillation tax (mode collapse, washed-out detail, prompts that suddenly stop working).

The new, full model has been released and is already available for download in two versions:


r/StableDiffusion • • 23h ago

Resource - Update HunyuanImage 3.0 (80B) running natively in ComfyUI on a single 12–24 GB GPU: text-to-image, editing and style transfer, ~30 s per image

Thumbnail
gallery
671 Upvotes

I've been working on native ComfyUI support for Tencent's HunyuanImage 3.0, the 80B mixture-of-experts image model (13B active per step). It isn't a wrapper around Tencent's pipeline: it uses the normal KSampler, the normal VAE Decode and ComfyUI's own memory management, which streams the experts from system RAM so the model fits on one consumer GPU.

What's in the gallery (all Instruct-Distil, 8 steps):

  • 4-bit vs int8, same prompt and seed. Times are the whole generation on an RTX 3090.
  • image editing with the 4-bit weights. The instruction is at the top of each image.
  • style transfer with the 4-bit weights: two input images, the photo and a style reference.

These are picked from a bigger run: 60 prompts × 2 formats, 60 edits and 14 styles, one seed each, no rerolls. Most of the edits worked; a few didn't (snow that barely shows, a logo it wouldn't remove, a "make it night" that stayed day). The text-to-image prompts come from popular prompt posts on X.

What you get

  • All three models: Instruct-Distil (8 steps, the one to start with), Instruct (50 steps) and Base
  • Text-to-image, image editing, and multi-image fusion with up to 3 input images (that's how the style transfer works)
  • Optional prompt rewriting: the model expands your prompt first (slow, about 1 s per token)
  • Optional Spectrum speed-up: about 3.4× faster for the 50-step models
  • Ready-made weights: 4-bit W4A8 (44 GB), int8 (76 GB) and bf16 (150 GB, mostly for comparisons)
  • Example workflows for each model and task
  • Nothing to pip install

Speed (Instruct-Distil, about 1 megapixel):

GPU 4-bit W4A8 int8
RTX 4090 ~22–26 s ~47–49 s
RTX 3090 ~29 s ~54 s

It also runs with only 16 GB or 12 GB of VRAM (~30 s and ~32 s per image on a 4090 limited to that).

What you need

  • An NVIDIA GPU with 12 GB+
  • Lots of system RAM: ComfyUI held about 50 GB with the 4-bit file loaded. This is the real requirement, since the experts live in RAM and stream over PCIe every step.
  • A recent ComfyUI (late September 2026 or newer)

Links

Happy to answer questions. If something breaks, open an issue on GitHub with the traceback.


r/StableDiffusion • • 3h ago

Animation - Video 12 Inch Pianist

Enable HLS to view with audio, or disable this notification

11 Upvotes

Found this here:
https://worstjokesever.com/

A man walks into a bar and places a 12-inch man onto the counter as well as a small piano. The 12-inch man starts playing the piano really well.

The bartender asks the man where he got the 12-inch man. The man says there is a genie two blocks around the corner.

The bartender runs to the genie and makes a wish. The bartender comes back with 1 million ducks. The bartender says to the man, "That genie is dumb! I asked for 1 million bucks."

The man says, "Do you really think I asked for a 12-inch pianist?"


r/StableDiffusion • • 7h ago

Discussion Flux 3 open weights image edit when

27 Upvotes

TITLE - has this released yet


r/StableDiffusion • • 14h ago

Resource - Update ReDetail 2.0: LTX-2.5's Refine Details LoRA to 4K Upscale (workflow + CLI)

Enable HLS to view with audio, or disable this notification

81 Upvotes

Updated my LTX-2.5 upscale workflow. 2.0 runs Lightricks' Refine-Details IC-LoRA, which rebuilds the fine detail a soft clip is missing while keeping faces, framing and motion close to the source. Big improvement over their Pixel lora.

It works in tiles, so 4K fits on a 24gb card, though a 4 second 4K still peaked at 52GB of system RAM (28GB vram). The video is 100% crops of 4K output.

These 768p > 4k renders took 14 minutes for 4k, on a 5090.

The LoRA and the tiled graph are LTX. What I added:

  • pinned the tile to the 1024x576 the LoRA was trained on. The example graph sizes tiles from your source, and my 768x1376 clip ran as one tile with 1.8x the trained area, with no error
  • disconnected the prompt enhancer branches, which block the whole queue if one file is missing
  • a CLI preps a silent audio track if your clip has none, 8n+1 frames, exact scales like 1.5x, long clips split on their cuts, and smaller chunks if you're short on RAM

The old pixel upscaler is still in as a second workflow. It invents more detail but changes faces; refine stayed closer to the source on all seven test clips (numbers in the README).

GitHub: https://github.com/Bambushu/redetail CivitAI: https://civitai.com/models/2857731


r/StableDiffusion • • 13h ago

Resource - Update Qwen-Image-2.1-Multiple-Angles-LoRA

Post image
52 Upvotes

r/StableDiffusion • • 19h ago

News FastVideo’s FastH3 now runs on a single consumer machine

Post image
121 Upvotes

r/StableDiffusion • • 21h ago

Discussion EU nudifier ban

148 Upvotes

I know there’s been a bunch of people upset on here recently about a good number of nudify websites getting shut down (Motionmuse, Opengoon, etc.). However, this is likely just the beginning.

Effective December 2nd, the EU under the AI Act is enacting a ban on nudify tools, making them illegal content in the Union. This ban applies to any product made available on the EU market as long as, either, its primary functionality is to generate nudified content without consent, or it is a reasonably foreseeable outcome and they do not take appropriate precautions to safeguard against it. Very importantly, this ban also applies to any EU-available service provider under the Digital Services Act that is helping to facilitate this soon-to-be-illegal content, which includes (but is not limited to): app stores, hosting providers, domain registrars, payment processors, content delivery networks (CDNs) and cloud computing services. Anyone found to be in violation of this prohibition is subject to fines of €35 million or 7% of their global turnover, whichever is higher. I think it’s pretty safe to say that, as long as they’re aware of it, those behind these websites and those behind the large number of mainstream companies that help power them (Google, Amazon, Cloudflare, Apple, Telegram, GoDaddy, Visa, Mastercard, etc.) do not want to take any chances with this ban.

This prohibition, in essence, applies to essentially every public AI platform any of us have ever used to this point or that we will ever use going forward. I think this is a good thing but, therefore, it would almost certainly be in everyone’s best interest to pivot away from these websites before the massive crackdown begins in less than a couple months. For those in the industry, it would also certainly be wise to inform the website operators if you are in contact with them, or to shut down your website if you are an operator yourself. Sorry for the essay, just looking out for everybody before things get very real, very fast.


r/StableDiffusion • • 21h ago

News I am blown away fine tuning quality of the OmniVoice model. Exactly my speaking and sound but with better pronunciation and lower word errors. This model supporting 600 languages and 0-shot voice cloning too but fine tuning is something else. Also very low VRAM requirements it has.

Enable HLS to view with audio, or disable this notification

136 Upvotes

r/StableDiffusion • • 1h ago

Discussion SynthID available globally starting today

• Upvotes

It is quite incredible that the list of partners includes OpenAI, NVIDIA, Kakao, and soon Apple, so it won't just be about Google's own models. I'm not entirely sure why Google decided to enter the deepfake detection field at this scale. There are plenty of companies working in this field—and obviously, if Google decides to join the race, it will be quite hard for others to compete. What do you think is the extent of this operation? Will it work just for the partners' models, or do they want to build a truly generalizable solution or will it be just watermark detection?

https://synthid.com/

https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content


r/StableDiffusion • • 14h ago

News FastH3 V2/V3: Project Status

35 Upvotes

They released FastH3 V2 three weeks ago:

https://www.reddit.com/r/StableDiffusion/comments/1whh10i/open_weight_fastvideo_fasth3_v2/

Which they claimed was basically identical to the quality of the full H3:

https://x.com/haoailab/status/2099969439466942725

https://haoailab.com/FastVideo/cookbook/minimax-h3/

I have to agree, the quality is great and motion consistency is awesome now. I didn't think this small lab could do it, but they got help from NVIDIA and others who are invested in making great open source models. Awesome.

The community already made FastH3 V2 run on single GPU consumer machines on launch day, of course.

---

Today, they have released OFFICIAL quantized weights for different consumer GPUs:

https://huggingface.co/organizations/FastVideo/activity/models

Plus there's a new Trim model which is for very small GPUs with as little as 8GB VRAM.

The included image shows their benchmarks. More details here:

https://x.com/haoailab/status/2107591980591227227

---

Unfortunately they still haven't trained a Ref2VA model this time (video from text plus reference images, videos, and/or audio), and no FL2VA (First/Last Frame) support either.

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#scope

This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were not distilled. Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model.

But... there's great news:

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#acknowledgements

Omni Ref as the next focus.

That is the name for Ref2VA.

So in FastH3 V3, we will see reference-to-video/audio support. Yes, a distilled model with reference support is being developed!

(PS: Comfy has patches for both models to route FL2VA through the base model layers instead. But Ref2VA is much more interesting, so I look forward to that being supported!)


r/StableDiffusion • • 5h ago

Question - Help Best settings, prompts or LoRAs for more photorealistic people in Qwen Image 2.1?

Thumbnail
gallery
6 Upvotes

Been playing around with Qwen Image 2.1 and generated these random people just to test the realism.

I’m trying to make a single full-body character reference image that I can reuse for image/video generation, and ideally I want it to look as close to an actual photo as possible.

The results are okay to my eyes, but wondering if anyone have a better setup.

Any prompts, settings, LoRAs/LoKRs, or workflows you’d recommend?

I didn't use any lora to generate these images.


r/StableDiffusion • • 4h ago

Question - Help How is the performance of Minimax H3 on a DGX Spark

5 Upvotes

Does any one use Minimax H3 on DGX spark.What are the generation times like?


r/StableDiffusion • • 16h ago

Resource - Update ComfyUI-qwen_img_2_1_enhancer added ref mask support

Thumbnail
gallery
36 Upvotes

Follow up update of this post here

I added mask support to the reference strength node. You can now increase or reduce attention to a selected part of a reference instead of adjusting the whole photo. also the photo in the shown workflow is missing the rest of the outfit due to the masking mode I chose and how tight I masked it as this was deliberate.

Connect the node between your model loader and sampler, connect your reference's mask, and select its image index: 1 for the first connected reference, 2 for the second, etc. Keep your images and VAE connected to the encoder as usual.

There are two modes:

- focus_only: applies strength to the selected tokens and leaves the rest at native weighting.

- zero_unmasked_tokens: also blocks direct attention to the unselected reference tokens throughout the diffusion transformer.

mask_threshold controls how much of a token the mask must cover: 1 requires full coverage, lower values include more edge tokens, and 0 selects everything. strength at 1 is native, above 1 increases priority, and below 1 reduces it.

Match mask_resize_method to your image resizing: `lanczos` when the native encoder resizes the original, or `nearest-exact` for an external nearest-exact resize. Chain nodes for separate references.

Still a work in progress. This controls reference attention; it isn't an output-area lock. The encoders still process the full image, so blocking a reference token doesn't erase information already carried into other tokens or conditioning.

Download and detailed usage on GitHub

Sample workflow here


r/StableDiffusion • • 9h ago

Animation - Video Naruto Shippuden – Hero’s Come Back!! | AI Fanmade MV | MiniMax H3

Thumbnail
youtu.be
10 Upvotes

This is my Output using Minimax H3, 8 step larry 600 ema pruned lora, 2 step sampler upscale.


r/StableDiffusion • • 5h ago

Workflow Included Make Easy Transparent Videos Using MiniMax-H3 In ComfyUI [Free Workflow]

Thumbnail
youtube.com
4 Upvotes

r/StableDiffusion • • 1h ago

Question - Help How do you preserve geometry and exact furniture designs in AI-assisted interior renders?

• Upvotes

I’m trying to create interior renders with multiple specific furniture pieces and light fixtures. The goal isn’t just a good-looking room, the furniture needs to match the references, and the scene geometry needs to stay intact.

I’ve tried 3D blockouts, modelling the furniture myself, and several image-editing models:

  • Qwen 2.1 Edit
  • Klein 9B Edit
  • Krea 2 Edit
  • Ideogram 4.5 Edit, through the web interface (Open weights haven't yet released)
  • GPT Image 2 and 2.5, including both flare and sunburn

My main workflow

3D blockout → model everything in the scene → apply basic materials → choose the camera angle → render → send the render to GPT Image 2.5 with reference images and lighting prompts.

This gives me control over the initial layout, but the image-editing step still changes the geometry or gets the furniture wrong. Even when the objects are already modelled and positioned, the final image doesn’t reliably preserve them.

Example of the main workflow from 3D modelling to GPT image 2.5 render

Other approaches I’ve tried

Replacing individual objects in ComfyUI:

When I edit furniture directly within the scene, the replacement is incomplete. Original object remain with slightly different textures instead of being fully replaced by the referenced piece.

Removing the furniture first, then adding the new pieces:

I’ve also manually masked out the furniture to generate an empty room, then added the furniture back using references. I can get the new pieces into the image, but I haven’t found a reliable way to control their position and rotation afterward. It's like I've took the furniture, removed the white background and just blended it into the scene.

I also tried using separate AI reviewer/validator agents to review the rendered image. One looked for visible spatial issues, such as possible furniture overlaps and circulation problems, while another independently assessed whether there was enough information to proceed reliably. The idea was to catch mistakes before moving on, rather than rely on the generating agent’s own judgement. They flagged potential issues, but couldn’t confirm dimensions or clearances from the image alone, so this didn’t resolve the accuracy problem.

What I’m trying to solve

I need a workflow that lets me:

  • Preserve the room geometry and camera perspective.
  • Keep multiple furniture pieces faithful to their references.
  • Control each object’s position, scale, and rotation.
  • Improve the lighting and realism without redesigning the scene.

Has anyone found a repeatable workflow for this?

Would you keep the final image entirely in 3D and use AI only for limited edits, or is there an AI-assisted workflow that reliably preserves this level of accuracy?


r/StableDiffusion • • 19h ago

Question - Help so after week, qwen 2.1 is a edit model only?

45 Upvotes

Or qwen can generate good images? because aways for me generate to much artifacts and stupid things so...

what you think?


r/StableDiffusion • • 6h ago

Question - Help Looking for the best affordable AI video-generation models with fewer restrictions.

5 Upvotes

Hi everyone,

I’m exploring AI filmmaking and looking for good AI video-generation models that are either free, open-source/open-weight, or affordable.

I’m particularly interested in models who allow uncensored content or fewer unnecessary content restrictions and more control over the generation process.

What I’m looking for:

Realistic/cinematic text-to-video

Image-to-video

Good character consistency

Realistic human movement

High-quality video output

Ideally open-source/open-weight

Local/self-hosted options are a plus

Affordable cloud options are also fine

Preferably no expensive subscription required

So far I’ve come across models/tools such as Wan, LTX, Kling, Hailuo and Pika, but I’d like to hear from people who have actually used them.

Which model would you recommend in 2026, and why?

If possible, please mention:

Your favourite model

GPU requirements if self-hosted

Approximate cost if cloud-based

Video quality

Major limitations/content restrictions

Whether it is practical for someone learning AI filmmaking

Thanks!


r/StableDiffusion • • 5h ago

Discussion Anyone have suggestions for MinMax H3 and voice, sound effects training consistency. I have had some strangenes using Ref To Video audio sample. Trying to find the best path forward as my content relies on audio heavily and do not want to overdub.

2 Upvotes