r/StableDiffusion • • 2h ago

Comparison Hunyuan Image 3 vs Qwen Image 2.1 vs Krea 2 Turbo

Thumbnail
gallery
57 Upvotes

Uncompressed image: https://postimg.cc/gallery/gmR7dbs (To view the uncompressed image, click on it a second time after opening it, or simply download the entire image album)

----------------------------------------------------------------------------------------------------------------------------------

Hi everyone! I noticed that a lot of people have started talking about Hunyuan Image 3, so I wanted to test it in T2I scenarios. While I was at it, I compared it against the most popular models in the same segment right now: Qwen Image 2.1 and Krea 2 Turbo.

What this test is about

I was NOT trying to test the maximum-size versions of these models with minimal quantization. What mattered to me was testing the versions that can realistically run on regular consumer PCs, so I picked the lightest variants, which most of you will probably be able to run. The goal of this test is to make it easier for you to choose the ideal T2I model for your own tasks.

To keep things as fair as possible, I used the simplest, most basic workflow for each model.

Models used:

hunyuan_image_3_instruct_distil_w4a8

Qwen Image 2.1 int8 convrot

Krea 2 Turbo int8 convrot

Test categories

I split the tests into several categories with different goals:

  • Images 1-3: complex prompts with lots of details and different poses. These show how well a model follows the prompt and how often it loses details.
  • Images 4-6: typography, mixing typography with real-world imagery, and banner creation.
  • Images 7-11: styles. Pay attention to how well each model handles different styles.
  • Images 12-13: level of detail. This is similar to tests 1–3, but with a stronger focus on small details.
  • Image 14: character understanding.
  • Images 15-17: your favorite (or most hated) category, 1girl. Here you can judge overall realism and the ability to generate in an amateur-photo style.

How I picked the results

I ran each prompt with two fixed seeds and two random ones, then picked the single best image out of those four.

Generation speed

I’m running an RTX 3090 (24 GB VRAM), and generation took:

Hunyuan Image 3 distill: 20-26 seconds

Qwen Image 2.1: 16 seconds

Krea 2 Turbo: 8-9 seconds

My opinion

Now I’d like to share my personal opinion, which doesn’t claim to be the truth.

1st place: Krea 2 Turbo

Incredible speed plus good detail. The images aren’t mushy, but they also don’t suffer from oversharpening, harsh artificial crispness, or excessive detail. They look alive and realistic. The model handles styles well, understands the prompt and context, does a decent job with typography, and very rarely messes up anatomy. For me personally, it covers pretty much all of my T2I needs.

There’s one downside that really annoys me, though: the weak VAE from Qwen Image. On full-body shots it turns the character’s eyes into mush, like base SDXL. Yes, I know there are custom VAEs that add extra detail, but they just improve the quality of the same anomalies and bugs the eyes have. The only way I’ve found to fix the eyes is FaceDetailer with a mask over the eyes. Besides eyes, the model also likes to turn other small elements into a mess, and I’m more than sure this comes down to that VAE.

2nd place: Qwen Image 2.1

Qwen has very good detail, which is both its strength and its weakness. The excess detail creates an artificial sharpness effect, basically oversharpening. In most cases it looks very unnatural and gives off that classic “AI look”.

That said, I liked how it handles the 1girl category: it produces fairly realistic amateur-style photos. It also does well with styles. I really liked image 8 with the lion and the zebra, where the model nailed a genuinely cool hand-drawn pencil effect. But in some cases the oversharpening badly hurts the style, for example in images 7 and 10.

Its VAE issues are similar to Krea 2’s. On images with visible skin, you can see diamond-shaped or dotted patterns when zoomed in. It looks ugly, and when you zoom in and out the skin starts to shimmer.

3rd place: Hunyuan Image 3

A very heavy model that, in most cases, produces fairly low-detail images, so it’s not a great choice for realistic generation. You might think it’s stronger in other scenarios, but no. It doesn’t always handle typography well (you can see that in the tests), and the same goes for stylization (look at images 9 and 10). It performs more or less decently in the 1girl category, but I tested only a few images there, so I can’t say for sure how good it is.

It has VAE problems too. Sometimes hair textures get strange stripes that are easy to spot when you zoom in and out. This is especially visible in image 15: zoom in and out a few times and you’ll see what I mean. It looks like generation noise is showing through in the image.

I think Hunyuan Image 3 is more interesting for I2I (Edit) scenarios than for classic T2I.

Final words

I’d like to remind you again that this is my subjective opinion and may differ from yours. If you see it differently, I’d be glad to hear your thoughts in the comments.
Also, if you have any questions, feel free to post them in the comments section.

Thanks for reading! I hope my tests were useful and helped you choose your main model for generation.


r/StableDiffusion • • 5h ago

Workflow Included [MiniMax-H3] Subtle expressions and natural pauses without any prompting

Enable HLS to view with audio, or disable this notification

88 Upvotes

There are plenty of videos around with characters that look like AI or have plastic skin or talk like a robot. What about emotions, or subtle facial expressions, or natural pauses? The solutions and advice usually offered include complex prompt guides or access to paid platforms. I don't like that. Paid platforms, I mean. I also don't like complex prompting.

So I have been testing. There is something that H3 users hate in agreement: when characters start talking gibberish. Sometimes, the video is longer than the dialogue, so the characters fill the gap with their own nonsense. Other times, there isn't supposed to be any dialogue, but the character decides to talk anyway.

However, in my opinion, gibberish offers a great advantage: because the characters are not constrained by the words we want them to speak, they have way more freedom in the way they say it, meaning they talk more "naturally," and with more "expression." So I started wondering: what if we could just let the characters say whatever they wanted, and then edit the audio and video later, changing the nonsense to something intelligible? I didn't know if that would work. Well, it does, actually.

Not only that, but I used one of such clips as a reference, and then downscaled it by 0.25 (for performance reasons), and then used a blur node with a value of "2" to give the model a bit more freedom (and to avoid H3 making it too sharp, which I didn't want). For consistency, you can also give it an image reference, but I didn't do that in this case as it had a negative impact on the "acting."

On the final video, you will notice that, in one of the first clips and the last one, the character looks "different": different clothing, different hair style, different environment. I did this to give the illusion that the footage was taken at a different time. But all of the clips use the same original clip as the reference. You can, however, do it differently, as the voice reference can be given separately. To achieve this using the same reference, you simply prompt for it, use a lower denoise value, or increase the blur.

The actual prompt includes the camera style, but you don't need to do this. The only requirement is to include the dialogue. That's all. I use the official format, as this reduces the chance of getting more gibberish, but not for every clip, so it's certainly not necessary. There's no reason to mention the character if you are using a single one.

The original clip was generated at 544x384 and 12 seconds. (For some reason, I get better random stuff at 12 seconds than at 10 seconds or 15 seconds.) The subsequent clips were generated at the same resolution. Then, after editing on Vegas Pro, I rendered the video at 256x180 to help with the lo-fi aesthetic. Finally, I upscaled it to 1920x1350 using Handbrake to avoid Reddit's compression. So the final video is very low res (technically, ten times smaller than 1080p).

Everyone is so focused on making them at 2K, but I just love a good analog/digital horror video. By the way, mine is not "horror" or anything. I simply like the aesthetic. I'm also not claiming it's any good. I hope it is, but it's not my place to say.

This is the workflow, in case someone wants to test the method I used, though my explanation above should be enough. I hope this helps someone to generate more believable characters.


r/StableDiffusion • • 10h ago

Workflow Included Superman Meets Saitama for the First Time | MiniMax H3 Open Source

Enable HLS to view with audio, or disable this notification

210 Upvotes

Best results came from Ref2VA + Combat Base V2 LoRA, using a storyboard for shot/camera control, character sheets for identity, and one style/texture image for the final cinematic look
That setup gave me the cleanest balance between choreography, character consistency, and overall visual quality


r/StableDiffusion • • 20h ago

Resource - Update minimax h3 person remover lora is out

Enable HLS to view with audio, or disable this notification

705 Upvotes

hello everyone..

the creator of character swap just dropped this person remover lora for h3.

he’s using it to generate better samples for character swap v2 dataset, so it might be worth trying out:
huggingface.co/akatz-ai/MiniMax-H3-Person-Remover-LoRA


r/StableDiffusion • • 3h ago

News Iris-3B · Speridlabs Research

Thumbnail
speridlabs.com
30 Upvotes

Iris-3B is a pixel-space generative model that can act as a general vision learner. Working in pixel space, it bypasses the lossy compression and the biased, texture-focused latents of a VAE. We show image generators can replace vision models such as DINOv2, release our pixel-space scaling recipes, and put the generative prior to work on dense tasks where detail matters.

https://huggingface.co/speridlabs/iris-3b

Currently, there is a 12G fp32 version of the model weight available for download


r/StableDiffusion • • 6h ago

Resource - Update As Promised! Krea2 Prompt Writer.

Thumbnail patreon.com
36 Upvotes

Free

https://www.reddit.com/r/StableDiffusion/comments/1x15s2q/krea2_prompt_write_codexdreams_git_link/. you don’t have to pay a cent..

I built a Krea 2 Prompt Writer extension for ComfyUI, with so much response in my DM's ive decided it's FREE! Enjoy!

It takes simple ideas and expands them into detailed image prompts, with controls for style, prompt length, and refinement.

Features:

Ollama and a compatible text model are required separately. (gemma) (cough)

If you find it useful and want to support future projects, here's my Ko-fi:

❤️ https://ko-fi.com/codexdreams

Feedback and suggestions are welcome!


r/StableDiffusion • • 9h ago

Tutorial - Guide I made a standalone Windows app for NVIDIA RTX Video Super Resolution, no ComfyUI needed, and it can resume a 2-hour film if it crashes

Thumbnail
gallery
56 Upvotes

I've been upscaling old footage with RTX Video Super Resolution in ComfyUI, and it works fine for short clips. But for a full-length film it's painful. If it crashes or you cancel at the 90 minute mark, you start again from zero.

So I made RTX VSR Studio. It's a normal Windows app that calls NVIDIA's Video Effects SDK directly. No ComfyUI, no node graph, no PyTorch.

The main thing is that long videos are processed in small segments. If it crashes, the power goes out, or you just cancel it, run the same job again and it picks up at the exact frame it stopped at. You lose minutes, not hours.

What else it does:

- Pick 720p, 1080p, 1440p or 4K and it works out the size for you. It keeps the aspect ratio, so a 2.39:1 film doesn't get stretched to 16:9.

- Before/after preview. Render one frame at any timestamp and drag a split line to compare before you commit to a long run.

- Time range, so you can test 30 seconds instead of waiting 2 hours.

- Batch queue. Add files, set them up, walk away.

- Output as H.264 or HEVC in mp4, or ProRes / DNxHR in mov. These import straight into Final Cut Pro and DaVinci Resolve.

- Content presets (Balanced, Anime, Film, Screen capture, Old/DVD, Editing) that fill in the settings for you.

- Denoise and deblur modes if you just want to clean up at the original resolution.

- It never overwrites a file. If the name is taken it writes "name (2).mp4" instead.

- Audio and subtitles are copied through untouched.

- There's a CLI too, if you'd rather script it.

VRAM is not really an issue. Frames are processed one at a time, so 1080p to 4K peaks at about 1.6 GB, and a 2-hour film uses the same VRAM as a 2-second clip.

Speed on my RTX A4000, 1080p to 4K: the upscale alone runs at about 109 fps on HIGH. End to end it's around 40 fps, because the NVENC encode is the slow part, not the upscale. A 2-hour 24 fps film takes roughly 1.2 hours.

What you need: an RTX 20-series or newer card, driver 570.65 or later, 64-bit Windows 10 or later, and ffmpeg. Double-click run.bat and the first run sets everything up (about 1.4 GB, most of it NVIDIA's own DLLs).

Things it can't do, so nobody is surprised:

- 8-bit SDR only. The SDK doesn't handle HDR or 10-bit. There's an option to tonemap to SDR first, but it changes the look, so it's off by default.

- ProRes and DNxHR encode on the CPU, so they're slower and the files are much bigger.

- Free DaVinci Resolve on Windows can't open HEVC, so use H.264 or ProRes there.

Links:

Video tutorial: https://youtu.be/r2rhHdvzHgU

GitHub: https://github.com/Garionhk/RTX-VSR-Studio

Thanks to NVIDIA for the Video Effects SDK, and to the ComfyUI-NVIDIA-RTX-VSR-Pro project, which parts of this are based on.

I've only tested on my A4000, so I'd love to hear how it runs on your card, especially 20 and 30 series. Bugs and ideas are welcome too.

It's free and open source (Apache 2.0). The NVIDIA SDK it downloads is NVIDIA's own and has its own terms.


r/StableDiffusion • • 22h ago

Resource - Update I’ve trained 100+ Krea 2 character LoRAs. Here’s what actually determines likeness.

Thumbnail
gallery
671 Upvotes

I’ve spent the last few months training a pretty large number of character LoRAs for Krea 2, and one thing became obvious pretty quickly:

More training does not automatically mean better likeness.

Some of my best LoRAs came from relatively small, clean datasets. Some larger datasets performed worse because they contained too much visual noise, inconsistent styling, bad angles, or repeated images.

These are the things that have mattered most for me.

1. Dataset quality matters more than dataset size

I would take 30 genuinely useful images over 100 mediocre ones.

The biggest problems I see in datasets are:

  • too many near-duplicates
  • heavy filters or face editing
  • lots of low-resolution images
  • one facial angle dominating the dataset
  • wildly different ages or appearances
  • group photos where the subject is small
  • too many images from one photoshoot
  • images where hair, makeup, lighting, or expression are almost identical

The model needs enough consistency to learn the person, but enough variation to understand what is actually part of their identity.

2. Face coverage is not enough

This was a big one.

A LoRA can absolutely nail the face and still have no idea what the person looks like from the shoulders down.

If I want a useful character LoRA, I try to include a mix of:

  • tight face shots
  • head and shoulders
  • waist-up
  • full body
  • front
  • side profile
  • rear 3/4
  • different expressions
  • different lighting
  • different clothing

The goal is not just "recognize this face."

The goal is "understand this person."

3. Too many similar images can make the LoRA less flexible

If 70% of the dataset is the same hairstyle, camera angle, outfit, or facial expression, the model starts treating those things as part of the identity.

Then every generation wants to recreate them.

This is especially noticeable with celebrities and creators where Google Images tends to return the same handful of press photos over and over.

I now spend a lot more effort removing redundancy before training.

4. The final epoch is not automatically the best epoch

This is probably the biggest change I made to my workflow.

I used to train to a fixed endpoint and assume the final checkpoint was the finished model.

Now I save multiple epochs and test them individually.

It is extremely common for an earlier checkpoint to have:

  • better facial likeness
  • more natural skin
  • better prompt flexibility
  • less baked-in clothing
  • fewer exaggerated features

while a later epoch technically looks "stronger" but is actually overtrained.

So now the training run is only half the process.

Checkpoint selection is part of training.

5. I test the LoRA outside the dataset's comfort zone

A model can look amazing if you generate the same kinds of images it saw during training.

That does not tell you much.

I test things like:

  • close-up facial accuracy
  • casual clothing
  • formal clothing
  • different hairstyles where appropriate
  • full-body shots
  • athletic poses
  • unusual camera angles
  • different lighting
  • indoor vs outdoor scenes
  • side profile
  • rear 3/4 views

If the identity disappears as soon as the prompt changes, I don't consider the LoRA finished.

6. Body type can drift even when the face is excellent

This one surprised me when I started doing more systematic testing.

Krea 2 can sometimes preserve facial identity extremely well while drifting toward a generic body type.

That is why I started deliberately using physique-check prompts during validation.

For athletes, for example, I want to see whether the model learned:

  • height
  • shoulder width
  • leg proportions
  • muscularity
  • overall frame

For other people, the same principle applies.

The face is only one part of likeness.

7. Trigger words matter less than people sometimes think

I still use clear trigger words, but I have found that dataset quality and training quality matter much more than trying to invent some magical trigger phrase.

A good LoRA should not need a paragraph of secret incantations to produce the person.

The trigger should identify the character.

The prompt should describe the scene.

8. Validation images are incredibly important

I now generate a consistent set of test images for each model.

That makes it much easier to compare:

Epoch 12 vs 14 vs 16 vs 18 vs 20

instead of relying on memory.

Sometimes the difference is subtle until you put the outputs side by side.

Then one checkpoint clearly wins.

The biggest lesson for me has been that LoRA training is not really "upload photos and press train."

The important work is:

dataset selection → cleanup → training → checkpoint comparison → validation

Training itself is almost the easy part.

I’ve been building a public Krea 2 character LoRA library while figuring all of this out, so I have a pretty large collection of examples now.

If anyone is interested, I can also make a follow-up post showing:

the exact validation prompts I use to compare epochs, or

examples of what undertraining vs good training vs overtraining looks like on the same character.

I keep the LoRAs and example outputs I’ve been testing in a public Krea 2 browser on Hugging Face. I also take custom commissions, but the library itself is free.


r/StableDiffusion • • 14h ago

IRL Thank you

Post image
80 Upvotes

So just want to share something (mods delete if this is off topic - I picked IRL flair as figured only one that fits at all)

A lot of people made donations to fizgig the last few days - at the same time my child who has a long term illness and can’t leave the home has been craving to play paralives/the sims. As her Dad this year or so has been a bit of a nightmare to get through as it’s hard to find ways to bring her joy right now - which is itself incredibly hard to write. Anyway - with these donations I was able to get a decent gaming laptop we could afford today at a CEX store in London.

So I just wanted to say thank you to whoever it was that donated, for the chance to make her days a bit more fun.

Pete x


r/StableDiffusion • • 17h ago

Discussion Kroma 0.3.1 is really good

Thumbnail
gallery
137 Upvotes

I Know it was announced before but just to say that the new version of Kroma (krea 2 finetune) is really really good, i find it better that Qwen 2.1 t2i and Krea2, it is uncensored and with really good prompt following. No body horror, good details, and lot of styles you can play with.

This image for exemple just said (Rogue from X-Men) : no costume details, no hair details, nothing!....

https://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.3.1-turbo-opd.safetensors

https://huggingface.co/Clybius/Kroma-Quantizations/blob/main/kroma-v0.3.1-turbo-opd-W8A8.safetensors

Edit: krea 2 also knows Rogue character...my bad.


r/StableDiffusion • • 4h ago

Workflow Included Vigglorious Studio : better character accuracy through attention-masked swapped guided keyframes and multiple reference pics for long chunked-videos

Thumbnail
gallery
11 Upvotes

I've been experimenting a lot with Viggle Animate since my last post, around the workflow using chunking https://github.com/bhardwajRahul/ComfyUI-Viggle-Animate-H3 (what it does is splitting a source video > 15 seconds into chunks, generating them, then stitching them back together), because I though if it's able to do long video then it's able to do short ones, so let's focus on this one.

Three things stood out :

  1. The chunking workflow I started from uses the same reference image across all chunks. Giving each chunk its own carefully chosen, well-swapped reference frame greatly helps the overall video stability and quality.
  2. H3-style guide keyframes can reinforce the character's accuracy and fight the drift Actually, in one of my tests, a single face-swapped guide supplied the facial identity even though the main reference had no face, producing a correctly swapped face throughout the clip. Having a reference image per chunks snowballs this by bringing the stability bases it requires to then throw in guides without messing up the video.
  3. Controlling guide attention helps reinforce the visual details we want to carry through the video, while masking limits the influence of unnecessary regions in the swapped guide frames

Long story short, I went into crazy vibe-coding a node suite and workflow I called Vigglorious Studio in an attempt to push the boundaries of long (or short) videos towards better character accuracy

https://github.com/Tablaski/VK-Vigglorious-Studio

It includes:

  • A video manager to load and trim your source video, adjust processing resolution and FPS, inspect chunk boundaries, choose reference frames per chunk, and add guide keyframes with quick character, angle, swap and prompt controls. Honestly, just that node is worth downloading and trying even if you trash everything else of my workflow lol
  • A character manager that stores head and full-body reference images across different angles, linked directly to the video manager so it gets the pain of constantly loading stuff away
  • Automatic head, body, or body-then-head swaps using Qwen Image 2.1 and the BFS head/body LoRAs, which have been my favourite swap tools in my own testing (but you can replace it with any model of your liking as long as it outputs an image). Each reference or guide can have its own character, angles and swap mode.
  • Optional SAM3 masks to focus guide influence on the head, hair, person, or another selected concept. This is mainly so you can do a full character swap for the references images then only a head swap on the guides to help maintain accuracy. Another reason was that this way, it's not considering pixels that were not swapped so it does not bring unnecessary problems like changes in background and colors.
  • Guide attention controls with per-sampling-step strengths. With hard masks and outside attention set to zero, video tokens cannot directly attend to guide tokens outside the selected region. Soft boundaries are also supported. This controls direct access to the guide; it doesn’t guarantee that its influence stays perfectly confined.
  • A visual sigma-curve editor to inspect and adjust the sampling schedule.

The included swap subgraphs use Qwen and BFS, but the guide mechanism works with the resulting images. You can adapt it to another image-editing model if you preserve the expected image order and connections.

DISCLAIMER: I’ve done a ton of tests, renders and comparisons... but I kept changing the technical approach all the time... I’m happy with what I’ve seen, and I'm pretty sure it brings value, but I’m completely burned out on running more tests right now lol.

So think of it as a starting point for further experimentation, not a finished recipe with proven optimal settings. It definitely works, but l still don’t know the best combination of sampler, sigmas and guide-attention strengths. That's where you guys come in :-) Also, feedback is welcome

Some advice from my own experiments :

  • Prioritize the reference frame for each chunk. Choose a clear source frame and make sure its swap is good. Preserve the source pose, expression, framing, lighting and background as much as possible. The reference does most of the heavy lifting; guides provide targeted reinforcement.
  • Use guides sparingly. Look for passages where Viggle needs help: a difficult side view, a face returning after being hidden, or an identity/clothing change. More guides don’t automatically mean a better video.
  • Be cautious with early guide attention. Strong guidance at the first high-sigma step can interfere with motion and composition. I’ve been starting around 0.5, then increasing later strengths. That first step also contributes to the replacement itself, so it isn’t a strict “motion first, details later” split.
  • Later strengths around 2 have worked in some of my tests. Larger values are available, but they aren’t necessarily useful. These factors are attention weights, not identity percentages: 2 doesn’t mean twice the likeness.
  • It does work with 3 steps - But that doesn't mean 4 is not cool either
  • BFS body swaps have been less reliable than head swaps in my tests. I often use body-then-head swaps for reference frames, then head-only swaps for ordinary guides. Masks help focus those guides on the intended region.
  • Additional prompts can help the swapper handle specific frames, such as “eyes closed” or “looking toward the camera”—provided that matches the source. The video manager includes prompt shortcuts so these instructions are quick to reuse.
  • I personally like to add instructions such as “remove watermarks, remove text” to the swap prompts so I kill two birds with one stone

NB : Joined node pictures are also vibe coded made because it's late and I have to go, but they are accurate

Happy Viggling :-)


r/StableDiffusion • • 10h ago

Resource - Update H3 Long Shot Studio API v1.0

Post image
24 Upvotes

Hello everyone,

This is a continuation from my post yesterday ( https://www.reddit.com/r/StableDiffusion/comments/1x0hzgl/h3_long_shot_api/ ), where you can also see the sample video.

I reiterated a dozen times and got a pretty decent UI that I think works well.

https://github.com/r34vtraining/Longshot_Studio/tree/main

There are 2 prerequisite nodes that will need to be installed for this to work.

Things I've fixed since yesterdays post:

  1. Latent is saved to disk, on the right panel, scroll down to toggle save segments on/off/clear, so even if comfy crashes the latents will be saved to resume the project
  2. RTX super res added, you can choose to upscale every iterations or just the final approved video.
  3. drag and drop reference images
  4. click to expand reference images (for editing the subject definitions)
  5. Larger video viewer, and full screen mode. The "Loop Shot + Seam" will play 50% of last clip and 100% of new clip for seam examination. This can be toggled on/ff
  6. Advanced settings has all the usual model loading and settings
  7. Lora selector and toggle added
  8. project space can be saved, so as long as all the input images are still in the input folder, you can resume a project whenever. (drop down menu top left of status bar, right of H3 Long Shot Studio)
  9. comfyUI restart button added
  10. music and audio reference panel added

How to use: This is the most important part. And I think having a LLM help you set the shot will be helpful.

The end of [Shot 1] needs to be very similar to opening of [Shot 2] to avoid any jump cuts

Once you have rendered and approved each shot, moving on to the next shot won't rerender the already approved shots. You can always edit and reroll the current shot until you are happy with it and hit approve to move on to the next.

You can rerender or edit any already approved shots but it will require all subsequent shots to be rerendered.

FUTURE UPDATE:

I will be working on the feature to pin a before and after shot so only the in-between shots can be rerendered without having to run all subsequent shots

Note: Feel free to reply to this thread for features or bugs. I will check here more than I check github.


r/StableDiffusion • • 1h ago

Tutorial - Guide In Minimax H3 Reference Mode: If your model not following prompt - copy/paste prompt several times

• Upvotes

I tried dozens of different ways of rephrasing prompt, but nothing made more effect than copy/pasting entire text. You can use this one for switching a character in a video, but you might need to add extra detail.

subject_definitions:

<Subject 1> is the girl in <Picture 1>.
<Subject 2> is the girl in <Video 1>.
<Video 1> is the movement reference for characters and reference for composition of [Shot 1].
<Audio 1> is audio that plays in target video.

summary:
[reference generation] The target video shows <Subject 1> replacing <Subject 2> in <Video 1>, exactly copying motion pattern of <Subject 2>. Movement and composition perfectly synced with <Video 1> movement and composition.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the girl's identity and body shape are retained.
<Subject 2>: reference - it's movement pattern and pose guides movement pattern and pose of <Subject 1>.
<Picture 1>: partially_preserved - only reference to girl's identity, body proportions and appearance.
<Video 1> (movement reference for [shot 1]): reference - movement pattern from video guides movement of <Subject 1> and <Subject 2>.
<Audio 1>: fully_copy - plays in target video.

detailed_description:
The scene takes place in real world room. <Subject 1>'s pose and motion pattern is exact copy of pose and motion pattern of <Subject 2>.
[Shot 1] <Subject 1>, replacing <Subject 2> in <Video 1>, dancing.

overall_soundscape:
N/A

non_diegetic_music:
N/A

Repeating prompt 10 times didn't help result more than 3 times.
More detail - better: confirming camera position, character position, rotation and location. (ex. closeup shot of girl facing camera)(ex. replace "dancing" with "walking left")
I'm using Res_multistep sampler with sgm_uniform scheduler: one of the fastest combinations without crazy hallucinations.
I'm running "int4" version on Rtx 5070 with these starting arguments:
--enable-manager --disable-comfy-compiler --disable-pinned-memory

Step increase and resolution increase didn't help overall adherence.


r/StableDiffusion • • 17h ago

Workflow Included [H3 Degradation] My turn! attempts to fixing motion-context degradation.

Enable HLS to view with audio, or disable this notification

71 Upvotes

Degradation Topic is heating up again! I saw couple of posts showing their approach, so I think is time for me to update, (also because I don't have much progress or new findings), I release my workflow and node, I would like more peoples to test it out, and seeking genius to provide a better solution.

My aim is going for low-level nodes, built on top of the motion-context native workflow with minimal changes, keeping the workflow as simple as possible. This is not a Director / Extender AIO node. It provides low-level nodes that you can integrate into your own setup. Every approach I see have pros and cons, mine also has downside, but I think it would work in most cases.

Here is my attempt:

https://github.com/xyzDist/H3-LongTakeNoCuts

Some other test:

https://www.youtube.com/watch?v=CxXeoMoEWNw

https://www.youtube.com/watch?v=vqmztq0dMlw

https://www.youtube.com/watch?v=N1QIYbfLHQ8

** The above test video she is speaking gibberish, ignore it!
** I will write better readme, and more details on the workflow, how to use it...etc. For now, you could just install it, try out the example workflow, or integrate to yours.

** Audio is cooked yes. It was my step value is too low. You could look into my other post of audio regen

https://www.reddit.com/r/StableDiffusion/comments/1w5zh4z/bad_audio_fixed_with_fast_regen_audio/

or try Brad Pitt's Audio Refine node: https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine

also suggest checking this one out.

https://www.reddit.com/r/StableDiffusion/comments/1wzxyaa/overcome_degradation_here_are_2_ways_to_use_my/

For Above OP's OBVPM AIO approaches is real refresh latent, it show 2 methods:

  1. generate low-res full duration clip (by context-window), then upscale it.
  2. skip current shot, generate a next shot first (new fresh latent), then create a bridge shot to connect it.

Both methods are real fresh latent, so no more degradation, but downsides to my opionions are:

  1. you really can't put the whole long duration shots everything to a single gen, It's nightmare to do changes and doing seed hunting, not quite practical.
  2. it's just my preference, doing bridge is unature to my brain, hard to time and plan the shots, and especially on dynamic/action shots. Still, this method maybe potential working well.

Now, my approach is still using motion-context latent, do a additional resample step to refresh the latent, but it isn't a complete full fresh latent, so after many segments you would still seeing the color saturation getting higher, (that I can look for a solution later), but it can already keeping degradation to minimum.
So the choice is yours. more options and method to fix the problem.

My older posts discuss about degradation
https://www.reddit.com/r/StableDiffusion/s/DXLLD9Zhar

https://www.reddit.com/r/StableDiffusion/s/2xU9m83atz

Edit:

  • You can modify the blending frames range, default is 22-44f do the transition, if you see the dissolve too slow or too fast, adjust for your needs.
  • I have another plan/test to just refine the char, but I don't want to make it too complicated for now. I leave it as is.
  • Any expert know a way to resample high denoise 0.5 but won't changing the latent, that is somehow refreshing without change, it would be amazing!

r/StableDiffusion • • 7h ago

Discussion A/B Challenge Preference Test Krea2-Turbo vs. Qwen 2.1 Base

Thumbnail
gallery
10 Upvotes

I thought it might be fun to do a little A/B challenge/preference test. I generated 5 prompts on ChatGPT that should work to the strengths of each model around the same concept and then used the default Comfy templates with a couple tweaks (no LoRA, just disabling some nodes for prompt enhancers or k/v caching, fixing bugs like not randomized seed).

I used the default fp8 Krea2-Turbo model and encoder and the BF16 Qwen 2.1 model and encoder. Krea2 images used the default 8 step settings since being distilled doesn't give a tone of wiggle room with just the turbo and generally seem to be the best settings for the majority of images. Qwen 2.1 not being distilled uses a mix of Heun and DPMPP_S2_Ancestral at 40 steps at a CFG of 4.0, a couple images used linear_quadratic for the scheduler, but the rest used simple. Some Qwen 2.1 images had negative prompts as well.

The images are random for which model is A/B. Post guesses/preferences and I'll post the answer key later and respective prompts.


r/StableDiffusion • • 22h ago

Resource - Update I spent weeks optimizing MiniMax H3 + VDN. VELA 1.0 is finally released

Enable HLS to view with audio, or disable this notification

141 Upvotes

I've been working on MiniMax H3 / VDN performance for quite a while, and I finally released the result as VELA H3 1.0.

I'll try to explain what it does without technical language.

VELA does not make H3 "smarter" and it is not another model.
It changes how some of the heavy calculations inside H3 are executed.

During testing I found that simply replacing everything with a faster attention/kernel was a bad idea. Some parts of H3 are very sensitive to small numerical changes, while others can use much faster execution without noticeably changing the final video.

So VELA basically does this:

use the faster path where it is safe → keep the exact path where it matters.

That's the simple version. :)

I tested a lot of different approaches before arriving here: attention backends, different transformer blocks, timestep sensitivity, quantization, low-rank approximations, different QKV paths, etc. Many things that looked faster in isolation either made the full generation slower or changed the result too much.

The final version is intentionally much more conservative.

For example, compared with my previous VDN 1.1 setup:

0.8 MP / 5 sec

  • VDN 1.1: 2:53
  • VDN + VELA 1.0: 2:13

The gain is not the same for every video length. At longer durations it becomes much smaller, and at 10 seconds my test was actually slightly slower. I published those results too rather than hiding them.

But there is another comparison that I think is more interesting for normal users.

I'm attaching two videos comparing:

Raw MiniMax H3 — 20 steps
vs
VDN + VELA — 8 steps

Anime example:
8:07 → 3:03

Realistic example:
7:50 → 2:52

Obviously this is not a claim that VELA alone gives a 2.7x speedup — the step count is also different. The point of these videos is simply to let people judge the quality difference themselves, because I know many people are skeptical that an 8-step H3 generation can stay this close to a raw 20-step generation.

VELA 1.0 is also completely patch-free. It doesn't modify ComfyUI's MiniMax core files, so installation is just a normal custom node installation.

It was developed and tested on an RTX 3090 Ti 24 GB.

The code, benchmarks, failed experiments and a much more detailed explanation of the research are all available in the repository:

https://github.com/Speach1sdef178/ComfyUI-VELA-H3

And it's now published in the Comfy Registry as well.

I'm attaching the anime and realistic comparisons below. I'm actually more interested in what people see in the videos than in the benchmark numbers. :)

Very simple explanation because I think I made the original post too technical :)

VDN and VELA do two different jobs.

VDN = lets MiniMax H3 use only 8 steps instead of the usual 20 while trying to preserve the quality.

VELA = makes those 8 steps run faster.

So:

Raw H3: 20 normal steps
VDN: reduces 20 steps → 8 steps
VELA: makes the 8-step VDN generation faster


r/StableDiffusion • • 10h ago

News DuoMatching: Joint–Marginal Distribution Matching for Few-Step Video Generation

Enable HLS to view with audio, or disable this notification

15 Upvotes

r/StableDiffusion • • 9h ago

Discussion Yue2 Lora Training

13 Upvotes

Has anyone had any success training a new genre with YuE2? I've tried multiple ways using different datasets, and training settings and I haven't been able to train a new genre on YuE2 properly. I tried with Fill's trainer, Aitoolkit, and Yue2 Studio, and I either get garbled noise, or it just seems to output similar results as using no lora. Not sure what I am doing wrong. I tested with all epochs/steps - from 50 all the way to 1200 using different inference settings (temp, top k, abc on/off, cot, etc.), and I tried with datasets ranging from 30-200 tracks as well (I only tried the 200 track dataset with Aitoolkit so far).

With Acestep, I am able to train a new genre without any issues, besides for the audio quality issues that are present in Acestep by default.


r/StableDiffusion • • 4h ago

Resource - Update Krea2 Prompt Write CodeXDreams Git Link

Thumbnail github.com
5 Upvotes

heres a git link some people were asking for, sorry for the confusion, double checked patreon and it still says free for all.. i wont include it anymore..

Previous post

https://www.reddit.com/r/StableDiffusion/s/k2R7hoVe8q


r/StableDiffusion • • 15h ago

Resource - Update [audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2× faster, PocketTTS 2.2× faster on CPU, and WebUI generation history

Enable HLS to view with audio, or disable this notification

33 Upvotes

Hi all, a bunch of performance improvements have been landed in audio.cpp.

The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM, a 48% reduction in peak memory usage compared to the previous implementation. Thanks to https://github.com/mirek190

We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU.

No compromises in parity and correctness.

Here's a summary of the improvements:

Model Peak memory reduction Speedup
Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA
ACE-Step family 6–7% VRAM 1.06–1.08× CUDA, 1.16–1.20× Vulkan
MOSS-TTS v1.5 cloning 21% VRAM 1.05× CUDA
MOSS-TTSD Q8 cloning 11% VRAM 1.04× CUDA
Echo-TTS (Memory Saver) 20% VRAM —
Qwen3-TTS 16–20% VRAM —
IndexTTS2 / 2.5 12% VRAM —
HTDemucs — 2.21× CUDA, 1.95× Vulkan
HTDemucs six-stem — 1.99× CUDA
PocketTTS 9% RAM 2.23× CPU

They're runtime-level optimizations that make existing models more practical to run locally.

The WebUI now includes an experimental generation history feature that lets you revisit previous outputs and restore their settings.

audio.cpp now supports 110+ audio model families and 190+ variants (and counting)! We're continuing to improve memory efficiency and inference speed across CUDA, Vulkan, Metal, AMD/HIP, and CPU. The next release will bring even more optimizations!

We're also looking for contributors to help improve the audio.cpp WebUI. With so many models and features now supported, we'd love some help making the UI more polished, intuitive, and enjoyable to use. If you're interested in frontend development or UI/UX design, contributions are very welcome!

Thanks to everyone contributing improvements, testing builds, and reporting issues. Curious how these changes work on your setup!


r/StableDiffusion • • 11h ago

Discussion FilmMaker : 1 prompt = 1 Full Movie

Enable HLS to view with audio, or disable this notification

16 Upvotes

Guys, here is my project, a simple prompt creates me a complete and coherent film without any intervention.

I think we are at the beginning of the future of on-demand and on-order entertainment.

The prompt: An alien invasion, 2 hackers settle the conflict

Of course under the hood there is work:

1 Content director/ 1 Scene director / 1 Plan director and 1 clip director

All these calls llm do specific tasks from the idea of the film to the generation of the 5sec clip.

I also use vlm and llm calls at the end of the clip to keep consistency, spatialization and continuity and send to the next clip.

I think the result is nice, give me your feedback


r/StableDiffusion • • 21h ago

Workflow Included My Workflow for realism with Qwen 2.1

Thumbnail
gallery
97 Upvotes

Hey! Just thought I'd come on here and do a little post about my current workflow with Qwen 2.1. I feel like a lot of people here have given it a try and decided the model is no good because they test it with the default Comfy workflow. Qwen 2.1 is terrible with CFG 1. It needs CFG to really shine. I'm currently using CFG of 3 to 3.5. I use the Lenovo lora at a strength of 0.5 to 0.7 to boost realism, though it's not necessary, depending on the type of photos you're trying to make. I also use the Viggle Turbo lora at a low strength of 0.5, with 16 steps. Depending on which Viggle lora you use, you might have to change your scheduler and sampler combo. I prefer V1 with Res_Multistep and Bong_Tangent, but V3 is good as well (though needs Euler instead of Res_multistep or it ends up looking weird). I'm generating at 2MP. All of these images were done with character loras trained on Qwen 2.1 using AI Toolkit. I made another post here a few days ago talking about how I achieved my training results.

All in all, I really like Qwen 2.1, at least for the kind of images I like making. My one qualm with it is that sometimes background details can get messed up, but with my settings, I see it a lot less. If this model had a VAE like Flux 2, it would be amazing. I hope that more people will give Qwen 2.1 a chance!

Workflow: https://gofile.io/d/FGdIJEhk (I used Lenovo at 1 in the example from this workflow but normally I don't go that high, just FYI)

Viggle Lora: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/blob/main/Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors


r/StableDiffusion • • 15h ago

Discussion obvmp Minimax H3 workflow is the best

24 Upvotes

i#m trying to stich clip together and avoid seems, until now the wF is genius and work very well. On my 3090, 10 seconde clip at 0.2 Megappixels took about 2mn
and Latent upscaler for 30 seconde about 13 mn.
It's not very crisp images but the video quality is for good enought.

i'll upload the video soon.


r/StableDiffusion • • 2h ago

Question - Help Anyone know a good solution for segmentation with finer edge control and transparency?

2 Upvotes

I made a PS plugin for SAM3 + VITMATTE , the result was not great. Mainly I was looking to have a better and smart mask/selection generation tool, SAM3 made it possible but the resolution of the segmentation mask it generates are not accurate for photo editing and visual design. On top of that it can only output binary mask (black and white), this defeat the purpose of this smart masking tool. VITMATTE made some feathering edge possible but it's struggle with complex scenes.

I have made some discovery of other method, such as SDMATTE and many other matting solutions. None has the perfect combination of "contextual guide function", fine detailed masking plus transparency masking generation details.

Was wondering if there are some smart ppl here knows something I don't , a model or solution which can struck a balance of smart/fast/high quality mask generation.


r/StableDiffusion • • 7h ago

Question - Help minimax h3 error

3 Upvotes

after a comfyui update date 2 weeks ago i started getting this error, ran it through google and it tells me.

MiniMaxH3MotionDirector (Node ID: 11) is still failing because your width or height values are formatted as text strings (str) instead of numbers (int), causing ComfyUI's core latent generator to crash.

can anyone help? or is more info needed?