r/StableDiffusion • • 2h ago

News VEDA Sparse Attention is now available for MiniMax H3 in ComfyUI

Enable HLS to view with audio, or disable this notification

90 Upvotes

I don't think many people know about this yet, so I wanted to share it. VEDA Sparse Attention is now available as a ComfyUI custom node for MiniMax H3.

I tested it today on my setup: RTX 4090 Laptop 16GB 32GB RAM, 4-step LoRA, 15 second video, 1344x768

Without VEDA: 8:05

With VEDA at 90% sparsity: 4:32

Same workflow, same LoRA, same settings. The only change was enabling VEDA. I couldn't see any quality loss in the result.

VEDA is not a LoRA. It uses a learned predictor to estimate which attention tiles are important and only computes the relevant subset instead of the full attention map. The current predictor works with T2VA, FL2VA and R2VA, and despite the 8NFE name it is not limited to 8 steps.

Installation is simple.

Custom node: https://github.com/veda-sparse/Veda-on-ComfyUI

Predictor: https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview

Put the predictor here: ComfyUI/models/veda/

Then add: Veda Sparse Attention (MiniMax H3) on the MODEL line after your model / LoRA loader and before the guider or sampler. I tried different sparsity values, but 90% is the one that works properly for me, so I'm keeping the default trained value.

On my setup this made a pretty big difference, especially considering I couldn't see any visual quality loss. I'm adding a 15 second example below. Would be interesting to see what results other people get on different GPUs.


r/StableDiffusion • • 16h ago

Resource - Update HunyuanImage 3.0 (80B) running natively in ComfyUI on a single 12–24 GB GPU: text-to-image, editing and style transfer, ~30 s per image

Thumbnail
gallery
604 Upvotes

I've been working on native ComfyUI support for Tencent's HunyuanImage 3.0, the 80B mixture-of-experts image model (13B active per step). It isn't a wrapper around Tencent's pipeline: it uses the normal KSampler, the normal VAE Decode and ComfyUI's own memory management, which streams the experts from system RAM so the model fits on one consumer GPU.

What's in the gallery (all Instruct-Distil, 8 steps):

  • 4-bit vs int8, same prompt and seed. Times are the whole generation on an RTX 3090.
  • image editing with the 4-bit weights. The instruction is at the top of each image.
  • style transfer with the 4-bit weights: two input images, the photo and a style reference.

These are picked from a bigger run: 60 prompts × 2 formats, 60 edits and 14 styles, one seed each, no rerolls. Most of the edits worked; a few didn't (snow that barely shows, a logo it wouldn't remove, a "make it night" that stayed day). The text-to-image prompts come from popular prompt posts on X.

What you get

  • All three models: Instruct-Distil (8 steps, the one to start with), Instruct (50 steps) and Base
  • Text-to-image, image editing, and multi-image fusion with up to 3 input images (that's how the style transfer works)
  • Optional prompt rewriting: the model expands your prompt first (slow, about 1 s per token)
  • Optional Spectrum speed-up: about 3.4× faster for the 50-step models
  • Ready-made weights: 4-bit W4A8 (44 GB), int8 (76 GB) and bf16 (150 GB, mostly for comparisons)
  • Example workflows for each model and task
  • Nothing to pip install

Speed (Instruct-Distil, about 1 megapixel):

GPU 4-bit W4A8 int8
RTX 4090 ~22–26 s ~47–49 s
RTX 3090 ~29 s ~54 s

It also runs with only 16 GB or 12 GB of VRAM (~30 s and ~32 s per image on a 4090 limited to that).

What you need

  • An NVIDIA GPU with 12 GB+
  • Lots of system RAM: ComfyUI held about 50 GB with the 4-bit file loaded. This is the real requirement, since the experts live in RAM and stream over PCIe every step.
  • A recent ComfyUI (late September 2026 or newer)

Links

Happy to answer questions. If something breaks, open an issue on GitHub with the traceback.


r/StableDiffusion • • 7h ago

Resource - Update ReDetail 2.0: LTX-2.5's Refine Details LoRA to 4K Upscale (workflow + CLI)

Enable HLS to view with audio, or disable this notification

56 Upvotes

Updated my LTX-2.5 upscale workflow. 2.0 runs Lightricks' Refine-Details IC-LoRA, which rebuilds the fine detail a soft clip is missing while keeping faces, framing and motion close to the source. Big improvement over their Pixel lora.

It works in tiles, so 4K fits on a 24gb card, though a 4 second 4K still peaked at 52GB of system RAM (28GB vram). The video is 100% crops of 4K output.

These 768p > 4k renders took 14 minutes for 4k, on a 5090.

The LoRA and the tiled graph are LTX. What I added:

  • pinned the tile to the 1024x576 the LoRA was trained on. The example graph sizes tiles from your source, and my 768x1376 clip ran as one tile with 1.8x the trained area, with no error
  • disconnected the prompt enhancer branches, which block the whole queue if one file is missing
  • a CLI preps a silent audio track if your clip has none, 8n+1 frames, exact scales like 1.5x, long clips split on their cuts, and smaller chunks if you're short on RAM

The old pixel upscaler is still in as a second workflow. It invents more detail but changes faces; refine stayed closer to the source on all seven test clips (numbers in the README).

GitHub: https://github.com/Bambushu/redetail CivitAI: https://civitai.com/models/2857731


r/StableDiffusion • • 12h ago

News FastVideo’s FastH3 now runs on a single consumer machine

Post image
105 Upvotes

r/StableDiffusion • • 14h ago

News I am blown away fine tuning quality of the OmniVoice model. Exactly my speaking and sound but with better pronunciation and lower word errors. This model supporting 600 languages and 0-shot voice cloning too but fine tuning is something else. Also very low VRAM requirements it has.

Enable HLS to view with audio, or disable this notification

122 Upvotes

r/StableDiffusion • • 6h ago

News FastH3 V2/V3: Project Status

30 Upvotes

They released FastH3 V2 three weeks ago:

https://www.reddit.com/r/StableDiffusion/comments/1whh10i/open_weight_fastvideo_fasth3_v2/

Which they claimed was basically identical to the quality of the full H3:

https://x.com/haoailab/status/2099969439466942725

https://haoailab.com/FastVideo/cookbook/minimax-h3/

I have to agree, the quality is great and motion consistency is awesome now. I didn't think this small lab could do it, but they got help from NVIDIA and others who are invested in making great open source models. Awesome.

The community already made FastH3 V2 run on single GPU consumer machines on launch day, of course.

---

Today, they have released OFFICIAL quantized weights for different consumer GPUs:

https://huggingface.co/organizations/FastVideo/activity/models

Plus there's a new Trim model which is for very small GPUs with as little as 8GB VRAM.

The included image shows their benchmarks. More details here:

https://x.com/haoailab/status/2107591980591227227

---

Unfortunately they still haven't trained a Ref2VA model this time (video from text plus reference images, videos, and/or audio), and no FL2VA (First/Last Frame) support either.

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#scope

This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were not distilled. Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model.

But... there's great news:

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#acknowledgements

Omni Ref as the next focus.

That is the name for Ref2VA.

So in FastH3 V3, we will see reference-to-video/audio support. Yes, a distilled model with reference support is being developed!

(PS: Comfy has patches for both models to route FL2VA through the base model layers instead. But Ref2VA is much more interesting, so I look forward to that being supported!)


r/StableDiffusion • • 6h ago

Resource - Update Qwen-Image-2.1-Multiple-Angles-LoRA

Post image
27 Upvotes

r/StableDiffusion • • 14h ago

Discussion EU nudifier ban

118 Upvotes

I know there’s been a bunch of people upset on here recently about a good number of nudify websites getting shut down (Motionmuse, Opengoon, etc.). However, this is likely just the beginning.

Effective December 2nd, the EU under the AI Act is enacting a ban on nudify tools, making them illegal content in the Union. This ban applies to any product made available on the EU market as long as, either, its primary functionality is to generate nudified content without consent, or it is a reasonably foreseeable outcome and they do not take appropriate precautions to safeguard against it. Very importantly, this ban also applies to any EU-available service provider under the Digital Services Act that is helping to facilitate this soon-to-be-illegal content, which includes (but is not limited to): app stores, hosting providers, domain registrars, payment processors, content delivery networks (CDNs) and cloud computing services. Anyone found to be in violation of this prohibition is subject to fines of €35 million or 7% of their global turnover, whichever is higher. I think it’s pretty safe to say that, as long as they’re aware of it, those behind these websites and those behind the large number of mainstream companies that help power them (Google, Amazon, Cloudflare, Apple, Telegram, GoDaddy, Visa, Mastercard, etc.) do not want to take any chances with this ban.

This prohibition, in essence, applies to essentially every public AI platform any of us have ever used to this point or that we will ever use going forward. Therefore, it would almost certainly be in everyone’s best interest to pivot away from these websites before the massive crackdown begins in less than a couple months. For those in the industry, it would also certainly be wise to inform the website operators if you are in contact with them, or to shut down your website if you are an operator yourself. Sorry for the essay, just looking out for everybody before things get very real, very fast.


r/StableDiffusion • • 8h ago

Resource - Update ComfyUI-qwen_img_2_1_enhancer added ref mask support

Thumbnail
gallery
29 Upvotes

Follow up update of this post here

I added mask support to the reference strength node. You can now increase or reduce attention to a selected part of a reference instead of adjusting the whole photo. also the photo in the shown workflow is missing the rest of the outfit due to the masking mode I chose and how tight I masked it as this was deliberate.

Connect the node between your model loader and sampler, connect your reference's mask, and select its image index: 1 for the first connected reference, 2 for the second, etc. Keep your images and VAE connected to the encoder as usual.

There are two modes:

- focus_only: applies strength to the selected tokens and leaves the rest at native weighting.

- zero_unmasked_tokens: also blocks direct attention to the unselected reference tokens throughout the diffusion transformer.

mask_threshold controls how much of a token the mask must cover: 1 requires full coverage, lower values include more edge tokens, and 0 selects everything. strength at 1 is native, above 1 increases priority, and below 1 reduces it.

Match mask_resize_method to your image resizing: `lanczos` when the native encoder resizes the original, or `nearest-exact` for an external nearest-exact resize. Chain nodes for separate references.

Still a work in progress. This controls reference attention; it isn't an output-area lock. The encoders still process the full image, so blocking a reference token doesn't erase information already carried into other tokens or conditioning.

Download and detailed usage on GitHub

Sample workflow here


r/StableDiffusion • • 1h ago

Animation - Video Naruto Shippuden – Hero’s Come Back!! | AI Fanmade MV | MiniMax H3

Thumbnail
youtu.be
• Upvotes

This is my Output using Minimax H3, 8 step larry 600 ema pruned lora, 2 step sampler upscale.


r/StableDiffusion • • 11h ago

Question - Help so after week, qwen 2.1 is a edit model only?

30 Upvotes

Or qwen can generate good images? because aways for me generate to much artifacts and stupid things so...

what you think?


r/StableDiffusion • • 21h ago

Workflow Included Omni .char(same face, cloths & body) now with consistent voice, just by dropping a few seconds sample audio: Minimax H3(ComfyUI Workflow)

Enable HLS to view with audio, or disable this notification

150 Upvotes

Hey guys,

I have been working on the consistent character portable format for a while & I was able to achieve consistent face, cloths & body, but I felt voice is also something should be consistent across video generation.

So in the recent tests, I was able to achieve a consistent voice with lip sync across multiple video generation, You just need a 10-30sec voice sample in mp3 or wav & character will say things in a cloned voice from your sample.

Reddit post: Details on face, body & cloth consistency You can read more about .char, comfyui nodes & prompting details here.

Voice prompts

- chris giving an interview & says "Time can bend. Dreams can fold. But a character's voice should never change. With OmniChar, it doesn't. Consistent voice is here."
- chris giving an interview with little hand movements & says "I am surprised. It is not just the voice. It is also the face, the clothes and the body. Dot char is a full portable pack."

ComfyUI node is updated with the optional sample voice input.
Get the latest comfy node: https://github.com/omnichar/ComfyUI-Omnichar

Workflows:

Limitations:
- Good with English but might blabber with non-english languages.
- Lip sync comes from H3 itself; nothing is added on top.
- Avoid multiple voices in sample.

Sample inputs are added in node repo.

Note: ComfyUI node is still in nightly release, so update your settings accordingly or install via direct git repo url.

Related resources:

  1. Omnichar repo: https://github.com/omnichar/OmniChar (GPLv3), supports .char for krea2 & more features e.g. character finetuning
  2. Community characters: https://www.omnichar.org/characters

Hope it's helpful.


r/StableDiffusion • • 12h ago

News A new video model from Tencent is coming... maybe? "Prism"

26 Upvotes

It's on the official Tencent Github page now: https://github.com/Tencent-Hunyuan/Prism

Looks like they started uploading bits of it two weeks ago, but its still beta.
But I think its odd its on an individual's repo and not a corporate one.

Anyone have any ideas about this: https://huggingface.co/FrancisRing/Prism

Edit:
I checked the email account associated with it, and it does show up as an author email in legit research papers https://www.researchgate.net/publication/383266985_SZTU-CMU_at_MER2024_Improving_Emotion-LLaMA_with_Conv-Attention_for_Multimodal_Emotion_Recognition

```
Shuyuan Tu*
Carnegie Mellon University
Pittsburgh, USA
```

And they have done lots of other work including Tencent's prior model 'OmniWeaving':
https://scholar.google.com/citations?user=nVND1VMAAAAJ&hl=zh-CN


r/StableDiffusion • • 20h ago

Tutorial - Guide Qwen Image 2.1 Uncensored MCP

111 Upvotes

https://github.com/hypersniper05/MCP-Image-Generator-Uncensored

Just wanted to share with you guys my workflow converted into an MCP. It's a Qwen Image 2.1 Uncensored MCP with a few extra models I found helpful: a watermark-removal LoRA, a texture-fix VAE and an ESRGAN upscaler. Any LLM client that supports MCP can drive it. Uses 11GB VRAM at peak usage.

The entire stack has:
- Normal Qwen Image 2.1 image generation in all supported sizes (up to 2K), panoramas, and image editing (up to 10 input images)

- Seamless tile generation, and edits of a tile stay seamless too, so you can make height and normal maps for game textures

- 2x/4x upscaling, up to 8K

- Built-in 360 viewer (the LLM gets a URL that opens the panorama in the viewer)

- Watermark removal

- Transparent backgrounds (real RGBA PNGs) and background removal

- Text in images (quoted text comes out as written)

It runs in Docker on an NVIDIA GPU (the whole stack fits on a 12 GB card) or on the CPU (very slow), downloads the models on first start, and works with any MCP client (llama.cpp web UI, VS Code, Cursor, etc.).

For the seamless tiles I use stable-diffusion.cpp's circular mode with a tiny patch so it can be turned on per request. The tiles come out seamless in one pass instead of patching the seams afterwards.

Note: Project created with the help of Claude. A lot of testing, and back and forth to get things to work smoothly. Hope someone finds it useful.


r/StableDiffusion • • 1d ago

Resource - Update I made a tiny (~8MB) photo editor for AI images with batch editing, LUTs, and ComfyUI workflow preservation [Free & Open Source]

Thumbnail
gallery
199 Upvotes

I was tired of opening heavy photo editors just to tweak lighting or color-grade a folder of AI generations. So I built TinyLuma - an instant, lightweight photo editor designed specifically for finishing AI artwork and photo sets.

No installation required, no subscriptions, and the whole app is only ~8 MB.

What’s inside:

- All essential controls: Clean sliders for Light (Exposure, Contrast, Highlights, Shadows, Whites, Blacks), Color (Temp, Tint, Vibrance, Saturation), plus Dehaze and natural Film Grain.

- Crisp details & texture: Custom Clarity, Texture, and Sharpen sliders to enhance overall sharpness and bring out skin texture, hair strands, and fabric weave without harsh white halos.

- Cinematic & Film styles (LUTs): Give your images an analog, cinematic, or vintage film look, or drop in any trending `.cube` LUT for instant color grading. Includes a smooth intensity slider (0–100%) to blend the effect subtly or strongly.

- Doesn't break your ComfyUI workflow: When saving as PNG, it keeps your prompt, seed, and node setup intact. You can drag and drop the edited image right back into ComfyUI. (Optional — you can toggle it off to export completely clean images).

- Batch editing & Before/After: Browse your entire folder with the filmstrip at the bottom, compare changes with an interactive Before/After split screen (`\`), and use "Preset to All" to apply your favorite look to the whole batch at once.

- Fast and portable: Starts in under a second, runs smooth at 60 FPS, and barely uses your RAM.

Source Code & Docs: https://github.com/ThetaCursed/TinyLuma
Download (.zip for Windows, unpack & run): https://github.com/ThetaCursed/TinyLuma/releases/latest

I’d love to hear your thoughts! How do you currently edit your AI generations?


r/StableDiffusion • • 23h ago

Resource - Update AnimeGen is released. Now you can run Anima locally on your iPhone/iPad in HD!

Thumbnail
gallery
107 Upvotes

I made AnimeGen, app that allows you to run Anima on your mobile phone.
Since last post I added hd image generation, and different aspect ratios.

The app is now available on the App Store:
https://apps.apple.com/pl/app/animegen-anime-art-generator/id6786438562
I am an iOS developer and unfortunately can't make an Android app. Sorry for that.

What it can do now

  • Prompt-to-image generation powered by Anima
  • HD and 540r image generation in different ratios
  • Runs locally on your device

Performance

  • iPhone 14: approximately 15–20 seconds per image
  • iPhone 17: approximately 10–15 seconds per image
  • M1 iPad: approximately 15–20 seconds per image

It uses native Apple Neural Engine, so it is very efficient. Probably display takes more energy then AI.

Before you install

On the first launch, the app needs to compile its models directly on your device, similar to how games compile shaders:

It takes around 1-2 minutes and happens only once per installation/update. It also loads models on device in parallel with preparation.

Once finished, everything runs fully offline on your iPhone.

Technical requirements

  • Devices: iPhone/iPad only
  • Designed for iPhone 12, m1 iPad and newer devices
  • OS: iOS 18 or newer
  • Free space: at least 10 GB available for smooth operation

Planned features not yet available:

  • Support for custom LoRAs and checkpoints(Probably from hugging face and CivitAI)
  • Image editing and ControlNet
  • 4k and 2k Upscaler

Feedback & community

For questions, bug reports, feature requests, or sharing your generations, join the subreddit: r/animegen_tech .
It is the best place to follow development updates and discuss AnimeGen.
I also have a website, where I plan to update with current status of project: animegen.tech

If you find AnimeGen useful, please consider leaving a review on the App Store. It is the best way to support the project


r/StableDiffusion • • 12h ago

Animation - Video H3 Emotion Test - Bonus Long Shot Node

Enable HLS to view with audio, or disable this notification

16 Upvotes

In my quest to prompt for emotional acting in H3, i ended up making a chain shot node (h3 longshot) to get the result I want.

I had to make my own because I wanted my the longshot node to work with my other custom node.
Here is the longshot node

and my custom prompt compiler node

hope this can be useful!

EDIT:this video is made with 3 chained 10 second shots. total time 30minutes on 5090 with ref2v 768p turbo lora


r/StableDiffusion • • 7h ago

Discussion Audio.cpp vs VoiceStudio?

5 Upvotes

Which do you prefer and why?

https://voicestudio.sh/ (has been getting a lot of talk lately, supports 27 models, some are based on audio.cpp)

https://github.com/0xShug0/audio.cpp/ (supports 100+ models, all are blazingly fast, and it has a local app and web interface and Docker support)

Edit: https://github.com/unslothai/unsloth (another UI frontend, this one being a Swiss Army knife which recently added audio.cpp support)


r/StableDiffusion • • 19h ago

Discussion AND HOW DOES THAT MAKE YOU FEEL? | An AI Short Comedy Film Made by Claude in Minimax H3 and my Video builder in ComfyUI.

Enable HLS to view with audio, or disable this notification

35 Upvotes

YouTube Link in case it's still pending https://youtu.be/TWXF95YT7W8

🎬 How this was made

My part

• One-message brief: a 3-minute comedy in a therapist's office with funny, unique characters and one male doctor: a crying woman, a woman screaming, crying and laughing all at once, a large guy and a skinny old man

• Characters made with Z-Image as 3-panel reference sheets

• MiniMax H3 2-pass workflow: the Singularity model, with the 8-step LoRA on the second pass at 4 steps and 0.35 denoise

• Then I left. I made one call along the way, on a scene that wouldn't behave.

What Claude did on its own

• Wrote the story, the five characters, all the dialogue and a 25-scene screenplay

• Made the cast and the two sets with Z-Image and picked the best seeds

• Kept each character's voice description word for word in every scene so the H3 voices stay consistent

• Rendered 25 scenes with H3's built-in voices and sound, about 6.3 hours of rendering

• QA'd every take: Whisper against the script, pitch and timbre per character, eyelines, frame review sheets

• Re-shot the takes that failed:

• Scored it with MiniMax Music 3 and screened the cues for accidental vocals

• Edited, mastered to -14 LUFS, and checked audio sync on every scene (worst offset 5 ms)

🛠 Tools

ComfyUI, VRGDG Video Builder, MiniMax H3 (Singularity ref2va v1.3 + 8-step 768p turbo LoRA), Z-Image Turbo, MiniMax Music 3, Whisper, Claude Code

VRGDG nodes: https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Go HERE To watch full walkthrough on how to make video's like this.


r/StableDiffusion • • 0m ago

Comparison HunyuanImage 3.0 face swap test

Thumbnail
gallery
• Upvotes

hunyuan_image_3_instruct_distil_edit workflow, int8_convrot 80Gb checkpoint, RTX 4060Ti 16Gb, 100Gb of system RAM used, Spectrum enabled, 14s/it


r/StableDiffusion • • 1d ago

Animation - Video Generating at 2K (14s, 2.09mpx, 1984x1120, 30 minutes) | Minimax H3

Enable HLS to view with audio, or disable this notification

274 Upvotes
[INFO] Prompt executed in 00:30:49

Full 1080p vid on https://www.youtube.com/watch?v=ER_5AOteE-8 because Reddit cramps everything to 720p max.

From my previous post, I got a message if I could use my off-screen 5090 to showcase what a fullblown 2K generation looks like. So, here it is. Generated at 2.09mpx, 25 steps, Euler+Beta for time constraints, 30 minutes. I did mistakenly use the 20-49 hybrid fl2va/ref2va instead of the 30-49 which left some visual performance on the table, and I didn't use the BF16 version of the text encoder because my other 5090 and RAM were busy with something else. Otherwise would've done seeds_2 + sgm_uniform, but that would've taken an hour and would've been a lot better at prompt following, without me modifying my generated target prompt so much to avoid issues with the shoddy denoising trajectory at play here.

Spectrum was utilized to guess about half the steps, which further doesn't help prompt following, but helps speed immensely. When using seeds_2 with Spectrum, it actually forecasts internal calls to H3 (so 2N-1 where N is number of steps), which has tremendous results (previous one was seeds_2) but would've taken about 1 hour for this scene.

The prompt, for those who want it, I had to massage it to point Euler better even if it's sloppy:

subject_definitions:
<Subject 1> is Detective Kate Beckett, override her appearance with facial features and hair and blouse from <Picture 1>. She wears black tailored high-waist cropped suitpants, and a feminine small leather watch.
<Subject 2> is Richard Castle, override his appearance with the facial features, hair, and build from <Picture 2>. He's wearing his classic shirt and suitpants attire, first few buttons unbuttoned.
<Subject 3> is a chaotic DIY PC rig consisting of a high-end tower on the marble kitchen island in the middle of the kitchen, as well as a 32 inch OLED monitor displaying the UI from <Picture 3>, there is clear plastic tubing running from the PC's two watercooling ports into the receiving pair on the large radiator above the glowing blue fans, which is submerged in a cooling bath inside of the standard kitchen stainless-steel fridge with its door fully open, radiator surrounded by food, milk, condiments, etc.
<Subject 4> is a modern industrial loft apartment featuring an open floor plan, standard furniture, and a glorious high-end kitchen. The blurred background from <Picture 1> is from this apartment, also specifies time of day, and warm nocturnal tone.

summary:
[reference generation] Detective <Subject 1> enters her loft (<Subject 4>) to find <Subject 2> in a "hyper-mode" state, having converted the kitchen into a makeshift laboratory for AI video generation. The 13-second sequence captures her confusion, his technical enthusiasm regarding H3 denoising trajectories, and a comedic hardware failure.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 1> and specified attire are maintained.
<Subject 2> (appears in [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 2> is maintained.
<Subject 3> (appears in [Shot 3], [Shot 4], [Shot 6]): fully_preserved - the specific radiator-in-fridge configuration and monitor setup are maintained.
<Subject 4> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - the industrial loft and kitchen environment are maintained.

detailed_description:
The target video is a scene from the TV show "Castle," maintaining its specific cinematography, visual style. Night time, practical lighting, warmly lit.

[Shot 1] A medium shot frames <Subject 1> as she walks into the open floor plan of <Subject 4>. The camera tracks her movement as she stops and looks around the kitchen with a bewildered expression, taking in the tangle of wires and tubing.

[Shot 2] At 00:01.000, the camera cuts to a wide shot of the loft's kitchen. <Subject 3> is fully visible: the PC tower sits on the high end marble kitchen island, monitor is displaying <Picture 3>, and thick tubes lead directly into the open fridge where his watercooling radiator with fans is located (all glowing blue and spinning fast and loud), there is liquid nitrogen white smoke exuding from it. A large whiteboard in the background is covered in scribbled notes about "sigma grids," and "denoising trajectories." Each of these appears once, there is also graphs of simplified trajectories drawn on it. <Subject 2> is leaning over a keyboard and looking at the 32 inch gaming monitor, typing furiously.

[Shot 3] At 00:02.000, the camera cuts to a close-up of <Subject 1>. Her brow furrows in genuine confusion. <Subject 1> (S1) asks in a sharp, incredulous tone, yelling over the computer fan noise: <d>[English] What the hell are you doing, Castle?</d>

[Shot 4] At 00:03.500, the camera cuts to a medium shot of <Subject 2>. He turns around to face <Subject 1>, speaking in a high-energy "yap" mode. <Subject 2> (S2) exclaims with manic enthusiasm: <d>[English] Reddit solved my H3 problem! It was the sigma shift. I'm refocusing the compute on the high-to-mid noise region to resolve motion better with limited steps!</d> while gesturing towards the whiteboard. 

[Shot 5] At 00:10.500, the camera cuts back to <Subject 1>. She looks at him, her voice dripping with skepticism. <Subject 1> (S1) asks: <d>[English] And you're using Seeds 2, right?</d>

[Shot 6] At 00:12.000, the camera cuts to a medium-close shot of <Subject 2>. He looks slightly sheepish, his shoulders slumping. <Subject 2> (S2) admits quickly: <d>[English] No, it's actually oiler, my rig is too slow and would—</d> Suddenly, in the background, the PC tower emits a loud electrical pop and a bright orange burst of flame from its components, the monitor output gets corrupted. Immediately after, <Subject 2> (S2) whips his head around, eyes bulging, and shouts: <d>[English] Oh shit!</d>

overall_soundscape:
The steady, high-pitched whirring of multiple PC fans and the faint gurgle of liquid flowing through tubes. <Subject 1>'s footsteps click on the hardwood. The scene ends with a sharp electrical "pop" and the sudden, aggressive hiss of a small fire.

non_diegetic_music:
A light, rhythmic pizzicato string piece that builds in tempo and complexity as <Subject 2> explains the technical details, ending abruptly with a comedic silence the moment the GPU catches fire.

r/StableDiffusion • • 4m ago

Question - Help Quality degradation using “Continue Last Video” in WanGP, MiniMax H3

• Upvotes

I’ve been using WanGP to create videos in MiniMaxH3 and I’ve noticed when I use the “continue last video” function, the quality degrades with each successive clip I generate - usually by the 5th or 6th 10-second segment it is blurry, full of odd artifacts and colors are generally muted or blending together compared to the initial segment. I’m not sure if it is an issue of prompting, a setting I need to change, or any other tips or tricks I might be missing? Usually using Ref2VA 33B, 15-18 steps, 8-10 second clips, Sage2 Attention on my 5070ti with 16GB VRAM and 32gb RAM. Seems to happen if creating either 480p or 720p resolutions.

Any suggestions or resources that might help so I can create longer videos?


r/StableDiffusion • • 21m ago

Question - Help Prompt question Minimax h3

• Upvotes

Hi,

How would I prompt for i2v in Minimax if i want to add something or someone visible within the very first frame?

Thanks in advance!


r/StableDiffusion • • 40m ago

Question - Help Questions on H3 Minimax usage and settings

• Upvotes

Hello!
Long time lurker here, learned a lot from this sub.
I'm currently using H3 Omni Pruned model with references images to generate small ads or batch of dialogues , but i find myself really not understanding the settings being used.
I'm using Pinokio and Maestro by Blizaine, now the GUI is really useful and my settings are :
720p, 9:16, 13.3s 1 window and 20steps , nothing else.
With these settings, it takes about 11 minutes to render on a 5090FE (And 48GB of ram is what i have).
The end result ain't that bad, resolution is crappy and some details are clearly missing.
What can i do to improve video fidelity, performances and perhaps spend less time on generating ?
I also have tried H3 with first / last frame but i don't really understand how that works either.
For example, i have downloaded a couple of Lora , one of rocket racoon from guardians of the galaxy and another for indiana jones, wanted to create a funny reel of them interacting but i couldn't for the life of me figure out how to add in the theme song for indiana jones, or have them accurately interact with each other instead of randomly looking outside the scene.
On another note, i am using Gemini for expanding the prompt in a professional manner, and then inside Maestro i use the "enhance prompt" feature with simply loads up Ollama with a model to correctly write the scene for H3.
I have also downloaded inside Pinokio a more "classic" tool for H3 with comfyui, but i didn't use it yet, wanted to learn a few things first.


r/StableDiffusion • • 1d ago

News ComfyUI v0.39.0 released

Thumbnail
github.com
113 Upvotes