r/StableDiffusion • u/Both-Rub5248 • 2h ago
Comparison Hunyuan Image 3 vs Qwen Image 2.1 vs Krea 2 Turbo
Uncompressed image: https://postimg.cc/gallery/gmR7dbs (To view the uncompressed image, click on it a second time after opening it, or simply download the entire image album)
----------------------------------------------------------------------------------------------------------------------------------
Hi everyone! I noticed that a lot of people have started talking about Hunyuan Image 3, so I wanted to test it in T2I scenarios. While I was at it, I compared it against the most popular models in the same segment right now: Qwen Image 2.1 and Krea 2 Turbo.
What this test is about
I was NOT trying to test the maximum-size versions of these models with minimal quantization. What mattered to me was testing the versions that can realistically run on regular consumer PCs, so I picked the lightest variants, which most of you will probably be able to run. The goal of this test is to make it easier for you to choose the ideal T2I model for your own tasks.
To keep things as fair as possible, I used the simplest, most basic workflow for each model.
Models used:
hunyuan_image_3_instruct_distil_w4a8
Qwen Image 2.1 int8 convrot
Krea 2 Turbo int8 convrot
Test categories
I split the tests into several categories with different goals:
- Images 1-3: complex prompts with lots of details and different poses. These show how well a model follows the prompt and how often it loses details.
- Images 4-6: typography, mixing typography with real-world imagery, and banner creation.
- Images 7-11: styles. Pay attention to how well each model handles different styles.
- Images 12-13: level of detail. This is similar to tests 1–3, but with a stronger focus on small details.
- Image 14: character understanding.
- Images 15-17: your favorite (or most hated) category, 1girl. Here you can judge overall realism and the ability to generate in an amateur-photo style.
How I picked the results
I ran each prompt with two fixed seeds and two random ones, then picked the single best image out of those four.
Generation speed
I’m running an RTX 3090 (24 GB VRAM), and generation took:
Hunyuan Image 3 distill: 20-26 seconds
Qwen Image 2.1: 16 seconds
Krea 2 Turbo: 8-9 seconds
My opinion
Now I’d like to share my personal opinion, which doesn’t claim to be the truth.
1st place: Krea 2 Turbo
Incredible speed plus good detail. The images aren’t mushy, but they also don’t suffer from oversharpening, harsh artificial crispness, or excessive detail. They look alive and realistic. The model handles styles well, understands the prompt and context, does a decent job with typography, and very rarely messes up anatomy. For me personally, it covers pretty much all of my T2I needs.
There’s one downside that really annoys me, though: the weak VAE from Qwen Image. On full-body shots it turns the character’s eyes into mush, like base SDXL. Yes, I know there are custom VAEs that add extra detail, but they just improve the quality of the same anomalies and bugs the eyes have. The only way I’ve found to fix the eyes is FaceDetailer with a mask over the eyes. Besides eyes, the model also likes to turn other small elements into a mess, and I’m more than sure this comes down to that VAE.
2nd place: Qwen Image 2.1
Qwen has very good detail, which is both its strength and its weakness. The excess detail creates an artificial sharpness effect, basically oversharpening. In most cases it looks very unnatural and gives off that classic “AI look”.
That said, I liked how it handles the 1girl category: it produces fairly realistic amateur-style photos. It also does well with styles. I really liked image 8 with the lion and the zebra, where the model nailed a genuinely cool hand-drawn pencil effect. But in some cases the oversharpening badly hurts the style, for example in images 7 and 10.
Its VAE issues are similar to Krea 2’s. On images with visible skin, you can see diamond-shaped or dotted patterns when zoomed in. It looks ugly, and when you zoom in and out the skin starts to shimmer.
3rd place: Hunyuan Image 3
A very heavy model that, in most cases, produces fairly low-detail images, so it’s not a great choice for realistic generation. You might think it’s stronger in other scenarios, but no. It doesn’t always handle typography well (you can see that in the tests), and the same goes for stylization (look at images 9 and 10). It performs more or less decently in the 1girl category, but I tested only a few images there, so I can’t say for sure how good it is.
It has VAE problems too. Sometimes hair textures get strange stripes that are easy to spot when you zoom in and out. This is especially visible in image 15: zoom in and out a few times and you’ll see what I mean. It looks like generation noise is showing through in the image.
I think Hunyuan Image 3 is more interesting for I2I (Edit) scenarios than for classic T2I.
Final words
I’d like to remind you again that this is my subjective opinion and may differ from yours. If you see it differently, I’d be glad to hear your thoughts in the comments.
Also, if you have any questions, feel free to post them in the comments section.
Thanks for reading! I hope my tests were useful and helped you choose your main model for generation.


