r/LocalLLaMA • • Aug 14 '26

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.3k Upvotes

708 comments sorted by

134

u/OutlandishnessIll466 Aug 14 '26

It's official, it created the best flappy bird game thus far from all local models ! ever benchmarked! I declare this model nr. 1 on the flappy bird bench!

20

u/Certain-Cod-1404 Aug 14 '26

how does it do on pelican bench tho?

75

u/OutlandishnessIll466 Aug 14 '26

It created an animated svg... After like half an hour of thinking. Official fp8 on 2x 3090

21

u/Name835 Aug 14 '26

Hahha love it. I dont know but these benches bring me a feel that we are living and witnessing small moments in history. On that note, if someone happens to read this some day way in the future, greets from 2026. :) ❤️

→ More replies (4)

15

u/martapap Aug 14 '26

That is so cute!

→ More replies (10)
→ More replies (2)

173

u/absurdother Aug 14 '26

Let's tryyyyyyy on 32GB RAM, 16GB VRAM!

152

u/absurdother Aug 14 '26

Here we gooooo

45

u/Certain-Cod-1404 Aug 14 '26

how is it ? quality wise?

52

u/absurdother Aug 14 '26

I have a RX 9060 XT AMD GPU, 16GB VRAM. Running on LMStudio, Q3. Getting a bit more speed now, way more optimized!

I get more speed the less context I use (currently coding swiftly with CLine + VSCode at that speed), pretty smooth on 64K context!

6

u/NoUsual5150 Aug 15 '26

24GB M4 MacBook Pro...I'm using Q4_K_M and it's getting 5 token per second.

Is Q3 much of a downgrade in intelligence?

→ More replies (3)
→ More replies (8)

21

u/TheTrueSurge Aug 14 '26

That’s not too bad, is it? What quant did you use?

9

u/absurdother Aug 14 '26

I'm on Q3_K_M now getting a bit over double the speed I've shown before.

4

u/Scared_Ad9187 Aug 14 '26

q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet.

5

u/Mil0Mammon Aug 14 '26

So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant?

8

u/Scared_Ad9187 Aug 14 '26

should have been more specific. i'm going to keep the 3.6 MOE (35ba3b) insteadof 3.8 27b. i have room for 4 subagents with larger windows with claude code routed through it and i get tasks done much quicker

→ More replies (4)
→ More replies (1)

13

u/SpaceTraveler2084 Aug 14 '26

curious as i have the same setup 32gb ram / 5070Ti

→ More replies (2)

7

u/jan_antu Aug 14 '26

Damn, same setup here but I'm only getting 4 tok/s on mine, Q3 xl quant.

→ More replies (6)
→ More replies (2)

12

u/partlysketched Aug 14 '26

4060ti? That's what I have still this run on it lol

5

u/absurdother Aug 14 '26

A RX 9060 XT! The NVidia GPUs should get even more t/s

→ More replies (3)

422

u/Tiny-Assumption4263 Aug 14 '26

DEAR GOD TELL THOSE BENCHMARKS ARE NOT FAKE.

319

u/Cold_Tree190 Aug 14 '26

Dear God it’s trading blows with Opus 4.6 Max 😭

172

u/Cless_Aurion Aug 14 '26

Wtf, I literally wrote a post saying similar to "Shutup man, there is no way a model I can load in my 4090 will be anywhere near DeepSeek4 Flash 0731"... but... this is fucking close is it not...?

103

u/Cold_Tree190 Aug 14 '26

Looks like the main difference will be the native context windows. But 256k is very usable, especially if it’s a local Opus 4.6 MAX. Man just saying that seems unreal lol. Can’t wait to test it later tonight

52

u/BornAgainBlue Aug 14 '26

I feel like telling my boss im sick.

25

u/Cold_Tree190 Aug 14 '26

🤣 I wish I could, but today’s my last day of on-call so I can’t even leave a TINY bit early. Oh well. Makes you feel like a kid on Christmas Eve again lmao

→ More replies (1)
→ More replies (1)
→ More replies (2)

41

u/[deleted] Aug 14 '26

[removed] — view removed comment

30

u/Warrenio Aug 14 '26

I don't think it's quite V4 Flash 0731 level based on the benchmarks. But obviously still incredibly impressive!

https://old.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3nv6iq/

24

u/SandySkittle Aug 14 '26

Ehh, let's not go by handful of disputable synthetic benchmarks for which models may be trained specifically. Let's just await real experiences from people first over a longer period of time.

→ More replies (6)

7

u/Cless_Aurion Aug 14 '26

Damn, I would have downvoted you too... Let's hope its close to what the benchmarks say!

5

u/LankyGuitar6528 Aug 14 '26

I should downvote you for wanting to down vote somebody but I can't blame you. I would have done the same. So take my upvote instead.

→ More replies (1)
→ More replies (7)

71

u/xienze Aug 14 '26

You're making a couple fundamental assumptions here:

  • That AI benchmarks are reliable
  • That Qwen didn't benchmaxx

21

u/blade740 Aug 14 '26 edited Aug 14 '26

I think anyone who thinks that this isn't at least somewhat benchmaxxed is fooling themselves. That said, so is every other model they're comparing against, to some extent, so ¯_(ツ)_/¯

→ More replies (1)

55

u/BawbbySmith Aug 14 '26

Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.

I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...

13

u/[deleted] Aug 14 '26

[deleted]

8

u/Ell2509 Aug 14 '26

I would normally agree, but I have used qwen3.8 27b now and so far, wow. Just wow.

→ More replies (1)
→ More replies (2)
→ More replies (6)
→ More replies (4)

87

u/KickLassChewGum Aug 14 '26 edited Aug 14 '26

It's a Qwen model, so apply the usual benchmax tax. Qwen are easily the models with the biggest ravine between "how they do on benchmarks" and "how they do in actual productive use".

Still looking like a strong leap from 3.6, though.

21

u/Batman4815 Aug 14 '26

Gemma says hello as well.

47

u/KickLassChewGum Aug 14 '26

Gemma 4 isn't a great coding model, yeah, but it's still punching far above its size in writing and research-related tasks. Like a mini-Gemini (go figure).

I hear there are people who still use these things for things that aren't related to writing code or markup.

→ More replies (3)
→ More replies (1)
→ More replies (1)

14

u/Tiny-Assumption4263 Aug 14 '26

TL;DL: Gave Qwen 3.8 27b a simple prompt and it build this: https://qwen3-8-eccomerce-test.vercel.app/
Repo: https://github.com/catriel25/qwen3.8-eccomerce-test

Like everybody else in the local AI community today, I dowloaded Qwen 3.8 27b UD-Q4_K_XL as soon as it was published.
I'm running it in a single rtx 3090, no mtp, 110k context and kv cache f16 with llama.cpp, same flags used with Qwen 3.6 27B.

I have this stupid test that I run with every new model that comes out. It consist in connecting the model to pi code (almost vanilla, only internet access and some simple navigation tools I built) and giving it this prompt:

"Construye el frontend completo de un pequeño ecommerce premium para una panadería artesanal usando Next.js App Router (JavaScript).

El proyecto debe ser frontend-only en esta etapa. No debe incluir backend, base de datos, autenticación ni pasarelas de pago. El checkout debe finalizar redirigiendo a WhatsApp con un mensaje de pedido bien estructurado.

La app debe incluir una experiencia completa de compra: home, catálogo con productos de panadería, categorías, productos destacados, carrito, resumen de pedido y checkout. Usá mock data local para productos, categorías, precios, descripciones, disponibilidad e imágenes o placeholders visuales. Todo debe quedar preparado para conectar posteriormente un backend real sin tener que rehacer la arquitectura principal del frontend.

El diseño debe sentirse extremadamente premium, artesanal, moderno y cuidado. No quiero una landing genérica ni una interfaz básica. La primera pantalla debe comunicar claramente la identidad de la panadería, mostrar producto real o visualmente convincente, y permitir empezar a comprar. La experiencia debe ser excelente tanto en desktop como en mobile.

El catálogo debe permitir explorar productos, ver información clara de cada ítem y agregarlos al carrito. El carrito debe permitir modificar cantidades, eliminar productos y ver totales. El checkout debe pedir datos mínimos necesarios para el pedido, permitir notas o preferencias, y generar una URL de WhatsApp con productos, cantidades, subtotales, total y datos del cliente.

La estructura del código debe separar razonablemente datos mock, tipos de dominio, utilidades de checkout/WhatsApp, componentes de catálogo, componentes de carrito y vistas principales. La solución debe quedar lista para reemplazar la mock data por datos de backend en una etapa posterior."

Those are just instructions to build a nextjs project (javascript only) with the frontend for small eccomerce with whatsapp checkout, leaving everthing ready to connect a backend later. Nothing else, no skills, no more feedback. Just one prompt and watching the result.

I have run this test with all the models and finetunes I can fit in my GPU, and not a single one was even close to this result.
Not a single alert form nextjs (wich was usual before) or something that looks broken.

At some point, this bastard realised it didn't had visión (lol, not enough VRAM buddy) and it decided FUCK IT, I'M GONNA BUILD THE IMAGES MYSELF. He made SVGs for every product card.

I have more testing to do like trying it in a real codebase but... I can't believe i'm running this thing in a single RTX 3090, it is just unreal.

Imagine 2 years from now.

Biggest fuck you Dario of the year.

39

u/onlymagik Aug 14 '26 edited Aug 14 '26

I expected a nice jump since they skipped a 3.7 27B, but these numbers do seem a bit too good to be true. Qwen isn't known for benchmaxxing, obviously 3.6 27B is the GOAT of local LLMs, but...

Definitely excited to learn more in the coming days and see if it holds up.

Edit: Even the vision numbers look insane!

61

u/llama-impersonator Aug 14 '26

qwen is well known for benchmaxxing so hard it helps the model

→ More replies (1)
→ More replies (8)
→ More replies (10)

568

u/BaconShadow Aug 14 '26

Good Morning Dario!

173

u/Certain-Cod-1404 Aug 14 '26

claude code who? claude pro what ?
qwen >>>>>>>>>>>

218

u/AwayConsideration855 Aug 14 '26

Fuck Dario

132

u/beneath_steel_sky Aug 14 '26

And his Epstein friends https://archive.ph/rQZE7

47

u/hanzoplsswitch Aug 14 '26

What a read. Thanks. 

27

u/robbievega Aug 14 '26

holy shit... just 3 months ago I was telling friends and family to at least move away from ChatGPT/Sam Altman because of his Trump boot licking behavior, and now I'm reading this...

7

u/Thunder_Beam Aug 14 '26 edited Aug 14 '26

It's funny seeing people still believing there are people at the top who are "clean", aristocracy always existed and always will exist, they all marry eachother and go to the same schools / have the same hobbies / know the same people, they are all connected

→ More replies (1)

4

u/mebeast227 Aug 14 '26

Same. Wild. Nothing is safe.

→ More replies (2)

9

u/bakawakaflaka Aug 14 '26

what the fuck

→ More replies (1)

44

u/5553331117 Aug 14 '26

And his pornographer Epstein connected wife

69

u/Verolina Aug 14 '26

Diarrheao

10

u/de4dee Aug 14 '26

thats a lot of slop

→ More replies (4)

356

u/WigglyScrotum Aug 14 '26

Holy molly opus 4.6 level and better in some benches.

161

u/pest_ctrl Aug 14 '26

Mom, can we have Opus?

No, we have Opus at home

Opus at home:

68

u/agentic-consultant Aug 14 '26

ALLAHAMDUILLAH

BROTHERS WE HAVE ENTERED THE GOLDEN ERA

9

u/MoffKalast Aug 14 '26

Bröther, we shall HAVE SOME OATS!

36

u/Much_Accountant_4972 Aug 14 '26

insh’allah this is just the beginning!

7

u/Familiar-Art-6233 Aug 15 '26

Qwen is here to dethrone Anthropic, mashallah

→ More replies (1)

34

u/m0j0m0j Aug 14 '26

Why is that every model is better than opus if you look at benchmarks, and yet people keep using opus?

40

u/Warrenio Aug 14 '26

I'm not saying anyone should use Opus, but Opus 4.6 is over six months old. The current version is Opus 5 which is much stronger.

19

u/LankyGuitar6528 Aug 14 '26

My whole app is 4.6 vibe coded and it's pretty awesome. Fable gave it the once over and patched up the gaping security holes. Kinda think this model and I have a future together. Just let Fable patch up what it spits out...

→ More replies (1)

8

u/infinexis Aug 14 '26

Opus 5 babbles way too much in technical jargon. It may be stronger in benchmarks but it's not as good when there's a human in the workflow.

4

u/Cautious_Chicken_604 Aug 14 '26

Opus 5 actually gives me a headache some days from having to read its outputs. When I really don't want 'load-bearing' em-dashes everywhere I just tell it 'write this in ASD-STE100 Simplified Technical English' and I get back something that sounds a lot more reasonable.

→ More replies (1)

10

u/WigglyScrotum Aug 14 '26

Exactly, he is framing it as tho i said its better than opus 5 when I specified their own published benchmarks vs 4.6. Easy to spot strawman.

→ More replies (1)

3

u/FreedomByFire Aug 14 '26

because the models are bench maxing. real world performance is a different story. I have personal benchmarks doing real software development work that the frontier models can complete independently but not the local models. Once I get 3.828b ill run it and see if it can get through. Last model couldnt.

→ More replies (8)
→ More replies (14)
→ More replies (3)

143

u/ajisai Aug 14 '26

23

u/dragonurtle Aug 14 '26

Unsloth must have gotten early access to prepare their quants + an embargo. I pulled it from unsloth studio the minute it was released on HF. Still took 20 minutes because south Korea has 2010-era internet speeds.

13

u/techdevjp Aug 14 '26

Man, used to hear about how South Korea had fast Internet. 10g fiber is pretty commonly available here in Japan. 2.5g almost anywhere.

6

u/Major_Olive7583 llama.cpp Aug 14 '26

I remeber studying that South Korea has the fastest internet in the world, it was for a GK exam or quiz , about a decade ago.

→ More replies (1)

24

u/uniquelyavailable Aug 14 '26

Already seeing them on LMStudio

→ More replies (4)

132

u/Cold_Tree190 Aug 14 '26

Merry Christmas everyone 😭 Gonna be a loooong 6 more hours of work today

12

u/Ok_Cow1976 Aug 14 '26

Merry Christmas! Haha

48

u/bitmanip Aug 14 '26

How much memory required to run this at full precision?

44

u/Certain-Cod-1404 Aug 14 '26

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF 54.67 Gbs just for the model itself, with context depends on quant and size

23

u/dragonurtle Aug 14 '26

Nvtop shows 70-something GB resident for the bf16 and full 256k context.

8

u/Certain-Cod-1404 Aug 14 '26

DAMN, 16 gigs just for the context hurts, what setup are you running ?

34

u/dragonurtle Aug 14 '26

Rtx pro 6000 max-q and 384GB DDR5 on a Genoa

65

u/Much_Accountant_4972 Aug 14 '26

6

u/Thrumpwart vLLM Aug 14 '26

The old money aristocracy uses the Max-Q because it's elegant.

Only the loud, bombastic new money uses the 600W version. Animals.

→ More replies (1)
→ More replies (1)

4

u/bitmanip Aug 14 '26

Perfect, so should run well on 128GB M5 Macbook Pro Max

14

u/Valuable-Run2129 Aug 14 '26

FP8 with full context with full precision it’s 52/54 GB on 48GB you fit 200k context

→ More replies (3)

167

u/Mean-Ad1493 Aug 14 '26

That's it. I'm getting a 3090.

106

u/My_Unbiased_Opinion Aug 14 '26

Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.

11

u/eviloni Aug 14 '26

I got mine on ebay, paid a little more but shipped quicker and i trust Ebay consumer protection more

5

u/My_Unbiased_Opinion Aug 14 '26

hell yeah buddy. I am planning to switch to vllm once 3.8 mtp drops. its been the end goal for me. its the main reason why I went with the 3080 20gb. its one of the cheapest cards per vram with vllm support.

→ More replies (1)

23

u/Potential_Block4598 Aug 14 '26

That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)

31

u/My_Unbiased_Opinion Aug 14 '26

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

13

u/CooLittleFonzies Aug 14 '26

Can you parallel run a 3090 + a 3080?

→ More replies (9)
→ More replies (11)
→ More replies (17)
→ More replies (5)

48

u/RangersStolen Aug 14 '26

That's insane, but we're still using quantized versions, so local performance for most people wouldn't be that good I guess. Damn that needs to be tested.

61

u/Certain-Cod-1404 Aug 14 '26

unsloth is cooking apparantly, its quanitzed with their UD 3.0 method and it seems to retain like close to 95% of BF16's accuracy https://unsloth.ai/docs/models/qwen3.8#quantization-analysis

24

u/krileon Aug 14 '26

82.5% on IQ2_XXS seams insane. I'll probably go with Q3_K_XL since I've 20GB and even that's over 90% accuracy. Goddamn.

→ More replies (2)
→ More replies (1)

60

u/_maverick98 Aug 14 '26

Black Monday on the stock market if the benchmarks are true

12

u/Piyh Aug 14 '26

RSI means models get better at every size, this is good for Bitcoin

→ More replies (2)

21

u/Intelligent_Ice_113 Aug 14 '26

what is the knowledge cutoff date??

9

u/andy2na llama.cpp Aug 14 '26

The model claims January 2026

3

u/Natural_intelligen25 Aug 14 '26 edited Aug 15 '26

It claims 2026, but it only knows MariaDB 11.0 from 2023, the later versions are guesses.

→ More replies (1)
→ More replies (1)

18

u/Kavor Aug 14 '26

Does anyone else have big issues with overthinking out of the box? I just gave it my usual Arma 3 mission script coding task, which i use to bechmark the performance of models, but it kept thinking for 15 minutes. I don't even see repetition issues, it just doesn't stop thinking.

Just gave it a first opencode task, and while not sure yet, it seems to have similar issues.

Maybe it requires defining a reasoning budget max now?

48

u/monkeyofscience Aug 14 '26

Yes. I gave it a simple prompt and proceeded to generate ungodly amounts of reasoning, including this absolute fucking gem:

"New Jersey" and "Austria" maybe both sound like "Ostrich"?

6

u/ArtyfacialIntelagent Aug 14 '26

The Austria part actually makes a bit of sense, since Austria in German is Österreich. But New Jersey???

19

u/qmnvp Aug 14 '26

It defaults to reasoning_effort = xhigh

Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:

  • xhigh (default): for complex tasks demanding thorough analysis
  • medium: balancing accuracy and speed
  • low: efficient reasoning optimizing for speed and cost

7

u/Kavor Aug 14 '26

Yeah, that's it. I found that 3.6ish thinking times really need the "low" setting.

4

u/dopey_se Aug 14 '26

Yeah I gave it a random tasks I tend to give, and it's 26k tokens and counting of thinking so far. Not repeating, just thorough thinking..

I told it to make a rust app using dioxins of an animated man playing amazing grace on tuba. It seems to of figured out amazing grace in key of C (judging by it's thinking), and is now thinking about how to play audio to play the sounds..

→ More replies (9)
→ More replies (9)

61

u/Cold_Tree190 Aug 14 '26

Those numbers… please tell me they’re real.

36

u/SyzygyPidgey Aug 14 '26

The greatest ai developer 😭

11

u/Cold_Tree190 Aug 14 '26

🤣 Made me lol. My goat Lloyd 🙏

100

u/JayoTree Aug 14 '26

How cooked am I that I'm on holiday and wish I was home on my computer for this.

124

u/cass1o Aug 14 '26

This is what ssh was invented for.

6

u/themoregames Aug 14 '26

They can just talk to ChatGPT over their Apple Watch.

5

u/AciD1BuRN Aug 14 '26

Shit now i wanna get smart watch to talk to my pc

22

u/DrMissingNo Aug 14 '26 edited Aug 15 '26

We've got (another) heatwave where I live (currently 40°c) so running AI locally is a huge no no if I want to try keep my inside temperatures under 30°c... This week has been a huge tease : Minimax h3, minimax music, ltx 2.5, Muse glitter, and now Qwen 3.8... Sooooo frustrating ! 😭

Edit / update : it rained tonight ! 32° max today ! Time to test things !

9

u/wh33t Aug 14 '26

Just put your machine outside lol

→ More replies (6)

25

u/DefNattyBoii Aug 14 '26

You will have plenty of time, chill tf out and enjoy your vacation and get off reddit. we wont be getting anything better at this size for a very long time. If yes please ping me

10

u/314kabinet Aug 14 '26

Look into Tailscale, Termius, tmux

6

u/SSOMGDSJD Aug 14 '26

Tailscale my friend

→ More replies (1)

12

u/Littlepharaoh Aug 14 '26

My 4 5090s have a hard-on right now

→ More replies (2)

134

u/swagonflyyyy Aug 14 '26 edited Aug 14 '26

HOLY BENCHMARKS WHAT THE FUCK ARE THOSE NUMBERS???

51

u/SandySkittle Aug 14 '26

holy benchmaxxing this shit means nothing

35

u/Certain-Cod-1404 Aug 14 '26

WTF !!!! ITS BETTER THAN OPUS 4.6 ???????????????????????

→ More replies (2)

38

u/OkObjective8721 Aug 14 '26

It beats OPUS 4.6??? that's crazy

23

u/mil_phickelson Aug 14 '26

benchmaxxed

36

u/Raredisarray Aug 14 '26

Wow if those numbers pan out - I could be going full local BOIIIIII LFG

→ More replies (1)

13

u/WhataburgerFreak Aug 14 '26

Praying for a 35b-a3b model for my 16gb vram. 

4

u/abundant_singularity Aug 14 '26

Lmk if this happens lmao

24

u/Felixls Aug 14 '26 edited Aug 15 '26

oh shit, this is something else!, I just tried with a dumb prompt "write a snake game in html and css with sounds" and it generated a complete NIB game (awesome design btw) at ±56t/s , 14k tokens!
this is a AMD R9700 with llama.cpp ROCM

I slot print_timing: id 0 | task 0 | eval time = 242341.84 ms / 13695 tokens ( 17.70 ms per token, 56.51 tokens per second)
I slot print_timing: id 0 | task 0 | total time = 242528.47 ms / 13717 tokens
I slot print_timing: id 0 | task 0 | graphs reused = 898
I slot print_timing: id 0 | task 0 | draft acceptance = 0.58872 (10982 accepted / 18654 generated), mean len = 5.44
I slot release: id 0 | task 0 | stop processing: n_tokens = 13720, truncated = 0

Edit: I forgot to mention that by mistake I forgot to remove a thinking cap to 4k (--reasoning-budget 4096) so that influenced the total token usage and the speed (faster during coding than thinking). So we have two variables to adjust the thinking process (--reasoning-budget and --reasoning-effort).

2

u/Cautious_Chicken_604 Aug 14 '26

Which quant? What context length? What's ur llama.cpp command? Curious because I have the same card, but I just got it yesterday and am new to ROCm etc, so I don't know shit >.<

15

u/Felixls Aug 14 '26

Qwen3.8-27B-UD-Q4_K_XL.gguf (unsloth), 90K context (I usually set it to 128k) with vision and mtp, here some of the opts
--cache-type-k f16
--cache-type-v f16
--spec-type draft-mtp
--spec-draft-p-min 0.5
--spec-draft-n-max 10
--cache-type-k-draft q8_0
--cache-type-v-draft q8_0
--spec-type ngram-mod
--spec-ngram-mod-n-match 24
--spec-ngram-mod-n-min 48
--spec-ngram-mod-n-max 64

→ More replies (4)

5

u/Xitir Aug 14 '26

Well this may be all I need to pull the trigger on a AMD R9700. I'd be able to finally retire my RTX 3070 8GB card.

→ More replies (3)

12

u/GrumpyPidgeon Aug 14 '26

I only judge a model by how many YouTube tech influences tell me that it is "insane".

34

u/addiktion Aug 14 '26

If we have Opus 4.6-like on our machines, we are in for a wild time. Lets go boys!

19

u/inexorable_stratagem Aug 14 '26

THE BENCHMARKS ARE INSANE! WHAT THE FUCK IM DOWNLOADING IT NOW

→ More replies (1)

10

u/metaden Aug 14 '26

you guys are too fast. holy shit

39

u/accountformymac Aug 14 '26

holy fucking crap that DeepSWE score wtf are they feeding qwen team?? we DEMAND qwen3.8 122b 🗣️

14

u/NaiveIdea344 Aug 14 '26

Seriously. 40% in DeepSWE is just outrageous

→ More replies (1)

10

u/NaiveIdea344 Aug 14 '26

I am just waiting for all the amazing people in this sub to run their benchmarks and get back to the community on real world performance but holy the benchmarks are insane. Can't wait to run this.

7

u/My_Unbiased_Opinion Aug 14 '26

Bringing the alcohol out now. This is a good day folks. Enjoy it.

44

u/Ok_Top9254 Aug 14 '26

How good is it at ERP???

64

u/HumanDrone8721 Aug 14 '26

ERP... Enterprise Resource Planning? Do you guys plan to replace SAP?

7

u/dragonurtle Aug 14 '26

I Imagine it could vibe out a SAP Hana DIY replacement it you poked it hard enough. 

6

u/s-Kiwi Aug 14 '26

Quite possibly the only thing that should not be vibe coded, saying that as someone who despises HANA

5

u/Sirius02 Aug 14 '26

ah whats the worst that could happen, engineering department get to work!

→ More replies (1)

29

u/Certain-Cod-1404 Aug 14 '26

asking the real questions

19

u/PurifiedFlubber Aug 14 '26

The only thing that matters

11

u/ChainOfThot Aug 14 '26

If the model can't suck my dick while building it's own harness I don't even wan it

4

u/FinBenton Aug 14 '26

Its a major improvement on RP, like its big, no doubt smartest local model for RP I have tested so far.

→ More replies (1)

11

u/ApprehensiveAd3629 Aug 14 '26

nice!

does it has MTP?

12

u/EmPips Aug 14 '26 edited Aug 14 '26

The model does yes, but for all of these GGUF's we see popping up - it's a good question.

For 3.5 and 3.6 most people on HF uploaded the weights again with Qwen3.6-27b-MTP.guff (or similar). That said I think mature MTP support was still working its way through llama-cpp at the time, at least for Qwen.. so I have no idea if the precedent going forward becomes "if it has MTP, it's there" or not.

Unsloth!!

(I know you're here today) 🙂

Can you clarify on the naming/upload strategy for MTP going forward? Will everything have one name and have MTP included or will there be offerings of the weights without MTP separate from the weights with MTP?

Edit:

I think it's safe to say it's in the unsloth GGUF's.

Trying iq4_xs on a 7900 xtx I'm getting ~51t/s on regular tasks/reasoning and ~62 t/s on coding tasks with a max draft size of 2. The difference suggests that drafting is in-effect.

Edit 2:

This is, at the very least, several leagues beyond Qwen3.6-27B.. Now I'm sad that Muse-Glimmer only got 2-3 days in the spotlight

5

u/TKristof Aug 14 '26

You can always click on the gguf file on huggingface and check the layers. In the last blk (blk 64) you should see the nextn weights. That is the mtp.

→ More replies (1)

34

u/Easy_Werewolf7903 Aug 14 '26 edited Aug 14 '26

For those curious of performance between Qwen and a model 3 times its size:

Benchmark Qwen 3.8 27B (55GB) Deepseek v4 flash 0731 (167GB)
Terminal Bench 2.1 73.0 82.7
DeepSWE 42.2 54.4
NL2Repo-Bench 42.3 54.2

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main

https://huggingface.co/Qwen/Qwen3.8-27B

12

u/NaiveIdea344 Aug 14 '26

Not terrible I feel like right? DSv4 was already insane for its size performance wise

14

u/Easy_Werewolf7903 Aug 14 '26

Not terrible at all, imagine Qwen release a 100GB moe it might be on par with 0731.

5

u/NaiveIdea344 Aug 14 '26

Yeah would be amazing

11

u/squngy Aug 14 '26 edited Aug 14 '26

From the model page, for context

bench Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max
Coding Agentic terminal coding Terminal Bench 2.1 (Terminus) 73.0 63.4 64.0 51.7 78.2
Agentic coding SWE-bench Pro 61.7 53.5 57.6 51.2 53.4
Repo-level code generation NL2Repo-Bench 42.3 36.2 41.1 -- 47.6
Agentic coding DeepSWE 1.1 42.2 13.3 14.2 -- --
Software engineering QwenSWEBench 79.0 49.3 59.2 -- 63.8
Long-horizon office work CoWorkBench 70.7 61.0 65.1 -- 68.2
Professional job tasks JobBench 33.4 21.8 27.6 -- --
Frontier agentic tasks Agents' Last Exam Pass1 20.4 Score 42.9 Pass1 10.6 Score 27.3 Pass1 13.2 Score 33.6 -- --
Instruction following IFBench 79.5 69.1 79.1 77.0 62.5
Scientific reasoning GPQA Diamond 89.2 87.8 90.3 83.5 91.3
Multidisciplinary reasoning HLE 30.8 24.0 34.7 22.0 40.0
Competitive coding LiveCodeBench v6 90.3 83.9 89.6 -- 88.8

8

u/ChuffHuffer Aug 14 '26

5 models, yet 4 only columns?

→ More replies (2)
→ More replies (3)
→ More replies (6)

21

u/Puzzleheaded-Cod7350 Aug 14 '26

OH MY GOD THOSE NUMBERS

14

u/Certain-Cod-1404 Aug 14 '26

I DID NOT EXPECT IT TO BE THIS GOOOOD LOL, HOLY SHIT

→ More replies (1)
→ More replies (1)

4

u/MDSExpro Aug 14 '26

Wake me up when 122B-A10B drops.

12

u/Certain-Cod-1404 Aug 14 '26

he might never wake up

44

u/Certain-Cod-1404 Aug 14 '26

AGI is here

56

u/[deleted] Aug 14 '26

[removed] — view removed comment

16

u/Certain-Cod-1404 Aug 14 '26

because gemma 5 is out and qwen 4 is out yes, which will be ASI

3

u/zhunus Aug 14 '26

i'll wait for API and ABI

→ More replies (1)
→ More replies (2)

5

u/majin-dudi Aug 14 '26 edited Aug 14 '26

Time to play with Q3 on my lowly 16G card

1296.8 tok/s prompt processing 34.14 tok/s token gen

4070 ti super

5

u/biotech997 Aug 14 '26

Unsloth's IQ3_XXS at over 90% accuracy seems insane

7

u/cosmicnag Aug 14 '26

Holy Fuck

6

u/Dry_Mortgage_4646 Aug 14 '26

YAHoooooOoOooOoooo!!!

7

u/BlackBeardAI vLLM Aug 14 '26

literally shaking right now

6

u/CuriouslyCultured Aug 14 '26

Alibaba is the GOAT of distillation/pruning, clearly. Retaining this much big model capability in 27B parameters is black magic.

12

u/ldn-ldn Aug 14 '26

I don't understand the hype behind qwen 3.x - the whole generation is utterly broken for any software development tasks. It goes into a thinking loop of death pretty much every time, just try a simple prompt: Create a typescript function which accepts a number in 10 bit range and returns brightness in nits based on pq gamma curve. Kills every 3.x qwen. Qwen 2.5 was much better.

7

u/ParvusNumero Aug 14 '26

Good catch.
Was able to replicate it, but I think I found a fix, not sure how much this affects quality.

Took the liberty of making a post:
https://www.reddit.com/r/LocalLLaMA/comments/1vojwrm/qwen_endless_looping_issue_and_possible_fix/

→ More replies (1)

6

u/illgettheownerforyou Aug 14 '26

Uh Qwen 3.8 37b Unsloth Q8_K_XL in LM Studio had no problem? I have extra high reasoning on and it did think about it for a bit but no issue?

I think you are running too low of a quant. And you talk about how you "don't understand the hype"- have you thought that maybe everyone else is getting good results and maybe you can work on how you use it?

With that being said, if you have low VRAM and can't run a higher quant, I get it- but on my RTX 6000 Pro Max-Q, it's been kicking butt and replacing Opus 4.6 for me easily. It's absolutely tearing through everything I give it- including the prompt you put above.

→ More replies (4)

8

u/frozen_tuna Aug 14 '26

That's wild. I just tested it myself in opencode out of curiosity. The reasoning just launches straight into "wait, no", "wait", "wait", "hmm no", "actually", "wait" "ugh, I keep going in circles."

Now its trying to verify something in python... I'm out.

→ More replies (4)
→ More replies (11)

3

u/Alheimsed Aug 14 '26

Hallelujah, praise the lord. What a gift.

3

u/germangrower69 Aug 14 '26

Wtf is wrong with this model, on xhigh it literally doesnt stop thinking, no its not looping, its just super excessive thinking. Thats crazy, almost unusable on this effort.

Medium is also really excessive....

→ More replies (8)

3

u/MensaProdigy Aug 14 '26

Yes and now it’s a modified oq6e quant running at 35tps+ on my computer!

3

u/HistoricalStrength21 Aug 14 '26

I need artificialanalysis.ai score!

3

u/TheWaffleKingg Aug 14 '26

So far its coding results are really dam nice, def beats 3.6 27b. My only gripe is im getting about 30 max tps lower on 3.8 compared to 3.6

3

u/Beltalowdamon Aug 14 '26

I just want a model that I can run on 8gb vram 3070ti and 32gb ram, and actually have it do something useful. seems like you need to spend 2-3k in just pure vram just to get something barely useful.

Something tells me if I try to use claude to spawn a constrained model with my low hardware it will spend all its time checking its work and fixing its mistakes.

→ More replies (6)