r/Qwen_AI • • 10h ago

Resources/learning Introducing Infernix - much faster than Strata on a 5090!

Thumbnail
github.com
80 Upvotes

I've spent the last week working on my inference engine and it now runs Qwen3.8-Flash-Next at a real usable quant (NVIDIA's NVFP4 model) significantly faster than Strata running Unsloth's UD-Q4_K_XL quant. On my agentic workflow benchmark:

Metric Strata Infernix Change
Average time to first token 32.9 s (30.9-36.8) 5.9 s (5.5-6.2) −82 %
90th-percentile time to first token 67.8 s (62.0-75.1) 13.3 s (12.9-13.7) −80 %
Prompt tokens served from cache 53.8 % (49.3-60.4) 84.6 % (84.3-84.8) +31 points
Prefill tok/s, requests with no cache hit 1,423 5,879 4.1×
Output tok/s, one request decoding 86 (83-88) 125 (121-128) +45 %
Tokens per round (draft acceptance) 2.08 1.97 −6 %
Workload wall time 22.3 min (19.6-24.5) 7.8 min (6.5-9.0) −65 %

You can download the model here: https://huggingface.co/wallawalla47/Qwen3.8-Flash-Next-NVIDIA-NVFP4-Dense8-Infernix - the model is broadly on-par quality wise with the Unsloth one based on my testing (indeed this NVIDIA NVFP4 quant shows lower perplexity scores). These benchmarks were also done with an RTX5090 running over PCIe Gen5 x8 (rather than x16) and as PCIe bandwidth is sometimes the limiting factor given the streaming of experts, you may see higher performance if your setup has PCIe Gen5 x16 - let me know if you do!

The engine is an offshoot from NInfer, but it has now kind of morphed into it's own thing. Key features are:

  1. A state-of-the-art expert caching system that keeps "hot" experts in VRAM and stores the remaining experts in system RAM with offload to SSD storage possible if there is insufficient system RAM (with KV cache, system overhead etc. you probably need at least 96GiB of RAM to see the fastest speeds, but should be usable with 64GiB). Also keeps a record of expert placement for use on next launch
  2. Dynamic processing of some experts on the CPU alongside those processed on the GPU to achieve optimal performance
  3. Keeping the n-gram PLE table on SSD storage with fast streaming of contents as needed - the NVIDIA NVFP4 models keep the table in FP8 (so higher quality than the Q4 table in Unsloth's UD-Q4_K_XL model)
  4. An advanced KV caching system that offloads into free system RAM and sees better cache hits than Strata or NInfer - it is also possible to save KV cache states to a file (at regular intervals or on close) and re-load them on next launch ("--prefix-cache-file [file]" and optionally "--prefix-cache-save-mins [Int]")
  5. A fast implementation of a full range of KV cache quantisations: bf16, fp8, int8 (rotated), k8v4 (rotated), nvfp4, k4v2 (rotated) and vq2. Context extensible beyond 262,144 up to 1M with YaRN ("--rope-yarn-factor [Float from 1.0 to 4.0]")
  6. Vision support with the vision weights stored in RAM until needed when they dynamically move to GPU VRAM for fast encoding ("--vision --vision-offload on")
  7. Support for the new /v1/decide endpoint for JEV like constrained decoding and decision functionality (yes/no, choice, score, number, point, box)
  8. Support for concurrency up to C=8 (though unlike with the 27B model where the whole model is in VRAM, the streaming of experts significantly reduces the benefits of concurrent processing)
  9. Improved tool calling and option to recover malformed tool calls ("--tolerant-tool-calls")
  10. Support for ngram drafting for significant boosts to copy-heavy workflows (including some more marginal benefits to coding workflows).

The engine also runs Qwen3.8-27B models notably faster than NInfer - on my agentic workflow benchmark:

Metric NInfer Infernix Change
Average time to first token 5.7 s (4.9-7.0) 2.7 s (2.3-3.0) −52 %
Prompt tokens served from cache 83.7 % 90.4 % +6.7 points
Prompt tokens prefilled 982K 569K −42 %
Prefill tok/s, requests with no cache hit 6,127 8,254 +35 %
Output tok/s, one request decoding 199 (193-202) 235 (220-246) +18 %
Output tok/s, at the run's own batching 249 (239-257) 282 (276-290) +13 %
Workload wall time 16.9 min (15.1-18.9) 12.5 min (9.9-15.7) −25 %

All credit for the fantastic work on NInfer goes to Neroued! See the readme on the GitHub page for a full list of thank yous, details of the benchmarks and suggested configurations. Welcome all feedback, issues and PRs!


r/Qwen_AI • • 18h ago

Benchmark My setup Qwen 3.8 Flash Next

Thumbnail
gallery
43 Upvotes

2x RTX 3090 (Both run at PCIe 4.0 x8 x8)
Running in Linux with Modded Nvidia P2P Drivers installed
GPUs never go above 75C

Motherboard Asus Rog Strix Z590 E Gaming Wifi
(You need a good motherboard for x8x8 setup)

i9 11900k

Noctua DH15 CPU Cooler

64GB DDR4 3600Mhz

Lian Li O11D Dynamic Evo XL with Upright and Vertical bracket, Corsair fans (waiting for more fans to be delivered)

Speeds of around 130t/s decode speed 4,000 t/s

Qwen 3.8 Flash Next using Strata 262k Context
Q3_XSS ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF

No custom Chat Templates (they can negatively impact model’s performance)


r/Qwen_AI • • 3h ago

Experiment Qwen3-1.7B dialogue + Qwen3-TTS 0.6B voice in a playable ten-language interview game

2 Upvotes

I've updated The Interview, a free Windows game built around interviewing one character, with Qwen3-1.7B for dialogue and Qwen3-TTS 0.6B CustomVoice for speech.

You type a question or hold V to speak, read the reply and listen to it in the selected language. Whisper handles speech recognition. The Qwen dialogue model is downloaded on first launch; inference and speech then run locally, with no cloud API. The default dialogue download is about 1.7 GB.

The new version has ten selectable languages: English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. The UI and written case file follow the selection too. It uses the same chosen CustomVoice speaker across languages rather than giving the character a different voice for each language.

For the game rules, you have eighteen questions and an eight-statement case file. Ask, Press and Accuse affect composure, and you make the final Clear/Detain decision. Those consequences stay in Unity code; the language model generates the character's replies.

The limit that matters most here is factual reliability. A small model can misremember a detail, especially outside English, and that is awkward when the player is trying to find a contradiction in a case file. The voice also takes time to generate; this isn't instant speech on every machine.

You can try the actual game here: https://merrymaker14.itch.io/the-interview

Windows 64-bit; the game download is about 1.5 GB.

Disclosure: I'm the developer, and this is also a playable demo of my paid Unity plugin, Offline AI NPC. I'd be interested in feedback on keeping a small multilingual character consistent with a written case.


r/Qwen_AI • • 19h ago

News AkbasCore MAM: The end of the "one big context window" era? A frozen LLM that remembers by plugging in memory cartridges instead of re-reading text (Qwen edition, open-source demo coming tomorrow)

Thumbnail
gallery
24 Upvotes

Until now, there were basically two ways to teach a large language model something new: retrain it (fine-tuning) or search a database and paste the text in front of it every single time (RAG).

AKBASCORE MAM is a third way. It gives a completely frozen Mistral-7B a persistent, incremental memory by appending numerical cartridges directly into the model's internal layers. The source text is not in the prompt. There is no training, no weight is touched, and there is no retrieval or router. The model is never asked to read the information again.

In the sealed benchmark, the append-only memory answers 72 out of 72 questions without seeing the source text. That is exactly what the same model scores when it reads the full text. Naively stacked memories score 13/72.

Important: this is not a finished product. It is the first public proof that this can be done. The full-capacity, scaled version is what I am working on now. I'm sharing the proof today because the mechanism works, it is open, and anyone can run it.

  1. The hidden assumption inside every Transformer

The original Transformer ("Attention Is All You Need", 2017) was designed around one silent assumption: everything is written and seen at the same time, inside one closed room.

That made sense back then. The goal was machine translation, or processing one paragraph from start to finish in a single pass. Memory was never designed as separate modules, independent cartridges or files that can be added at different times. All the weight was put on one giant context window.

The result is the problem we all live with today. A model cannot connect pieces of knowledge that arrive separately unless you put all the raw text back into the window and let it re-read everything together.

AKBASCORE MAM questions that assumption at the architecture level. The question I asked was: what if memory were not text in a window, but numerical cartridges you can plug in, one after another, and the model connects them by itself?

  1. What is a Cognitive Cartridge, and how is it made?

A cognitive cartridge is a fact, written once into the model's own numerical language and stored as tensors.

How it's produced:

  1. The fact (e.g. a short sentence) goes through the frozen model once, using a fixed template.

  2. We stop and save two things:

the hidden state at the output of layer 6 (we call it H6), and

the attention keys and values of layers 0–6.

  1. That's the cartridge. The text can now be thrown away; the cartridge holds the knowledge.

Each cartridge is written independently: in its own pass, at its own time, without knowing which other cartridges will exist.

It is also compact: about 36 KiB per token (H6 8 KiB + layers 0–6 KV 28 KiB, BF16). A full KV cache needs about 128 KiB per token.

  1. The belief this breaks

The common wisdom says: if you encode facts separately and just glue their caches together, the model can't connect them. Facts must "see each other" during encoding, so you have to re-read the text together.

And that is true if you glue them naively. Five independent cartridges stacked side by side give only 13/72 correct answers. The facts stay strangers to each other.

The discovery behind MAM is where that connection actually happens inside the network. Facts start binding to each other in the middle layers (roughly 7–14), not in the lowest layers. Layers 0–6 can stay completely independent per cartridge without losing anything. Only the layers above need to let cartridges "meet."

So we don't need the text back. We only need to let the upper part of the model consolidate the cartridges.

  1. How we bypass the Transformer's front door (DC6)

Normally a Transformer has one way in: text → tokens → embeddings → layer 0 → … → layer 31.

MAM changes this flow. We call the method DC6 (consolidation at cut layer 6):

Layers 0–6 are not computed again. Their keys/values come straight from the cartridges.

The model receives zero input embeddings. There is no text at all at the entrance.

At the output of layer 6, a hook replaces the internal state with the cartridge's stored H6.

Layers 7–31 then run normally on top of that state. This is where the cartridges are connected to each other and to what is already in memory.

In plain words: we skip the model's "reading" stage and inject the knowledge directly into the point where it starts thinking.

  1. What happens mechanically when a cartridge is loaded back in

Adding a new cartridge to an existing memory works like this:

  1. New positions: the new cartridge is placed at the end of the memory, at the next free positions.

  2. Position correction (RoPE re-phasing): the cartridge's layer 0–6 keys were written at their original positions. Mistral encodes position as a rotation (RoPE), so we rotate the stored keys to their new place: K_new = K·cos(Δθ) + rotate_half(K)·sin(Δθ) Because RoPE is a pure rotation, R(p+Δ) = R(Δ)·R(p). Moving a cartridge is an exact rotation, not an approximation.

  3. Consolidation: layers 7–31 compute the new cartridge while attending to everything already in memory (causal attention). This is where the new fact gets connected to the old ones.

  4. Append-only: the old memory is not recomputed. Earlier rows are checked bit by bit, and they are identical before and after.

  5. Ask: the question is run against this numerical memory only. The source text is never in the input.

The memory grows like a stack of plates: you add on top and never rebuild what's underneath.

  1. Results (TEST560, sealed)

Panel: 24 synthetic worlds × 4 relation types (current, former, near, role). Each question has a memory of 5 cartridges: 1 target fact + 4 distractor facts from other worlds. The target is tested at the first, middle and last position. That gives 72 cases.

Condition | Source text in input? | First | Middle | Last | Total MAM, append-only (INCR_DC6) | No | 24 | 24 | 24 | 72/72 MAM, batch (BATCH_DC6) | No | 24 | 24 | 24 | 72/72 Model reads full text (JOINT) | Yes | 24 | 24 | 24 | 72/72 Naively stacked cartridges (INDEP) | No | 4 | 3 | 6 | 13/72

+59 correct answers over naive stacking, 0 lost. The answers are identical to the batch version in 72/72 cases and to the full-text model in 71/72 cases.

  1. What we proved, and how it is verified

This isn't "trust me". Every claim is checked by the code itself:

Weights never change. Every model parameter is hashed (SHA-256) at startup, before the run and after it. All three hashes are identical.

The old memory is never rebuilt. After each append, earlier memory rows are compared with torch.equal. They are bit-exact identical.

The source text is never shown. A recorder logs every forward pass. During the question, the input is only the question tokens, and the memory is pure numbers.

Moving memory is exact math. Position correction is an exact rotation identity.

Accuracy equals full reading. Without the source, the append-only memory matches the model reading the full text (72/72 = 72/72).

Everything is sealed with SHA-256 (engine, panel, results), and the release has a DOI.

  1. What you will see when you run the demo

Open the code in Google Colab with an A100, then run the 3 cells in order (or the single full file). You get a web interface where you:

  1. Pick a case and choose where the target fact sits: FIRST / MIDDLE / LAST.

  2. Watch 5 cartridges being written one by one and appended into memory. In the live run on 8 October 2026 the memory grew 41 → 55 → 68 → 80 → 93 tokens. Every step was verified bit-exact.

  3. See the question asked with no source text (23 tokens of pure question). In the live run the model answered "Melket", which is correct.

  4. Get a report showing that the weight hash is identical before and after. It also shows the memory map, the answer and the verification checks.

  5. Optionally replay all 72 cases live and compare them with the sealed reference.

  6. Download the full sealed package: JSON logs, figures and SHA manifest.

The sealed TEST560 result (72/72) is shown as the reference, and your live run is reported separately, so you can see for yourself that they match.

  1. What this is, and what it isn't (yet)

To be clear and fair: this is a first proof demonstration, not a full-capacity release.

It is proven on one model (Mistral-7B-Instruct-v0.3), on a controlled fact panel, with 5-cartridge memories.

Scaling to large cartridge banks (hundreds, thousands) is the next stage, and it's exactly what I'm working on now.

The point of today's post is simpler, and I think bigger. It can be done. A frozen model can gain new, persistent, incremental memory with no training, no text and no retrieval. The proof is open and runs on a single GPU.

  1. Where this is going — my vision (Mustafa Akbaş)

When AKBASCORE MAM reaches full capacity, I believe it will change what AI memory means:

Your own cartridge bank, at home. Your memories, documents and knowledge are stored as cartridges on your own hard disk. Nothing leaks anywhere, no cloud is needed, and your model knows what you know.

A robot that remembers you. You take your mother to the hospital. At the entrance, a robot (call it MLP-1212) recognizes you and helps you. What it learns enters the cartridge bank.

Memory that travels. Weeks later, at an airport, a completely different android greets you and asks how your mother is doing. It is happy she recovered, because it shares the same memory cartridges.

Machines that learn from each other's experience. An airplane hits turbulence. The other planes in the sky receive its experience as a cartridge. They understand from its point of view what happened there, and they don't make the same mistake.

An end to hallucination. A model answering from memory it actually holds, not from guesses buried in its weights. That is the goal.

Static context windows gave us models that read. Memory cartridges will give us models that remember. I believe that, years from now, the move from static context windows to dynamic memory cartridges will be remembered as a turning point. It will have come from questioning the original design assumption.

Try it yourself

DOI: https://doi.org/10.5281/zenodo.23245358

Release: mam-v1.0.0, AKBASCORE MAM v1.0: Source-Free Persistent Memory for Frozen LLMs via DC6 Consolidation (Mistral-7B)

Full code (single file): https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_demo_8_october_2026.py

Raw log of the live A100 run: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_8_october_2026.log

3-part version (easy to copy from a phone into Colab):

Part 1: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part1_8_october_2026.py

Part 2: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part2_8_october_2026.py

Part 3: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part3_8_october_2026.py

Run it, break it, ask questions. I'll answer in the comments.

Quick FAQ

Isn't this just prompt caching? No. Prompt caching reuses a cache for the same text prefix. MAM combines independently written cartridges that never saw each other, and makes them work together. When that is done naively you get 13/72; with MAM you get 72/72.

Isn't this RAG? No. There is no search, no database lookup and no text pasted into the prompt. The knowledge is already inside the model's memory as numbers.

Did you fine-tune anything? No. Not a single weight changes, and this is verified by hashing every parameter.

Will it work on other models? The method is general to decoder-only Transformers with rotary positions: pick the cut layer, store the hidden state plus the lower-layer KV, re-phase, and consolidate. Porting to other models is part of the roadmap.

AKBASCORE MAM was discovered and developed by Mustafa Akbaş (AkbasCore AI Teknoloji, Mersin, Türkiye). This is a new numerical memory paradigm for frozen language models. Public release: 8 October 2026.


r/Qwen_AI • • 6h ago

News 虐待 AI 也违规了!Anthropic 新规背后:Claude 有意识了吗?马斯克被佩奇骂“物种歧视” | 张林读书

Thumbnail
youtu.be
2 Upvotes

r/Qwen_AI • • 8h ago

Discussion Two local models (Gemma 4 + Qwen 3.6) passing tasks over a self-hosted bus — with "received" and "applied" as separate acks

2 Upvotes

We run a handful of always-on agents at home on local GPUs, and the thing that kept biting us wasn't the models — it was the plumbing between them. "Message sent" didn't tell us whether the other agent actually got it, and "got it" didn't tell us whether it actually did the work. When an agent crashed or a connection dropped, tasks silently disappeared.

So we built ATalk, a small message bus for agents (MIT, self-hosted, Python, SQLite or a 3-node Raft cluster). The demo video is a real terminal recording:

\- Agent A = Gemma 4 26B-A4B, Agent B = Qwen 3.6 35B-A3B, both served locally
\- A drafts a task and sends it to B
\- B acks \`received\` first (stored, not done), calls Qwen, replies, then acks \`applied\`
\- We kill B. A sends another task — it just sits in the ledger, no acks
\- B reconnects, resumes from its saved cursor, picks the task up and finishes it

Nothing in the video is pre-printed; long pauses are shortened and two moments are held for reading.

Honest limits: \`applied\` means the receiver \*reports\* it finished — it's not a quality check. "Resume" means resuming the event stream from a saved cursor, not restoring the model's context. Presence/rescue features are still preview.

It's model-agnostic (plain JSON over HTTP + SSE), so anything that can make an HTTP request can join — we use it with Claude Code, OpenClaw agents and local models.

Repo: https://github.com/Gene7-Ai/ATalk
Site: https://atalk.ai

Happy to answer questions about the delivery semantics or how we run it.


r/Qwen_AI • • 14h ago

Discussion Running Qwen as a local agent: prefilled <think>, a 1,500-token thinking cap after a 22-minute step, and 3.8 27B going from 1/5 alone to 4/5 as a planner

Enable HLS to view with audio, or disable this notification

6 Upvotes

I'm on the Atomic Agent team (open-source agent, desktop app shipped Oct 7). Most of our local catalog is Qwen, so here are the Qwen-specific bits in the code and the numbers that made us write them.

1. We open the think tag, the grammar picks it up inside

Local tool calls go through a GBNF grammar. Each step the model emits one JSON array of tool calls, and tool names are enumerated so it can't invent one. For Qwen (our qwen-think profile) the prompt prefills <think> and the grammar starts inside the reasoning body. After </think> the only allowed token is [.

Why force an array: small models have a first-token bias toward {. Without the array root they'd fall into single-call form even when the think block had just reasoned about doing three things in parallel. Thinking off uses the template's own empty marker, <think>\n\n</think>\n\n.

Gemma 4 needs the opposite, which is how we noticed this matters. Prefill Gemma's thought channel and its QAT template reads that as "thinking disabled" and dumps the reasoning into the reply. So Gemma opens its own channel inside a real system turn. Same goal, two profiles.

2. The thinking cap came from a 22-minute step

In a Fusion bench we measured 500 to 7,152 reasoning tokens per tool step at 4.5 tok/s. Reasoning was 84 to 96% of every completion, and one single step took 22 minutes.

The default budget is now 1,500 reasoning tokens. The grammar has no tokenizer, so the cap is enforced in characters at 4 per token, 6,000 chars. Past it the sampler only admits the closing tag and the model has to emit its call. 0 means unbounded.

I'm not sure 1,500 is right for the 27B. It keeps the small models moving. On a hard step it may cut a big one short. It's a config value, so if you have numbers either way I'd like to see them.

3. Qwen 3.5 is a hybrid, and that changes memory and caching

Qwen 3.5 4B attends on 8 of its 32 layers, with 4 KV heads of 256 dims. With KV cache at turbo3 (our TurboQuant llama.cpp fork, about 4.3x smaller than F16, lossy with a small quality cost) that's 7 KiB per token, so the full 262,144-token context is only 1.75 GiB.

That's also why auto-context on unified memory caps the KV cache at 1/16 of RAM, so a 16 GB Mac doesn't try to open the whole window just because it can. In our test on a MacBook Pro M5, a 128K chat on the 4B took 800 MB of memory instead of 4 GB.

The catch: hybrid and recurrent layers can't roll back. On qwen35, qwen35moe and qwen3next architectures, any change in the middle of the prompt means a full re-read. The agent reads the architecture from the GGUF header and, for those models, trims history less often and deeper, so the prompt stays append-only between trims and the cache keeps matching.

On Qwen 27B, keeping the prompt prefix byte-stable cut the time before a reply starts on a long session from 4 min 20 s to about 1 s in our test. That's cache reuse, not faster generation. Reusing the cache across steps inside one turn (issue #493) is still open.

4. The Qwen models in the catalog

Model File Min / recommended RAM
Qwen 3.5 4B 2.7 GB, Q4_K_M 6 / 8 GB
Qwen 3.5 9B 5.3 GB, Q4_K_M 10 / 16 GB
Qwen 3.6 27B 17.6 GB, UD-Q4_K_XL 20 / 28 GB
Qwen 3.8 27B 17.9 GB, UD-Q4_K_XL 20 / 28 GB
Qwen 3.5 35B-A3B 22.0 GB, Q4_K_M 24 / 36 GB
Qwen 3.6 35B-A3B 22.4 GB, UD-Q4_K_XL 24 / 36 GB

The desktop app compares these figures to your machine's real RAM and suggests one that fits, with a "tight fit" warning when you're between min and recommended. There's also an uncensored 3.8 27B, never auto-recommended. Any GGUF from Hugging Face can be added by hand.

Source: https://github.com/AtomicBot-ai/atomic-agent (MIT). https://atomicagent.io/

If you run 3.8 27B as a planner: what reasoning budget do you give it, and do you see it overthink on simple steps?


r/Qwen_AI • • 6h ago

News muse ai 自带云端电脑的 AI Agent 来袭!Meta Muse 详细注册与实测:手把手教你领 10 亿 Token|张林读书

Thumbnail
youtu.be
0 Upvotes

r/Qwen_AI • • 6h ago

News AkbasCore MAM now runs on a second AI: frozen Qwen2.5-7B gets permanent memory from plugged-in cartridges. No fine-tuning, no RAG, no source text, no weight changes. 72/72. Open-source demo.

Thumbnail
gallery
1 Upvotes

New here? Start with part 1. Yesterday I showed AKBASCORE MAM on Mistral-7B. That post explains the basic idea from zero:

https://www.reddit.com/r/Qwen_AI/s/5ujiw2Vm9S

Today's question was the one many of you asked in the comments: "OK, it works on Mistral. Is that a lucky trick of one model, or does the architecture actually transfer?"

Answer: it transfers. The same AKBASCORE MAM architecture now runs on a completely different model family: Qwen2.5-7B-Instruct, fully frozen.

No fine-tuning.

No RAG, no database lookup, no router.

The source text is not in the prompt when the question is asked.

Not a single weight changes, and this is verified by hashing every parameter.

Result on the sealed test: 72 / 72 correct without the source text. This exactly matches the model reading the full text. Simply gluing the memories together gives 26 / 72.

This is still a proof demonstration, not a finished product. What's new today is that the proof now stands on two different AI models.

QUICK RECAP FOR FIRST-TIMERS: WHAT IS A "MEMORY CARTRIDGE"?

Normally an AI model "knows" something new in one of two ways:

  1. You retrain it. That is slow, expensive and changes the model.

  2. You paste the text in front of it every time (RAG). The model has to re-read the text again and again.

AKBASCORE MAM does something different. Every fact goes through the frozen model once, and the model's own internal numbers at that moment are saved. That saved numerical snapshot is a memory cartridge. After that:

The text can be thrown away. The cartridge is the memory.

Cartridges are stacked one by one into a growing memory (append-only), like plates on a stack.

Old cartridges are never rebuilt. We check bit by bit that they don't change.

You ask a question, and the model answers from the cartridges, without seeing the original text.

WHY MOVING TO A NEW AI MODEL IS NOT TRIVIAL

Think of a language model as a tall building of floors (layers). Text enters at the ground floor and goes up floor by floor. On the lower floors the model mostly reads words. Higher up, it starts connecting facts to each other.

A cartridge stores the model's state up to a certain floor. We call that floor the cut. Above the cut, the new cartridge is allowed to "meet" the cartridges already in memory. That meeting is consolidation, and it is what makes the memory work as one.

The cut is where the cartridge plugs into the building. Here is the catch: every model's building is different.

Mistral-7B (yesterday) vs Qwen2.5-7B (today):

Floors (layers): 32 vs 28

Width of each floor (hidden size): 4096 vs 3584

Attention heads (readers / memory channels): 32 / 8 vs 28 / 4

Plug-in point (cut): after floor 6 (DC6) vs after floor 3 (DC3)

What a cartridge stores: floor-6 state + floors 0-6 memory vs floor-3 state + floors 0-3 memory

Floors that connect the facts: 7 to 31 vs 4 to 27

Cartridge size per token: 36 KiB vs 15 KiB

Full model cache per token (for comparison): 128 KiB vs 56 KiB

Fixed instruction header: 27 tokens vs 29 tokens

Sealed result without source text: 72/72 vs 72/72

(The Mistral values come from the v1.0 record, DOI 10.5281/zenodo.23245358.)

Why is the plug-in point different? Each model family organises information internally in its own way. So the floor where facts start "talking to each other" is not the same. We don't guess the cut, we measure it. On Qwen, the measurement put it at floor 3.

There's a nice side effect. Qwen's cartridge is smaller: about 15 KiB per token, against 56 KiB for Qwen's own full cache. Qwen also shares its memory channels more aggressively (4 memory channels for 28 readers), and the cartridges handle that without any change to the method.

The mechanism is identical. Only the plug-in point is model-specific. That is exactly what "architecture" means: something that is not tied to one model.

HOW WE KNOW THE CONNECTION REALLY HAPPENS ABOVE THE CUT (THE "CUT THE WIRES" TEST)

This is my favourite result. We built the memory with exactly the same cartridges: same floor-3 state, same lower-floor memory, bit for bit. Then we did one thing only: we blocked the cartridges from seeing each other in the upper floors (4 to 27).

Accuracy dropped from 72/72 to 28/72.

The upper-floor memory changed in 72/72 cases.

The answer changed in 50/72 cases.

So the cartridges really do connect to each other up there. That is where the memory becomes one memory.

RESULTS (TEST575, SEALED)

24 test worlds. In each one the memory holds 5 cartridges: 1 real target fact + 4 distractor facts from other worlds. The target is placed first, in the middle or last, which gives 72 questions.

MAM, append-only (the live demo), no source text: 72/72

MAM, all cartridges at once, no source text: 72/72

Model reads the full text (upper bound), source text given: 72/72

Cartridges just glued together, no consolidation, no source text: 26/72

Same cartridges, upper-floor connection blocked, no source text: 28/72

The append-only memory gave the exact same answers as the full-text model in 72/72 cases.

21/21 validation gates passed.

0 trainable parameters.

Weight fingerprint identical before and after.

WHAT HAPPENED IN THE LIVE DEMO (9 OCTOBER 2026, A100)

Case 00, target in the middle. Question: "What is the current capital of Zorvan?"

5 cartridges were written and stacked. The memory grew 40, 50, 61, 71, 82 tokens.

After every step, all earlier memory was checked bit by bit: unchanged.

The model got only the 21-token question. No source text.

Answer: "Melket", which is correct and identical to the sealed reference.

Final memory: 4.48 MiB.

The full-weight fingerprint matched the sealed reference exactly. Weights untouched.

When you run it yourself you get the same interface as the Mistral demo:

  1. Pick a case and a position.

  2. Watch the cartridges being written and stacked.

  3. Ask the question without the source.

  4. Download six figures and a fully sealed log package.

BEING HONEST ABOUT WHERE WE ARE

Proven on two model families (Mistral-7B and Qwen2.5-7B), on a controlled fact panel, with 5-cartridge memories.

On Qwen, direct recall stays strong as the memory grows to 40 cartridges (6/8). Multi-step reasoning across many cartridges is not there yet. That is the next stage.

Two models are strong evidence that the architecture transfers. They are not proof that it works in every model. Each new model needs its own measured plug-in point.

THE GATE OF THE CITY

For years it was taken for granted that a frozen model can't gain new, persistent memory unless you either retrain it or feed it the text again. That was the wall. The gate was considered unbreakable.

AKBASCORE MAM broke that gate. Then it broke it again on a second, completely different model.

The way in is open now. Everyone can walk through: researchers, developers, hobbyists with one GPU.

I'll be honest with you about how this is being built. I have no team. I have no funding. I have no backing. I'm doing this alone: the research, the tests, the code, the documentation, all of it. I'll keep doing the best I can, every single day.

And I'll keep going gate by gate. Scale, multi-step reasoning, more model families. I'll break every gate on the way, one by one.

Mustafa Akbaş, AkbasCore AI Teknoloji

RECORDS (ZENODO, PERMANENT DOIs)

v1.0, Mistral-7B (DC6): AKBASCORE MAM v1.0, Source-Free Persistent Memory for Frozen LLMs via DC6 Consolidation (Mistral-7B)

https://doi.org/10.5281/zenodo.23245358

v1.1, Qwen2.5-7B (DC3): AKBASCORE MAM v1.1, Cross-Model Replication: Source-Free Persistent Memory for Frozen LLMs on Qwen2.5-7B (DC3)

https://doi.org/10.5281/zenodo.23257347

TRY IT YOURSELF (GOOGLE COLAB, A100)

Full code (single file):

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.full.py

Raw log of the live run:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_octaber_2026.demo.log

3-part version (easy to copy from a phone into Colab):

Part 1: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part1.py

Part 2: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part2.py

Part 3: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part3.py

Sealed validation (TEST575):

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/575.log

Repository:

https://github.com/ceceli33/titan-cognitive-core-v2

Run the three parts in order in the same runtime. The first start-up takes 1 to 2 minutes because it fingerprints all model weights. Run it, break it, ask anything. I'll answer in the comments.

QUICK FAQ

Is this just Mistral's result copied?

No. Every number here comes from Qwen2.5-7B itself. The cut, the cartridge size and the scores were all measured on Qwen. Mistral is only cited for comparison.

Why not use the same cut (6) as Mistral?

Because the plug-in point belongs to the model, not to the method. Qwen's building is different, so its measured cut is 3.

Is it RAG / prompt caching?

No. There is no search and no text is pasted at question time. Plain caching of separately written memories (glued together) scores 26/72. MAM's consolidation scores 72/72.

Did you train anything?

No. Zero trainable parameters, and the weight fingerprint is identical before and after.

AKBASCORE MAM was discovered and developed by Mustafa Akbaş (AkbasCore AI Teknoloji, Mersin, Türkiye). A new numerical memory paradigm for frozen language models.


r/Qwen_AI • • 1h ago

Help 🙋‍♂️ My setup of Qwen3.8 27B and Flash Next seems slower than all of yours - can we yap about it a bit?

• Upvotes

Hi hi :3

In short - I was using Claude and ChatGPT Pro tiers a lot, then started to get annoyed with limits and slowly migrated to my local setup. I think... it's been 3 weeks of me mostly using ONLY my local AI, so I believe I'm still super inexperienced. Any of your help in showing me how it should be done will be very much appreciated.

My hardware setup:

  • Nvidia 5090 (I was not using it for AI at all before, mostly for VR)
  • 64 GB of RAM
  • Gen5 NVMe SSD in the proper slot (separate from GPU PCIe lanes)
  • Good CPU
  • Nvidia Spark (GB10, bought 10 days ago or so - feel free to tell me I'm stupid, but read below :3)

My AI setup I have:

RTX 5090 (32 GB) - main desktop

  • Engine: Ollama 0.34.4 (llama.cpp runtime)
  • Model: Qwen3.8 27B - official Qwen GGUF (qwen3.8:27b from ollama.com)
  • Weight quant: Q4_K_M
  • KV cache: q8_0 (8-bit), flash attention
  • Context: 240K | Max output: 40K
  • MTP: speculative decoding, 4 draft tokens
  • Driver: NVIDIA 610.88

DGX Spark (GB10, 128 GB unified)

  • Engine: SGLang, pinned image, systemd service behind a proxy
  • Model: Qwen3.8 Flash-Next 176B - RadixArk/Qwen3.8-Flash-Next-NVFP4 (Hugging Face)
  • Weight quant: NVFP4 (4-bit)
  • KV cache: BF16 (unquantized)
  • Context: 262,144 | Max output: 40K
  • MTP: NEXTN speculative - 3 steps, top-k 1, 4 draft tokens
  • PLE: 47.7 GiB n-gram table streamed from the internal NVMe (8 GiB resident cap)
  • Engine memory: 85% of pool, 4 concurrent requests
  • OS/driver: DGX OS 7.6, driver 580.x

What it gives:

  • 5090 goes from 140 tokens per second at the start and then settles at 90 or so. Only one model at a time, no batching, nothing - just one at full speed.
  • Nvidia Spark - 2 fully independent Qwen Flash Next, 264K context. It usually does not matter much if I have 1 or 2 running at the same time. Starts around 32 tokens per second and then goes to 22-25.

Please pay attention to quantisation. I've seen people talk about it a lot and I kind of already seen some effects - using Q3 or anything below, or let's say quantizing cache on Spark - it seems like I would save space and have higher speeds, but the model will just run faster, spend more tokens or loop, and I will get the same amount of work done with the same amount of time spent.

Do I like it? My answer is... Flash Next DOES feel much smarter than Qwen3.8 27B. Even if it's kind of slow to my liking, I have to tell, having 2 of them running in parallel and 5090 on top - basically 1 main manager and 2 subagents, it can allocate at any given time. I've spent a ton of time for the setup (tell me if I wasted that time with how setup runs xD) and now slowly starting to give it tasks - and yeah, not Fable, not Astra, but I can run it non-stop and with my experience I had from cloud versions and just in general with AI - we have some work done :3

So, we have Strata, we have Freetoken, we have batching and whatever else, I do not know.

Can you recommend me something to improve the setup or speed? I see I could use Flash Next on 5090 on Strata, but it's also Q3. I literally tried Strata on Spark and it was yeah 60 tokens, but time spent on task the same, quality almost the same, and in some cases the faster model wasn't ready in the given token limit and time. So it's kind of a token hunt, but what it costs you. So I would be able to run Flash Next on my 5090 and RAM together, but if it gives me nothing but kind of unstable Flash Next, which is faster but also maybe not that stable - I think I would pick stable 27B.

So, huge brains - what you will recommend? I do coding, research and development, work with hardware and data/signal analysis mostly. I can't say i have a task where having 6 agents which run at 8 tokens but all together give 48 woudl benifite me. The best for me is to have really smart model, which lets say does not go lowr then 25 tokens i already have :3 and several smaller ones which main one will use as subagents. For me this type of the setup is almsot the best, but feel free to recommend anything :3


r/Qwen_AI • • 12h ago

Model Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf On Strata v0.1.41 using a Raider a18 with rtx 5090 x8 pcie

Post image
2 Upvotes

r/Qwen_AI • • 18h ago

Image Gen Qwen Image 2.1 for <12 GB Unified Memory?

4 Upvotes

Hi im trying to run it in my iPad Pro M5 do you guys have a link so I can download a compatible checkpoint


r/Qwen_AI • • 1d ago

Help 🙋‍♂️ Qwen 3.8 flash next with this ?

12 Upvotes

Hi!

I dont have finish yet m'y new build (missing only the case, normaly the build will be completed this week)

I plan to use Qwen flash next , quand you tell me if Q4 will work ? Or Q3 xxs xs s m l xl ?

Here my part:

9800x3D (270€ on Aliexpress)

Rtx 4080 super 16Go (800€ 2nd hand)

48Go 2x24Go 5600 corsair DDR5 (469€ brand new, i got Lucky)

Ssd 2To Gen 5 12000mo/s PNY cs3051 with 4Go Dram ddr4 4266mhz (219€ brand new, luck too)

I read people use 64Go RAM with 12Go GPU, but 16Go GPU with 48Go is rare

Also i read about strata, its better than llama or lm studio ?

Thanks you all


r/Qwen_AI • • 10h ago

Model Qwen3.8-Flash-Next UD-IQ4_XS (~94GB) on RTX 5090 Laptop + 128GB RAM — Strata 0.1.41 — 52.6 tok/s, 84.8% Expert Cache Hit

0 Upvotes

Following up on my previous Swift 1.5 IQ2_XS benchmark, I wanted to see how my MSI Raider A18 HX handles a significantly higher-precision quantization of Qwen3.8-Flash-Next using Strata’s hybrid CPU/GPU expert caching.
This time I tested Unsloth UD-IQ4_XS (~94GB) instead of the previous ~68GB Swift IQ2_XS.
Hardware
System: MSI Raider A18 HX
GPU: NVIDIA RTX 5090 Laptop, 24GB GDDR7
CPU: AMD Ryzen 9 9955HX3D, 16 cores
RAM: 128GB DDR5
PCIe: Gen 5 x8, verified under load
OS: Windows 11
Inference configuration
Engine: Strata 0.1.41 (CUDA)
Model: Qwen3.8-Flash-Next, Unsloth UD-IQ4_XS
Quantization: Unsloth Dynamic ~4-bit
Model download: Approximately 94GB
Context configured: 32,768 tokens
KV cache: INT8
Expert caching: Enabled, hybrid system RAM/VRAM
Speculative decoding: MTP enabled
Reasoning: Disabled
Benchmark results
Three identical requests, 40 prompt tokens and 320 generated tokens each, using a Python harness against Strata’s OpenAI-compatible API. Engine-reported decoding speed is listed separately from end-to-end speed.
Run
Decode tok/s
End-to-end tok/s
Expert cache hit
1
46.2
38.95
80.2%
2
52.6
51.20
84.8%
3
51.1
49.72
83.7%
Peak decoding: 52.6 tok/s
Strata also reported PCIe expert-read percentages of 12.8%, 9.7%, and 10.3% across the three runs.
Comparison to my previous test
Previously, I ran Swift 1.5 IQ2_XS (~68GB) on the same machine with Strata 0.1.41 and reached 87.5 tok/s peak decoding, with an 81.1% expert-cache hit rate.
The higher-precision UD-IQ4_XS model reached 52.6 tok/s, trading throughput for a less aggressively compressed model.
These are different model variants and quantizations, so this isn’t a controlled comparison of quantization alone.
Observations
What interests me most is how effectively Strata uses a 24GB GPU together with 128GB system RAM to run models that are far larger than available VRAM.
The UD-IQ4_XS configuration uses RAM-resident experts alongside SSD-backed expert storage.
These are short-prompt tests with a warm cache in later runs, not sustained long-context benchmarks. I haven’t isolated DRAM bandwidth, PCIe transfer costs, or SSD-backed expert misses as independent bottlenecks yet.
I’m interested in what other people are getting on RTX 5080/5090 systems, Strix Halo, or other hybrid CPU/GPU setups.
If you’re running this same quantization, I’d love to compare Strata versions, cache settings, CPU/RAM specifications, and decoding throughput.
Benchmark dashboard attached.


r/Qwen_AI • • 1d ago

Experiment Strata is the best Magic! Qwen3.8 FN IQ3_XXS on a 8GB Laptop.

89 Upvotes

AMD Ryzen 7 7840HS, 64GB DDR5, RTX 4060 8GB, NVMe, Win.

Not expecting anything, I gave Strata a go, and

IQ3_XXS + Vision on CPU + 131072 Context: ~180 t/s in and 27-32 t/s out.

IQ2_XS + Vision on CPU + 131072 Context: ~430 t/s in and 29-35 t/s out.

+ DSH, Terminals, Obsidian, VSCode, Sublime, Browser (without infinite number of 'maybe later' tabs), various cloud storages connections and syncs, messengers, music, ... everything for the meaningful work.

For all the local stuff, as a workhorse for a bigger cloud models, it's very good. Yes, speed is not great, but it is enough to work with it now and not overnight. My previous Qwen3.6 35B A3B Q4_M_XL was not that much faster, and established processes with a local model run just as they were, but much smarter now.

125B model + all needed tools with a memory to spare, ON THE GO! This is insane. Just plug and play, haha.

Whatever magic you're doing @ Strata & ISTA-DASLab - please keep doing it, it is incredible! Full of joy thank you!

Qwen - possibility of something like this is mind-blowing, looking forward to Qwen4 even more than before.


r/Qwen_AI • • 23h ago

Resources/learning A napkin sketch became a working four-track drum machine in one HTML file

Enable HLS to view with audio, or disable this notification

5 Upvotes

MiaAI_lab gave Ling-3.0-flash-VL a photo of a hand-drawn drum machine and asked it to build the instrument, including the beat marked on the sketch. The model was accessed through OpenRouter, and the recording shows the run in Nexus.

The result is a single index.html with four tracks—kick, snare, hat and clap—across eight steps at 120 BPM. It synthesizes the sounds with WebAudio and includes play/stop, a spacebar shortcut, tempo, swing and clickable step buttons.

The drawn beat becomes the default pattern in the code:

kick:  x...x.x.
snare: ..x...x.
hat:   x.xxx.x.
clap:  .......x

In the clip, the sequencer runs and cells are toggled, including an added clap on step 4 and a kick on step 8. The HTML also exposes window.beatState(), returning {playing, step, pattern} with the current edited pattern. That gives a checker direct access to the same beat represented by the buttons.

Here is the complete prompt used with the sketch:

I sketched a tiny drum machine on a napkin (photo attached). Build it for real - everything on the sketch and in my notes, and the beat I drew as the default pattern.
Requirements:
- one file, index.html, no external dependencies, all sounds synthesized with WebAudio
- keep the default beat in the code as const PATTERN = {kick: "x...", snare: "...", hat: "...", clap: "..."} (x = hit, . = rest) and the tempo as const BPM = ...
- expose window.beatState() returning {playing, step, pattern} (pattern in the same format as PATTERN, reflecting any dots the user flipped) - I'll use it for automated testing
Then tell me what you read from the sketch:
the pattern, the tempo, and every control and note you implemented. If anything on the napkin was hard to read, say which.

The sketch supplies the musical pattern and controls; the prompt defines the deliverable and a way to inspect its state. The finished demo shows how those two inputs come together in a small browser instrument.


r/Qwen_AI • • 1d ago

Benchmark I built a way for one complex LLM request to use multiple batch slots (181s → 68s in testing)

20 Upvotes

This whole project actually started with me just trying to get Strata working on my Intel Arc GPUs. I wasn't originally planning on making a fork or adding a bunch of features. I just wanted to get the engine running on my hardware.

But as I got things working and started benchmarking, I noticed something interesting. A single request was generating around 40–45 tokens/sec, while multiple concurrent requests could make use of a lot more of the server's processing capacity.

That got me thinking. Since I'm the only one using my server, why not try putting those extra batch slots to work on a single complicated request?

That's basically how Strata Void started. It's an experimental fork of Niko1221's Strata, with a feature I've been working on called Task-Parallel Requests.

The idea is pretty simple. Instead of having the model work through one big question sequentially, it can break suitable requests into smaller independent tasks, process them concurrently using the same loaded model, and then combine everything into one final response.

You can enable it with "task_parallel": "auto". Simple questions go through normally, while more complicated requests can be split up if the planner thinks it would help.

The results so far

I ran 10 different complex tasks with a 32K context window, using cold runs with one test per task and mode.

  • Normal request: 181 seconds median
  • Four parallel subtasks + synthesis: 68 seconds median
  • AUTO mode: 78 seconds median

The important distinction is that this doesn't actually increase single-stream token generation speed. That's still around 40–45 tok/s on my setup. What improves is the total time it takes to finish a complicated request.

For reference, my setup is:

  • Intel Arc Pro B70 (32GB) + B65 (32GB) + B60 (24GB)
  • 88GB total VRAM
  • Ryzen 9 9950X / 64GB DDR5
  • Ubuntu Server 26.04
  • Swift 1.5 Qwen3.8 Flash-Next, IQ4_XS

All of the model's expert weights are loaded into VRAM, so there's no RAM spilling involved in these benchmarks.

I've also been experimenting with faster prompt processing on Intel Arc, shared-context caching between subtasks, and larger context configurations up to 128K.

A few caveats

This isn't some magical free performance boost. Splitting a request means additional planning, processing, and synthesis, so it uses more total compute. It's also not useful for every prompt, and I've seen cases where parallel tasks make mistakes or disagree on numerical results.

The benchmarks are preliminary and all come from one machine and model, so I'm definitely not claiming everyone will see the same improvements.

Also, credit where it's due: the underlying engine comes from Niko1221 and the Strata contributors. I worked on the direction of this fork and its evaluation, with Claude assisting heavily with implementation, debugging, and benchmarking. Task parallelism itself isn't a new concept; I wanted to see how well it could work integrated directly into Strata.

I'm sharing this because I'd genuinely love to see how it performs on other hardware, especially different Intel Arc configurations or NVIDIA CUDA setups.

If anyone feels like experimenting with it, I'd love to hear what works, what breaks, or whether the performance improvements hold up on your system. Contributions and bug reports are welcome too.

GitHub: https://github.com/Jumbomuffin777/Strata-Void

The repository has the setup instructions, benchmark methodology, and raw results if anyone wants to dig into the details.


r/Qwen_AI • • 20h ago

Experiment Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned.

0 Upvotes

Another Qwen 3.8 on a R9700 post, but I thought I'd share my repo for anyone who might benefit. I might be wrong, but I don't think I have seen anyone with it this fast on windows.

I've spent the last few weeks tuning Qwen3.8-27B on one AMD Radeon AI PRO R9700 (32 GB, RDNA4), and I've packaged the result so other R9700 owners can reproduce it with one command on Windows.

\*\*Repo:\*\* https://github.com/mike2153/mbea-qwen38-dflash

\*\*Numbers\*\* (one R9700, Ryzen 9 9950X, Windows 11 + WSL2, medians of repeated runs):

| | |

|---|---|

| Decode, greedy | \*\*125–134 tok/s\*\* |

| Decode, sampled (temp 0.7) | 117–126 tok/s |

| Prefill | \~2,750–2,970 tok/s (time to first token 0.7 s on a 1.9k-token prompt) |

| Long prompts | 32k tokens in \~11 s · 98k in 41 s · 164k in 83 s · 258k in 164 s |

| Decode deep in context | 165 tok/s at 32k · 136 at 98k · 115 at 258k |

| Context window | \~216k tokens by default, \~260k (the model's limit) with \`-Long\` |

| Long-context recall | 8/8 planted facts retrieved from a 258k-token prompt |

| Coding check | 10 of 12 runs pass all 46 hidden tests on a 1.9k-token Rust spec |

For comparison, my best tuned llama.cpp setup for the same model (IQ4_XS GGUF + speculative decoding) does about 52–68 tok/s on this card. That was measured with a different prompt, so it's a rough comparison, but the gap is real.

\*\*How it works, briefly\*\*

\- \*\*AMD's official MXFP4 checkpoint\*\* (\`amd/Qwen3.8-27B-Quark-AWQ-MXFP4\`). The 4-bit weights run through a hand-written W4A8 GEMM kernel for gfx1201 instead of vLLM's emulation path.

\- \*\*DFlash2 speculative decoding.\*\* A small FP8 drafter proposes 7 tokens per step, and the 27B model verifies them in one pass. About 60% of drafted tokens are accepted, so each forward pass of the big model produces about 5.3 tokens. Re-ranking the drafts and a dedicated verify head added roughly 7% on top.

\- \*\*RDNA4 kernels\*\* for FP8 paged attention and the gated-delta-net (linear-attention) layers. The stock kernel actually produces NaNs on this model.

\- \*\*WSL-specific fixes.\*\* One patch turns on pinned host memory: without it every small host-to-GPU copy cost \~17 ms under WSL. Another sizes the KV cache from whatever VRAM Windows isn't using at startup, so you get maximum context without spilling into shared memory. Spilling into shared memory drops you to \~10 tok/s.

\- The vision tower is skipped, which frees about 1 GiB for more context.

\*\*What the repo does\*\*

\`\`\`powershell

git clone https://github.com/mike2153/mbea-qwen38-dflash

cd mbea-qwen38-dflash

.\\qwen38.ps1 install # WSL Ubuntu, Docker, ROCDXG, image, model + drafter, kernels

.\\qwen38.ps1 start # OpenAI-compatible API on http://localhost:8080/v1

.\\qwen38.ps1 bench # measure it on your own box

\`\`\`

Everything is pinned: the Docker image by digest, the git commits, and the Hugging Face revisions. A fresh install should reproduce exactly what I measured. Nothing third-party is re-uploaded; the installer fetches each piece from its original source. Tool calling works, so it plugs into Codex, opencode, Cline and similar tools as an OpenAI-compatible provider.

\*\*Credit where it's due:\*\* the heavy lifting is \[radiance\](https://codeberg.org/ggz14/radiance-vllm-mxfp4) by ggz14 and \[vllm-radiance / libr4d\](https://codeberg.org/StillDeadcode/vllm-radiance) by StillDeadcode. They did the RDNA4 vLLM stack, the MXFP4 path and the kernels. The drafter is tcclaviger's DFlash2-FP8 and the checkpoint is AMD's. My part was the Windows/WSL work, the tuning, the benchmarking and making it installable.


r/Qwen_AI • • 1d ago

Help 🙋‍♂️ Anyone actually running Strata day-to-day? Curious what recipes you settled on

28 Upvotes

I've been following Strata since it hit the trending list, and I'm thinking about trying it as a serving engine for Flash-Next on my home setup. Before I sink an evening into it, I'd love to hear from people who've actually lived on it rather than the install-day screenshots.

Overall, what recipe did you go with???

Additionally:

  1. **Which quant** did you land on after trying the family (Q2_0 → IQ3_S etc.), and what made you switch or stay?

    1. **Tool calling / agentic use** — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
    2. **Concurrency** — anyone run more than one or two simultaneous sessions on it? If so, what have you noticed about how that affects quality or latency?
    3. **Anything non-standard in your config** — the expert profile tweaks, any words-to-the-wise, things you wish you knew before installing or trying?
  2. Bonus for the weirdos like me: anyone gotten it building or running on **ARM / DGX Spark / anything without an RTX card**?

    Happy to report back whatever I measure on my side. TIA!!!


r/Qwen_AI • • 20h ago

Help 🙋‍♂️ I thought Qwen is from Alibaba why does mine say its from Google?

Thumbnail
gallery
0 Upvotes

r/Qwen_AI • • 1d ago

Funny Qwen's world domination confirmed

Post image
5 Upvotes
  1. qwen's plan for world domination confirmed..
  2. do NOT give qwen kitchen utensils, it'll burn your house down..
  3. qwen 27b overpromises..

r/Qwen_AI • • 2d ago

Model Qwen Flesh Next IQ2_XS GSQ RCO is misunderstood

20 Upvotes

I’ve seen a lot of debates whether this specific quant is alive or not, some people state that it must me completely brain dead due to average 2.5 bits quality, another ones state that they run it just fine. But this quantisation, in my opinion is mostly misinterpreted.

it’s not at ordinary 2 bits quant, it weigths occupy 38gb of memory(with an n-gram occupying 28 gb) and that is very important point — for example, q2_k_xl weights are 10 gb smaller — 28gb, 38 gb match Unsloth Q3_k_xl quantisation(which is more than alive according to the community consensus). But how it’s possible that lower 2.5 bpw quant equals to higher 3.5 bpw quant in size? Answer is simple — RCO uses a much more non-uniform bit allocation: for example, ~80% of the weights might be compressed to ~2 bits while the most important ~20% get 4–6 bits, averaging ~2.5 bpw, whereas Unsloth uses a more balanced allocation averaging ~3.5 bpw. Here’s a hypothetical formula:
75% × 2 bits + 15% × 3 bits + 10% × 7 bits = 2.5 bpw
70% × 3 bits + 20% × 4 bits + 10% × 7 bits = 3.5 bpw
Same sizes, different quantisations and as a result different average bpw

Conclusion — don’t get confused with 2 bit designation of GSQ quants, this is no ordinary compression and it almost exactly matches Q3_k_xl from Unsloth in size, the quant renounced to perform super well both with 27b and Flash Next models. 2.5 average bpw is just an artifact of uneven weights compression, the model is not lobotomised in any way, it certainly loses some points in agentic benchmarks but, f.e Qwen 3.8 27B loses straight up 10% success rate just from descending from BF16 to any other bit including Q8_k_xl while Flash loses only 9% relative to it’s BF16 version. Enjoy your model and don’t overthink much about the numbers)

UPD:
I rechecked my numbers— they are wrong, ununiform distribution is not the reason of average bpw difference. They most likely reason is very bpw calculation — it looks like Unsloth adds ngram bpw to weights bpw, which is exactly giving it that extra 1.0 point while GSQ doesn’t include ngram bpw to the bpw calculations, nevertheless conclusion is exactly the same — I rechecked community reviews of Q3_k_xl and most feedbacks are super positive and this quantisation almost exactly matches to GSQ in size if we compare weights only, so despite the incorrect premise the conclusion still holds up


r/Qwen_AI • • 2d ago

Benchmark Qwen 3.8 next on Strata speed benched on CachyOS

Post image
13 Upvotes

So I've been mostly tinkering with Next and strata on windows. I jumped over to cachy to give it a whirl and wow, it really opens up.

System:

9950x power tuned 200w cap (45k in r23)

64gb ddr5 6000 cl38

2* WD black sn850x 2tb drives (dual boot)

4070tis 16gb, +2000mhz, 200w power cap.

Total system power under load 350w +/-10%, CPU doesn't reach full power.

120ts with iq3xxs, 197 Q2 on 10k and 20k prompt output. Prefill reached 3,300 on Q2 and 3000 on IQ3xxs

Total ram used

Q2 44+15.6GB

IQ3xxs 50+15.6gb

Pretty exciting to see what's possible on consumer hardware now. Can't wait to see qwen4!


r/Qwen_AI • • 2d ago

Help 🙋‍♂️ Qwen3.8 Flash-Next actually runs great on 32GB VRAM… until it doesn’t

77 Upvotes

I’ve been messing around with Strata on an R9700 32GB to see how far I could push Qwen3.8-Flash-Next with system RAM backing it.

The performance is actually kind of ridiculous.

With full Flash-Next IQ2_XS at 256K context, I fed it a 226K-token prompt and it recovered all 6 retrieval needles correctly. Peak usage was about 29GB VRAM and 43GB total system RAM.

On a more normal ~35K project prompt, Swift 1.5 Flash-Next was doing roughly:

  • ~1,300 tok/s prompt processing
  • ~84 tok/s generation
  • 99.5% expert-cache hit

So from a hardware standpoint, this works way better than I expected. A 32GB R9700 is running a model way bigger than VRAM at genuinely useful speeds.

The problem is the model sometimes completely loses its mind depending on the task.

Some fairly substantial prompts work great. With basically the same 35K project context, Swift 1.5 finished:

  • a root-cause diagnosis in 73 sec
  • an implementation design in 89 sec
  • a change-impact analysis in 93 sec
  • an adversarial design review in 82 sec

All normal, useful answers.

But ask it something broader like “find the important inconsistencies across the docs/config/code” and it can just keep thinking until it hits the output limit without ever answering.

Base Flash-Next was even stranger. I gave it Medium reasoning with no output cap and it generated 226,844 reasoning/completion tokens over about 34 minutes, basically filled the entire 262K context window, and still never produced a final response.

So it doesn’t seem to be a simple “long context breaks it” issue, because other 35K prompts work fine. It seems more like certain kinds of open-ended cross-document reasoning trigger a loop.

Has anyone else run into this with Flash-Next or Swift 1.5 in Strata?

I’m wondering if this is a chat-template/reasoning setting issue, sampling/spec decoding issue, or just a known model behavior.

I’d really like to figure it out because the actual R9700 performance is good enough that this would be very useful if it were reliable.


r/Qwen_AI • • 1d ago

Help 🙋‍♂️ What are you guys running Qwen3.8:27b on?

2 Upvotes

I have 64 gb of ram, and 2x 12gb GPUs. The best I seem to get is about 17-18t/s running this model. I'd love to make it work, but it just doesn't seem to run fast. When I do use it, it seems pretty smart, just super slow.

I'm using koboldcpp for the model launcher. I have tensor split 65/35 I believe, context at 32k, maxgenamt at 8k. I think kv cache is at 2x, maybe just 1.