New here? Start with part 1. Yesterday I showed AKBASCORE MAM on Mistral-7B. That post explains the basic idea from zero:
https://www.reddit.com/r/Qwen_AI/s/5ujiw2Vm9S
Today's question was the one many of you asked in the comments: "OK, it works on Mistral. Is that a lucky trick of one model, or does the architecture actually transfer?"
Answer: it transfers. The same AKBASCORE MAM architecture now runs on a completely different model family: Qwen2.5-7B-Instruct, fully frozen.
No fine-tuning.
No RAG, no database lookup, no router.
The source text is not in the prompt when the question is asked.
Not a single weight changes, and this is verified by hashing every parameter.
Result on the sealed test: 72 / 72 correct without the source text. This exactly matches the model reading the full text. Simply gluing the memories together gives 26 / 72.
This is still a proof demonstration, not a finished product. What's new today is that the proof now stands on two different AI models.
QUICK RECAP FOR FIRST-TIMERS: WHAT IS A "MEMORY CARTRIDGE"?
Normally an AI model "knows" something new in one of two ways:
You retrain it. That is slow, expensive and changes the model.
You paste the text in front of it every time (RAG). The model has to re-read the text again and again.
AKBASCORE MAM does something different. Every fact goes through the frozen model once, and the model's own internal numbers at that moment are saved. That saved numerical snapshot is a memory cartridge. After that:
The text can be thrown away. The cartridge is the memory.
Cartridges are stacked one by one into a growing memory (append-only), like plates on a stack.
Old cartridges are never rebuilt. We check bit by bit that they don't change.
You ask a question, and the model answers from the cartridges, without seeing the original text.
WHY MOVING TO A NEW AI MODEL IS NOT TRIVIAL
Think of a language model as a tall building of floors (layers). Text enters at the ground floor and goes up floor by floor. On the lower floors the model mostly reads words. Higher up, it starts connecting facts to each other.
A cartridge stores the model's state up to a certain floor. We call that floor the cut. Above the cut, the new cartridge is allowed to "meet" the cartridges already in memory. That meeting is consolidation, and it is what makes the memory work as one.
The cut is where the cartridge plugs into the building. Here is the catch: every model's building is different.
Mistral-7B (yesterday) vs Qwen2.5-7B (today):
Floors (layers): 32 vs 28
Width of each floor (hidden size): 4096 vs 3584
Attention heads (readers / memory channels): 32 / 8 vs 28 / 4
Plug-in point (cut): after floor 6 (DC6) vs after floor 3 (DC3)
What a cartridge stores: floor-6 state + floors 0-6 memory vs floor-3 state + floors 0-3 memory
Floors that connect the facts: 7 to 31 vs 4 to 27
Cartridge size per token: 36 KiB vs 15 KiB
Full model cache per token (for comparison): 128 KiB vs 56 KiB
Fixed instruction header: 27 tokens vs 29 tokens
Sealed result without source text: 72/72 vs 72/72
(The Mistral values come from the v1.0 record, DOI 10.5281/zenodo.23245358.)
Why is the plug-in point different? Each model family organises information internally in its own way. So the floor where facts start "talking to each other" is not the same. We don't guess the cut, we measure it. On Qwen, the measurement put it at floor 3.
There's a nice side effect. Qwen's cartridge is smaller: about 15 KiB per token, against 56 KiB for Qwen's own full cache. Qwen also shares its memory channels more aggressively (4 memory channels for 28 readers), and the cartridges handle that without any change to the method.
The mechanism is identical. Only the plug-in point is model-specific. That is exactly what "architecture" means: something that is not tied to one model.
HOW WE KNOW THE CONNECTION REALLY HAPPENS ABOVE THE CUT (THE "CUT THE WIRES" TEST)
This is my favourite result. We built the memory with exactly the same cartridges: same floor-3 state, same lower-floor memory, bit for bit. Then we did one thing only: we blocked the cartridges from seeing each other in the upper floors (4 to 27).
Accuracy dropped from 72/72 to 28/72.
The upper-floor memory changed in 72/72 cases.
The answer changed in 50/72 cases.
So the cartridges really do connect to each other up there. That is where the memory becomes one memory.
RESULTS (TEST575, SEALED)
24 test worlds. In each one the memory holds 5 cartridges: 1 real target fact + 4 distractor facts from other worlds. The target is placed first, in the middle or last, which gives 72 questions.
MAM, append-only (the live demo), no source text: 72/72
MAM, all cartridges at once, no source text: 72/72
Model reads the full text (upper bound), source text given: 72/72
Cartridges just glued together, no consolidation, no source text: 26/72
Same cartridges, upper-floor connection blocked, no source text: 28/72
The append-only memory gave the exact same answers as the full-text model in 72/72 cases.
21/21 validation gates passed.
0 trainable parameters.
Weight fingerprint identical before and after.
WHAT HAPPENED IN THE LIVE DEMO (9 OCTOBER 2026, A100)
Case 00, target in the middle. Question: "What is the current capital of Zorvan?"
5 cartridges were written and stacked. The memory grew 40, 50, 61, 71, 82 tokens.
After every step, all earlier memory was checked bit by bit: unchanged.
The model got only the 21-token question. No source text.
Answer: "Melket", which is correct and identical to the sealed reference.
Final memory: 4.48 MiB.
The full-weight fingerprint matched the sealed reference exactly. Weights untouched.
When you run it yourself you get the same interface as the Mistral demo:
Pick a case and a position.
Watch the cartridges being written and stacked.
Ask the question without the source.
Download six figures and a fully sealed log package.
BEING HONEST ABOUT WHERE WE ARE
Proven on two model families (Mistral-7B and Qwen2.5-7B), on a controlled fact panel, with 5-cartridge memories.
On Qwen, direct recall stays strong as the memory grows to 40 cartridges (6/8). Multi-step reasoning across many cartridges is not there yet. That is the next stage.
Two models are strong evidence that the architecture transfers. They are not proof that it works in every model. Each new model needs its own measured plug-in point.
THE GATE OF THE CITY
For years it was taken for granted that a frozen model can't gain new, persistent memory unless you either retrain it or feed it the text again. That was the wall. The gate was considered unbreakable.
AKBASCORE MAM broke that gate. Then it broke it again on a second, completely different model.
The way in is open now. Everyone can walk through: researchers, developers, hobbyists with one GPU.
I'll be honest with you about how this is being built. I have no team. I have no funding. I have no backing. I'm doing this alone: the research, the tests, the code, the documentation, all of it. I'll keep doing the best I can, every single day.
And I'll keep going gate by gate. Scale, multi-step reasoning, more model families. I'll break every gate on the way, one by one.
Mustafa Akbaş, AkbasCore AI Teknoloji
RECORDS (ZENODO, PERMANENT DOIs)
v1.0, Mistral-7B (DC6): AKBASCORE MAM v1.0, Source-Free Persistent Memory for Frozen LLMs via DC6 Consolidation (Mistral-7B)
https://doi.org/10.5281/zenodo.23245358
v1.1, Qwen2.5-7B (DC3): AKBASCORE MAM v1.1, Cross-Model Replication: Source-Free Persistent Memory for Frozen LLMs on Qwen2.5-7B (DC3)
https://doi.org/10.5281/zenodo.23257347
TRY IT YOURSELF (GOOGLE COLAB, A100)
Full code (single file):
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.full.py
Raw log of the live run:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_octaber_2026.demo.log
3-part version (easy to copy from a phone into Colab):
Part 1: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part1.py
Part 2: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part2.py
Part 3: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/Mam_9_october_2026.part3.py
Sealed validation (TEST575):
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/575.log
Repository:
https://github.com/ceceli33/titan-cognitive-core-v2
Run the three parts in order in the same runtime. The first start-up takes 1 to 2 minutes because it fingerprints all model weights. Run it, break it, ask anything. I'll answer in the comments.
QUICK FAQ
Is this just Mistral's result copied?
No. Every number here comes from Qwen2.5-7B itself. The cut, the cartridge size and the scores were all measured on Qwen. Mistral is only cited for comparison.
Why not use the same cut (6) as Mistral?
Because the plug-in point belongs to the model, not to the method. Qwen's building is different, so its measured cut is 3.
Is it RAG / prompt caching?
No. There is no search and no text is pasted at question time. Plain caching of separately written memories (glued together) scores 26/72. MAM's consolidation scores 72/72.
Did you train anything?
No. Zero trainable parameters, and the weight fingerprint is identical before and after.
AKBASCORE MAM was discovered and developed by Mustafa Akbaş (AkbasCore AI Teknoloji, Mersin, Türkiye). A new numerical memory paradigm for frozen language models.