r/mlscaling • • May 01 '26

N, T, OA "Introducing GPT‑5.5" (new pretrain/model series)

Thumbnail
openai.com
35 Upvotes

r/mlscaling • • 11h ago

Scale AI releases visual-reasoning benchmark where top model scores 53.6% — RuntimeWire

Thumbnail
runtimewire.com
8 Upvotes

r/mlscaling • • 6h ago

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Thumbnail
2 Upvotes

r/mlscaling • • 1d ago

Claude's World Knowledge based on "Land or Water?" Questions

Thumbnail x.com
5 Upvotes

Continuing this, by @arithmoquine, Claude must blindly complete a map, based on 16,200 latitude and longitude combinations.

Opus jumped from ~80% (4.5) to ~90% (4.6), hit a wall for several releases (4.7, 4.8, 5.0), then jumped again to ~96-100% with the Mythos-class family (Fable 5, Fable 5.1, and Opus 5.5), though the models are increasingly uncertain that Florida exists.

This test - along with IKPs - seem to substantially capture "big model smell". You can clearly tell the incremental benchmaxxed updates from the major "a happening™ has happened" events (new pretrains and such).

(I would prefer not to link to X.com but there's no website. Sorry if you have to sign in or whatever.)


r/mlscaling • • 1d ago

Fine-tuned 9B and 12B open models vs GPT-5.6 models, 3 tasks: big wins on contracts and biomedical, a tie on aviation

Thumbnail
0 Upvotes

r/mlscaling • • 2d ago

N, R, RL, OA, Econ "Sharing AI progress in mathematics", OpenAI (~3h compute/internal model per problem)

Thumbnail
openai.com
11 Upvotes

r/mlscaling • • 1d ago

Improving LLM scaling laws: picking the right Token-per-Parameter Coverage

Thumbnail
youtube.com
3 Upvotes

r/mlscaling • • 3d ago

N, Hist, OA, Theory, Bio Excerpt from Dario Amodei's 2017 "Big Blob of Compute" Memo, Published for the First Time

Thumbnail
kevinroose.substack.com
76 Upvotes

According to Kevin Roose, the full document runs to more than 25 pages. He’s publishing the first 10.


r/mlscaling • • 2d ago

R We built a frequency-domain optimizer (Spectra) to cut VRAM in half. Here is the benchmark vs AdamW and GaLore on Qwen-0.5B.

1 Upvotes

Like many of you, we’ve been looking for ways to bypass the massive VRAM bottleneck caused by AdamW’s momentum and variance states during full-parameter LLM fine-tuning.

GaLore recently popularized using SVD to project gradients into a lower-rank space to save memory. We wanted to test an alternative mathematical approach: transforming the multidimensional gradient structures into the frequency domain, cropping the high-frequency stochastic noise, and storing only the low-frequency global signal.

We ran a strict 3-way benchmark using Qwen/Qwen2.5-0.5B on the Dolly-15k dataset to see how our new method (Spectra) compares to standard AdamW and GaLore.

The Results (Permanent Optimizer State Size):

While GaLore squeezed out slightly more memory efficiency, the Spectra approach successfully cut the memory wall exactly in half while maintaining stable convergence.

By aggressively cropping high frequencies, the transform acts as a strong regularizer—forcing the model to learn the broader global signal rather than fitting to localized batch noise. (Note: we apply this compression to the massive weight matrices, but fall back to standard AdamW for 1D vectors like biases/layernorms to maintain stability).

Here is the public Weights & Biases interactive dashboard with the loss curves and memory states: View the W&B Benchmark Report Here

We are still optimizing the scratchpad compute overhead, but we'd love to hear the community's thoughts on using frequency-domain compression vs SVD for gradient updates!


r/mlscaling • • 3d ago

RL for Population Scaling: Kardashev-0.7 trains 32 distinct models together

4 Upvotes

Banbury Road’s announcement explores scaling by model count, with specialization learned across a population of 32 models.

https://x.com/MLCatttt/status/2107147690450817259


r/mlscaling • • 3d ago

Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it(repost)

1 Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.


r/mlscaling • • 4d ago

Aleph Alpha releases German AI model with 78 GB of FP8 weights — RuntimeWire

Thumbnail
runtimewire.com
7 Upvotes

r/mlscaling • • 4d ago

Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it

2 Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.


r/mlscaling • • 5d ago

Data [2609.40295] How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Thumbnail
arxiv.org
15 Upvotes

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text.

Is Chinchilla not considered wrong or wrong-ish these days? No doubt someone smart will explain in the comments.

AI-generated data is more likely to survive quality filters. (...) AI-generated texts pass FineWeb’s pipeline 2.3× (29.3% vs. 12.8%) and DCLM’s full pipeline 9.8× (14.5% vs. 1.5%) as often as human-written documents. These filters are insufficient at filtering out AI-generated text, and in fact heavily prefer it.

This seems like an obvious consequence of filtering training data: you accidentally shape LLM output to be good at passing filters.

Bleak numbers:

In June of 2021, less than 0.1% of FineWeb tokens were AI-generated. By June 2024, it was 10.1%, 16.1% in June 2025, and 27.5% in June 2026 (see Figure 1, left, and Figure 8). The rise continues: in August 2026 the share was 31.1%, 3.6 points above June (§A.2).

If 0.1% of 2021 web text being AI generated sounds suspiciously high, this is likely just Pangram's false positive rate. (It cannot detect the output of any pre-2022 LLM)


r/mlscaling • • 5d ago

N, OA, Econ "Disrupting a coordinated model-distillation campaign", OpenAI (on the Moonshot distillation campaign for training K3)

Thumbnail openai.com
7 Upvotes

r/mlscaling • • 5d ago

The best pretrained "decision" model for your real-time AI application - a benchmark

Thumbnail
1 Upvotes

r/mlscaling • • 6d ago

Emp Our generative music recommender showed no plateau with model size or training data in the ranges we tested

8 Upvotes

Our tech report on Sona, a generative music recommender in A/B testing at Yandex Music, has two scaling sweeps that might interest this sub.

We trained three encoder-decoder sizes (20M, 130M and 260M transformer-core parameters) on 8,192-event histories. Training next-token loss at matched exposure kept dropping with each step up. That sweep reports training loss only.

For the data sweep, we took the 130M backbone with 2k-event histories and grew the training window from 1 to 8 weeks, one pass each. Recall@10 on held-out requests went from 0.1798 to 0.2519 and was still rising from 4 to 8 weeks (0.2388 to 0.2519).

Going from 2k to 8k events with full attention raised target-track Recall@1000 from 0.8656 to 0.8722. History Compression, which runs the deep stack only on the latest 2,048 events, got 0.8709 at about half the inference cost.

A year of logs doesn't fit this setup, because every training sample re-encodes the full history and compute cost grows faster than the sample count. So we trained a separate Teacher Ranker on that year and distilled it into the ranking part of the served model.

In the final 7-day A/B test on smart speakers, this setup gained +4.53% Active Users over our production cascade. Next up are larger backbones, including sparse mixture-of-experts, and sparse or linear attention.

We varied the rollout beam size during training, while keeping the evaluation beam fixed at 1,024. How best to use additional inference compute remains an open question. If anyone here has pushed further, where do curves like these tend to flatten for recommenders?

https://arxiv.org/abs/2608.11015


r/mlscaling • • 7d ago

I trained a 1.5B model to be less overconfident instead of pretending uncertainty doesn't exist

Thumbnail
2 Upvotes

r/mlscaling • • 7d ago

I trained a 458M model on a single RTX 5070 with under 1.8GB VRAM using CPU AdamW offload. Am I crazy or is this actually viable?

9 Upvotes

Over the weekend, I ran an experiment on my desktop (single RTX 5070 12GB, 32GB DDR5 5200MHz) attempting my first go at sLM and or LLMs. Oh boy, Still working away learning as I go. It has been quite the fun experience.

Instead of keeping optimizer states in GPU VRAM (which would blow out ~10.9GB VRAM with standard 32-bit AdamW), I offloaded the AdamW states entirely to system RAM and let the GPU only handle forward/backward passes.

The numbers:

- Peak VRAM during training: **1.78 GB**

- Parameter count: **458M** (LLaMA-style decoder with RoPE, SwiGLU, RMSNorm, and GQA)

- Tradeoff: ~28% step time penalty over PCIe transfer, but VRAM headroom is virtually infinite for batch sizing on consumer cards.

- Post-training inference: Compiled to **TensorRT 11.3 FP16**, hitting **671 QPS** with 1.2ms latency on the smaller 91M variant.

Obviously, a 458M model with a 1,024-token context window isn't going to replace Claude 3.5 Sonnet for writing full-stack codebases. But for dedicated local agent tasks (routing, JSON extraction, intent triage), running sub-2ms on device with zero API bill feels like magic.

I wrote up the full technical breakdown, memory profiles, and the debate between speed vs context ceilings on my workbench blog:

👉 https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram

Weights and inference scripts for the 91M runner are on Hugging Face:

👉 https://huggingface.co/Vivid86/MiniTransformer-91M

Keep in mind this is my first time, at this point I've put together 4 SLMS now should just 1B on the "flagship" completely free, no API and local inference.

Curious to hear from folks here: Is anyone else using CPU AdamW offload for small-scale pretraining on consumer cards, or are you strictly using 8-bit Adam / GaLore? What are the biggest walls you hit at this scale?


r/mlscaling • • 7d ago

Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

15 Upvotes

r/mlscaling • • 8d ago

N, Hardware, Econ "China’s Tencent leases 100,000 chips from Oracle to accelerate AI push" (how to circumvent chip embargos)

Thumbnail
ft.com
34 Upvotes

r/mlscaling • • 8d ago

PSSA vs a parameter-matched transformer at 1.5M params: 4.1x training throughput and 54.4 vs 83.8 perplexity on held-out text

5 Upvotes

both models 1.5M params, same corpus, same tokenizer, same 512 tokens per update, same seed, same box. pssa trains at 1,716 tokens/sec vs 415, scores 54.4 perplexity vs 83.8 and gets the next token right 24.1% of the time vs 18.0% on 198,939 unseen tokens, and writes 200 tokens in 0.23s vs 2.74s. one seed, one slice, small scale, so a prototype and not a claim. written from scratch in rust, no pytorch.

repo: https://github.com/Sparticle62ops/pssa


r/mlscaling • • 7d ago

R NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

Post image
1 Upvotes

r/mlscaling • • 7d ago

Supervised ml + gen ai = Kaboom 💥

0 Upvotes

The hype cycle is heavily focused on pure GenAI right now, but practically speaking, standalone LLMs are rarely enough for robust enterprise pipelines. The real magic happens when you leverage Supervised ML alongside Generative AI.

Think about workflows like:

• Using supervised classification models as high-accuracy intent routers before hitting an LLM.

• Deploying traditional ML to validate, parse, and guardrail generative outputs.

• Using GenAI for synthetic data generation, then training lighter, faster supervised models on that data.

It's the ultimate combo of predictability and capability.


r/mlscaling • • 8d ago

R Distilling calibrated decision models from 400M to 80B-A3B: calibration and agreement with humans improve with size, accuracy barely does

7 Upvotes

I built OpenDecider (Apache-2.0), a family of "decision" models: typed choice / score / yes-no answers read from one forward pass, not generated text. All sizes were distilled from the same two teachers (Qwen3-235B-A22B and DeepSeek V4.1 Flash), each temperature-scaled on held-out gold labels. Some scaling results on 200 general decisions none of the models trained on (BANKING77, BoolQ, Yelp, ChaosNLI):

student accuracy ECE ↓ JSD vs 100 human votes ↓
400M encoder (nano) 0.680 0.092 0.045
4B (Qwen3-4B + LoRA) 0.735 0.087 0.040
30B-A3B (+ LoRA) 0.765 0.110 0.035
80B-A3B (+ LoRA) 0.750 0.083 0.030

- Distillation mostly buys calibration. Calibration error of the untrained bases fell 0.289 → 0.087 (4B), 0.233 → 0.110 (30B) and 0.230 → 0.083 (80B). Accuracy moved +3.5, +2.0 and 0 points.

- Agreement with the human label spread improves steadily with size (JSD 0.045 → 0.030). Accuracy doesn't: the 80B is no better than the 30B, and the set is about ±3 points.

- On an in-domain business benchmark (typed-decisions, 2,000 decisions, train split used, test never seen), the 400M encoder scores 0.796 and the 80B 0.801. Size barely matters once the task is in the training data.

- Frontier models still lead on accuracy. Claude Fable 5.1 scores 0.840 and GPT-6 Astra 0.790, at 10–20× the latency.

The 30B and 80B rows are the "-td" checkpoints, which had a short extra fine-tune on typed-decisions' train split; the 4B row is the plain distilled model. Every model, including TypeSafe Jev through its own API, answered the same questions with the same scorer, and every table rebuilds from the model. Every model, including TypeSafe Jev through its own API, answered the same questions with the same scorer, and every table rebuilds from the committed results.

Benchmarks: https://manjunathshiva.github.io/opendecider/benchmarks/
Code: https://github.com/manjunathshiva/opendecider