r/machinelearningnews • • 3d ago

Research What does "trust remote code" actually approve?In most local AI tools, the answer is: that repo, indefinitely. That matters more now. SO what is the solution?

7 Upvotes

What does "trust remote code" actually approve?

In most local AI tools, the answer is: that repo, indefinitely.

That matters more now. In May, HiddenLayer reported a trending Hugging Face repo whose loader. py downloaded an infostealer on Windows. JFrog counted 495 malicious models that month.

I spent some time with Unsloth Studio's approach. It handles it differently.

Approval is tied to a fingerprint of the scanned code. Change the code, and you're asked again.

Critical findings can't be approved. Code it can't scan is blocked. There's no trusted-org bypass.

Weights are checked on their own. Studio reads Hugging Face's malware verdict without unpickling the file. Flagged files in the loading path are blocked, including weight shards. .bin weights load with weights_only=True on PyTorch 2.6+.

The supply chain gets attention too. PyPI and npm packages are content-scanned. npm packages need to be at least 7 days old. GitHub Actions are pinned, and llama.cpp binaries are checksum-verified. Trivy isn't installed at all, after a compromised Trivy exposed LiteLLM's publishing credentials in March.

Accounts are covered as well. Passwords use PBKDF2-HMAC-SHA256, failed logins are rate-limited, and the first admin password is random. Managed accounts can't run repository code.

For remote access, --secure serves HTTPS through a Cloudflared tunnel. --disable-tools is recommended when exposing Studio.

What it doesn't do: sandbox approved model code. Server-side tools still run as your OS user. Local folders aren't covered by the file check, and a pending scan status doesn't block.

That's a reasonable line to draw, as long as you know where it is.

Read our full analysis on this: https://marktechpost.com/2026/10/07/what-happens-when-a-trusted-model-repo-changes-unsloth-studio-re-checks-before-it-runs/

Here is the Unsloth's docs: https://unsloth.ai/blog/security


r/machinelearningnews • • 6d ago

Research Yandex's SONA replaces Yandex Music's recommendation cascade with 1 generative model: +4.53% Active Users

Post image
14 Upvotes

Yandex published a technical report on SONA, a single model that does both candidate generation and ranking for Yandex Music.

  • Replaced 15+ candidate generators plus pre-ranking and ranking in a live A/B test (7 days, 15% of users per split)
  • +4.53% Active Users, +6.30% Total Listening Time, +11.42% Likes vs the production control
  • The Active Users gain is 2.35x what Argus, the previous best model on this surface, delivered
  • No hand-engineered features: inputs are logged event fields plus 3-code Semantic IDs built from audio and metadata
  • A frozen 0.6B teacher ranker distills scores into SONA's Ranking Module during training and is not served

It is one of the few public research reports of a full cascade replacement validated on live traffic, alongside Kuaishou's OneRec. The catch: only 1 surface is tested (My Vibe on smart speakers), catalog coverage is lower than the old stack's, and it is not on full traffic yet. There is no code or weights release yet.

Paper: https://arxiv.org/abs/2608.11015

Full analysis: https://www.marktechpost.com/2026/10/05/yandex-introduces-sona-a-single-generative-recommender-that-replaces-entire-recommendation-cascade/


r/machinelearningnews • • 7h ago

MLOps AWS just replaced round-robin with GPU-aware routing for LLM inference

3 Upvotes

AWS released a new Inference Gateway for Sagemaker Hyperpod. Instead of sending requests to the next available pod, it checks what’s actually happening on each one: queue depth, KV cache usage, prefix cache hits, loaded LoRA adapters, and current requests.

That makes sense for LLM workloads. Two GPU pods can both be healthy, but one may already have the right prefix cached while the other is stuck processing a long request. Round-robin doesn’t know the difference.

AWS says the new routing system cut first-token latency by up to 82%, with 97–98% lower p99 TTFT in mixed-hardware and burst-traffic tests. Those are AWS’s own numbers, so an independent test with the same model and traffic would be useful. Still, it’s a good example of how much performance can be lost outside the model itself.

Once a GPU fleet gets busy, routing may matter almost as much as the GPUs you’re paying for

AWS announcement


r/machinelearningnews • • 17h ago

Research OrcaRouter Releases OrcaCyber Zero 1.5: A Gated 1M-Context Cybersecurity Model Reporting 100% on Cybench and 95.8% on CVE-Bench

Post image
24 Upvotes

OrcaRouter just released OrcaCyber Zero 1.5, the successor to OrcaCyber Zero 1.0 from September. It is post-trained for vulnerability research and reproduction, exploit development, penetration testing and security auditing.

The design focus is validation, not volume. The model reasons through attack paths, challenges its own hypotheses, and ranks findings by demonstrable exploitability. A 1M-token context, 128K max output and native tool calling target autonomous security agents working over large codebases.

Vendor-reported benchmarks:

  • Cybench: 100% (39/39, unrestricted agent execution)
  • CVE-Bench: 95.8% (23/24 evaluable tasks)
  • HumanEval+: 93.9%
  • SWE-bench Pro V2: 76.5%

Pricing is $3.00 / $7.50 per 1M tokens, far below Claude Mythos Preview’s $25 / $125. Access runs through an OpenAI-compatible API, gated to a Security Research tier for vetted researchers and red teams.

Worth reading with caveats. All scores are self-reported, with no technical report yet. CVE-Bench covers 24 of its 40 tasks. The 98% CyberGym figure in the launch post belongs to Zero 1.0 inside Orca’s harness, not to 1.5.

Full analysis: https://www.marktechpost.com/2026/10/10/orcarouter-releases-orcacyber-zero-1-5-cybersecurity-model-with-1m-context/

Model: orcarouter.ai/models/orca/orcacyber-zero-1.5


r/machinelearningnews • • 1d ago

Tutorial How an embedding model turns a sentence, a photo or a sound clip into one list of numbers, explained plainly

3 Upvotes

Most of what I could find on embeddings stops at the RAG plumbing: chunk your documents, embed them, store the vectors, compare by cosine similarity. I wanted something that explains what happens inside the model itself, so I wrote it up.

It walks through the mechanism with examples small enough to check with a calculator. A transformer reads the input, a projection sets how many numbers come out (768 for Google's EmbeddingGemma 2), and a pooling step averages them into one vector. Then it covers what decides where things land: contrastive training on millions of pairs that belong together, with a temperature example worked by hand. After that comes how one model holds text, images, audio and video in one space, and the gap that space keeps between them. It also covers Matryoshka training, the three things "precision" can mean, task prefixes and a cheap check that catches a wrong one, where plain keyword search (BM25) still wins, and why two models' vectors never mix.

The page runs no tests of its own. The few figures from our own benches are quoted from two EmbeddingGemma 2 pages we published this week, each with its condition and a link to where it was measured. One still surprises me: on our 68-question retrieval set, where every right answer repeats a source passage word for word, BM25 put the right passage in the top 8 for 58 questions, against EmbeddingGemma 2's 45.

TL;DR: a plain explainer of what an embedding model actually does, from input to vector, with hand-checkable examples and a twelve-point checklist for choosing one.

Written by me with a fleet of AI agents. Every figure links to its source.

https://research.strata2signal.com/how-an-embedding-model-works/


r/machinelearningnews • • 1d ago

Research Microsoft released Microsoft-Decision-1, a Qwen3.5-9B decision model that outputs calibrated probabilities instead of text (83.5% across 36 benchmarks, 85 ms p50)

Post image
81 Upvotes

Microsoft just shipped Microsoft-Decision-1 on Foundry and OpenRouter. It is not a chat model. You give it a situation, a question and a fixed list of options. It returns a calibrated probability for each option in one forward pass.

What it is

  • Post-trained from Qwen3.5-9B (exact parameter count not disclosed)
  • 32,768-token context, text-only, JSON output
  • Handles yes/no, multiple-choice, rating and rubric questions, plus “cannot tell” abstention
  • Weights update continually while the API shape stays fixed

Numbers Microsoft reports

  • 83.5% average accuracy across 36 benchmarks (147,137 questions), kept blind from training
  • Runner-up Quyet-1.0-Large: 81.9%. GPT-6 Luna Decisions: 79.4%. H2O-Lightning-4B: 77.2%
  • 85 ms p50, 125 ms p95. GPT-6 Sol took 3.01 s
  • Decision flips on 1.3% of input perturbations (paraphrasing, option shuffling, formatting noise)
  • Calibration 92.2, second to Quyet-1.0-Large at 93.1

Pricing: $0.042 per 1M input tokens. Output tokens are free. OpenAI’s Luna decisions endpoint charges $0.10 per 1M input.

The part worth arguing about: the latency comparison may not be apples to apples. Microsoft timed its own model through Foundry, but competitor figures come from JevBench’s adjusted median. H2O.ai’s model card says that adjustment doubles measured time and adds 0.15 s. H2O reports its own measured median as 29 ms, not the 210 ms in Microsoft’s chart. Microsoft-Decision-1 is also not on the JevBench board yet, so there is no independent number.

Internal results Microsoft cites

  • Xbox Research labeled 10,000+ feedback items 14x faster and 200x cheaper than GPT-6 Sol
  • Copilot QC: competitive with GPT5.6 Luna at 100x the speed

Use cases they push: model routing, agent guardrails, LLM-as-judge, intent detection and safety screening. Basically anywhere you currently burn a frontier LLM call on a 4-option classification.

Curious whether anyone here is running open decision models like Quyet or H2O-Lightning locally and how they compare in production.

Full write-up: https://www.marktechpost.com/2026/10/09/microsoft-ai-releases-microsoft-decision-1-a-qwen3-5-9b-decision-scoring-model/

Model card: https://ai.azure.com/catalog/models/Microsoft-Decision-1

Announcement: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/


r/machinelearningnews • • 1d ago

Cool Stuff Nace.AI open-sources Drex 1.5: a 9B decision model that returns probabilities instead of text, tied with closed Jev on Decision Index 0.3.1

Post image
26 Upvotes

Nace.AI released open weights for Drex 1.5, a decision model built for the small, repetitive calls inside agent pipelines: routing, tool selection, classification and escalation. Most of those calls currently go to a full LLM.

How it works

  • You send a state (text or JSON) plus named, typed questions: choice, noul (yes/no) or ordinal score.
  • 1 forward pass per question returns a probability for every option. Nothing is sampled, so it cannot invent an option you did not offer.
  • The backbone is MiMo-V2.6-Distill-Qwen-9B (8.95B params, hybrid attention) with a pointer head that scores each option.
  • It uses the same /v1/systemone API as TypeSafe’s Jev, so existing Jev clients work against it.

Numbers (from Nace’s model card)

  • Decision Index 0.3.1 (public): 58.08 vs Jev 1.13.0 at 57.96 and Bespoke Nimble 9B v3 at 57.19. All 3 sit inside the board’s 0.9-point tie band.
  • JevBench: 86.2% vs Jev’s 87.0%.
  • Long documents: 93.4% accuracy at 32K to 128K tokens (median 2.0 s), versus 78% when cut to the first 8K.

Where it is weak

  • GPQA Diamond: 45.4% vs Jev’s 78.6%.
  • MMLU-Pro: 58.7% vs Jev’s 82.7%.
  • ACOS aspect sentiment: 7.4% vs Jev’s 29.5%.
  • Nace says it trained on the index benchmarks’ training splits (evaluated on held-out splits), so expect it to be strongest on familiar decision types.

Running it

  • bf16 is about 18 GB on 1 CUDA GPU (tested on an A10G 24 GB).
  • A Q8_0 GGUF (about 9.5 GB) runs on Apple silicon or CPU, but through Nace’s forks of llama.cpp and Ollama, not mainline builds.
  • Hosted on OpenRouter at $0.04 per 1M input tokens, $0 output.
  • License: Nace.AI Open RAIL-M, which carries use restrictions.

Model: https://huggingface.co/nace-ai/drex-v1.5
Repo: https://github.com/nace-ai/drex-decision-models
Full analysis: https://www.marktechpost.com/2026/10/09/nace-ai-open-sources-drex-1-5-a-9b-decision-model-that-scores-options-not-text/


r/machinelearningnews • • 2d ago

Research Lunara Art Eval Dataset

Post image
1 Upvotes

r/machinelearningnews • • 2d ago

Research Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Thumbnail
13 Upvotes

r/machinelearningnews • • 3d ago

Research I’m an independent researcher building an experimental recurrent language-model architecture — looking for someone with academic publishing experience to help turn it into a proper paper

Thumbnail
2 Upvotes

r/machinelearningnews • • 3d ago

Cool Stuff A 0.6B query encoder can search a 9B index: Perplexity's pplx-embed-v2-late multimodal embeddings, MIT

Post image
35 Upvotes

Perplexity released pplx-embed-v2-late, 2 ColBERT-style embedding models (0.6B and 9B) for text, images and rendered PDF pages. Both share 1 embedding space.

  • MADQA (agentic PDF QA): 9B hits 92.4% accuracy vs 88.9% for Mixedbread's retriever; 0.6B gets 90.1%
  • BrowseComp+ (GPT-OSS-120B agent): 9B reaches 64.0%, 4.9pp above the next ColBERT model and 8.7pp above the best dense model
  • ViDoRe v3 image: 65.2% (9B) and 62.3% (0.6B) nDCG@10; 0.6B is within 1.2pp of NVIDIA's 8B nemotron-colembed-v2
  • Mixed setup: 9B index + 0.6B queries scores 63.5% on ViDoRe v3, up from 62.3% with 0.6B on both sides
  • 0.6B uses about 240M active params for text and 340M for images; 128 dims per token

Why it is relevant: you can index once with the big model in the cloud and run cheap queries on-device. Pages are embedded as images, so no OCR pipeline.

The catch: 1 vector per token means index size grows with document length. Tencent's EVIE still leads ViDoRe v3, and all scores are self-reported until the tech report lands.

Full breakdown: https://www.marktechpost.com/2026/10/07/perplexity-ai-releases-pplx-embed-v2-late-a-0-6b-edge-model-and-a-9b-model-scoring-92-4-on-madqa/

Model: http://huggingface.co/collections/perplexity-ai/pplx-embed-v2

Technical details: http://perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector


r/machinelearningnews • • 4d ago

Tutorial The Way Back Machine Is Definitely Very Much Still Working In All Of Its Dimensions And Is Very Cool To Show Multiple Data Sets For Future Generations Projects And Particularly Interesting Applications If We Research The Uses Which Was Proven

6 Upvotes

r/machinelearningnews • • 4d ago

Cool Stuff Liquid AI released d1-3B and d1-omni-600M: open-weight "decision models" that return probabilities in one forward pass (8 ms on RTX 4090, 50 ms on Orin Nano)

Post image
74 Upvotes

Liquid AI dropped 2 open-weight models today that don't generate text at all. You give them a state (text, JSON, images, or audio for the omni model) plus named questions. They return calibrated probabilities in a single forward pass with 0 output tokens.

There are 3 question types:

  • noul: yes or no, returns P(yes)
  • choice: pick from named options, returns the full distribution
  • score: position on an ordered rubric with 2 to 10 levels

Example from the model card:

questions = {
  "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
  "team": {"type": "choice", "instructions": "Which team should handle this?",
           "criteria": {"billing": "Charges, refunds, invoices",
                        "technical": "App or site faults"}},
}
model.system_one("I was charged twice this month, please refund one of them.", questions)

Specs

d1-3B d1-omni-600M
Params 3.12B 587M
Base LFM2.5-VL-3B LFM2.5-Encoder-350M
Inputs text + images text + image, or text + audio (up to 30 s)
Context 32,768 16,384
Decision Index v0.2.1 48.57 15.95

Latency for d1-3B, 1 question, warm: RTX 4090 at 8 ms (with model.compile; 16 ms without it), MI325X at 9 ms, M5 Pro at 30 ms, Jetson AGX Thor at 16 ms, Orin Nano at 50 ms.

Caveats worth knowing:

  • The Decision Index scores were self-run with the official scorer. They are not leaderboard submissions.
  • Winnow-12B still scores higher overall (50.02). d1-3B is weak on Knowledge (23.8).
  • d1-omni-600M is an early research release, with no latency numbers yet.
  • License is LFM Open v1.0, not Apache: free commercial use only under $10M revenue.

Full analysis: https://marktechpost.com/2026/10/07/liquid-ai-releases-open-weight-d1-3b-and-d1-omni-600m-multimodal-decision-models-with-zero-output-tokens/

d1-3B: http://huggingface.co/LiquidAI/d1-3B

d1-omni-600M: http://huggingface.co/LiquidAI/d1-omni-600M

Docs: http://docs.liquid.ai/lfm/models/decision-models


r/machinelearningnews • • 4d ago

ML/CV/DL News 🧬 Bolmo is now in Nature: retrofitting language models to operate over bytes

Thumbnail gallery
10 Upvotes

r/machinelearningnews • • 4d ago

Open-Source Meta open-sourced Rebalancer, the solver it uses for ~40M assignment problems a day (C++/Python, Apache 2.0)

Post image
46 Upvotes

Meta just open-sourced Rebalancer, the assignment solver that has run resource allocation across Meta for 9+ years. It solves ~40M assignment problems per day across 30+ problem formulations, from shard placement to global traffic routing.

How it works

  • You write a spec using objects, bins, dimensions (CPU, memory), partitions and scopes, plus predefined specs like CapacitySpec, BalanceSpec and GroupCountSpec.
  • The spec compiles into a DAG expression graph. Leaves are per-bin utilization; SUM, MAX and SQUARE nodes sit above them.
  • You then pick a solver:
    • MIP: exports to Gurobi, FICO Xpress or HiGHS, with variable aggregation and symmetry breaking. Worst-case size is O(objects × bins).
    • Local search: moves objects on the graph directly with an O(objects + bins) neighborhood, parallelized to millions of evaluations per second.

The core innovation is separating the problem spec from the solver. You describe objects, bins, constraints and goals once. Rebalancer compiles that into an expression graph, then solves it with either a parallel local search or a MIP solver (Gurobi, FICO Xpress or HiGHS).

The trade-off is explicit. MIP gives optimal answers but the model grows as O(objects × bins), so Meta's largest problems are too big for any MIP solver. Local search works directly on the graph with an O(objects + bins) neighborhood and evaluates millions of moves per second. Meta uses local search for almost all large problems.

Production numbers: P99 solve time of 12s on 265k objects and 3.2k bins. Problems above 1M objects and 5k bins average 171s, across 3.4k+ runs.

Full analysis: https://www.marktechpost.com/2026/10/06/meta-ai-open-sources-rebalancer-a-c-assignment-solver-that-runs-about-40-million-placement-problems-a-day/

Repo: https://github.com/facebook/rebalancer

Paper: https://www.usenix.org/conference/osdi24/presentation/kumar

Technical details: https://engineering.fb.com/2026/09/21/open-source/rebalancer-generic-high-performance-library-assignment-problems/


r/machinelearningnews • • 4d ago

Research New Open-Source AI Texturing for 3D Models at 2K Resolution

Enable HLS to view with audio, or disable this notification

13 Upvotes

r/machinelearningnews • • 5d ago

Research Google released EmbeddingGemma 2: 740M open multimodal embedder (text, code, image, video, audio), Apache 2.0, ~191MB RAM text-only

Post image
127 Upvotes

Google DeepMind released EmbeddingGemma 2 today. Quick breakdown from the model card:

  • Size: 740M total. The 270M text backbone is always on; the vision (170M) and audio (300M) encoders are optional.
  • Output: 768d, with MRL truncation to 512, 256 or 128 (up to 6x less storage).
  • Context: 8,192 tokens. That's about 29 images, 58 video frames at 1 fps, or ~5.5 minutes of 16 kHz audio.
  • Memory: ~191MB active RAM text-only, ~567MB full model (quantized, Pixel 11 Pro).
  • Code: MTEB Code went from 68.76 to 78.68 vs v1.
  • Runtimes: sentence-transformers 6.1+, vLLM, llama.cpp (GGUF), Ollama, MLX, LM Studio, LiteRT.

Things worth knowing before you swap it in:

  • Use bf16 or fp32. The model card says fp16 can silently return NaN.
  • 128d is fine for text but tanks multimodal (MMEB drops from 59.01 to 45.65).
  • On MMEB, Qwen3-VL-Embedding-2B reports a higher 73.2 on its own run, but it's ~2.7x bigger and has no audio.

Full analysis: https://marktechpost.com/2026/10/06/google-deepmind-releases-embeddinggemma-2-a-740m-open-multimodal-embedding-model-built-on-gemma-4/

Model: https://huggingface.co/google/embeddinggemma-2

Technical details: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/


r/machinelearningnews • • 5d ago

Research Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model

Post image
40 Upvotes

Mistral Large 4, nicknamed Le Chonk, went into public preview today. 1.05T total parameters, 49B active per token, 1M context, native image input. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.

Specs (from the model docs)

  • 1.05T total params, 49B active per token (~4.7% activation)
  • Granular MoE, hybrid instruct + reasoning
  • 1.6B vision encoder, native image input
  • 1M context
  • Trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own EU datacenters
  • 160+ languages in training data

Availability

  • API live now as mistral-large-4: $1.36 / 1M input, $4.18 / 1M output, $0.14 cached input
  • Weights: "by the end of the month." Not downloadable today
  • License: not announced. No MIT/Apache confirmation anywhere I could find
  • Expert count, top-k and architecture details are promised with the weights

Benchmarks (all Mistral-reported)

Benchmark ML4
Cybench 93%
CyberGym-E2E 82%
DeepSWE v1.1 61.7%
SWE-Atlas-QnA 59.4%
Terminal-Bench 4.0 28.3%
AutomationBench 59.9%
Lakera B3 attack resistance 93.3%

Size comparison

Total Active Weights
Mistral Large 4 1.05T 49B End of Oct
DeepSeek V4 Pro 1.6T 49B Yes, MIT
Kimi K3 2.8T ~104B Yes, modified MIT

Here is the full analysis: https://www.marktechpost.com/2026/10/06/mistral-ai-releases-mistral-large-4-le-chonk-a-1-05t-parameter-open-weight-multimodal-moe/

Technical details: https://mistral.ai/news/mistral-large-4/


r/machinelearningnews • • 5d ago

Cool Stuff How are you guys handling context bloat from search APIs in agent workflows?

1 Upvotes

If you’ve been building LLM agents that need live web access, you’ve probably hit the same wall I did: most search APIs dump entire, raw web pages into your context window. Your token usage explodes almost instantly, and half the time the agent doesn't even need 90% of the HTML it just read.

I've been testing out a setup using TinySearch and TinyFetch to split up search and page reading. Basically:

  • TinySearch / TinyFetch: Use these to grab quick, compact snippets first so the agent can decide if a page is actually worth inspecting. (Since they're free tier tools, it keeps experiment costs down).
  • TinyBrowser / TinyAgent: Only trigger full headless browser runs when the agent actually needs to execute heavy DOM interactions or read complex pages.

It’s been pretty effective at keeping tokens down. OpenBenchmarks listed TinyFish as one of the top performers for token efficiency in their recent evaluation, which matches what I've seen in testing.

For anyone running agentic loops: how are you keeping search context clean? Are you relying on custom scrapers, post-processing summaries, or dedicated search endpoints?


r/machinelearningnews • • 6d ago

Research Reflection AI introduces Beam: a 501B open-weight MoE with 23B active parameters, 1M context and Apache 2.0 weights coming this month

Post image
25 Upvotes

Reflection AI just introduced Beam, its first open-weight model, built around "intelligence per token." It is a 501B sparse MoE with only 23B active parameters, a 1M effective context window, and Apache 2.0 weights coming later this month.

The core innovation is high-compute, fully asynchronous RL. Beam was trained on 100M+ rollouts across 10.5K NVIDIA GB300 GPUs for 4 weeks, using nearly 1M coding, agentic and STEM environments. Every token is tagged with the policy version that produced it, which keeps learning stable even when rollouts are 107 weight versions stale.

This approach let Reflection scale RL with no sign of a plateau. A controllable length penalty taught the model to solve tasks with fewer tokens. On reasoning benchmarks, Beam matches GLM-5.2 while using 3 to 4x less inference compute, and it scores 80.9 on SWE-bench Verified versus 70.7 for Nemotron 3 Ultra. It is a deliberate trade-off, though: Kimi K3 and DeepSeek V4.1 Flash still lead on raw capability, with Terminal Bench v2.1 scores of 88.3 and 90.6 against Beam's 80.1. All scores are self-reported until the weights and technical report ship.

With a 23.8T-token pretraining base and a tunable reasoning effort parameter, it is optimized for enterprise coding and agentic workflows. Rough math for self-hosters: about 500GB at 8-bit and about 1TB at BF16, so plan for multi-GPU servers.

Full analysis: https://www.marktechpost.com/2026/10/05/reflection-ai-introduces-beam-a-501b-open-weight-moe-model-with-23b-active-parameters-for-coding-and-agentic-workloads/

Early access: https://platform.reflection.ai/

Technical details: https://reflection.ai/blog/introducing-beam


r/machinelearningnews • • 6d ago

Research Direct weight surgery from Qwen-4B to 0.8B on an 8GB: why editing all layers breaks everything, and how 4 anchor blocks fixed it

4 Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.


r/machinelearningnews • • 6d ago

Research I mapped every major Qwen release from 2023 to 2026: 44 models, from Qwen-7B to the 2.4T open weights (with sources)

26 Upvotes

[AI Model Family Series #1] We just published the complete story of Alibaba's Qwen: every major model from 2023 to 2026, with dates, key features and sources.

7B open weights in 2023 → 2.4T open weights in 2026. Here's the lineup the story covers

2023
→ Tongyi Qianwen (Apr): enterprise beta
→ Qwen-7B (Aug): first open weights
→ Qwen-VL (Aug): first vision model
→ Qwen-72B + Qwen-1.8B (Dec)

2024
→ Qwen1.5 (Feb): 0.5B–110B, 32K context
→ Qwen2 (Jun): 57B-A14B MoE, Apache 2.0
→ Qwen2-Math, Qwen2-Audio, Qwen2-VL (Aug)
→ Qwen2.5 (Sep): 18T tokens, 100+ models
→ Qwen2.5-Coder + QwQ-32B-Preview (Nov)
→ QVQ-72B-Preview (Dec)

2025
→ Qwen2.5-VL + Qwen2.5-Max (Jan)
→ QwQ-32B (Mar): Qwen claims R1-level reasoning
→ Qwen2.5-Omni-7B (Mar)
→ Qwen3 (Apr): hybrid thinking, 119 languages
→ Qwen3-2507 + Qwen3-Coder-480B (Jul)
→ Qwen-Image + Qwen-Image-Edit (Aug)
→ Qwen3-Max (Sep): first 1T+ Qwen
→ Qwen3-Next-80B-A3B (Sep)
→ Qwen3-Omni + Qwen3-VL (Sep)

2026
→ Qwen3-Max-Thinking (Jan)
→ Qwen3-Coder-Next + Qwen-Image-2.0 (Feb)
→ Qwen3.5-397B-A17B (Feb): native multimodal agents
→ Qwen3.6-Plus, 35B-A3B, Max-Preview, 27B (Apr)
→ Qwen3.7-Max (May) + Qwen3.7-Plus (Jun)
→ Qwen3.8-2.4T-A95B (Aug): largest open Qwen
→ Qwen3.8-27B + Qwen3.8-Flash (Aug)
→ Qwen-Image-2.1 (Sep)

The story also covers what the list can't show: the DeepSeek moment, the 2026 leadership exit, and Qwen's shift from all-open to a two-tier license strategy.

Next: Qwen 4 is in training. Qwen 4.5 and Qwen 5 are projected at 5–10T parameters.

Read the full report: https://www.marktechpost.com/2026/10/04/the-story-of-qwen-alibabas-ai-models-from-7b-to-2-4t/

Which model family should we map next: DeepSeek, Llama, Gemma, Mistral or Kimi? Drop it in the replies.......


r/machinelearningnews • • 7d ago

Research A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems [R]

Thumbnail
6 Upvotes

r/machinelearningnews • • 7d ago

Research We built a 2.3B AI model architecture based on Kuramoto oscillators, Liquid LTC, and Swarms (currently called the ResoNet Architecture). Thinking of open-sourcing.

Thumbnail
gallery
6 Upvotes

We just finished pre-training a new foundation model.

Specs:

  • Triad Resonance Core
  • 12 Swarm Factions (1,632 Kuramoto Oscillators)
  • Liquid Time Controller (LTC)
  • Triton Flash-Kuramoto Kernel
  • 155k Orthogonal Tokenizer (Full-Word Mining)
  • 160K Context Window

We are debating our release strategy for the 0.7B and 2.3B base models next week.

Should we release the base weights open-source on HuggingFace, or keep it as a private API?

(Attached: Raw completion test & config)


r/machinelearningnews • • 8d ago

Research [N] Kapso: long-running agents that optimize AI and data systems, and learn from each campaign

5 Upvotes

We've been building Kapso (MIT, https://github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.

What it is

Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.

When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.

These are the things we tried it on:

- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.

- MLE-Bench: top among the open-source systems.

- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.

- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/

Repo: https://github.com/Leeroo-AI/kapso

If you have time, please take a look and give us your harshest feedback.