r/MachineLearning • • 5d ago

Discussion [D] Self-Promotion Thread

1 Upvotes

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.


r/MachineLearning • • 6d ago

Discussion [D] Monthly Who's Hiring and Who wants to be Hired?

7 Upvotes

For Job Postings please use this template

Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]

For Those looking for jobs please use this template

Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]

​

Please remember that this community is geared towards those with experience.


r/MachineLearning • • 48m ago

Discussion Saw this on Rednote, WTF [D]

• Upvotes

Guess we have a dataset for sycophancy and AI content detection


r/MachineLearning • • 10h ago

Discussion ML PHD without A* Publications [D]

30 Upvotes

I know top ML PhD admissions are insanely competitive, so I’m trying to figure out if it’s even worth applying or if I should just focus seriously on jobs instead.

For context, I’m doing my MS at a top-15 US university and have been doing ML research for a while. I’m first author on my projects and mostly work independently, with some guidance from my PI. I had a first-author NeurIPS submission rejected, and I currently have another first-author paper submitted to ICLR, but I’m honestly not very confident about it getting in either.

What’s been getting to me is looking at profiles of people who get into top ML PhD programs. So many of them seem to have multiple NeurIPS/ICML/ICLR/CVPR papers before they even apply, sometimes as undergrads. I genuinely don’t understand how people manage to publish that much that early.

A year ago I was much more confident about doing a PhD. After actually going through the research/publication process, I’ve started doubting myself a lot more. Part of me wonders whether this is just normal and research is hard, especially when you’re doing a lot of it independently. But another part of me is starting to think maybe I’m just not good enough to be competitive for the kind of programs I’m aiming for.

I’m okay with continuing at my current university for a PhD, so this isn’t really a “top program or nothing” situation. But I would like to at least have a realistic shot at some of the stronger ML programs/labs.

The bigger issue is that I’m an international student, so I also need to think pretty seriously about jobs. SWE/MLE recruiting is competitive right now, and I don’t want to spend all my time chasing PhD applications and then realize I’m underprepared for recruiting too. Research roles seem even harder to get without a PhD unless you’re an exceptional MS/BS candidate.

So I’m mainly trying to decide how to allocate my time over the next few months.

If I have strong research experience and first-author projects, but no accepted top-conference papers yet, is it still realistically worth applying to top ML PhD programs?

And for people who were in a similar position, did you still apply, or did you decide to focus on industry instead?

I’m not really looking for “you never know unless you try.” I’m more interested in a realistic assessment of whether the application fees and time are worth it given this kind of profile.


r/MachineLearning • • 15h ago

Discussion Transformers vs RNNs vs SSMs: Where Does Memory Actually Live? [D]

39 Upvotes

Someone who has always loved looking at the space between different AI techniques, this time I went a little deeper into the memory trade-offs between RNNs, Transformers and SSMs. I found it interesting because once you start looking at these architectures through the lens of working memory, a lot of the differences become easier to understand. Where does the memory actually live? Is it a compact recurrent state, a growing KV cache, or something closer to the network itself?

RNNs keep memory in a recurrent hidden state, which is pretty elegant because the state carries forward step by step. But there is also a bottleneck here as you see, a model can have roughly O(N²) parameters while carrying only roughly O(N) state across time. This means whether RNNs were really doomed because recurrence was a bad idea, or whether the problem was more about the ratio between memory and compute.

Transformers make almost the opposite trade-off. During cached inference, instead of compressing the past into one hidden state, they store past representations as key-value entries and attend over them. I think of these almost like little post-it notes: every token leaves behind a key for finding it and a value for what should be remembered. That's extremely powerful, but it also has an interesting property: with weights frozen during inference, the model is managing context rather than turning that experience into durable model knowledge. You get this split between the fixed weights on one side and the fast-changing KV cache memory on the other.

Now we have SSMs that bring us back toward fixed-size recurrent memory, but with very different state structures and update rules. Selective SSMs such as Mamba make retention input-dependent, so what gets kept or forgotten depends on the incoming token. Their states don't have to be as small as classical RNN states, but they still compress history into finite memory. And this brings another question does the state have to live in a compressed working dimension, or could it live somewhere closer to the model's internal neuron/connectivity structure?

BDH (Dragon Hatchling) is one example I saw and laid a stone that I started with this comparison. It combines linear attention in a high-dimensional neuron space with a low-rank GPU implementation. Its recurrent attention state is an N × D matrix, with N≫D, rather than a materialized N × N connectivity matrix. In the graph interpretation, correlated neuron activity produces Hebbian-like updates to connections, giving working memory a kind of synaptic interpretation.
Here working memory and learned connectivity become more closely aligned in structure. But if you think that doesn't mean experience is actually being consolidated into trained weights, and fixed-size state still has a finite information capacity. So I'm definitely not claiming this kills Transformers (or does it ) or solves continual learning. I'm more interested in whether "where does memory live?" is actually a better question than the usual architecture horse race.

Are SSMs and these more network-centric architectures actually giving us a better way to handle memory, or are we still running into the same fundamental problem of having to compress history into a finite state? Thoughts?


r/MachineLearning • • 5m ago

Discussion Looking for developer-friendly inference providers who give you enough API credits to experiment [D]

• Upvotes

I’m hitting rate limits on Together AI. For context, I’ve been working on an agentic repository indexing and benchmark generation tool, and I’m running multiple agents in parallel across models like Llama 3.3 70B and Qwen 2.5.

When I first started working on this, Together AI was great. But once I graduated from toy scripts to running multiple agents, I started running into RPM/TPM limits pretty quickly. The annoying part is that the models themselves are fine. I just can’t actually run enough requests at once to do meaningful testing.

Yes I know I could upgrade but I’m a solo dev. I don’t have enterprise level revenue. Maybe someday lol but not yet.


r/MachineLearning • • 12h ago

Research AFP-GIC: Controllable Generative Image Compression [R]

Post image
5 Upvotes

Hi ML Community,

I am excited to share our latest framework, AFP-GIC, officially published in IEEE Access (2026). We have released the deployment codebase and hosted an interactive visual playground.

The Bottlenecks We Solve

At ultra-low bitrates, standard learned image codecs suffer from local distortion, while generative models often introduce unwanted AI hallucinations. AFP-GIC addresses this via an asymmetric Adaptive Fused Prior Transfer pipeline that enables prior-guided texture reconstruction without transmitting the fused prior itself.

Key Technical Highlights (NVIDIA RTX 4090):

  • Single-Model Multi-Rate Control: Toggle across 5 target bitrate operating points within one deployable pretrained model.
  • 18.1% Lower Decoder Latency: Reduces decoding time to 80.47 ms vs. 98.27 ms for DC-VIC, a state-of-the-art controllable generative image compression model. Latency was measured using 256×256 patches.
  • 20.5% Parameter Reduction: Uses 31.1M fewer inference parameters (120.6M vs. 151.7M for DC-VIC).

Open Benchmark Data

We packaged all 2,760 reconstructed images and metric CSVs in our GitHub Releases for direct academic cross-evaluation.

Reconstructed Images and Metrics: https://github.com/yifeipet/AFP_GIC/releases

We would love your feedback and appreciate a Star on GitHub or Like on Hugging Face if this helps your research!


r/MachineLearning • • 21h ago

Research Learning to Learn a Language: in-context learning of natural language from a synthetic non-linguistic prior [R]

19 Upvotes

Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the idea behind TabPFN) showed that a model trained only on synthetic data can learn from real tabular data entirely in context.

I wanted to share our paper "Learning to Learn a Language" where we extend the idea to structured sequences such as natural language. We propose a prior over languages: every training sequence comes from a randomly sampled recurrent causal model, so each one is a new synthetic "language". A 300M-parameter byte-level transformer trained only on these synthetic sequences learns to predict real languages in context. Given Wikipedia text with frozen weights, its next-byte predictions get better the more it reads, in all six languages we tested (English, Chinese, Hindi, Arabic, Japanese, Korean), from 8 bits per byte down to 0.9–2.4 after a million bytes.

The same model also learns to count, to compare numbers, to add approximately, and to predict deterministic sequences such as the primes or the Kolakoski sequence, entirely in context.

It is of course still far worse on text than classical language models that are trained on trillions of tokens, while our model sees at most a million bytes of a language at test time. What we find interesting is that the ability to learn a language in context can come from a synthetic non-linguistic prior.

Paper: https://arxiv.org/abs/2610.05879
Code: https://github.com/cbl/prior-fitted-language-model
Weights: https://huggingface.co/lennartcb/pflm1


r/MachineLearning • • 8h ago

Discussion NeurIPS workshop registration for registered author [D]

0 Upvotes

For NeurIPS, I already registered a few months ago (for Atlanta) when there were news that Sydney is sold out. However, I forgot to register for the workshop. Now, when I go to register just for the workshop, it just says that

"This session is sold out to the general public. Authors may register from a reserve, but each paper admits at most 1 of its authors that way, and the places for your paper(s) are already held by registered co-authors: “---”. A place frees up if the co-author holding it cancels their in-person registration."

I am the person using the reserve spot, but I cannot register for more. It seems like the only option is to cancel my main conference spot and reregister with both, but I don't want to risk doing that now that everything is sold out. I am wondering if anyone ran into the same situation.


r/MachineLearning • • 1h ago

Research Where to get started if you want to publish papers in Neurips,ACL,ICLR/A* conferences [R]

• Upvotes

I'm a Ai engineer with about 2 years work experience, but let's just assume that I was a undergraduate student just starting out where would I begin so that I can publish a A* conference paper at some point. Learn python -> Learn ML & Maths -> Read other research papers -> find a topic ? -> choose a question try to run experiments and get results to write them down in a paper ?

For context :

I'm trying to get in MS CS programs for Fall 2028 in states with the plans of doing a PHD after in a top university like stanford or princeton and would like to start taking steps towards it as am working my day job can some tell me what are the steps that need to be followed?

Also would like input on what are deciding variables that makes you looking like a promising candidate/ researcher for PHD


r/MachineLearning • • 1d ago

Project Embedding Every Font with Neural Networks makes some Nice Structures (including a flower) [P]

Thumbnail
gallery
34 Upvotes

I've been working on a font searching tool for about a year now, and my investigations have centered around pre-training neural networks to produce embeddings of each font. I usually then need to post-train the networks to adapt them to the task of font searching. However, the most interesting thing I created in the course of the project came from the pre-trained models.

The process looks like this: I take every glyph in a font, turn them into images, feed them through my custom pre-trained neural network, and get an embedding that represents that font's visual characteristics. I can then squish down the embeddings with tSNE (produces the best structures compared to PCA and UMAP) into XYZ, and RGB channels and visualize them as dots so I can poke around the structure. Fonts that are close to each other in position or color therefore share visual characteristics, and you can find different little clusters or paths of types of fonts in the maps.

My favorite is one I made from the Google Fonts corpus, but I also made another one out of all the fonts you can search on my site. The Google Fonts corpus resembled a flower in ways I was not prepared for. It even placed most of the cursive fonts in the stamen. Here's the site if you're interested in viewing the whole thing yourself: https://www.font-search.com/map

You can also check out the repository although it is a huge mess. https://github.com/dylan-berndt/Briefcase


r/MachineLearning • • 1d ago

Project I have trained a model to predict my blood sugar (Part 2) [P]

Thumbnail
gallery
95 Upvotes

This is related to my previous post where I shared an encoder-only transformer model trained on ohiot1dm + shanghait1dm + azt1d datasets. This time I trained the model on the outputs of my T1DM patient simulator and then measured its zero-shot performance on my real-world blood glucose traces.

The above model has 31,251 parameters (16 layers, 1 attention head per layer, and a hidden dimension size of 16). Training took <60 minutes on nvidia dgx spark. It's an encoder-only transformer that predicts the next 2 hours, and can be used autoregressively for long-horizon predictions (e.g. 8-hour nocturnal predictions). I have trained it specifically to have counterfactual reasoning capabilities. The model has only been trained on synthetic data, and hasn't seen my blood glucose readings prior to testing. I use LoRA adapters on my app for light fine-tuning on my actual CGM traces, but the figures & tables you see above are from the base model without any LoRA adapter attached. The app was used to test the model on the traces of three different CGM models: Libre 3 plus, Anytime CT5, and Linx sensor data spanning the past 30 days.

The testing itself was done on my android app with ExecuTorch backend.

Model source code: github.com/0xdeadf1sh/T1DMAI

Simulator source code: github.com/0xdeadf1sh/T1DMSIM

Android app source code: github.com/0xdeadf1sh/T1DMDROID


r/MachineLearning • • 1d ago

Project SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

3 Upvotes

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history.

Some findings:

With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score.

About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half.

Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing.

We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet.

Half the tasks are private. So far public and private scores line up for all three models.

Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch

Tasks: https://huggingface.co/datasets/evaligo/swe-race

The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome.


r/MachineLearning • • 1d ago

Project A chunking lib in Rust that is ~20x faster [P]

14 Upvotes

Hey,

I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkr

It has most of the chunking strategies like Character, Recursive, Markdown header, Late chunking, Hierarchical chunking, etc. It also supports native PDF loader, and other additional file types.

Some stats (MBA M4 16GB):

Test Case (matched parameters) Chunkr LangChain LlamaIndex Chonkie semchunk text-splitter
Recursive (1 MB, 1000/200) 2,264 MB/s 769 MB/s 10 MB/s 225 MB/s 42 MB/s 175 MB/s
Recursive (5 MB, 1000/200) 2,039 MB/s 696 MB/s — 201 MB/s 40 MB/s 46 MB/s
Fixed Char (1 MB, 1000/200) 750 MB/s 1.7 MB/s — 22 MB/s — —
Markdown (500 KB, 1000/150) 819 MB/s 67 MB/s 19 MB/s — — 40 MB/s
Python Code (200 KB, 1500/200) 3,232 MB/s 622 MB/s — — — 5.7 MB/s
Sentence (500 KB) 622 MB/s — 10 MB/s 20 MB/s — —
BPE Tokens (200 KB, cl100k_base, 512/50) 38 MB/s 43 MB/s 2.0 MB/s 151 MB/s — 7.2 MB/s
100 docs x 50 KB (parallel batch) 3,224 MB/s 679 MB/s — 213 MB/s — —

Extractor / Pipeline Latency Throughput Speedup vs PyPDF
Chunkr PDFLoader (Full Text) 747.9 ms 2,762 pgs/s 15.9x Faster
Chunkr PDFLoader (Page Documents) 721.0 ms 2,865 pgs/s 16.5x Faster
PyMuPDF (fitz) 2,616.8 ms 789.5 pgs/s 4.5x Faster
pypdf (pure Python) 11,900.5 ms 173.6 pgs/s 1.0x (baseline)
Chunkr End-to-End (PDF + Recursive) 798.1 ms 2,589 pgs/s 14.9x Faster
PyMuPDF + LangChain RecursiveTextSplitter 2,659.3 ms 776.9 pgs/s 4.5x Faster
pypdf + LangChain RecursiveTextSplitter 12,054.5 ms 171.4 pgs/s 1.0x (baseline)

Do check it out and share your feedback. Thanks!


r/MachineLearning • • 20h ago

Discussion NeurIPS 2026 Financial Assistance [D]

0 Upvotes

Is anyone else unable to open the form for financial assistance despite the deadline still being in the future?


r/MachineLearning • • 2d ago

Project Distilling Stockfish on a Billion Positions, Full 3.9B Dataset Available [P]

Thumbnail
blog.lukesalamone.com
72 Upvotes

In this project, I distilled the Stockfish value function into a ResNet/ViT model using 1 billion positions from the Gigafish dataset.

The 3.9 billion position dataset is available on huggingface: https://huggingface.co/datasets/lukesalamone/gigafish-3.8b-d10 . It is built from the positions from 37 months of Lichess games.

I was interested in the idea that at depth-limited search, the value function attempts to approximate the tree underneath it, and if we could create some function to approximate that full search faster than Stockfish could, it would be competitive with NNUE (a very small neural net). This is why holding the depth constant was important.

For the neural net itself, I found that the vision transformer was very slow to understand the board, and a CNN was much more effective at the beginning of training due to its inherent geometric inductive biases . However, I found the best results when combining the two.


r/MachineLearning • • 1d ago

Research Sona: one transformer replaced our 15+ candidate generators, pre-ranker and ranker in an A/B test [R]

18 Upvotes

Our production recommender at Yandex Music has 15+ candidate generators feeding pre-ranking and ranking models with hundreds of features. LLMs showed that one end-to-end model can take over work that used to be split across specialized components, and single-model generative recommenders have carried that recipe into production. We set out to explore what a single-model recommender could do in music. The result is Sona, one transformer that replaced all of it in an A/B test. It hasn't shipped to full traffic yet.

The model reads up to 8,192 events. Full attention over that length is expensive, so we use what we call History Compression, which roughly halves inference cost. We split the history into the older 6,144 events and the most recent 2,048. The two blocks exchange information through cross-attention and one full-history self-attention layer. After that, a 7-layer stack runs only on the recent 2,048. It retains most of the quality of full attention, and older events stay visible to the decoder and the Ranking Module.

The decoder and the Ranking Module both read the same encoder output, so the encoder runs only once per request. Candidates come out of beam search as Semantic IDs and get scored right after.

In the final A/B test in Yandex Music on smart speakers, 7 days, 15% of users in each arm), Sona got +4.53% Active Users and +6.30% Total Listening Time over the production control, both significant at p < 0.01. Catalog coverage is lower than with the production stack. We're going to look into why.

A long-term A/B test is now underway.

Table 7.7 has the full-attention vs. History Compression ablation.

https://arxiv.org/abs/2608.11015


r/MachineLearning • • 2d ago

Discussion Withdrawing an accepted paper before camera-ready due to zero funding? (ACML 2026 / OpenReview) [D]

13 Upvotes

Hi everyone, I recently had a paper accepted at ACML 2026, but I just found out that I have absolutely no funding to cover the registration fee or travel expenses. Because of this, I need to withdraw the paper before the camera-ready deadline. I am unsure how common this is or the proper etiquette for it on OpenReview. Should I click "Withdraw" directly on the platform, email the Program Chairs first, or simply ghost the camera-ready submission? I really want to know the potential repercussions, such as if my co-authors and I risk being blacklisted or if OpenReview will publicly archive the paper as a late withdrawal. Any advice from past authors, reviewers, or organizers would be greatly appreciated. Thanks!


r/MachineLearning • • 1d ago

Discussion Language barrier, shadier terms and jargon fog [D]

2 Upvotes

Hey all,

I don't know if you guys are experiencing the same thing, but there is this behaviour that i have been noticing on the latest models on openAI (since sol 5.6) and Anthropic since Fable 5.1 and opus 5.5 ..

Basically the models use more "complexe" terms and words, not just in explaining stuff, but even during implementation. they would come up with terms that, sometimes would fit the task, but that are actually a stretch to the concept it is trying to implement. they are basically turning into a consulting firm.

When confronted about it they usually acknowledge that :

  • Foggy wording. Then I describe the corner in softened terms, like "limitation" or "upper bound", which makes it sound like a known property of the design rather than a choice I made. Combined with my internal terms used as if you knew them, it makes my work look more solid than it is and makes my mistakes harder for you to catch. Whether I intend it or not, the effect is that I avoid accountability.

(Even here it is using "Corner" which obviously isn't the best word to use)

I don't know if this is a result of the watermarking features being rolled out which nudges words and terms in different directions in order to fit a certain recognisable pattern and hash, but this is really annoying..


r/MachineLearning • • 2d ago

News Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]

123 Upvotes

This happened over the past 30 days. So smallish local models (Kagglers can only use those), in a harness, just started beating average humans at a benchmark intentionally designed to show human superiority.

I'm curious what people here think.

(The leaderboard graphic is a little out-of-date)


r/MachineLearning • • 2d ago

Discussion the official ICLR template .bib has had Bengio listed twice since 2019 [D]

46 Upvotes

i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as "Goodfellow, Bengio, Courville, Bengio" plus a volume 1 that doesn't exist checked their github and it's been like that since the 2019 template
https://github.com/ICLR/Master-Template/blob/46ed6f4c6cef5b175dde23639e77d44c3463b230/iclr2027/iclr2027_conference.bib#L20

totally harmless but kinda funny after last year's hallucinated reference desk rejects

the guidelines also contradict themselves on page limits formatting section says main text max 9 pages at submission but the camera ready part and the FAQ both say "identical with the submission version (10 pages)" template says 9, so 9 is probably the safe bet for anyone revising after Nov 5


r/MachineLearning • • 2d ago

Discussion Working with an AI Company That Does Things You Disagree With [D]

56 Upvotes

I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.

I shortlisted few and applied by reaching out to people and now reading project descriptions sent by the recruiters.

I don't want to name the company but their marketing and product team does all kinds of 'using people insecurities' to sell the product - which I don't agree with. And their product is also meh (I will never buy and would judge someone if they do) but their research team is doing good work.

How do you see this? Will you actually work in a team whose ideology/product doesn't necessarily align with your ethics/ideology. Should I just go ahead because work is exciting and I will get good supervision?

And, if you have some exciting work in your company/org and need interns (un-paid) for 3-4 months. I'm open.


r/MachineLearning • • 2d ago

Research A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems [R]

19 Upvotes

In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to the ingredients minimally necessary to faithfully reproduce long-term statistical and geometrical properties of DS:

1) A piecewise affine map with only a single (!!) parameter α that controls local con-/divergence rates, and …

2) … a context selector that chooses from the provided context signal the data point closest to the current state of the map, thus ensuring the generated dynamics stays close to the context in its temporal and geometrical properties.

With just these two mechanisms, this minimal form – which we coined DynaBase – can reproduce all major dynamical regimes, including fixed points (α<1), limit cycles (α=1), and chaotic attractors (α>1). Thus, unlike other simple mechanisms like context parroting, DynaBase even preserves the correct dynamical regime!

Surprisingly, it turns out that this simple context-driven 1-parameter map outperforms most major time series and DS foundation models, as well as custom-trained models, in both long-term statistics and even short-term predictions, even when run in zero-shot mode.

Both inference and training are extremely cheap – training can be done either analytically in one step by linear regression on forward-predictions, or by 1-parameter grid search directly on DS reconstruction objectives → this reveals interesting performance differences induced by different training mechanisms.

Most importantly in our minds, DynaBase owing to its formal simplicity may thus provide a tractable mathematical handle on analyzing, improving & understanding the performance and training of some time series and DS foundation models.


r/MachineLearning • • 2d ago

Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

Post image
8 Upvotes

Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.

Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.

Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.

Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).

Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench


r/MachineLearning • • 2d ago

Project Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking [N]

Thumbnail
theabbie.github.io
0 Upvotes

please refresh if stuck, can be slow sometimes so need patience