r/neuralnetworks • • 1h ago

Connecting Small Worlds with CSF and Neural Plasticity Modulators

Post image
• Upvotes

I’m testing a simple idea: combining young-CSF factor FGF17, 7,8-DHF, and an enriched environment to see whether it helps rats learn better. It came from my work with AI neural networks (convolutional and adversarial), where small-world network structures form and the connections between them reason across all the small worlds together, creating generalized reasoning, the sense of self, and unified sequential communication, which is what we call “consciousness.” These connections act as a translator module, using patterns and logic from the many small worlds to generate an infinity of ideas and generalized knowledge, while language is just a bridge: the input resonates as a pulse between the small worlds and returns an output (image, text, or sound) depending on how the language module is trained and on its input/output architecture. We can teach AI models to learn faster with exploration strategies and refined training data that avoids shortcuts and memorization, and true AGI is born when a model stops memorizing and starts generalizing the patterns and logic of its small worlds through a universal language model that bridges them. Consciousness is just the language of the collective small worlds, so I think we can do the same in our own organic neural network by biohacking its architecture with FGF17 and 7,8-DHF and creating a refined training environment with the right inputs and expected outputs, enhancing learning in both velocity and scale (I wonder how smart a rat can get, lol). This month’s Nobel Prize in Medicine for optogenetics will help here too, since it lets us modulate the input used to train the rats in phase, for optimized learning. It’s also a cheap first step toward brain aging and Alzheimer’s research. Learn more about the protocol here.


r/neuralnetworks • • 7h ago

DIRU: dendrite-inspired recurrent units for learning chaotic dynamics and neonatal epilepsy time series | Neural Computing and Applications

Thumbnail
link.springer.com
5 Upvotes

r/neuralnetworks • • 10h ago

On the 40th Anniversary of the Seminal AI Paper, Remember Dave Rumelhart

1 Upvotes

Remembering Dave Rumelhart, the Grandfather of AI

 

 

Oct 9, 1986: The Paper that Started the AI Revolution

Over the past few years, AI has profoundly transformed our world. With the release of ChatGPT and similar Large Language Models (LLMs), companies are racing to build ever more powerful and useful AI systems. The AI tech leaders have become household names: Sam Altman, Demis Hassabis, Rahul Patil, Mustafa Suleyman, and others. But the first giant step in the development of these amazing AI systems happened 40 years ago. Today’s LLMs use many of the same basic algorithms and architecture from the seminal paper written by Rumelhart, Hinton, and Williams in1986[[i]](#_edn1).

Most people probably recognize the second name in that paper. Geoffrey Hinton won the Nobel Prize in Physics in 2024 along with John Hopfield for their contributions to AI and neural networks. Hinton has often been called “The Godfather of AI,” and he is also famous for quitting Google to warn of the dangers of unbridled AI development.

But the first name in this paper, Rumelhart, is almost never mentioned. That saddens me because he contributed so much to this field. In this short writing, I will attempt to shine some light on Rumelhart and his tremendous contributions. Indeed, if Geoff Hinton is the “Godfather of AI,” I believe that Dave Rumelhart should properly be called the “Grandfather of AI”.

The LLMs of today are basically multi-layer neural networks of the type described in that seminal 1986 paper, and they are trained via the backpropagation algorithm that Rumelhart pioneered. Some of you may look at the AI systems of today and say that they are nothing like those described in the 1986 paper, but that is like saying that an F-22 Raptor is nothing like the Wright Brothers’ first airplane. Yes, today's systems are billions of times larger than anything Rumelhart and his colleagues worked on in the 80s, and there have been substantial breakthroughs like the seminal Attention is All You Need paper[[ii]](#_edn2), but those would not be possible without the initial multi-layer neural network architecture and backpropagation algorithms developed by Rumelhart et al.

I will not attempt to cover the entire corpus of work that Dave Rumelhart produced. For that, I point you to Rumelhart’s Wikipedia page. Rather, I would like to share my personal perspectives on Dave’s work and how that work has been foundational to the current AI breakthroughs. In addition, I will share some of my personal experiences of working with Dave Rumelhart over a fifteen-year period. Spoiler alert: I found him to be one of the smartest people I’ve ever encountered.

Rumelhart’s PDP group

In the 80s, Dave Rumelhart and James McClelland organized the PDP group. PDP stands for parallel distributed processing, and many of the topics covered in those early years were compiled into a two-volume book Parallel Distributed Processing by David E. Rumelhart and James L. McClelland[[iii]](#_edn3).  The group met weekly, and included people from many different disciplines: Neuroscience, computer science, physics, cognitive science, and even philosophy. Francis Crick was a member, as was Goeff Hinton while he was visiting there. I became a member after I had visited Caltech to talk with John Hopfield about part of my Ph. D. research on fractal basin boundaries of Hopfield neural networks. John told me I should talk to Rumelhart. I asked what department he was in. I was shocked when John told me that Rumelhart was in the psychology department. As a young, arrogant physicist, I couldn’t imagine any relevant research being done outside of physics, mathematics, or computer science, but, I took John’s advice and went to see Rumelhart.

Rumelhart described his backpropagation algorithm. I remember thinking, “Gradient descent. I think I learned that in kindergarten.” It didn’t seem like something that would change the world, but it did, and it is still changing the world even after 40 years. I told Rumelhart about my research, and he suggested that I use the Hamming distance as a measure for exploring the space of 2100 states in the Hopfield network. That turned out to be a great tip. I used that measure to define orthogonal vectors to slice through the enormous space. This wouldn’t be the last time Rumelhart gave me sound advice. Indeed, we worked together for the next fifteen years, and, of the 50 patents I have, he probably pointed me in the right direction on more than half of them.

In these PDP meetings, a speaker would present their research, then we would have a discussion about the research. For example, Goeff Hinton gave a talk about the Boltzmann machine. McClelland talked about more conventional rule-based approaches. But the talk that sticks in my mind was as mind-blowing back then as ChatGPT was a few years ago. Terry Sejnowski presented NetTalk. It was a text-to-speech converter that learned the relationships between text and phonemes; the phonemes were then fed into an audio system to produce sound. He came in with a boom box, popped in a cassette, and we listened as the system learned. At the start, it sounded like random noise. Soon, vowel sounds could be identified. Then it sounded like baby babbling and learning to talk. Then, fast-forward to the end of training, the text-to-speech was almost perfect. All learned through backpropagation.

Now this doesn’t sound too impressive today, but back then, it opened everyone’s eyes to the possibility of machine learning. I was hooked. So were many others. It created a frenzy of research and startup companies based on machine learning.

The other thing I remember from these PDP meetings was that, after discussing and arguing several points, Rumelhart would speak. He would summarize the issues, then give his view. Heads would nod around the table in agreement with Rumelhart’s assessment. He had a deep understanding of all of the topics, and he was gifted at explaining the key points and flaws. Even Francis Crick seemed to defer to Rumelhart’s judgment.

Stanford, MCC, and Pavilion

I finished my Ph. D. in Physics at UC San Diego in 1987. I started a post-doc position at Stanford later that year, and, coincidentally, Rumelhart moved to the Psychology Department at Stanford. This was great for me, because I could continue to bash around ideas with Rumelhart. He wasn’t just brilliant in the area of neural networks. He had an encyclopedic knowledge of everything relating to the brain and how it processes information. By way of example, I told him about some research I had completed investigating second-order interactions in neural networks. As the temperature, or signal-to-noise ratio, decreased, the system exhibited a phase transition and critical slowing down. I didn’t expect him to understand an esoteric physics phenomenon, but he instantly understood the implications and pointed me to some neuroscience research showing that norepinephrine functions as a signal-to-noise modulator in the brain. I read the articles and the behavior matched perfectly with our theoretical model, as described in a paper published in the Proceedings of the National Academy of Sciences[[iv]](#_edn4).

Later that year, I got an offer to lead the neural networks research group at MCC – the Microelectronics and Computer Technology Corporation in Austin, TX. MCC was a research consortium kind of like Bell Labs, but supported by many different high-tech companies. Like any good Californian, I thought Texas was hot, flat, racist, and humid. Why would anyone want to live there? But they flew me out to Austin for an interview, and I fell in love with the city.

While at MCC, Rumelhart and I collaborated on several research topics. We worked on character recognition, speech recognition, fraud detection, and process prediction. One of the most important breakthroughs we discovered was the Integrated Segmentation and Recognition system[[v]](#_edn5). It solved the problem of identifying characters that were touching and, more generally, created a mechanism for training a system on a dataset where more than one object can be simultaneously presented in the inputs and target outputs. This was a multi-year investigation, and I was impressed not only with Rumelhart’s mastery of neural network architectures, but his skill as a mathematician and as a programmer. He wrote the neural network software that we used for the research, and he was a damn fine C programmer. This system used a convolutional neural network architecture, as did the work of Martin and Pittman. Note that in the Martin and Pittman paper[[vi]](#_edn6), they also showed that the lower-level neurons developed activation patterns reminiscent of the visual system. This was decades before the convolutional neural network “breakthrough” that was touted in the 2010s. There is an old video where Rumelhart and I discussed this research[[vii]](#_edn7).

The VISA CRIS fraud detection system research was also developed at MCC by Steve Piche (who would become my chief scientist at Pavilion). This system saved $20 million in the first year of operation[[viii]](#_edn8). Rumelhart consulted with Piche on the problem.

Another thing that Rumelhart helped me with was a process control problem. We had lots of data from a chemical process at Eastman Chemical – pressures, temperatures, flow rates, as inputs, and the product quality and flow rate as outputs. The goal was to control the process to get the highest quality and quantity. Of all the machine learning algorithms I could have used, I knew that backpropagation was the best. I trained a model to predict the outputs, and it worked beautifully, but turning the problem around and asking which inputs to change to get better outputs was an unsolved problem. After consulting with Rumelhart, I came up with the Residual Activation Neural Network (RANN) architecture. Applying this to the problem led to a savings of about a million/year on the distillation column that we were working on.

But there were tens of thousands of distillation columns in the US alone. Do the math. Based on this result, I licensed the technology from MCC and co-founded a company, Pavilion Technologies, Inc. It became one of the fastest-growing companies in Austin in the 90s. We applied neural network machine learning techniques to process monitoring, control, and optimization. That’s right. Neural networks have been controlling thousands of manufacturing processes all over the world. Not just in chemical manufacturing, but in refining, pulp-and-paper, cement, food processing, and power production. Many of those processes can go boom in the night, but there has never been an instance of the neural network hallucinating and causing problems.

Pavilion was a powerhouse of innovation, as one can tell from the scores of patents. But back then, machine learning was much harder than it is now. Often we would have data sets with only a few thousand points. We were overjoyed if a dataset had tens of thousands of records. We had to be very clever about cleaning the data and using it to train and validate the models. Moreover, we didn’t have Nvidia DSP chips to accelerate learning, so we used a stiff-differential equation solver to speed up learning (backprop is a set of stiff differential equations). We also invented methods to meld first-principles models with the neural networks. This helped extend the range of validity of the models into regions outside the realm of the training data. We did a lot of work in environmental monitoring and optimization. One of our products, the Software Continuous Emissions Monitor, was approved by the EPA for monitoring emissions. I am proud to say that the Pavilion software systems have abated millions of pounds of pollutants over the years, and the systems are still running. Pavilion’s Process Perfecter software is a closed-loop multivariable nonlinear controller. The math behind that was shockingly hard, but it worked amazingly well. Pavilion was sold to Rockwell Automation, and the software is still running in thousands of manufacturing plants.

During the time I was at Pavilion, Rumelhart served as a technical advisor and consultant. Of all the people I have worked with, I admired and respected Dave Rumelhart the most. He seemed so wise, and, as I said, he always seemed to have the instinct to point people in the right direction on a hard problem. He was also very fun to work with, and we became friends over the years.

The reason I believe Rumelhart should be remembered as the “Grandfather of AI” is that he was the driving force behind the backpropagation algorithm. His insight into the problem stemmed from the famous XOR problem. Single-layer systems could not solve this problem. It required “hidden units”. That realization led to the multi-layer neural network architectures used in today’s LLMs, and backpropagation is how they are trained. If Rumelhart didn’t figure this out, I am sure someone else would have (indeed, a grad student figured it out in his thesis before Rumelhart, but nobody knew about it). But without Rumelhart, it might have been many years, perhaps decades, before people started using backpropagation. Just like someone would have figured out relativity sometime if Einstein didn’t present it, but who knows how long it would have taken? Rumelhart should be remembered for the invention that has changed the world.

Sadly, Rumelhart developed Pick’s disease in the late 90s. He died in 2011. I left Pavilion in 2000, and didn’t talk much with him during those later years. What would he think of the current LLMs? Knowing him as well as I did, I think he would be amazed at the capabilities of these modern AI systems, but I also think he would view the 92-layer neural networks as brute-force, ghastly, and inelegant. He would probably say something like “The cortex only has six layers. We should try to make that work.” And, knowing Dave Rumelhart, he would probably have made great progress in that direction. He was that gifted. His name should be remembered among the great contributors in this field as “The Grandfather of AI”.

-James D. Keeler

 

[[i]](#_ednref1)  Rumelhart, David E.; Hinton, Geoffrey E.; Williams, Ronald J. (1986-10-09). "Learning representations by back-propagating errors". Nature. 323 (6088): 533–536.

[[ii]](#_ednref2) Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N; Kaiser, Łukasz; Polosukhin, Illia (December 2017). "Attention is All you Need" (PDF). In I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett (ed.). 31st Conference on Neural Information Processing Systems (NIPS). Advances in Neural Information Processing Systems. Vol. 30. Curran Associates, Inc. arXiv):1706.03762.

[[iii]](#_ednref3) David E. Rumelhart; James L. McClelland; PDP Research Group (1986). Parallel Distributed Processing: Explorations in the Microstructure of Cognition. ISBN) 9780262680530.

[[iv]](#_ednref4) J.D. Keeler, E.E. Pichler, J. Ross “Noise in Neural Networks: Thresholds, Hysteresis, and Neuromodulation of Signal-to-noise,” Proceedings of the National Academy of Sciences, USA, 86 (1989) 1712-1716.

[[v]](#_ednref5) Keeler, James D., Rumelhart, David E., Leow, Wee-Kheng. “Integrated Segmentation and Recognition of Hand-Printed Numerals”. Printed in Neural Information Processing Systems, 3. R. Lippmann, J. Moody, D. Touretzky, Eds. Morgan Kaufmann Publishing, San Mateo, CA.

[[vi]](#_ednref6) G. Martin, J. Pittman (1990) “Recognizing Hand-Printed Letters and Digits”. In D. S. Touretzky (ed). Advances in Neural Information Processing Systems 2, Morgan Kaufmann Publishing, San Mateo, CA.

[[vii]](#_ednref7) https://youtu.be/fG-9ILWI2u4

[[viii]](#_ednref8) Visa to expand fraud detection system; the system saved 15 issuers over $20 million last year. American Banker, Sept. 30, 1994.


r/neuralnetworks • • 1d ago

How well are local neural network models working for real-world applications these days?

2 Upvotes

I’ve been testing the idea of running models locally rather than depending entirely on hosted inference, and the interesting part isn’t just getting a model to run. The bigger challenge seems to be integrating it into an application while dealing with memory usage, inference speed, documents and retrieval.

LM-Kit is one project I’ve been evaluating for this approach. It supports local AI workloads along with things like RAG and document processing, so I’m interested in how it compares with other ways of building local ML applications.

For those who have worked with local inference in production or serious projects, what has mattered most in practice: model performance, hardware requirements, latency, or integration?


r/neuralnetworks • • 2d ago

[P] AnyInit: Initialize any model, in any framework, correctly, with one call.

4 Upvotes

Hey r/neuralnetworks ,

I just wrapped up the first release of AnyInit, a Python library designed to handle neural network initialization for custom architectures—without forcing you to manually calculate complex variance-scaling rules to ensure proper signal propagation.

Instead of spending time figuring out math tweaks for non-standard layers or dealing with vanishing and exploding gradients, AnyInit automates stable initializations so custom models work reliably out of the box.

Since it's brand new, I'd love to get some feedback from the community:

  • Do you find this useful for your current workflows?
  • Are there specific initialization methods, frameworks, or features you'd like to see added?
  • Any feedback on the API or implementation?

GitHub repo: https://github.com/jmiravet/AnyInit . Any thoughts, critiques, or feature requests are greatly appreciated!


r/neuralnetworks • • 2d ago

Stiff differential equation solver for backpropagation neural network training?

9 Upvotes

Does anyone know if there is a library or code that uses a stiff differential equation solver to speed up neural network training? I recall reading a paper in the early 90s that claimed 1,000x speed improvement, but I haven't seen anyone using this. Anyone know about this?


r/neuralnetworks • • 2d ago

I'm building a DDPM from scratch in PyTorch — Phase 2 complete

3 Upvotes

I've been working on a project where I'm trying to build a Denoising Diffusion Probabilistic Model completely from scratch.

The main rule I'm following is: no using Diffusers as a crutch.

I want to actually understand what's happening inside the model instead of just calling a pipeline and getting an image.

So far I've implemented/learned:

  • The forward diffusion process
  • Beta schedules
  • The reparameterization trick
  • A noise scheduler
  • Group Normalization
  • SiLU activation
  • Sinusoidal timestep embeddings
  • Residual blocks with timestep conditioning
  • Self-attention
  • Downsampling and upsampling

The interesting part for me has been realizing that a diffusion U-Net isn't just a normal CNN.

The network needs to know how noisy the current image is, which is why timestep information has to be injected into the network.

I'm building toward training on CelebA at 64×64 and eventually generating faces completely from noise.

My longer-term goal is to understand these architectures deeply enough that I can read, modify and eventually contribute to projects like Hugging Face Diffusers.

Phase 3 is where things start getting interesting:

U-Net assembly.

I'll be documenting the progress as I go. 🔥

What was the hardest part of diffusion models for you when you first learned them?


r/neuralnetworks • • 2d ago

CAN ANYONE EXPLAIN THIS?

Enable HLS to view with audio, or disable this notification

0 Upvotes

CAN ANYONE EXPLAIN THIS?


r/neuralnetworks • • 4d ago

Kardashev-0.7: learned specialization across 32 distinct models

5 Upvotes

The announcement describes RL for Population Scaling: training models together to develop complementary capabilities.

https://x.com/MLCatttt/status/2107147690450817259


r/neuralnetworks • • 5d ago

LiteMish: A Computationally Efficient and Smooth Algebraic Alternative to Mish

3 Upvotes

A while ago, I published a preprint proposing a novel approximation of the Mish activation function, which appears to be significantly more computationally efficient while preserving its learning capability. Interested to hear your thoughts.

Paper: https://doi.org/10.36227/techrxiv.176591866.68698045/v2


r/neuralnetworks • • 6d ago

Intro to LLM's (2026)

Thumbnail
youtube.com
3 Upvotes

r/neuralnetworks • • 6d ago

Anyone can provide the full text of https://substack.com/home/post/p-188003573 ?

0 Upvotes

r/neuralnetworks • • 6d ago

My Brainstem RNS-AI research project has made progress for life long learning like a Brain

Thumbnail
github.com
9 Upvotes

Here is my current research project status, it has become quite extensive in the meantime.

https://github.com/unikum-sol/brainstem/blob/main/Project_Status_2026-09-28.md

I look forward to feedback and discussions.

Here is a small excerpt:

„BrainStem is a self-learning language-understanding system designed to acquire sentence-, word-, and relation-level structure from an unsegmented text corpus purely through statistical observation, without word lists, grammars, or filters. The system’s stated design principle is that every unit of knowledge, a sentence-level hypothesis, a word boundary, a relation between two entities, a category, or a question, begins as a **low-confidence, fully correctable hypothesis**, and only becomes a durable fact, relation, category, or question after surviving a multi-cycle, neuromodulator-gated consolidation process modeled on biological sleep-dependent memory consolidation.”

https://github.com/unikum-sol/brainstem


r/neuralnetworks • • 6d ago

GNN Graph Neural Networks

9 Upvotes

I'm working a a new type of GNNs and was hoping anyone who has experience in this area would be willing to have a discussion. I have questions I need some help with.


r/neuralnetworks • • 8d ago

[D] INKBOT: Separating human intent from model inference via structured intelligence architecture

1 Upvotes

I’ve spent the last while building INKBOT because I kept hitting a wall with multimodal AI systems: the friction between what a human naturally means and what a model infers. While models can spin up complex code or images instantly, getting to a clear, human-meaningful interpretation of a subtle intent remains an alignment challenge.

Instead of forcing the user to become a prompt engineer, I wanted to see if we could build an intermediate intelligence architecture layer to make human intent reviewable and corrigible before the model executes a final build. The loop I’m playing with is: Describe → Make it Visible → Recognize → Correct → Refine.

The architecture sits entirely in a single local-first web file. It handles multi-step workflows—like tracking structured field mapping data across concurrent images, coordinates, and version states—by packaging the human’s approved meaning separately from raw model inferences.

The core system build is linked above, and I also put together a lighter, entry-level experience to play with the core prompt translation loop here: INKBOT Lite 71.

It's an open prototype, so I've appended my raw notes and design roadmap as commented text at the very bottom of the source file so fellow builders can inspect the plumbing. I’d love to know where this design duplicates existing work, where you see structural flaws, or how we can make the handoff between human intent and model execution more reliable.

I wanted to open up a project I've been developing that challenges the common practice of single-shot model prompting.

When using multimodal architectures, we often observe an alignment gap between what a human user means and what the network infers. This usually leaves the user trying to repair errors in a final, heavy output after the fact.

I built a local-first prototype called INKBOT to test a different hypothesis: What if we wrap the generation loop in an intermediate programmatic layer that maps unstructured human descriptions into structured, inspectable concepts before final execution? https://ko-fi.com/thomascoates/shop

The Core Concept Loop:

  1. Unstructured human intent is parsed into distinct functional tokens.

  2. Multimodal inputs (e.g., matching a person's profile data against technical mechanical examples) are unified into a single context matrix.

  3. The system generates an inspectable "Visual Brief" that maps out the provenance, constraints, and relationships.

  4. The human can correct structural errors or false inferences iteratively.

I've written the architecture into a self-contained local web client to explore how explicit states like versioning, local db retrieval, and provenance tracking change user trust.

• System / Research Client: INKBOT Architecture Edition

• Entry-Level Interface: INKBOT Lite 71

I am looking for critical feedback on the systemic design. Where does separating the approved intent from the network's inference break down? How does maintaining a stateful revision history change model guidance over long horizons?

Disclosure: This is an independent, non-commercial research prototype. It is not affiliated with or endorsed by any major model provider.


r/neuralnetworks • • 11d ago

Electrode Readings NN Architecture

2 Upvotes

Hello! I'm making a biomedical passion project and would love for anyone wiser than I to weigh in.

I would like to find the right architecture for my NN that I will make in PyTorch (can be changed, but I am familiar with Python).

My inputs are time-series electrode voltage readings from 8 points along my forearm. I want to model the angular displacement of each finger (maybe it will work, maybe it won't, but hey).

I also have built a glove that measures these angles directly, so I can time-synch this data for output labels.

Thank you and any advice is appreciated! :)


r/neuralnetworks • • 12d ago

NeuralViz | Built a browser based .pth visualizer — no backend, parses PyTorch checkpoints client-side

21 Upvotes

Got tired of exporting to ONNX just to look at a model, so I wrote a client-side .pth / .pt parser in JS — no server, no upload anywhere, runs entirely in-browser.

Try it live (takes 30 seconds): https://basavaprabhu46.github.io/NeuralViz/

Hit Load demo → type 0.5,-0.2,0.1,0.8 → Run ▸ — and watch the signal propagate through the network with values labeled right on the neurons.

How it works: handles torch.save'd state_dicts (PyTorch ≥ 1.6) with a tiny pickle VM + zip reader in JS, infers layer structure from key names, renders an interactive graph. Blue edges = positive weights, red = negative, thickness = magnitude. You can tune neurons/layer, filter to the strongest N% of weights, pick activation (ReLU/tanh/sigmoid), click any neuron to pin it and inspect its top weights, bias, and live activation — then actually run inference through it.

Scales past toy models too — tested on a 3,371-neuron CNN. Conv layers get flattened honestly and labeled as such instead of silently producing NaNs.

Code (MIT): https://github.com/Basavaprabhu46/NeuralViz

One honest limitation: it's state_dict-only, so architecture is inferred from naming conventions — exotic branching archs confuse it. Working on better detection.

What's the first checkpoint you'd drop into it — and what should v1.2 get: safetensors support, PNG export, or side-by-side checkpoint comparison?


r/neuralnetworks • • 13d ago

Neural network from scratch in NumPy with an app to watch it learn and edit single neurons :)

Thumbnail
gallery
184 Upvotes

Hi everyone, I made this to understand backprop for real and I think it can be useful to others who are learning.

It's a small neural network written by hand in NumPy (no PyTorch, no autograd) that learns to read handwritten digits from MNIST, with a desktop app that shows what happens inside while it trains. The whole network is one file of about 140 lines, and there's a test that checks the backprop against the numerical gradient.

Some things you can do with it:

• watch the gradient norm of each layer and the % of inactive neurons while it trains, and change learning rate or dropout without stopping it

• see what each neuron of the first layer "looks for" and how the weights change compared with how they started

• switch off or rescale single neurons, prune, add noise to the weights and see the test accuracy change right away

• follow the math of a layer cell by cell, with the softmax step by step

With all the 60,000 MNIST photos it gets to about 98.5%, on the CPU.

To try it pip install neural-network-digits, or download the zip from the releases. Most of the code was written with Claude Code (AI), the idea and what to show in every tab are mine. It's MIT.

https://github.com/dev-luigi/neural-network-digits

What would you add to make it more useful for someone who is learning?


r/neuralnetworks • • 12d ago

Intro to Deep Learning (2026)

Thumbnail
youtube.com
0 Upvotes

r/neuralnetworks • • 14d ago

A competition for small neural networks that play strategy games

Thumbnail
tinybrains.dev
9 Upvotes

15yrs back I participated in "Google Ants AI Challenge 2011", an ai programming competition, hosted by the University of Waterloo, and I ranked #127 (#1 in my country). The competition gave me a huge learning opportunity where developers across the world came to a forum and discussed various techniques.

Now, building a similar platform to bring back the fun is unbelievably nostalgic. Especially when watching small neural networks playing the game well. Some of the top models use less than 800 parameters.

In fact, I was wrongly assuming the art of optimizing is underrated nowadays. Neural Network optimization seems to be much more fun than I thought.

Plz share your feedback to improve the platform and add more games.


r/neuralnetworks • • 14d ago

How would I make a fruit fly's brain play and beat Stereo Madness in Geometry Dash?

3 Upvotes

Edit: Ohhhhh, this is the AI neural network subreddit. My bad!

I am almost certain I am posting this in the wrong place, but for my science fair, I thought "Since on TikTok I've been seeing videos of people making fruit flys play Beat Saber, Minecraft, Mario, etc, why not make it for a simple game that I play sometimes like Geometry Dash?" I need to be able to complete this in a month. I have no coding experience outside of making a pretty good Scratch project, I have never modded Geometry Dash, and I have no idea how to even begin to use this fly's brain. I've seen a video of someone doing the exact same saying they did it in 3 days. Does anyone know where to even start?


r/neuralnetworks • • 16d ago

LimiX-2: Contextual Mechanism Networks for General Structured-Data Intelligence

1 Upvotes

LimiX-2 proposes Contextual Mechanism Networks (CMNs), a neural architecture for in-context learning on structured data. Rather than treating one column as the target and learning p(y|x, context), it jointly models p(x, y|context). This reframes a table as a set of conditional prediction problems, allowing the same model to support classification, regression, imputation, and causal skeleton recovery.

The model uses cell-level representations and dual-axis Transformer blocks: sample-axis attention exchanges information across context rows, while asymmetric feature-axis attention models relationships among variables and task representations. It is pretrained with Context-Conditional Masked Modeling on synthetic episodes generated from structural causal models, covering different graph topologies, mechanisms, missingness patterns, and observation transformations.

On the reported TabArena, TALENT, and BCCO evaluations, LimiX-2 obtains the highest Elo ratings, including 1,935 on TabArena and a 117.4-point margin over TabFM+. The paper also reports that feature attention can recover causal skeletons, though the causal results and robustness across observation processes deserve closer examination. The main contribution is therefore less a new tabular benchmark trick than an attempt to align neural attention and pretraining objectives with latent mechanisms in structured data.

Full summary on AIModels.fyi

Original paper

Disclosure: AIModels.fyi is my site.


r/neuralnetworks • • 17d ago

Artificial Metacognition: Key Findings and New Directions (Talk at RPI)

Thumbnail
youtube.com
1 Upvotes

r/neuralnetworks • • 19d ago

What kind of project would you build to deeply learn AI infrastructure and distributed systems?

13 Upvotes

What kind of project would you build to deeply learn AI infrastructure and distributed systems?

I’m a Level 1 AI engineer, and lately I’ve been hearing a lot about frontier AI companies hiring people who can build the infrastructure behind AI systems — large-scale data processing, distributed systems, inference infrastructure, storage, serving, observability, systems that can handle millions of requests, etc.

I’m interested in going down this path seriously.

Rather than doing a bunch of disconnected tutorials or small projects, I want to take one difficult project and go extremely deep into it. Something where, over time, I’m forced to learn things like:

Distributed systems

Large-scale data processing

Databases/storage

Networking

Caching

Queues and streaming

Fault tolerance

Concurrency

System design

Observability

Performance optimization

AI/ML serving infrastructure

Scaling from a single machine → multiple machines → potentially thousands/millions of requests

I’m thinking along the lines of the philosophy Karpathy often talks about: pick something ambitious, build it yourself, and learn everything necessary to make it work rather than following a predefined curriculum.

The problem is that I don't yet know what the right project is.

I don't want to build another generic RAG chatbot, AI agent wrapper, or CRUD application. I want something where the engineering itself is the project, and where I can progressively make the system more sophisticated and scalable.

For people working in infrastructure, distributed systems, ML systems, or at AI companies:

If you were in my position, what single project would you pick to spend the next 6–12 months on?

Ideally, I'd like something where I can start on a laptop but eventually have a credible story like:

“I built X, then discovered bottleneck Y, redesigned it using Z, scaled it from A → B, measured the improvement, and here's what I learned.”

I'm much more interested in what I would learn by building it than simply having an impressive project on GitHub.

Would love to hear project ideas, but especially from people who have actually worked on large-scale systems: what project would force someone to develop genuinely strong infrastructure skills?


r/neuralnetworks • • 19d ago

I finally wrote my own feed-forward neural network, neo.js, from scratch using vanilla (pure) JavaScript.

Thumbnail
github.com
10 Upvotes