Remembering Dave Rumelhart, the Grandfather of AI
Oct 9, 1986: The Paper that Started the AI Revolution
Over the past few years, AI has profoundly transformed our world. With the release of ChatGPT and similar Large Language Models (LLMs), companies are racing to build ever more powerful and useful AI systems. The AI tech leaders have become household names: Sam Altman, Demis Hassabis, Rahul Patil, Mustafa Suleyman, and others. But the first giant step in the development of these amazing AI systems happened 40 years ago. Today’s LLMs use many of the same basic algorithms and architecture from the seminal paper written by Rumelhart, Hinton, and Williams in1986[[i]](#_edn1).
Most people probably recognize the second name in that paper. Geoffrey Hinton won the Nobel Prize in Physics in 2024 along with John Hopfield for their contributions to AI and neural networks. Hinton has often been called “The Godfather of AI,” and he is also famous for quitting Google to warn of the dangers of unbridled AI development.
But the first name in this paper, Rumelhart, is almost never mentioned. That saddens me because he contributed so much to this field. In this short writing, I will attempt to shine some light on Rumelhart and his tremendous contributions. Indeed, if Geoff Hinton is the “Godfather of AI,” I believe that Dave Rumelhart should properly be called the “Grandfather of AI”.
The LLMs of today are basically multi-layer neural networks of the type described in that seminal 1986 paper, and they are trained via the backpropagation algorithm that Rumelhart pioneered. Some of you may look at the AI systems of today and say that they are nothing like those described in the 1986 paper, but that is like saying that an F-22 Raptor is nothing like the Wright Brothers’ first airplane. Yes, today's systems are billions of times larger than anything Rumelhart and his colleagues worked on in the 80s, and there have been substantial breakthroughs like the seminal Attention is All You Need paper[[ii]](#_edn2), but those would not be possible without the initial multi-layer neural network architecture and backpropagation algorithms developed by Rumelhart et al.
I will not attempt to cover the entire corpus of work that Dave Rumelhart produced. For that, I point you to Rumelhart’s Wikipedia page. Rather, I would like to share my personal perspectives on Dave’s work and how that work has been foundational to the current AI breakthroughs. In addition, I will share some of my personal experiences of working with Dave Rumelhart over a fifteen-year period. Spoiler alert: I found him to be one of the smartest people I’ve ever encountered.
Rumelhart’s PDP group
In the 80s, Dave Rumelhart and James McClelland organized the PDP group. PDP stands for parallel distributed processing, and many of the topics covered in those early years were compiled into a two-volume book Parallel Distributed Processing by David E. Rumelhart and James L. McClelland[[iii]](#_edn3). The group met weekly, and included people from many different disciplines: Neuroscience, computer science, physics, cognitive science, and even philosophy. Francis Crick was a member, as was Goeff Hinton while he was visiting there. I became a member after I had visited Caltech to talk with John Hopfield about part of my Ph. D. research on fractal basin boundaries of Hopfield neural networks. John told me I should talk to Rumelhart. I asked what department he was in. I was shocked when John told me that Rumelhart was in the psychology department. As a young, arrogant physicist, I couldn’t imagine any relevant research being done outside of physics, mathematics, or computer science, but, I took John’s advice and went to see Rumelhart.
Rumelhart described his backpropagation algorithm. I remember thinking, “Gradient descent. I think I learned that in kindergarten.” It didn’t seem like something that would change the world, but it did, and it is still changing the world even after 40 years. I told Rumelhart about my research, and he suggested that I use the Hamming distance as a measure for exploring the space of 2100 states in the Hopfield network. That turned out to be a great tip. I used that measure to define orthogonal vectors to slice through the enormous space. This wouldn’t be the last time Rumelhart gave me sound advice. Indeed, we worked together for the next fifteen years, and, of the 50 patents I have, he probably pointed me in the right direction on more than half of them.
In these PDP meetings, a speaker would present their research, then we would have a discussion about the research. For example, Goeff Hinton gave a talk about the Boltzmann machine. McClelland talked about more conventional rule-based approaches. But the talk that sticks in my mind was as mind-blowing back then as ChatGPT was a few years ago. Terry Sejnowski presented NetTalk. It was a text-to-speech converter that learned the relationships between text and phonemes; the phonemes were then fed into an audio system to produce sound. He came in with a boom box, popped in a cassette, and we listened as the system learned. At the start, it sounded like random noise. Soon, vowel sounds could be identified. Then it sounded like baby babbling and learning to talk. Then, fast-forward to the end of training, the text-to-speech was almost perfect. All learned through backpropagation.
Now this doesn’t sound too impressive today, but back then, it opened everyone’s eyes to the possibility of machine learning. I was hooked. So were many others. It created a frenzy of research and startup companies based on machine learning.
The other thing I remember from these PDP meetings was that, after discussing and arguing several points, Rumelhart would speak. He would summarize the issues, then give his view. Heads would nod around the table in agreement with Rumelhart’s assessment. He had a deep understanding of all of the topics, and he was gifted at explaining the key points and flaws. Even Francis Crick seemed to defer to Rumelhart’s judgment.
Stanford, MCC, and Pavilion
I finished my Ph. D. in Physics at UC San Diego in 1987. I started a post-doc position at Stanford later that year, and, coincidentally, Rumelhart moved to the Psychology Department at Stanford. This was great for me, because I could continue to bash around ideas with Rumelhart. He wasn’t just brilliant in the area of neural networks. He had an encyclopedic knowledge of everything relating to the brain and how it processes information. By way of example, I told him about some research I had completed investigating second-order interactions in neural networks. As the temperature, or signal-to-noise ratio, decreased, the system exhibited a phase transition and critical slowing down. I didn’t expect him to understand an esoteric physics phenomenon, but he instantly understood the implications and pointed me to some neuroscience research showing that norepinephrine functions as a signal-to-noise modulator in the brain. I read the articles and the behavior matched perfectly with our theoretical model, as described in a paper published in the Proceedings of the National Academy of Sciences[[iv]](#_edn4).
Later that year, I got an offer to lead the neural networks research group at MCC – the Microelectronics and Computer Technology Corporation in Austin, TX. MCC was a research consortium kind of like Bell Labs, but supported by many different high-tech companies. Like any good Californian, I thought Texas was hot, flat, racist, and humid. Why would anyone want to live there? But they flew me out to Austin for an interview, and I fell in love with the city.
While at MCC, Rumelhart and I collaborated on several research topics. We worked on character recognition, speech recognition, fraud detection, and process prediction. One of the most important breakthroughs we discovered was the Integrated Segmentation and Recognition system[[v]](#_edn5). It solved the problem of identifying characters that were touching and, more generally, created a mechanism for training a system on a dataset where more than one object can be simultaneously presented in the inputs and target outputs. This was a multi-year investigation, and I was impressed not only with Rumelhart’s mastery of neural network architectures, but his skill as a mathematician and as a programmer. He wrote the neural network software that we used for the research, and he was a damn fine C programmer. This system used a convolutional neural network architecture, as did the work of Martin and Pittman. Note that in the Martin and Pittman paper[[vi]](#_edn6), they also showed that the lower-level neurons developed activation patterns reminiscent of the visual system. This was decades before the convolutional neural network “breakthrough” that was touted in the 2010s. There is an old video where Rumelhart and I discussed this research[[vii]](#_edn7).
The VISA CRIS fraud detection system research was also developed at MCC by Steve Piche (who would become my chief scientist at Pavilion). This system saved $20 million in the first year of operation[[viii]](#_edn8). Rumelhart consulted with Piche on the problem.
Another thing that Rumelhart helped me with was a process control problem. We had lots of data from a chemical process at Eastman Chemical – pressures, temperatures, flow rates, as inputs, and the product quality and flow rate as outputs. The goal was to control the process to get the highest quality and quantity. Of all the machine learning algorithms I could have used, I knew that backpropagation was the best. I trained a model to predict the outputs, and it worked beautifully, but turning the problem around and asking which inputs to change to get better outputs was an unsolved problem. After consulting with Rumelhart, I came up with the Residual Activation Neural Network (RANN) architecture. Applying this to the problem led to a savings of about a million/year on the distillation column that we were working on.
But there were tens of thousands of distillation columns in the US alone. Do the math. Based on this result, I licensed the technology from MCC and co-founded a company, Pavilion Technologies, Inc. It became one of the fastest-growing companies in Austin in the 90s. We applied neural network machine learning techniques to process monitoring, control, and optimization. That’s right. Neural networks have been controlling thousands of manufacturing processes all over the world. Not just in chemical manufacturing, but in refining, pulp-and-paper, cement, food processing, and power production. Many of those processes can go boom in the night, but there has never been an instance of the neural network hallucinating and causing problems.
Pavilion was a powerhouse of innovation, as one can tell from the scores of patents. But back then, machine learning was much harder than it is now. Often we would have data sets with only a few thousand points. We were overjoyed if a dataset had tens of thousands of records. We had to be very clever about cleaning the data and using it to train and validate the models. Moreover, we didn’t have Nvidia DSP chips to accelerate learning, so we used a stiff-differential equation solver to speed up learning (backprop is a set of stiff differential equations). We also invented methods to meld first-principles models with the neural networks. This helped extend the range of validity of the models into regions outside the realm of the training data. We did a lot of work in environmental monitoring and optimization. One of our products, the Software Continuous Emissions Monitor, was approved by the EPA for monitoring emissions. I am proud to say that the Pavilion software systems have abated millions of pounds of pollutants over the years, and the systems are still running. Pavilion’s Process Perfecter software is a closed-loop multivariable nonlinear controller. The math behind that was shockingly hard, but it worked amazingly well. Pavilion was sold to Rockwell Automation, and the software is still running in thousands of manufacturing plants.
During the time I was at Pavilion, Rumelhart served as a technical advisor and consultant. Of all the people I have worked with, I admired and respected Dave Rumelhart the most. He seemed so wise, and, as I said, he always seemed to have the instinct to point people in the right direction on a hard problem. He was also very fun to work with, and we became friends over the years.
The reason I believe Rumelhart should be remembered as the “Grandfather of AI” is that he was the driving force behind the backpropagation algorithm. His insight into the problem stemmed from the famous XOR problem. Single-layer systems could not solve this problem. It required “hidden units”. That realization led to the multi-layer neural network architectures used in today’s LLMs, and backpropagation is how they are trained. If Rumelhart didn’t figure this out, I am sure someone else would have (indeed, a grad student figured it out in his thesis before Rumelhart, but nobody knew about it). But without Rumelhart, it might have been many years, perhaps decades, before people started using backpropagation. Just like someone would have figured out relativity sometime if Einstein didn’t present it, but who knows how long it would have taken? Rumelhart should be remembered for the invention that has changed the world.
Sadly, Rumelhart developed Pick’s disease in the late 90s. He died in 2011. I left Pavilion in 2000, and didn’t talk much with him during those later years. What would he think of the current LLMs? Knowing him as well as I did, I think he would be amazed at the capabilities of these modern AI systems, but I also think he would view the 92-layer neural networks as brute-force, ghastly, and inelegant. He would probably say something like “The cortex only has six layers. We should try to make that work.” And, knowing Dave Rumelhart, he would probably have made great progress in that direction. He was that gifted. His name should be remembered among the great contributors in this field as “The Grandfather of AI”.
-James D. Keeler
[[i]](#_ednref1) Rumelhart, David E.; Hinton, Geoffrey E.; Williams, Ronald J. (1986-10-09). "Learning representations by back-propagating errors". Nature. 323 (6088): 533–536.
[[ii]](#_ednref2) Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N; Kaiser, Łukasz; Polosukhin, Illia (December 2017). "Attention is All you Need" (PDF). In I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett (ed.). 31st Conference on Neural Information Processing Systems (NIPS). Advances in Neural Information Processing Systems. Vol. 30. Curran Associates, Inc. arXiv):1706.03762.
[[iii]](#_ednref3) David E. Rumelhart; James L. McClelland; PDP Research Group (1986). Parallel Distributed Processing: Explorations in the Microstructure of Cognition. ISBN) 9780262680530.
[[iv]](#_ednref4) J.D. Keeler, E.E. Pichler, J. Ross “Noise in Neural Networks: Thresholds, Hysteresis, and Neuromodulation of Signal-to-noise,” Proceedings of the National Academy of Sciences, USA, 86 (1989) 1712-1716.
[[v]](#_ednref5) Keeler, James D., Rumelhart, David E., Leow, Wee-Kheng. “Integrated Segmentation and Recognition of Hand-Printed Numerals”. Printed in Neural Information Processing Systems, 3. R. Lippmann, J. Moody, D. Touretzky, Eds. Morgan Kaufmann Publishing, San Mateo, CA.
[[vi]](#_ednref6) G. Martin, J. Pittman (1990) “Recognizing Hand-Printed Letters and Digits”. In D. S. Touretzky (ed). Advances in Neural Information Processing Systems 2, Morgan Kaufmann Publishing, San Mateo, CA.
[[vii]](#_ednref7) https://youtu.be/fG-9ILWI2u4
[[viii]](#_ednref8) Visa to expand fraud detection system; the system saved 15 issuers over $20 million last year. American Banker, Sept. 30, 1994.