By

From The Perceptron To AGI

 

The AI Revolutions That Had to Fail Before Deep Learning Could Win

From Expert Systems to Deep Learning

 

 

The AI Revolutions That Had to Fail Before Deep Learning Could Win

Artificial intelligence did not leap from laboratory curiosities to modern recommendation engines in a single breakthrough. It passed through competing eras of logic, hand-written rules, decision trees, statistical learning and small neural networks before a combination of data, hardware and training methods made deep learning practical.

That history matters because the current AI boom is often described as if scale alone created intelligence. It did not. Each revival solved one problem while exposing another: rules were intelligible but expensive to maintain; trees were transparent but unstable; statistical models were disciplined but dependent on human-designed features; neural networks were flexible but once too difficult to train. The technologies that now appear obsolete were not detours. They were the engineering record of what machine intelligence could and could not do.

Before the boom, machines were taught to reason

The earliest influential AI systems were not trained on oceans of data. They were built around descriptions of a world supplied by people. The ambition was straightforward: identify a specialist’s method of reasoning, express it in a formal language, and let a computer apply that method at speed.

This was the age of symbolic programmes, logic programming and expert systems. In Japan, the Fifth Generation Computer Systems project became the most visible national attempt to push computing towards knowledge representation and logical inference. Prolog, a language associated with logic programming, offered a way to describe relationships and rules rather than write every computational step in the conventional sequence. Lisp, another foundational language, had already become closely associated with AI research because it made it convenient to manipulate symbolic structures.

The promise was attractive. A machine did not need to understand everything if it could reason reliably inside a narrow domain. A medical diagnostic system could work with symptoms, test results and known diseases. A configuration system could translate a customer’s requirements into a technical specification. A chemistry programme could infer molecular structures from observations. Within those boundaries, intelligence looked less like inspiration and more like disciplined bookkeeping.

Stanford’s historical teaching material lists DENDRAL, which inferred molecular structure from mass spectrometry, MYCIN, which diagnosed blood infections and recommended antibiotics, and XCON, which converted customer orders into parts specifications. The same material says XCON had saved Digital Equipment Corporation $40 million by 1986.1 Those examples are evidence that pre-machine-learning systems could have serious industrial value when the problem was bounded and the knowledge could be represented.

But there was a catch that would recur in every later AI cycle. Expertise was not sitting in a neat database waiting to be copied. It had to be extracted from people, translated into rules, tested against exceptions and revised whenever the real world changed. A doctor might say that a symptom “usually” suggests one condition, unless it occurs with another sign, except when the patient belongs to a particular risk group. Turning that practical judgement into a complete rule base was slow. Keeping it current was slower.

The machines were not defeated by a lack of cleverness. They were defeated by the size and instability of the world they were expected to describe.

Rules made expertise visible — and made its limits unavoidable

An expert system typically represented knowledge through production rules: if a set of premises was true, then a conclusion or action followed. The form was plain enough for a non-specialist to recognise. If the engine is turning but the vehicle does not move, inspect the transmission. If a patient has a particular collection of symptoms and laboratory results, consider a defined diagnosis. If a customer selects one configuration, add the compatible components.

The attraction was not only speed. It was inspectability. A user could ask why a system had reached its conclusion and receive a chain of rules. In a period when computers were often treated as mysterious calculators, that ability to expose the path from evidence to decision was a major advantage. It also separated the system from the human expert in a revealing way. The programme did not possess a general understanding of medicine or engineering. It possessed a carefully delimited map of decisions.

That map could include uncertainty. Real specialists rarely deal in absolute rules, so expert systems often attached confidence scores or probabilities to competing conclusions. Such numbers made the systems more realistic, but they did not remove the underlying maintenance problem. A confidence score still depended on the quality of the rule, the evidence behind it and the assumptions built into the domain model.

The historical limitation is sometimes described as the “knowledge acquisition bottleneck”. The phrase sounds technical, but the problem is familiar: someone must still do the work of asking the expert what they know. A programme may execute thousands of rules in seconds, yet the cost of eliciting, formalising and validating those rules can overwhelm the benefit of rapid execution.

The issue was also one of rigidity. A rule-based system could be impressive within its design envelope and helpless outside it. A mechanic can notice that an engine sounds wrong in a way that was never written into the manual. A doctor may recognise that the patient’s combination of age, history and appearance does not fit the textbook pattern. A rule engine cannot quietly import years of experience unless someone has already found a way to encode that experience.

This is why the history of expert systems leads naturally to machine learning. The question changed from “How do we write every rule?” to “Can the machine infer a useful set of rules from examples?” That shift did not make human judgement irrelevant. It moved human labour upstream, into the choice of data, labels, variables and objectives. The computer was no longer asked to recite a finished body of expertise. It was asked to discover a structure that could imitate decisions.

The decision tree was the bridge between hand-built knowledge and learned behaviour

The decision tree offered an unusually clear answer to the problem. Instead of presenting a dense mass of rules, it arranged questions in a hierarchy. At the top was a broad question about the data. Each answer sent the case down a branch. Further questions narrowed the possibilities until the system reached a leaf containing a classification or prediction.

The structure resembles ordinary troubleshooting because ordinary troubleshooting often has the same shape. A doctor does not ask every possible question in random order. A mechanic does not dismantle every component before checking whether the battery is charged. The specialist tries to ask the question that will eliminate the largest number of possibilities at the lowest cost.

A decision tree formalises that instinct. Suppose a system must decide whether a machine is likely to fail. It might first ask whether the temperature is abnormal, whether vibration has increased or whether maintenance is overdue. A poor first question leaves the cases mixed together, so the remaining branches remain uncertain. A useful first question divides the cases into groups with more consistent outcomes. The tree is not “thinking” in the human sense. It is arranging evidence so that each answer makes the next decision easier.

That arrangement gave decision trees two qualities that large neural models later struggled to preserve: readability and a direct relationship between input and output. An engineer could inspect the route to a prediction. A manager could see which variables had been given prominence. A regulator could at least identify the factors the model used, even if the model was not automatically fair or correct.

The method was not perfect. A tree that grows without restraint can memorise its training cases rather than learn a general pattern. Small changes to the data can produce a different structure. A variable with many possible values can appear more useful than it really is. Yet these weaknesses were visible, and visible weaknesses can be tested, pruned and debated.

IBM describes a decision tree as a hierarchical structure made of a root, branches, internal decision nodes and leaves. It also notes the central danger of allowing the tree to grow too far: data become fragmented, and the model can overfit. Pruning removes weak branches and reduces unnecessary complexity.2 The language is less glamorous than today’s talk of foundation models, but the practical problem is the same. A model can be made more elaborate than the evidence justifies.

The tree therefore became a kind of middle ground. It could be written by hand, learned from examples or assembled through a mixture of both. It looked like a decision procedure that a human could understand, while allowing data to determine which questions mattered most.

ID3 turned “what should we ask next?” into a calculation

The influential step was not the invention of the tree itself but the creation of a systematic method for growing one. Ross Quinlan’s ID3 algorithm, described in a 1986 paper, was designed to synthesise decision trees from examples. The paper’s abstract presents the goal with unusual directness: to build knowledge-based systems by inductive inference from examples, while addressing data that might be noisy or incomplete.3

ID3’s central idea was information gain. At each node, the algorithm considered candidate attributes and estimated how much each would reduce uncertainty in the remaining examples. The most informative attribute became the next question. The process was repeated for the resulting subsets, growing the tree from the top down.

The underlying concept, entropy, comes from information theory. A set containing a mixture of outcomes has higher uncertainty than a set in which nearly every example belongs to the same class. A good split takes a mixed set and produces cleaner groups. If the examples are applications and the outcome is approval or rejection, an effective question might divide the cases into branches where the results become much more consistent. A weak question merely rearranges the confusion.

This sounds abstract until it is placed beside an expert’s working method. The expert may not calculate entropy, but the expert is trying to find a question that changes the decision. ID3 replaces intuition about the next question with a numerical comparison among available attributes. It does not understand the domain. It identifies a feature that best separates the labelled examples supplied to it.

That distinction is important. ID3 did not discover truth in a vacuum. It discovered a compact decision structure in a particular dataset. If the examples were incomplete, biased or poorly labelled, the tree could make the same defects look official. If the target variable captured a bad policy, the algorithm would optimise the wrong policy with impressive efficiency.

Quinlan’s original work also recognised that the simple procedure had weaknesses. The Springer record says the paper discussed noisy and incomplete information, a reported shortcoming of the basic algorithm, and ways of overcoming it. Later methods such as C4.5 extended the family, while CART used a different splitting criterion based on Gini impurity. IBM identifies ID3 as the predecessor of C4.5 and explains that both information gain and gain ratio can be used when selecting splits.4

The achievement of ID3 was therefore more modest and more consequential than the mythology of AI usually allows. It did not create an artificial expert. It provided a repeatable answer to a practical question: given these examples, which question should be asked first?

Statistical learning won the 2000s by demanding less from the machine

By the beginning of the 2000s, neural networks had not disappeared, but they had lost their position as the obvious future of machine learning. Their advocates had to contend with limited computing power, relatively small labelled datasets and difficult training behaviour. The deeper the network, the harder it was to make learning signals reach the earlier layers. Gradients could become too small, weights could be badly initialised and training could consume resources that were not available to ordinary research groups.

Other methods looked more dependable. Support-vector machines offered a mathematically disciplined approach to separating classes, especially when a useful kernel could map the data into a more convenient space. Boosting combined many weak learners into a stronger ensemble. Random forests reduced the fragility of an individual decision tree by averaging across many trees built from varied samples or feature choices. These methods did not promise that a computer would learn every useful representation from raw data. They made a more limited promise and often delivered it.

The trade-off was clear. Traditional machine learning commonly depended on feature engineering: a human decided which measurements, transformations or descriptors should be presented to the model. In image recognition, that might mean designing features that captured edges, shapes or textures. In speech, it meant converting sound into a representation that exposed relevant frequencies. In commercial prediction, it meant deciding which customer behaviours counted as meaningful signals.

That human preparation could look like a weakness once deep learning became successful. In the 2000s, it was often a practical advantage. A carefully selected feature set could prevent a model from wasting capacity on irrelevant variation. A smaller model could train on smaller datasets. A method with clearer statistical behaviour could be easier to validate in an application where a wrong answer had a cost.

The neural network’s problem was not that it lacked expressive power. Even comparatively small networks could represent complicated relationships. The problem was finding useful internal representations and adjusting the parameters without getting lost. The machinery needed for modern deep learning required better ways to initialise and train many layers, larger quantities of data, faster numerical computation and architectures suited to the structure of the input.

A 2015 review in Nature described this contrast in plain terms. Conventional machine-learning systems had required careful engineering and domain expertise to create feature extractors, while deep learning aimed to discover representations from raw data through multiple processing layers.5 That division between human-designed features and learned representations became the central fault line in the decade’s argument over what neural networks could become.

The scepticism towards neural networks was not irrational. It reflected the conditions under which researchers were working. An algorithm that needs more data, more computation and more delicate training than its competitors is not a revolution. It is a research project waiting for its missing infrastructure.

The 2006 revival was a repair job, not a miracle

The return of deeper neural networks began with techniques that made the old ambition less fragile. Around 2006, researchers associated with the Canadian Institute for Advanced Research helped revive interest in deep feedforward networks by showing that layers could first be trained in an unsupervised fashion and later fine-tuned with backpropagation.

The approach addressed a blunt practical problem. Instead of starting a deep network with essentially arbitrary weights and asking labelled data to teach every layer at once, researchers could train one layer to model the patterns produced by the layer below. A restricted Boltzmann machine could learn relationships among units. Stacking such layers created a deep belief network, with each stage forming a more elaborate representation. Once the stack had been sensibly initialised, supervised training could adjust the whole system for the final task.

The value of pre-training was greatest when labelled data were scarce. Unlabelled examples could still help the network learn the general shape of the input, even when only a smaller subset had been assigned human labels. The machine was no longer forced to learn perception and classification from the same narrow supply of labelled cases.

The 2006 paper by Geoffrey Hinton, Simon Osindero and Yee-Whye Teh described a fast, greedy method for learning deep belief networks layer by layer. Its abstract says the method used complementary priors and trained the layers successively, with fine-tuning producing a generative model for handwritten digits and their labels.6 The important point was not the terminology. It was the demonstration that depth could be made workable.

The revival also changed the interpretation of neural-network failure. Vanishing gradients and unstable optimisation were no longer evidence that layered learning was fundamentally misguided. They were engineering problems that could be reduced by better initialisation, better activation functions, better architectures and better hardware. That was a much more dangerous conclusion for the established methods than a single benchmark victory would have been.

Yet pre-training should not be turned into a founding myth. It did not alone create modern deep learning. It helped bridge a period in which researchers had deep models but insufficient tools. Later progress would make the unsupervised stage less central in many applications, as improved activation functions, optimisation methods, regularisation and abundant labelled or weakly labelled data changed the balance again.

The lesson is broader than the fate of one algorithm. A technology can look obsolete when the surrounding system is immature. Neural networks were not waiting for one magical idea. They were waiting for several ordinary improvements to arrive at the same time.

GPUs supplied the missing scale

The second part of the repair involved computation. Neural networks perform large numbers of matrix operations, and graphics processing units were designed to carry out many similar calculations in parallel. Once GPUs became programmable enough for machine-learning research, training times could fall from an academic nuisance to a manageable engineering cost.

The Nature review records the practical effect. It says that the arrival of fast, programmable GPUs allowed researchers to train networks 10 or 20 times faster. It also links the combination of pre-training and GPU acceleration to record results in speech recognition, including systems that later reached Android phones.7

That change had consequences beyond speed. Faster training allowed researchers to run more experiments, compare more architectures and recover from failed choices without losing weeks or months. It made larger datasets usable and enabled the repeated tuning that complicated models require. A method that is theoretically possible but practically too slow will not win the field. Researchers select the methods they can test.

Data completed the triangle. A small neural network may contain thousands of adjustable parameters, but a modern deep-learning system can contain vastly more. The more flexible the model, the greater the danger that it will memorise the training examples. More data gives the model more opportunities to learn patterns that generalise, provided the data are relevant and the objective is sound.

This is why the usual story about deep learning is incomplete when it credits algorithms alone. The breakthrough depended on a relationship among algorithmic ideas, datasets and hardware. Remove any one of the three and the result becomes less dramatic. Better training without enough examples leaves the model starved. More examples without adequate computation make experimentation expensive. Faster hardware without a model that can use the data produces a large electricity bill and little insight.

The combination also altered the economics of research. A university group with access to specialised hardware could explore models that once belonged to theoretical papers. Companies operating online services possessed enormous streams of images, text, speech and user interactions. Recommendation systems benefited from this environment because they did not need to produce an explanation in the form of a visible tree at every step. They needed to predict what a user might click, watch or buy at a scale where even small improvements could matter.

That application exposed a new tension. The systems became more powerful as their internal representations became less legible. The decision tree made its questions visible. The deep network absorbed layers of features into numerical parameters that were difficult to translate back into human reasoning. The field gained performance but lost a familiar route to accountability.

The trade was not always unacceptable. A recommendation engine can be tested through outcomes, and a modest error may be tolerable. A system used in healthcare, employment, credit or public safety faces a different standard. The same engineering culture that celebrates a lower error rate must also ask which errors have been reduced, which remain and who bears them.

The neural network’s advantage was representation, not imitation of the brain

The word “neural” encouraged a misleading picture. Artificial neural networks are inspired by some abstract features of biological neurons, but their practical success did not depend on recreating the brain. Their strength came from composing many mathematical transformations so that raw inputs could be converted into increasingly useful representations.

In an image system, early layers might respond to edges. Later layers can combine edges into motifs, motifs into parts and parts into objects. In speech, the hierarchy can move from acoustic patterns to phonetic units, words and larger structures. In text, representations can capture relationships that are not visible in a simple list of tokens. The machine does not need a human engineer to specify every intermediate feature if the training process can discover features that improve the objective.

The Nature review calls this representation learning and describes deep learning as a stack of non-linear modules that transform raw data into increasingly abstract representations. It argues that this capacity is especially valuable for high-dimensional data, where a shallow classifier operating on raw inputs cannot easily distinguish meaningful variation from irrelevant change.8

That explains why neural networks could eventually displace methods that had been more reliable on modest datasets. The competition was not simply between a tree and a network, or between one classifier and another. It was between two ways of deciding who would design the representation. Traditional methods placed more responsibility on people. Deep learning placed more responsibility on the data and the optimisation procedure.

Convolutional neural networks made the arrangement especially effective for images and other structured arrays. Local connections let the system focus on nearby patterns. Shared weights let it detect the same motif in different locations. Pooling reduced sensitivity to small shifts. These choices were not decorative biological analogies. They were ways to impose useful structure on the learning problem while preserving the ability to learn from examples.

The advantage of deep learning, then, was not that it had escaped the need for assumptions. It had changed where the assumptions lived. A convolutional network assumes that local patterns and repeated structures matter. A decision tree assumes that sequential splits can separate useful categories. A support-vector machine assumes a particular geometric treatment of the data. No model is free of design choices.

This is the part of the AI story that fashionable language often obscures. “The machine learned” does not mean the machine learned without guidance. Someone selected the architecture, the data, the labels, the loss function, the hardware and the evaluation standard. The modern system may have fewer hand-written rules than an expert system, but it has not escaped human judgement. It has distributed that judgement across the pipeline.

Recommendation engines made the quiet revolution visible

The public encountered the new machine-learning era not first through philosophical claims about intelligence but through ordinary acts of selection. A video appeared in a queue. A product was placed beside another product. A search result was ranked. A news story or social-media post was pushed towards one user and away from another.

Recommendation engines were a natural home for the transition from older statistical models to deeper networks. Their raw material was abundant: clicks, purchases, viewing time, searches, ratings, pauses and returns. Their task was measurable: predict engagement, relevance or the probability of a future action. Their environment rewarded systems that could update and scale.

Earlier methods were not useless. Collaborative filtering, regression, decision trees and ensembles could all contribute. But deep models could learn representations of users and items, allowing the system to connect patterns that were difficult to encode through fixed rules. A user did not have to declare that they liked a particular category if their behaviour placed them near other users or products in a learned space.

The result was convenient, but it also made the model’s objective more consequential. A system trained to maximise clicks may learn to favour material that provokes attention rather than material that is accurate or beneficial. A recommendation engine can narrow a user’s choices by repeatedly showing what resembles the past. A platform can call this personalisation while quietly optimising for time spent or advertising value.

The old expert system at least announced its limits. Its rules were written for a domain and could be challenged line by line. A recommendation system may be more accurate in aggregate while giving the individual no clear account of why a particular item was selected. Its hidden representation can absorb historical preferences, commercial priorities and social patterns without presenting any of them as an explicit rule.

That does not make transparent models automatically fair. A visible tree can encode discrimination as easily as an opaque network. It does mean that interpretability is not a cosmetic preference. It is part of deciding whether the model can be governed. When a model affects a person’s access to credit, work, education or healthcare, the question is not only whether it predicts well. It is whether the institution can explain the decision, test its consequences and correct it when the target itself is wrong.

The Stanford lecture material makes this point in a different way. It warns that machine-learning systems are used in areas such as education, credit, employment, advertising, healthcare and policing, and states that an algorithm is not an acceptable excuse for mistakes or unwanted consequences.9 That warning belongs in any history of AI because the technology’s social importance is determined by where it is deployed, not by how impressive its benchmark score looks.

Every AI revival carried its own fashionable claim

The pattern is now familiar. A new method is presented as the answer to the failures of the previous method. Its early successes are treated as evidence of a general transformation. Investment follows. The limits become visible. Interest fades or changes shape. Later researchers recover the useful parts and combine them with new infrastructure.

The Fifth Generation project did not create general reasoning. Expert systems did not capture all of human expertise. Decision trees did not eliminate the need for judgement. ID3 did not solve noisy or biased data. Neural networks did not become dominant in the 1980s simply because they were flexible. Deep learning did not make assumptions disappear.

But calling these projects failures without qualification is lazy history. Logic programming contributed languages and methods for representing structured relationships. Expert systems demonstrated that narrow AI could deliver industrial value. Decision trees made learned behaviour inspectable and created a practical vocabulary for classification. ID3 showed how information theory could guide the construction of a decision procedure from examples. Statistical learning clarified the importance of generalisation, regularisation and the relationship between model complexity and data. Neural-network researchers preserved ideas that later became central when hardware and data caught up.

The real failure was usually the claim that one method would be enough. Intelligence is not one problem. A system that must calculate a tax liability, recognise a face, diagnose a rare disease, recommend a film and control a vehicle confronts different data, risks and standards. The model that is useful in one setting can be reckless in another.

This is why the history matters to the present debate. The current language of AI often collapses distinct tasks into one grand promise. It speaks of systems that reason, understand, create and act, then measures progress through a collection of benchmarks that may not capture reliability in the world outside the test set. The earlier eras are a warning against confusing a successful demonstration with a solved discipline.

The old methods also remind us that simplicity has a strategic value. A small rule set can be audited. A short decision tree can be challenged. A statistical model can be recalibrated. A neural network can be powerful enough to justify its opacity in one application and needlessly opaque in another. The right question is not whether a system is old-fashioned or state of the art. It is whether the system’s behaviour can be tested against the consequences that matter.

The next revolution will still inherit the old arguments

The progression from expert systems to ID3 and then to deep learning can look like a march from rules to statistics to scale. It is better understood as a continuing argument over where intelligence should reside.

In an expert system, intelligence resides in explicit rules supplied by specialists. In a decision tree, it resides in the sequence of questions induced from labelled examples. In a neural network, it resides in distributed numerical representations adjusted through training. Each arrangement changes the balance between human control, data dependence, computational cost and interpretability.

None of those balances is permanent. Modern systems already combine them. A recommendation platform may use neural embeddings, tree-based ranking, explicit business constraints and hand-written safety rules. A medical tool may pair a learned image model with a structured clinical workflow. An AI assistant may generate fluent language while relying on retrieval systems, deterministic checks and human approval. The boundaries between the old schools have become less important than the engineering problem of making them work together.

That synthesis also revives a question that expert systems never answered: what should the machine be allowed to decide? An algorithm can identify patterns, rank alternatives and estimate probabilities. It cannot settle the legitimacy of the objective by optimising it. If the objective rewards speed, the system may sacrifice care. If it rewards engagement, it may elevate provocation. If it rewards approval rates, it may learn to avoid difficult cases. No amount of mathematical sophistication turns a chosen target into a moral fact.

The history of AI offers a less exciting but more reliable definition of progress. Progress is not the disappearance of limitations. It is the ability to identify them earlier, measure them more honestly and build systems that fail in ways people can detect and correct.

The quiet architecture beneath the AI boom is therefore not one architecture at all. It is a layered record of discarded promises, repaired techniques and recurring trade-offs. The systems that once seemed too narrow, too brittle or too small supplied the questions that later models had to answer. The machine did not suddenly become an expert. Researchers kept changing the way expertise was represented, learned and tested until the hardware, data and mathematics finally made a larger ambition possible.

Sources

  1. Stanford University, CS221 Lecture 19: Conclusion and History of AI.
  2. IBM, “What is a Decision Tree?”.
  3. J. Ross Quinlan, “Induction of decision trees”, Machine Learning, 1986.
  4. IBM, “What is a Decision Tree?”.
  5. Yann LeCun, Yoshua Bengio and Geoffrey Hinton, “Deep learning”, Nature, 2015.
  6. Geoffrey E. Hinton, Simon Osindero and Yee-Whye Teh, “A fast learning algorithm for deep belief nets”, Neural Computation, 2006.
  7. Yann LeCun, Yoshua Bengio and Geoffrey Hinton, “Deep learning”, Nature, 2015.
  8. Yann LeCun, Yoshua Bengio and Geoffrey Hinton, “Deep learning”, Nature, 2015.
  9. Stanford University, CS221 Lecture 19: Conclusion and History of AI.

This post contains affiliate links. If you purchase through these links, I may earn a commission at no extra cost to you.

 

Leave a Reply

Discover more from Thoughts on Technology

Subscribe now to keep reading and get access to the full archive.

Continue reading