You’ve Got to be Wrong to be (Almost) Right

Gödel, Maimonides, Bayes (and AI)*

During World War II, Allied engineers faced a deadly puzzle: returning bombers were riddled with bullet holes, especially across the wings and fuselage. The obvious response was to reinforce those damaged areas. But statistician Abraham Wald saw what everyone else had missed. The planes being examined were the survivors. The aircraft hit in the places with few or no visible holes—around the engines, cockpit, and other vital systems—were the ones that never came home. Wald argued that armor should be added not where the returning planes were damaged, but where they were not, thereby saving the lives of multiple airmen.

For years, basketball coaches often judged players by the shots that went in: the spectacular fadeaway, the contested jumper, the last-second three. But modern analytics began paying just as much attention to the misses. By charting thousands of failed attempts, teams discovered that some shots – especially long two-pointers – produced poor results even for skilled players. The lesson wasn’t simply to “shoot better,” but to change where you shoot from.

Anecdotal as those may seem, they tell a story, one that goes against our intuition. We tend to imagine learning as the steady accumulation of information, as though ignorance were an empty container gradually being filled, a confirmation process of facts and figures, a series of successive experiences each charging our minds with additional knowledge about the world. While somewhat true, it is the least effective way to learn.

In Information Theory terms, a confirmation usually tells us what we already suspected. A mistake tells us something new. The most valuable signal is often the one we did not expect: the result that breaks the pattern, exposes a hidden assumption, or reveals where our model of reality fails. Success says, “this worked again.” Failure says, “your understanding is incomplete, and here is where.” That is why mistakes can be so disproportionately useful. They do not merely add another fact; they force us to redraw the map.

In layman’s terms, success confirms a hypothesis; failure compresses the hypothesis space.

With all due respect to Claude Shannon (the father of Information Theory), the idea that mistakes carry more information than confirmations goes back a long way.

In the philosophy of Nāgārjuna (sometimes referred to as the Second Buddha), going back 1800 years, wisdom is often approached not by adding ever more descriptions of reality, but by removing the ones that fail. A proposition is examined, its contradictions exposed, and another piece of conceptual territory is ruled out.

For the 12th-century Maimonides, this problem reaches its sharpest form when we speak about God. Positive descriptions are dangerous. To say that God is wise, powerful, merciful, or knowing seems innocent enough, but each word carries a human image. We know wisdom through human minds, power through human bodies and institutions, mercy through human emotion. The moment we use these words positively, we risk turning God into an enlarged version of something already familiar to us.

Maimonides therefore proposes something similar to Nāgārjuna, namely knowledge by negation. We advance not by saying what God is, but by stripping away what God cannot be.

A person may know for certain that a ‘ship’ is in existence, but he may not know to what object that name is applied, whether to a substance or to an accident: a second person then learns that the ship is not an accident; a third, that it is not a mineral; a fourth, that it is not a plant growing in the earth; a fifth, that it is not a living body whose parts are joined together by nature; a sixth, that it is not a flat object like a board or a door; a seventh, that it is not a sphere; an eighth, that it is not a solid piece of wood; a ninth, that it is not incapable of moving on water; a tenth, that it is not hollowed out of a single tree trunk.

The Guide for the Perplexed, Maimonides, Part I, Ch. 60

This famous ship parable makes the structure almost algorithmic. Imagine someone who has heard of a thing called a ship but has never seen one. At first, almost anything might fit the word. Then he learns that a ship is not an accident, not a plant, not spherical, and so on. Each negation removes possibilities. He may still be unable to say exactly what a ship is, but he is less wrong than before.

Knowledge, in this view, grows through subtraction.

The more contemporary 18th-century Presbyterian minister Bayes concluded that learning from mistakes is basically updating beliefs when reality disagrees with expectation. A correct, expected result may only slightly strengthen the hypotheses you already favored. But a surprising mistake can have a much larger effect, because it can suddenly make some hypotheses much less plausible and others more plausible. The learner does not merely record “I was wrong”; it redistributes probability across the whole space of possible explanations.

While Bayes approached empirical learning mathematically, historians believe his underlying motivation was deeply rooted in his role as a Presbyterian minister. A major philosophical debate raged during his life, sparked by the skeptic philosopher David Hume. Hume argued that believing in miracles (like the resurrection of Jesus) based on human testimony is irrational. Hume reasoned that because miracles are inherently low-probability events, any witness testimony is far more likely to be a lie or an error than a reflection of reality.

Bayes’ framework flips this by showing exactly how informative a “mistake” in reality can be. If you have an incredibly strong prior belief that something is impossible, it takes a massive shock to change your mind. But if a piece of evidence is so specific that it would be an absolute mathematical miracle for it to happen by pure chance or error, that “error” becomes highly informative. It forces the Bayesian calculator to shift from “the witness made a mistake” to “my fundamental understanding of the world made a mistake.”

Miraculously, Bayes’ theorem underpins a much more modern school of thought: Machine Learning, which has recently evolved into what we now call AI (Artificial Intelligence).

Machine learning can be imagined as a search through an invisible space of possibilities. Every point in that space is a hypothesis: one possible account of what the data means, one possible rule for distinguishing a cat from a dog, a signal from noise, a cause from coincidence. At first, the space is vast. Almost anything is still possible.

Then experience begins to cut into it. Each example, each correction, each failed prediction draws a boundary through that space, excluding whole regions of explanation. In geometric terms, these boundaries resemble hyperplanes slicing through an N-dimensional landscape. With every cut, the surviving region becomes smaller, sharper, more constrained – a polyhedron gradually carved out of remaining possibilities. A multi-faceted diamond encapsulating refined knowledge.

A 2D Visualization of Hypothesis Space Carved by Hyperplanes

Seen this way, learning is less like filling an empty vessel with facts and more like sculpting. Knowledge emerges through subtraction. What matters is not only what survives, but what has been ruled out. The model becomes more precise because the world has repeatedly told it where it errs.

This makes Maimonides’ account of knowledge, as well as the Bayesian one, unexpectedly modern.

At first glance, the worlds could hardly be further apart: ancient philosophers writing about God (Bayes was also interested in finding God’s work in the world through probability distributions), and contemporary systems learning from data. One speaks in the language of metaphysics, the other in probability distributions, hypothesis spaces, and optimization. Yet both are haunted by the same question: how do we know something we cannot grasp directly?

Note, though, that unlike Maimonides, Bayesian inference softens this picture. Bayes does not always eliminate a hypothesis outright. It changes its weight. Evidence makes some explanations more plausible and others less so. Instead of a world divided sharply into possible and impossible, we get a landscape of degrees of belief. The movement is similar, but gentler. A new observation does not necessarily tell us what the hidden cause is. It tells us which causes have become harder to believe.

In that sense, Bayes, and in a more philosophical form Maimonides, as well as modern Machine Learning, give way to an ancient intuition: knowledge may advance without ever becoming complete possession of the thing itself. What changes is not that the hidden object suddenly becomes “known”, but that our field of plausible interpretations narrows.

This is also where the analogy becomes most interesting, because it begins to fail.

Every machine-learning system assumes a space in which learning takes place. It includes hypotheses, parameters, labels, representations, and latent variables. Even Bayesian reasoning requires a prior: before any evidence arrives, we must decide what kinds of possibilities exist and how to represent them.

The system may not know which point in the space is correct. But it assumes that the truth is somewhere inside the space.

Maimonides is more radical.

His claim is not simply that God is an unknown point on a map. It is that the map itself may be inappropriate.

This changes everything.

If God does not belong to the same conceptual space as created things, then piling up negations does not gradually reveal a hidden object in the ordinary sense. We are not carving a finer and finer statue out of conceptual stone. We are disciplining language. We are discovering the limits of what our categories can (and cannot) say.

Centuries later, Wittgenstein would echo this exact boundary-drawing exercise with his famous “ladder” metaphor – propositions designed to be climbed, understood, and ultimately kicked away once the mind reaches the correct vantage point (how Buddhist of him). In the end, both the medieval theologian and the modern logician arrive at the same radical conclusion: the truest form of intellectual honesty isn’t building massive systems of thought, but mapping the precise borders of our own minds, turning silence into a profound monument to the transcendent.

But here, as the sharp-eyed reader may have already guessed, is where the distance between Maimonides, Wittgenstein, et. al. and machine learning becomes philosophically interesting.

Machine learning is extraordinarily powerful at uncertainty within a representational space. It can distinguish among possibilities, adjust confidence, infer latent variables, and discover structures that no human explicitly programmed. But it still depends on representation. It can learn only what the system can access as a hypothesis, a parameter, a vector, a class, a distribution, or a relation.

Maimonides asks a more disturbing question: what if the object of knowledge cannot be represented within the available space at all?

This is not merely a theological problem. It is also a problem for artificial intelligence.

When we say a model “understands,” “knows,” or “represents” something, we may be doing exactly what Maimonides warned against: taking concepts drawn from our own experience and projecting them onto something whose internal mode of being differs from ours. We observe behavior and then quietly convert behavioral similarity into claims about essence.

Maimonides distinguishes between what something is and what it does. That distinction becomes strangely relevant again in the age of AI. With a complex model, we may know its behavior in immense detail. We can measure inputs and outputs, test responses, map activations, compare representations, and analyze errors. But none of this automatically tells us what the system is in the stronger sense implied by words such as understanding, intention, or meaning.

We know the effects. We infer the hidden structure.

And then we are tempted to mistake inference for revelation. Perhaps this is where Maimonides and Bayes finally meet, and finally part ways.

Bayes teaches us how to reason carefully about what we cannot observe directly. Evidence, especially conceptual errors, changes what we should believe about hidden causes. Maimonides accepts the humility of this movement but pushes it further: what if the very categories we use to construct our hypotheses are inadequate to the thing we are trying to know?

Here lies an intriguing, though imperfect, echo of Gödel. Gödel showed that a sufficiently powerful formal system cannot settle every question expressible within its own language; there will always be propositions that its own rules cannot decide. Maimonides makes no such mathematical claim, but the philosophical resemblance is striking. Both warn against confusing a system’s power with its completeness. A framework may reason impeccably within its boundaries while remaining unable to capture everything that lies at, or beyond, them.

Bayesian uncertainty says: I do not know which hypothesis is true.

Maimonidean negation asks: how do you know that the truth is one of your hypotheses?

That may be the most important question machine learning inherits from medieval philosophy.

As one may remember, we opened by challenging the premise that learning is the steady accumulation of information, but we have since gone a long way. Not only does removing unlikely or wrong hypotheses help us know the world better, but we are also limited by design. We cannot truly grasp what Kant called the thing in itself (though I’ll admit the Kantian view does not exactly align with the concepts discussed so far), but rather a limited, approximated version of it.

Machine learning, in its own way, has rediscovered the power of this movement. It learns from error, exclusion, likelihood, boundaries, distinctions. But it also reveals the limit of that idea. Every act of learning takes place inside some representation of what is possible.

Maimonides, as well as Gödel, and in a softer way, Bayes, force us to confront the possibility that reality may exceed representation.

And perhaps that is the deepest connection between negative theology and artificial intelligence: both are best served with errors; however, they may only lead us to the edge of knowledge, to the point where better inference no longer solves the problem, because the problem may lie in the language of inference itself.

* In fond memory of Gödel, Escher, Bach: An Eternal Golden Braid by Douglas Hofstadter.