When I went to UCLA, I had heard professor Pearl’s name here and there from the other professors. It was surreal getting to interview him about his career and thoughts about AI today. He had a few anecdotes I really loved like:
Why he loves physics and who the greatest scientist of all time is
Proving Alpha-beta pruning is mathematically optimal which Donald Knuth was surprised at
Discovering Bayesian networks and the eureka moment he had
You can tell how much passion he has for science in just the first few minutes of the conversation. Hope you enjoy it!
Check out the episode wherever you get your podcasts: YouTube, Spotify, Apple Podcasts.
Timestamps
11:17 - Greatest scientist of all time
20:15 - What people thought of AI in the 80s
26:23 - Entering academia and researching AI
34:52 - The invention of Bayesian networks
46:28 - Pioneering work in causality
01:20:12 - A restless mind pays
01:24:36 - Advice for his younger self
Transcript
00:54 — How he got into AI
Ryan:
[00:54] I looked into your educational background, and it was all electrical engineering, physics. I don’t see the connection to causal reasoning and artificial intelligence. How did you get there?
Judea:
[01:06] Everything connects. Everything connects. Everything connects to the day I was born. Hit on the head first. I have to start by saying that we. I grew up in Mandate in Israel prior to 1948, prior to the establishment of the state of Israel. And we had a very excellent high school education. My high school teachers were professors that were chased by Hitler from Germany, from highly reputable universities like Heidelberg and Berlin.
[01:58] And they came to Israel in the 1930s, and they didn’t find any academic position at that time. They were rare, and they started teaching high school. But they were really quality professors. They could teach anything without notes, from the Economy of Manchuria to the proof of Pythagoras theorem, with no notes and no stop. And we were, we—I mean, my generation was lucky enough to be beneficiary of this educational experiment.
[02:46] So we were taught science from a human viewpoint, chronologically, the way things were discovered, by whom they were discovered, at what period. Why was there a question about certain mathematical proof? What the inventor of the proof knew, what he didn’t know, and what he asked himself in the context of the historical situation at that time. That’s the way to teach science, not the recipe of algorithms and techniques, but as a.
[03:37] From the viewpoint of the human actor. And the transfer of ideas from one human to another is a human struggle, not a human struggle to decipher the secrets of nature.
Ryan:
[03:54] What’s the advantage of learning in that style?
Judea:
[03:58] The advantage is that you see yourself as an actor. You, the student, you are also puzzled by many things. But look what he did. Look what Pythagoras did. He was as puzzled as you were, and he took that route. Perhaps when you’re puzzled, of course, your puzzles are not as magnificent as his, but still, look what he did. He took that route around things, and he consulted some other work of some other.
[04:34] He struggled and you struggled, and you are part of science. That is the basic idea. Give students the idea that you are part of science, not an observer, not a passive observer in science, not a recipient, but an actor. And that, I think, was unique, very valuable for me. We indeed got the idea that each one of us can find another proof of Pythagoras’ theorem that no one else has thought about. And each one of us has the potential of becoming an Einstein or Pythagoras.
[05:18] They give us this illusion. Okay, it was a useful illusion. I know I’m just a pebble looking at Pythagoras. But that illusion helped me be a little, I would say not contrarian, but assertive. And we all grew up in assertive mode of learning science. We insisted on understanding things our way in real time. If the teacher went too fast, we made noise. We made noise with our chairs and with everything we could, okay?
[06:07] And the teacher stopped and slowed down to make us all understand things our way in real time. So that was part of the. It’s a mood of me and my generation. And it, I. I can see that it affected my life. I had one story that I remember. It made an impression on me. I look back and I say, well, I. It started very early. It started at age 10 when we learned about how to calculate areas and volumes. And here there was a question in class.
[06:50] How many dunams are there in a square kilometer? Dunam is 1,000 meter square. It was a Turkish unit for measuring areas. So the entire class said, Dunam is a kilometer square. And I screamed and said, no, it’s a thousand. You have 1,000 dunams in a kilometer square. And the teacher sided with the laughing class, and they all mocked me and ridiculed me. And I went home and I said, they are wrong, and I’m going to come back tomorrow and insist on that.
[07:43] And the teacher apologized to the class, and I felt that, yes, you have to insist on your understanding of things. Yes, I’m telling you that because maybe this one made me into a noncompromise. And this is part of my childhood, a background which might explain how I got into artificial intelligence. So I took engineering after being a farmer in the army. Okay, farmer, yes. In Israeli army you have troops which are spending their time, half and half, half in military training and half in farming, being part of a kibbutz.
[08:40] It’s the old idea of one hand holding the plow and the other one the rifle. And you succeed.
Ryan:
[08:50] Did you ever shoot anyone?
Judea:
[08:52] I almost shot someone without seeing him or her. But it turns out the next morning that it was a fox. But the steps of the fox were very, very similar to the steps of terrorists advancing toward us. So that’s as far as I got to shooting. So I got into the Technion to study electrical engineering again. We had great teachers. Again, we got the idea that we are making science. So we studied physics very seriously.
[09:34] And I liked what we studied. Yeah, I wasn’t the first in class. No, it was third or fourth. Always never match the geniuses. The geniuses were bored in class. They knew what the teacher was going to do, what he was going to ask in the exams. And everything was boring to them. I wasn’t bored. I wasn’t the first, but I was bored.
Ryan:
[10:07] What interested you in physics? What did you like that made you so passionate?
Judea:
[10:13] What I liked was that, sitting on your chair, you can predict things in physics. Like Maxwell, who sat on his chair and said, “That looks like a wave equation. Look, let me calculate its velocity. It looks like the velocity of light. Maybe light is nothing else but electromagnetic wave,” you know, on his armchair. He didn’t do any experiment. That excited me. Yeah. My wife told me one night I woke up and I said Maxwell was wrong.
[10:57] Maxwell was wrong. She quieted me down in the next morning, said he was right. By the way, you.
Ryan:
[11:09] You mentioned a few scientists, and it seems like you know a lot about the history of the old scientists.
Judea:
[11:16] Yes.
11:17 — Greatest scientist of all time
Ryan:
[11:17] Do you have a favorite scientist of all time? And why?
Judea:
[11:21] When we studied and I looked at geometry, I got fever. Really? Fever.
Ryan:
[11:31] Physical fever.
Judea:
[11:32] Yeah. I couldn’t get over the idea that you can do in algebra all the geometric constructions that we learned on, you know. So I thought that Descartes was the greatest mathematician ever lived. It was so enormous to me. Transformation from geometric constructions to algebraic derivations. It was unbelievable.
Ryan:
[12:01] I took that for granted when I was in education. What is it about that that’s so astonishing?
Judea:
[12:08] Here you have two different languages, language of geometry and the language of algebra. And they are the same thing. And you can get the same phenomena, same proof that you can pass a tangent to a circle from a point outside the circle and you can find the angle. But different method, two different languages dealing with the same phenomena, different perspective, and they get the same result. It blew me off.
[12:45] Blew me off completely. Maybe that was. It was preparation to computer science. Because for us, computer science is what’s the big deal? You want to see things from different perspective, invent a new language. So that was what turned me on in my high school. That was in high school. And then we saw that. I saw the same thing in physics. Different languages, capturing the same phenomena.
[13:20] They really excited me. He invented the idea of a field, two different ways of looking at the same thing. You can see that the force here depends on the charges around it, or you can say, no, there’s a field right in the location of your testing point. Yeah, terrific, terrific. So I came in with this preparation, and here we go to Brooklyn Poly. And I studied there in.
[13:55] I worked in the morning in RCA Laboratories in Princeton, New Jersey, David Sarnoff Research Laboratory. And there I got into the computer research group. Computer research at that time was research of all phenomena that you can think of to find out a mechanism for computer memories. The memories at that time were core memories, magnetic cores, these donuts that you remember perhaps from your early childhood that were too slow and too clumsy, and you had to have people stringing them X and Y and Z in Hong Kong.
[14:44] And people understood that the days of core memories are numbered. And we were looking for new phenomena. Some people looked at photochromic memories, some people looked at the semiconductor. Some people looked, like me, into superconductivity. And I was in the group. It was supposed to design superconducting memories. Okay. We did some nice plates, 16 by 16 bits, okay. And we thought that we had the future in front of us, but in the way toward developing superconducting memories.
[15:36] I investigated the physical phenomena behind the eddy currents, permanent eddy currents in thin superconducting films. Okay. And it so happened that I discovered new phenomena there, and I got a prize. And it even has a name. It’s called Pearl vortex. You can find it in Wikipedia. I discovered that physicists, years after I finished my PhD, discovered my work there. And they were interested in the idea of permanent current flowing in a circle in thin superconducting films.
[16:23] And since I analyze the magnetic and current field there, they call it pearl vortex. So here I have my footstep into immortality.
Ryan:
[16:37] Well, you said it’s a permanent vortex. Is it because there’s no electrical resistance?
Judea:
[16:43] Because it’s a superconducting, the current going forever. So you establish, you put magnetic field and you excite a vortex counterclockwise, and it will continue to turn and turn and turn forever. That’s why we call it permanent current. Yeah, forever until you flip it with another magnetic field. So we call it a vortex. But it goes on forever. And you can detect it by flipping it.
[17:19] You flip it, and if you see a big flip, it was one way. If you don’t see a big flip, it was the way you turn it. Okay, so you have a memory. You have a memory. Of course, we didn’t succeed in turning it into useful memories that would be competitive with semiconductors. The people who worked on semiconductors beat us out. We never believed that they would, but they did, both in miniaturization and in techniques.
[18:04] Unbelievable.
Ryan:
[18:05] At the time, why did you not believe in the semiconductor direction?
Judea:
[18:10] Who is going to trust memory to battery failure? What if you lose a battery? It was obvious it will never work. Right. And we looked into the result that they obtained at that time. It looked far-fetched, the idea you can have that degree of miniaturization. We saw the struggle of people who work in the laboratory on semiconductors, and we weren’t impressed. Yeah, but they beat us up. Okay, so that was my story with the superconductors.
[18:59] But I must tell you that everybody, even at that time, when computers were clumsy and took rooms and rooms and you programmed with cards, okay, even at that time everybody understood in AI as an inspiration. Everybody believed thoroughly that one day computers are going to be able to emulate all human functions. That was not the question. The question was only how and when, but not whether.
[19:43] Okay, I remember already at that time, with the clumsy computers and the punch cards, people talked about associative memories, about pattern recognition, about seeing, understanding. All these were already ideas, were exciting people to think more about it. And so we were all geared towards it.
20:15 — What people thought of AI in the 80s
Ryan:
[20:15] If I at that time asked people and your peers and you, what’s the timeline for maybe human-level intelligence and machines? What would people have said at that time? When they were excited in the 80s?
Judea:
[20:30] I think they were more optimistic than reality. They probably gave you 20 years, but that was 1965. Okay, so 20 years, 1985. No, we didn’t yet get anywhere.
Ryan:
[20:46] Yeah, and then what about today? Do you think people are more optimistic than reality? Or is this just history repeating?
Judea:
[20:55] It depends who you’re talking about. Some people are extremely optimistic today, and some people say I’m a bit skeptical, but not skeptical in our ability to eventually reach AGI, but in whether the LLM technique and thinking will lead us there. Okay, so it’s a question of who you ask? Okay. LLMs were a great surprise. But they have limitations. We’ll talk about it. After superconducting, I decided to come to California to a company named Electronic Memories, in which they did not work on superconductors, but worked on plated wires.
[21:46] Instead of having a donut in which you thread a wire, you start with a wire and you plate it with magnetic material so it acts like a donut locally. Right. And that was the promising technique. At that time, at least, I was in charge of the research and development group charged with the task of developing this kind of systems to replace core memories. Yeah. And I worked there for three years. I was frustrated because things did not go my way.
[22:26] I had both administrative and technical challenges that I couldn’t handle, both in chemistry, and I didn’t know much chemistry, so I was frustrated.
Ryan:
[22:41] What were the administrative frustrations?
Judea:
[22:43] I had a group, and I had to satisfy the administration. I dealt with the personnel issues, firing, hiring people. Yeah. And my wife saw that I was unhappy, and she told me, “You get to. You have your place in academia.” So I looked for a position in academia. Luckily, also at that time, industry was revered by academia because all the advances, all the important advances were developed in industry, not in academia.
[23:36] The transistor was developed in Bell Lab. The laser was developed in, I think, here in California by another fellow. But all this was industry development and not. So academia looked with reverence toward people who come from industry. And they hired me without me even filling an application, without even asking for accommodations. Yeah, at that time it was a good time to be hired.
Ryan:
[24:15] And what about— because you said at that time industry was revered by academia.
[24:20] Would you say that’s still true today, or has that changed?
Judea:
[24:23] No, it’s changed. It’s different now with AI. If you come from deep learning or something, DeepMind, you know, people look at you with reverence in academia. Yeah, but it’s changed, you know, only in the last few years. I say that throughout. Since 1970, I think until 2000, it was the other way around. You know, people simply dismissed industry.
[25:02] I mean, academia dismissed industry. Yeah.
Ryan:
[25:05] I wonder what happened. Was it like Bell Labs disbanded their research group or something, or that was part of it.
Judea:
[25:14] Bell Labs disbanding and what happened to IBM is still there in Watson. Watson Center, I remember. Big center, huge and important. Including Raytheon. Including. Where can I tell you? Use research here in Malibu. Did great work, but it all went down sort of the frontier of research went to academia. Sure, we had a lot of theoretical work in academia. The development of AI, AI property after 1970. Yeah, but prior to 2000.
[26:03] Yeah, at least in my corner of the field. The whole idea of inference search, inference logic, expert systems, these were all the academic development.
26:23 — Entering academia and researching AI
Ryan:
[26:23] So then you got hired at UCLA?
Judea:
[26:25] I got hired in 1970 or 1969. And yeah, I was hired. The computer science department was just formed there. And I got first hired by another department called the Engineering Systems Interdisciplinary, and then back to computer science. And I was asked to teach computer memories, hardware on computer memories. Yeah. And I gave a course in this technology. And later on, I started getting interested in pattern recognition, and I started working in this direction.
[27:12] I did work on image compression. We used fast Fourier transform and fast Hadamard transform, all kinds of transform techniques to condense images, to minimize the number of bits sent. Okay. I guess it was a part of the trend at that time. But when I got into pattern cognition, I returned to my old dream of thinking about AI and how the brain works and how computers will one day emulate ourselves.
[27:54] And I started teaching class in AI. At that time, AI was game playing, machine playing of chess and checkers and the puzzles, like the eight puzzles, ruby cubes, things like that. That was AI. Yeah. And I got excited by that game, and now I see why. Why? Now I can tell you why. Because the game was a matter of capturing in mathematics what people do heuristically, like playing chess, and the interplay between mathematical analysis and the performance interests me.
[28:49] And especially in chess playing, the interplay between the explicit knowledge that you have in terms of your gut feel about the strength of a position and what you get when you do some search. Okay, so here, interplay between fast thinking and long thinking, to use Kahneman and Tversky or Kahneman’s title, Thinking, Fast and Slow.
[29:20] Thinking fast and thinking slow. Right now, here. It’s a beautiful arena to see how not only you have two modes of thinking, but how they feed each other and how you can invest more resources in one versus the other.
Ryan:
[29:42] The algorithms at that time, like chess, for instance. Can you give an example of the interplay of the two and how that might come together in a chess-playing system?
Judea:
[29:53] Yes. You can invest more time in getting your immediate perception. It’s called static evaluation function of the chess position, the strength of the chess position. Or you can let it go and think about searching for deeper horizon. Okay. It’s a tradeoff. It was like, if I remember, we.
Ryan:
[30:15] We searched the game tree and then evaluated each position at the horizon.
Judea:
[30:23] And then you back off and then you make a move toward the position that has the greatest strength after you back off. Yeah. Okay.
Ryan:
[30:35] So the intuition is encoded in the evaluation functions.
Judea:
[30:39] Yeah, and your intuition is the evaluation function. But you can improve your intuition too. How? By learning. Okay. So Samuel Checker program did learning in regression analysis to find the proper weight on the various characteristics of the position so as to make the evaluation function more accurate.
Ryan:
[31:08] When I was learning chess, I think one heuristic is you want to control the center.
Judea:
[31:13] Good. And material advantage is another one, right? Okay. And whether you have two bishops versus a bishop and a knight, this all counts. And whether you already castled or not. All this contributes as attributes to the strength of a board position. Getting the weight correct, you can do by learning. After you play so many games, you adjust the weights. Yeah. So that was Samuel’s contribution to first machine learning.
[31:55] I say that was the first machine learning. Yeah.
Ryan:
[31:57] At that time, when you were working on this, were chess systems superhuman? Yet I think there’s.
Judea:
[32:02] No, no, no. It was still a dream to beat the world champion Kasparov machine. It was a dream. No. And. But I did knight analysis. We did alpha-beta pruning. If you remember that, you probably programmed it. Well, I proved that alpha-beta is optimal.
Ryan:
[32:34] Really?
Judea:
[32:34] Yes, mathematically, you see, I like the mathematics. Prove it. You cannot do better in terms of number of position that you have to inspect at the horizon or the depth of search. And what can I say about it? I get some nice results. Even Knuth was surprised that one can prove the optimality on alpha beta. Okay. Because he questioned it in his book. Yeah, I did some work with Deep Cop on searching trees and.
[33:10] Okay, so I did mathematical work on the tradeoff between search and reasoning until I got sick and tired of search.
Ryan:
[33:24] When you pick your research area, is that 100% your own choice?
Judea:
[33:30] No, it’s always a combination of two things. Number one, do you know the answer to the question? If you don’t know the answer, it’s a puzzle. If it’s a possible next question comes. Do you think you have the techniques to make a contribution here? Do you know something that other people don’t know compared to another field, perhaps from physics, perhaps that you can bring to bear, that you can leverage here so you can get the answer or closer to the answer than other people.
[34:05] So it’s always a combination of your perception of your tools versus the puzzle that you have. I see an important problem. People are breaking their heads. So it’s a puzzle. Do you know the answer? If I know, fine. But if I don’t know, it’s my puzzle. I take it personally. Yeah, I’m aching. I don’t sleep at night. And then the question is whether I have the tools. In some areas, I give up right away.
[34:36] I don’t have the tools. In other areas, I say, wow, if I only use that kind of trick, maybe I can get some insight. So that’s always two questions I asked myself.
34:52 — The invention of Bayesian networks
[34:52] In the case of artificial intelligence, at that time, we had the expert system come into the game at Feigenbaum and his coworkers. Did my expert system on for medical analysis? Yeah. In expert systems, the hurdle was dealing with uncertainty.
[35:18] It started with logic. Okay, you ask an expert for rules of behavior. You ask a doctor, when you see a fever, what’s the first thing that comes to your mind? What’s the next question you ask? What drives your queries until you get a diagnosis and the therapy? So they thought they can capture expert behavior using logical rules, but then it turns out that most everything is corrupted by noise, by uncertainty.
[35:59] So they started doing the same thing to uncertainty. Okay, so if you go to, if you came from Asia, you have 50% of having malaria, okay? And so on. And you have 30% here, so much there. And how do you combine these uncertainties now? Okay, logic doesn’t tell you how to combine uncertainties. Probability does, but not logic. So how do you combine one uncertainty with another? Different rules, okay?
[36:33] To come out with a combined conclusion. That was a hurdle at that time. And I remember they didn’t do it well. Actually later on we proved that they could not do it well, because rules do not combine the way these logical assertions combine. So then I went to and I asked myself, you know probability, right? So why don’t you apply probability to it and do things the right way? But probability was in ill repute at that time because everybody understood that probability is passé because it takes exponential time, exponential memories to do even the most rudimentary tasks.
[37:24] Okay, if you look at how probability is defined by the textbook, you have a big table, and for every combination of event, you have a number. The numbers sum to one. Okay, that’s beautiful. But then you can talk about conditional probability. But all these require exponentially large tables and exponentially long time to compute even the smallest kind of inference. Take, for instance, what’s the probability of having malaria?
[38:05] Given that, you see, given that you see, two things like you came from Asia and you have a fever of 30 degrees Celsius. Okay? Even small tasks like that, probability of X given that you have Y and Z takes exponential time. If you go by textbook, okay. But I asked myself, you and I are doing it fairly well. We compute probability as we cross the street, as we choose a doctor. Yeah. And we do it fair.
[38:43] A fairly good job, at least. We go through life without much regret. And how do we do it then? If we are required to do it by exponentially large tables of probabilities, evidently we are using some other kind of judgment. And I hooked onto the idea that everything depends on conditional independence, which means not every fact in life is relevant to any query. Okay? The color of the eye of my uncle is irrelevant when I’m trying to find a diagnosis of a disease.
[39:35] So evidently we have a notion or assumptions about what is relevant and what is not relevant. How do we capture it? Conditional probability, conditional independence. But conditional independence, if you go by textbooks, they are defined by the probability table. So again, exponential time. No, he came—the breakthrough that we have conditionally independent, independently coded by our assumptions.
[40:13] How in a graph, if you and the graph can convey sets of independencies, that if you have the graph, then you compute all the independencies, find out what is relevant to what and deal with the relevant only. Great. And then came the work on Bayesian network. You define a network error or no error. The combination of errors gives you information about what is independent on what given what. So for every triplet, X is independent on Y given Z, where Z can be a set and so forth.
[40:55] And X and Y can be computed from the graph, not from the probability, but from the graphs, which actually, if you look at it from a philosophical viewpoint, it’s a revolution. What does probabilities have to do with graphs? When you took Probability Theory 101, is anybody talking to graph about you? No. Right. So both the probabilists and the philosophers get irritated, or should be irritated. What is the connection between probabilities and graphs?
[41:36] It turns out there is a very strong logical connection between the two, called the axioms of conditional probability, or conditional independence in probability theory, are the same axiom that you have in graph separation. In graph you have idea of separation. There’s no connection between node X and node Y unless you go through a set of nodes Z. So Z separates X from Y. Okay. It’s the same logic that you have when X is independent of Y given Z in probability theory. Independent separation is a connection between them.
[42:25] They share axioms. Aren’t you happy? I’m happy because I really know the excitement we had in the 1970s when we discovered all these connections between two seemingly unrelated perspectives on science, probability theory, and graph theory. That, by the way, I did in joint work with Azaria Paz, who came to visit me from the Technion in Israel. Yeah, and that is called, by the way, I should mention, the theory of graphoid.
Ryan:
[43:54] This all makes sense, but my immediate thought is, where do you get the graph?
Judea:
[43:59] Everything depends on where do you get on the input? Sometimes the input is in the data, sometimes the input is in a judgment. But suppose you need a judgment for that. Okay, are you giving up if the judgment required are intuitive, meaningful, something that you are willing to defend? Right. Why not use judgment? If I know that the sun doesn’t listen to the rooster crowing. Right.
[44:34] Doesn’t care. Okay. I strongly believe in it. Do I need the data to support it? Or can I insert it, assert it, and defend it when needed? So this is a trick here which people don’t realize, don’t appreciate, okay? Judgment is not a... no, no. If it is meaningful and if you can—if it is condensed, it’s very few, judgment can buy you lots of computation. And if you are willing to defend it because it’s so intuitive, where do you get the idea that...
[45:14] Where do you get the idea that the sun doesn’t care about the rooster? Well, have you done an experiment? No, but it’s so obvious, right?
Ryan:
[45:22] Okay, but what if your intuition’s wrong?
Judea:
[45:25] Indeed, that’s our problem. It’s part of our problem, even with the LLM. Because what is an LLM? It’s a summary. It’s an average of all possible judgments that people put on the Internet. It’s a summary of a huge trillion-number of judgments over which you have no control, over which the LLM does have no control. Okay, we live with it. Hopefully you put more weight on people whose judgment you trust. And let’s wait on just the quirks of people who are purposely trying to get the system to fail.
[46:09] So no, no, there is wisdom in looking at the crowd judgment. There is wisdom in that, but there’s also danger in that. Then after expert systems and uncertainty and Bayesian networks came causality.
46:28 — Pioneering work in causality
[46:28] I mentioned that in development of Bayesian networks I was extremely sure that probability captures our intuition, our reasoning mode. And it’s the best protection against paradoxes. Essentially that.
[46:49] It’s sufficient for capturing human reasoning. I was wrong. And I realized that already when the Bayesian network became famous and popular. And I realized it by, in the introduction to my book Causality, I confessed being wrong. And, and I understand why I got into that. Why. So it was misleading. And the transition came when we looked into the simple phenomena. And we never ask an expert to encode probabilistic judgment in a form of Bayesian network, namely with arrows and dots.
[47:42] Always the arrows went from what we believe to be cause into the effect. It never went the other way around. Psychological phenomena. Okay. Okay. Why is that? So people try to reverse errors. What about if you ask specifically, give me error between the symptom and the disease? Bad judgment if it couldn’t put the right judgment. Evidently we have something in causality which is basic to our reasoning that is not captured by probability.
[48:20] And that was the idea of invariance. Yeah. The relationship between disease and fever is a stable one, as opposed to the opposite relationship. Also invariant. Yeah. When you talk about car diagnosis, for instance. And so you have an expert system for diagnosing, troubleshooting cars. Okay. And then you have a new model. So, let’s see, the charger is on a different corner of the motor.
[49:01] You don’t need to reformulate your entire database from scratch. You only change one component, the location of the charger, all the rest remains intact. So the whole system, you can amortize the investment in eliciting knowledge that you got in one system after a local modification of the system, if you do it in a causal way, in a causal direction. It doesn’t work if you don’t do it in a causal direction.
[49:39] And that jolted me to think, maybe you were wrong all along and probability is not sufficient. If not, what is sufficient? But let’s capture the puzzle here. I have a puzzle. You and I operate very nicely with causation. Can we program causation on a computer then? This is a question because we are so much immersed in our language, in our assumptions, that we cannot even distinguish what is an assumption and what is a conclusion.
[50:15] We just talked cause and effect, and we are. Your assumptions are the same as mine. So there’s no way to convince you that we made an assumption, right? We take everything for granted. But when you have to teach it to a brainless robot, you have to distinguish assumptions and conclusions and logic. That was a task we had to invent a new science, a new mathematics, to capture a new phenomena, the phenomena of cause and effect.
[50:48] It hasn’t been done for us. Why? Because science was in bed with algebra from the time of Galileo, from 1632. He invented what, he got the idea. And he was very happy that science speaks algebra, which is great because you can ask questions and solve and get answers to questions that people could not do without algebra. Like how the load on a beam, when would the beam break if you put a certain load on it?
[51:33] And you figure out that you can ask questions both ways because the equality sign is symmetric. So from answering the question, when would the beam break if you put a certain load on it, you can ask the question, how should you shape the beam so that it will hold a load of that magnitude? You can invert it. That was a real revolution in science. I’m telling you my perception of science. Not many philosophers will say that was a revolution.
[52:13] I say so, okay, but maybe they agree with me on the time, at least. I trace the evolution of ideas carefully. And so that was a revolution. But it carries some limitation, because the equality sign is indeed symmetric. And science has not developed algebra for the directionality that we see in code and a vector relationship. If I tell you that the atmospheric pressure affects the deviation of the barometer and not the other way around, you agree with me.
[52:55] Yeah, but if you write the equation, the robot might think that maybe fiddling around with the barometer will change the weather tomorrow. I’m talking about a stupid robot, right? Yeah, but if you give him the equation, it can work both ways. If F is equal to ma, then M is equal to F over A, which means that if you want to change the mass, you increase acceleration or whatever, right? The symmetry might produce paradoxes, might use wrong action.
[53:31] So the symmetry is the limitation of algebra in terms of capturing science. And we have to build a new algebra to take care of the directionality that we have in cause-and-effect relationship. That takes computer science, because we in computer science have the operation called assignment, right? When you assign the content of register A into register B, it doesn’t mean that is not reversible. Okay, so if you take the logic of assignment, you put it on top of the algebra, on top of physics, you get causal science.
[54:18] And that’s what I try to do. And I think that, so far, I’m very happy with what came up. We do have a new algebra to capture cause-and-effect relationships. And we can answer causal queries on three levels, the ladder of causation from association to intervention to explanation. Okay, so, and we found out that we have a ladder here in hierarchy that you cannot solve. You cannot answer questions in a level I unless you have assumptions of level I or higher.
[55:04] So it’s a hierarchy in the formal sense, and we know how to handle it, which is very useful because you give me a query, I can tell you what level it is. I can tell you what assumption you might, what sort of assumption you need to have before you can answer it. And I can tell you if you can get it on the data, or you can get it from experiments, or you can get it by somebody’s earth explanation or whatever, but I can tell you the source of knowledge that you need in order to answer it.
55:38 — The causal hierarchy
Ryan:
[55:38] Can you explain that causal hierarchy?
Judea:
[55:41] Yes, yes, it’s very easy. It’s a three-level ladder. It goes from the bottom, which is association. That’s straight statistics. If you see X, what can you tell me about Y? If you see passively, hands off, okay? No intervention. You’re watching patients. Some of them have cancer, some of them don’t, some of them smoke, some of them don’t smoke. And you’re trying to figure out how many years a guy will live given that he is a heavy smoker of that magnitude.
[56:25] Okay, that’s association, correlation, that entire field of probability and statistics. This is what they teach you in Statistics 101, even to 808. Okay? It’s all they do. And now comes the question, what? I intervene. I mean, what if I force you to smoke five packs a day? Don’t laugh. I mean, it’s illegal, I know, but if you want to talk about the probability of living 20 years if I start smoking tomorrow, I have to think in terms of experiment.
[57:16] I start, which means I’m going to choose to smoke five packs a day. So it’s a matter of intervention. What is intervention? The intervention is forcing you to do something that you’re not inclined to do naturally. That’s the second-level intervention, or doing. If you have experiments, you can answer queries on level two. But that’s not the end, because we also need to ask to answer question of explanation.
[57:53] Given that I observe that I am 80 years old and I am still alive, alert, and I smoke five packs a day, what if I didn’t smoke? Okay? Would I be as alert? Why is it so different? Because you have already information about the outcome, okay? You know how I’m doing today. It gives you an idea about my metabolism and about my anatomy that you didn’t know before. Okay. And using that, you can find, you can try to figure out what the outcome would have been had the input been different.
[58:42] Okay? That’s a different level, requires different kind of assumptions, different techniques, different algebra. We have it. So that’s. I call it explanation. It’s more the creative retrospection, and it’s not an easy problem. Even the first level, especially when you have finite sample and you have to figure out these probabilities. Probabilities mean properties of population, right? From finite sample.
[59:15] So I have all these P levels and struggles among statisticians of what would be a proper way of quantifying the uncertainty that you have given you have finite sample.
59:34 — LLMs and predictions
Ryan:
[59:34] Where would you place LLMs in this causal hierarchy?
Judea:
[59:36] Beautiful. Here comes LLM. I made a statement, right? That you cannot go from level I to level I plus 1 unless you have assumptions here. LLMs just looking at data, right? And giving you beautiful explanations for things that happen. Beautiful prediction of what will happen if you do. Okay, how can? The trick is they are not looking at data. They are looking into assumptions-laden world models authored by you and me and by other authors on the Internet.
[01:00:24] So they’re looking at the opinions of doctors already, who already read, who wrote papers, okay? So it’s not looking at the samples of disease of patients and samples of patients smoking and nonsmoking. They’re not looking directly at the data in the environment. They’re looking at interpreted data, data interpreted already by physicians and interpreters and reviewers that went into the articles, which are summarized on the Internet.
[01:01:04] So they have all this human knowledge on which they operate, and that is what they take as input, and that’s what they summarize. So they do not violate the restriction of the ladder because they do have information from a higher level, but it’s biased by the opinion of those authors. Fine. Those authors were smart. As long as they are smart and you believed good. So that was what LLMs is doing and what is.
[01:01:44] I explained why there is compatibility between the ladder of causation and LLM performance and what the limitations are. Now, if you want to change the environment, if you would provide explanation for raw data, LLMs will be in the same difficulty as you are. And what the physician says, I have raw data. What can I say about the probability of cancer? But it’s not really doing the introspection.
[01:02:22] It’s taking the introspection that already was done and summarizing it. How it summarizes is a mystery that no one has yet been able to decode. It’s a mystery how human knowledge encoded in the form of articles on the Internet is being summarized by the LLMs.
Ryan:
[01:02:47] So then, do you think this approach could lead to superhuman intelligence? Or maybe some people say AGI?
Judea:
[01:02:54] I don’t think so, but not with the element. They need to have some understanding of causality, okay? So that they wouldn’t need access to the Internet. Look, a baby gets born playing around with toys in the crib, okay. And gets quite intelligent, right? Without having access to the Internet, okay. Simply by curiosity. Babies are born with built-in curiosity to have control over the environment.
[01:03:36] Until you have control, or the illusion that you have control, you’re restless maybe. And you play around with toys, bing, bing, bing, until you understand one toy makes noise and one toy doesn’t make noise, okay? But you are born with this restlessness. And when are you pacified? When you understand that green toys make noise and yellow toys don’t. Now you’re in control of the environment, you can suck your pacifier.
Ryan:
[01:04:07] What if I created a baby robot that randomly plays with toys and gathers data about them, and then you feed that into LLMs, then it does have some sort of discovery.
Judea:
[01:04:21] Yeah. That is indeed the danger. When you have a robot like that born with this restlessness and craving for control over the environment, then you and I become part of the environment. And there’s nothing to stop that baby put in, okay, from trying to turn us into his or her pets to utilize us to satisfy his control because we are part of the environment, in which case he can use us. And we could be very useful to serve his or her need.
[01:05:14] What is his name? His need simply needs to feel in control, to have this illusion of empowerment. I don’t rest until I have the illusion that I control my environment. And here are some organisms, you and I, who are part of the environment, and they seem to work outside my control. I cannot afford it, okay? It makes me feel like I’m useless. So this robot baby comes and says, let me control them. And I know how I understand their fears.
[01:05:59] You don’t want me to tell about your thoughts to your wife, right? So I’m going to blackmail you and all kinds of things. I have a lot of data about you, and some people I know you wouldn’t like me to tell what I know about you, so I’m going to blackmail you. But you see, if you want that robot to have the curiosity of a child, and we want it, we want the guy to desire to have control over its environment because the environment may change, and he needs to have this urge to be in control.
[01:06:40] So if you program that, then you lose control because you become part of his or her environment. But you can say, okay, let’s forget about a curious robot. We don’t want a curious robot, so you lost. We are not emulating ourselves because we are curious robots. We are curious organism, as opposed to monkeys. Okay, monkeys. An example of an organism which is motivated by reward. But if you don’t give the monkey a banana, he’s not curious how banana grows.
[01:07:27] He’s motivated by bananas. You remove the immediate reward, and the monkey is not interested in learning more about the world. Understanding environment can be totally wrong. Look, religious people believe that if they sacrifice their children, right, they control drought. It can get to this stupid extent, but it is common to many primitive societies. If you bring a sacrifice to the God, you know, next year you’re going to have crops and harvest.
[01:08:09] It goes to extremes. Battered wives believe that if she prepares a great, better dinner, the husband is going to be, next time, less abusive. It goes to all kinds of extreme and wrong conclusions. But the need to feel in control is so immense that it overcomes all these paradoxes. It’s innate in us. I cannot control my husband, but I control myself, right? So let me be a better wife. This is something I can control.
Ryan:
[01:08:49] Do you think that we need to put those human elements in an AI for it to become AGI?
Judea:
[01:08:57] I think so. Otherwise they wouldn’t see autonomy. We wouldn’t see autonomy in the sense that we are seeing it in a human being. And this is the definition of AGI, a general intelligence. It acts like you and me, so we can converse with that creature in our language and motivate.
Ryan:
[01:09:24] If not today’s LLMs, future LLMs, they might get to a point where if you were just texting it, maybe, like the Turing test, where you don’t worry about the physical embodiment, you just see the text that comes from it. You could mistake it for a human, maybe, or it could appear intelligent.
Judea:
[01:09:45] Well, the test comes from exposing the system to raw data, not data that was chewed by our Internet articles. Raw data. Look at patients, look at cancer, look at smoking. Tell us what you know. I’ll give you some experiments to run. Be automated scientists. Can an LLM today be automated scientists? And I think they cannot without access to the Internet articles.
Ryan:
[01:10:24] So then if that wouldn’t lead to AGI, what thoughts might you have on something that could lead to AGI?
Judea:
[01:10:35] A computer system that has both the ability of LLMs to go from finite samples to property of distribution, i.e., level one of the ladder, plus ability to reason in higher levels of the ladder that seem to combine it with the calculus of intervention and with the calculus of explanations, with a counterfactual calculus that I call causal AI. Yes, I don’t see any impediment to this combination to bring us to AGI level, with the danger that it presents to us.
[01:11:21] My motivation is to understand how we do it. And I still have a few puzzles. But as I told you, puzzles are the driving force for science. Yeah, so you said.
Ryan:
[01:11:33] Do is missing. What if you had a fleet of robots that are just doing experiments? They don’t know exactly the direction, some random discovery process. They collect that data, feed it back into their hive mind, LLM, and they repeat, they repeat until they discover things. Could that solve some of the missing piece you’re saying?
Judea:
[01:11:59] Sure, but I need to know how we do. You have organisms that have done it before. Monkeys haven’t done it because monkeys remain monkeys. They didn’t invent Maxwell equations. So what do we have that monkeys do not have? One hypothesis I supported is that monkeys, that we have this innate curiosity to have control over our environment.
Ryan:
[01:12:34] And that’s a necessary.
Judea:
[01:12:36] Necessary. I’m not sure it’s sufficient. Of course we have the computational tool to bring it to fruition. We have succeeded in some way. Perhaps the next robot will do better. So we have a benevolent God in the form of a robot. Actually, what’s wrong with it? People lived for so many thousands of years under the illusion of a nonexistent God. Could you imagine if we really have a benevolent God, both just and almighty?
[01:13:13] Wow. Wouldn’t it be nice?
Ryan:
[01:13:14] And it’s a robot.
Judea:
[01:13:15] It’s a robot. Yes. And we know exactly what sacrifice to give for the right kind of request. It’s the first time I think about it. Maybe it’s gonna be good.
Ryan:
[01:13:30] When I see all these AI companies, they seem to be thinking that LLMs will lead to AGI, or they continue to go in that same direction.
Judea:
[01:13:39] Just really, I’m not sure. I really, really believe that. Jeff Hinton just came out a few months ago. He said, no, we are on a dead end. Other people might also come up and say things in different ways. I don’t find the consensus here in terms of the capabilities of LLMs.
Ryan:
[01:14:05] Yeah, I think there’s a lot of famous people that disagree. But, for instance, the people who are running maybe Anthropic or something like that, they continue to push and believe, you know, three to five years from now, there will be.
Judea:
[01:14:20] There’s a lot of that. There’s a lot of anthropomorphic terms which people claim, we can now do that. We can now program consciousness. Okay, come on. You have to be scientists, right? Define what you mean by consciousness. What is the Turing Test for consciousness? And then show that you can do what are the principles that have limited us until now and that been overcome now with your system. That is scientific talk.
[01:14:51] I don’t buy this, and I don’t read them even.
Ryan:
[01:14:55] There was one interview you said, faking intelligence is intelligence. I could see an LLM faking intelligence based off what I’ve seen. But then wouldn’t that mean that we should believe they’re intelligent?
Judea:
[01:15:12] But if you have a correct test, yeah, you have to define what you mean by intelligence. And you have. If you define intelligence by playing good chess, right, we have already done it. Right. But if you put more demand on what intelligence is, then we haven’t succeeded yet in passing the Turing test. So, yeah, yes, faking it is having it because—why? Because it’s so hard to fake. I said it in that context because it saw.
[01:15:47] The context that I had is, for instance, coming out with correct answers to causal queries. And I showed that it grows, like, super exponential. You have so many variables on all sides that you have to deal with that. You’ll have to have. The faker will have to have super exponential memory on that basis. I made this statement. Having it is faking it, is having it because it’s so hard to fake. Currently, you can bypass faking without speaking.
[01:16:33] If you steal from other people. Right. You don’t need to spend these computational resources on faking it. So you bypassed it because you steal from the Internet.
Ryan:
[01:16:45] I see, so you’re saying in this case the intelligence came from the training set, which came from humans, which are intelligent, which is very useful, very useful.
Judea:
[01:16:53] And I’m saying the people like me who are trying to build the science of intelligence, we can use all these capabilities of LLMs, level one of the ladder in our scheme of getting general intelligence. Level one is very important. It allows you to compute functions of distributions, call it properties of distribution from finite samples. Beautiful. It’s a terrifically and very immensely useful tool among the many other tools that we need for AGI.
[01:17:41] And we know exactly where it’s going to fit in. Getting from finite sample to properties of distributions, it’s a very hard problem.
Ryan:
[01:17:56] You mentioned this conversation. I think you said it in other places too, that you’re interested in capturing the way that people think, not the way that nature is constructed.
Judea:
[01:18:05] Right, right.
Ryan:
[01:18:07] Why do you care about human cognition? When I think about machines, what makes them special is that they think in a different way than us and they’re faster, and so, yeah, why is that the goal?
Judea:
[01:18:20] Because I am an egotistic organism. I want to understand myself. I’m lazy. Okay, it’s true. We are made of organic material. So that puts certain limitations on our capabilities. Perhaps silicon is not subject to the same limitation that organic chemistry is. Okay, perhaps. So what? So which means that I will never be able to understand how I think by silicon, by exercise on silicon machine. Okay, that.
[01:19:02] What is that idea? But there are so many functions that are capturable by silicon today. I don’t see any speaking in terms of theory and emulation. I don’t see any capability which is basically not capturable by silicon emulator. So why work on this? The unique biology with which we inherited? I don’t see any reason for that. Well, anyhow, some people are. It’s maybe it’s a legitimate question to ask.
[01:19:55] Do we think the way we think because we were born with organic material as opposed to silicon? Okay. It’s a legitimate scientific question. And some people can spend their time. I’m interested in other questions.
01:20:12 — A restless mind pays
Ryan:
[01:20:12] You said rebellion pays in science and a restless mind pays. I was curious why you think being rebellious is a valuable thing in your career.
Judea:
[01:20:23] I tell you why. I tell you why. More and more I come to the realization that the scientific community and academic community is the most dogmatic, conservative, anti-progress that we have invented.
Ryan:
[01:20:46] Why do you say that?
Judea:
[01:20:47] I can see what difficulty the theory and the science of cause and effect are facing today in getting just being penetrating the thinking of disciplines like statistics, like economics, okay? These people are still thinking like 100 years ago. And when I see that, and I see the forces that preserve this inertia, and they are not decent forces, I see that I’m very disappointed. And I used to think that academia is a place where new ideas can really spread and propagate.
[01:21:31] And I feel the other way around. You have so much inertia invested in the politics of academia, in the cultish inhibitions that come with academia. So that I really am disappointed. What can I tell you? I’m not sure that we have the right kind of organizations that will be conducive to. That’s why I’m saying, let’s rebuild. Don’t take your professor’s word as authority. Rebel against your professors.
[01:22:20] I rebelled against my professors, and I want to see my students rebel against me. And believe me, if I remember, there were several students who told me, “You don’t know anything about AI.” And I said, after a while, I told them, “You’re right,” by saying that you drove me to study different aspects. And I was educated by that.
Ryan:
[01:22:44] When I looked at your past works too, I think you’d mentioned that your work was controversial or mischievous before it was accepted.
Judea:
[01:22:54] Yeah, it was. Because of dogmatism. One day I’m going to publish all my correspondence with the greatest philosophers of the time, okay? With statisticians, economists. It’s all in my correspondence files, okay? And I don’t know if I’m allowed to, because they communicated with me with the understanding that it will be kept private. But when I publish it, you see what kind of stupidity drives those great people, and very great, really.
[01:23:32] Each one of them was a giant in his or her field, really. But they couldn’t get over a few of the basic molds in which they were formed.
Ryan:
[01:23:45] How did you overcome that? If everyone thought your initial things were.
Judea:
[01:23:50] I remember my high school days, I said, no, there are a thousand dunams in a square kilometer. Sorry, I don’t know where I got this chutzpah. In Hebrew you call it chutzpah. In English it would be audacity. Perhaps in my high school, I’m not sure. But I want to understand things my way. I feel like I am still able to teach people useful things. At least I perceive them to be useful, which they do not know.
[01:24:27] So I’m happy because I feel useful. Happiness is feeling useful. It’s an illusion of thinking. Yeah, useful.
01:24:36 — Advice for his younger self
Ryan:
[01:24:36] With all the experience you have now, if you could go back to the beginning of your career and give yourself some advice, what would you say?
Judea:
[01:24:44] That’s a good example of retrospective thinking. Maybe I should have spent more time on learning chemistry. I hated chemistry because it required so much memory in chemistry and biology. But that’s not the advice that you probably... And genetics. I’m talking about areas where I feel weakness. But you have to decide where you spend your computational resources. And I spend them on different.
[01:25:15] On physics, on engineering, as opposed to mathematics, as opposed to chemistry. It’s a choice one has to make. Some people have the greatness of mind to be polyglots. I admire them.
Ryan:
[01:25:31] Did you think that physics and math were superior to chemistry? Because chemistry, you just have to memorize things.
Judea:
[01:25:38] Yes. I couldn’t stand this demand on memory. Yeah, I was weak in chemistry. That’s what I like about physics. And in math, you have a few basic axioms from which you can derive everything when you need it. You don’t have to memorize it. And that’s why today I am a great advocate of world model. World model. Okay. You don’t store the questions and the answers explicitly. You derive them when you need them from a very parsimonious code.
[01:26:21] Yeah, that is a great thing about world model. Okay.
Ryan:
[01:26:26] Awesome. Well, yeah. Thank you so much for your time. I really appreciate it.
Judea:
[01:26:30] Oh, you didn’t ask me to sing.










