<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Peterman Post]]></title><description><![CDATA[Sharing learnings from the transparent career stories of technical people, written by an ex-software engineer @ Instagram

]]></description><link>https://www.developing.dev</link><image><url>https://substackcdn.com/image/fetch/$s_!645l!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5381e1ae-726f-45da-ae2c-21f0a14a3336_1280x1280.png</url><title>The Peterman Post</title><link>https://www.developing.dev</link></image><generator>Substack</generator><lastBuildDate>Tue, 15 Sep 2026 16:59:33 GMT</lastBuildDate><atom:link href="https://www.developing.dev/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Ryan Peterman]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[ryanlpeterman@gmail.com]]></webMaster><itunes:owner><itunes:email><![CDATA[ryanlpeterman@gmail.com]]></itunes:email><itunes:name><![CDATA[Ryan Peterman]]></itunes:name></itunes:owner><itunes:author><![CDATA[Ryan Peterman]]></itunes:author><googleplay:owner><![CDATA[ryanlpeterman@gmail.com]]></googleplay:owner><googleplay:email><![CDATA[ryanlpeterman@gmail.com]]></googleplay:email><googleplay:author><![CDATA[Ryan Peterman]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Casey Muratori: Surprises In Computer History And Where Bad Code Comes From]]></title><description><![CDATA[In this episode, my goal was to record a conversation that was completely free of any &#8220;AI doom&#8221; content.]]></description><link>https://www.developing.dev/p/casey-muratori-surprises-in-computer</link><guid isPermaLink="false">https://www.developing.dev/p/casey-muratori-surprises-in-computer</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 14 Sep 2026 13:03:29 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/215590461/52520a275d4e9410e72908df9a174eb5.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>In this episode, my goal was to record a conversation that was completely free of any &#8220;AI doom&#8221; content. There&#8217;s so much of it out there, I hope that this can be a nice timeline cleanser.</p><p>I had <a href="https://x.com/cmuratori?lang=en">Casey Muratori</a> on who is a video game developer and programming creator also known for several talks (<a href="https://www.youtube.com/watch?v=wo84LFzx5nI">ex1</a>, <a href="https://www.youtube.com/watch?v=hpj6r6CjJf8">ex2</a>) on the history of computing. In putting together those talks he dug deep into computing history. Enough where he had interesting stories and tidbits to share. </p><p>For instance, Dijkstra wrote about his depression and there were also several famous flame wars that happened between legends in the industry. They didn&#8217;t have social media back then so they wrote letters to each other calling each other out which we discussed.</p><p>Lastly, we talked about his career and the video game industry. I was curious to compare it with big tech and hear more about his career path. He&#8217;s a great story teller so I hope you enjoy the episode!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/jHLbL1Eg4gM">YouTube</a>, <a href="https://open.spotify.com/episode/4ywz6rOWmFardgZ04tuBTs?si=w-toSznlQFCn6qd6S_r1_g">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-jHLbL1Eg4gM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;jHLbL1Eg4gM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/jHLbL1Eg4gM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/215590461/0108-digging-into-computer-science-history">01:08 - Digging into computer science history</a></p><p><a href="https://www.developing.dev/i/215590461/0408-what-shocked-him">04:08 - What shocked him</a></p><p><a href="https://www.developing.dev/i/215590461/1623-dijkstra-was-depressed">16:23 - Dijkstra was depressed</a></p><p><a href="https://www.developing.dev/i/215590461/2617-the-personal-side-of-goto-considered-harmful">26:17 - The personal side of goto considered harmful</a></p><p><a href="https://www.developing.dev/i/215590461/3509-the-anatomy-of-a-35-year-mistake">35:09 - The anatomy of a 35 year mistake</a></p><p><a href="https://www.developing.dev/i/215590461/5012-clean-code-horrible-performance">50:12 - Clean code horrible performance</a></p><p><a href="https://www.developing.dev/i/215590461/5833-how-to-write-high-performance-code">58:33 - How to write high performance code</a></p><p><a href="https://www.developing.dev/i/215590461/010118-where-bad-code-comes-from">01:01:18 - Where bad code comes from</a></p><p><a href="https://www.developing.dev/i/215590461/010637-why-design-docs-before-code-is-a-bad-idea">01:06:37 - Why design docs before code is a bad idea</a></p><p><a href="https://www.developing.dev/i/215590461/010920-the-only-unbreakable-law-in-software-engineering">01:09:20 - The only unbreakable law in software engineering</a></p><p><a href="https://www.developing.dev/i/215590461/011557-how-he-got-into-programming">01:15:57 - How he got into programming</a></p><p><a href="https://www.developing.dev/i/215590461/012116-why-he-didnt-work-in-big-tech">01:21:16 - Why he didnt work in big tech</a></p><p><a href="https://www.developing.dev/i/215590461/013044-should-you-work-at-a-startup-early-on">01:30:44 - Should you work at a startup early on</a></p><p><a href="https://www.developing.dev/i/215590461/013452-what-video-game-engineering-is-like">01:34:52 - What video game engineering is like</a></p><p><a href="https://www.developing.dev/i/215590461/013957-why-preventing-recursion-is-reasonable">01:39:57 - Why preventing recursion is reasonable</a></p><p><a href="https://www.developing.dev/i/215590461/014324-is-vibe-coding-bad-for-the-industry">01:43:24 - Is vibe coding bad for the industry</a></p><p><a href="https://www.developing.dev/i/215590461/015017-technical-reading-recommendation">01:50:17 - Technical reading recommendation</a></p><p><a href="https://www.developing.dev/i/215590461/015427-advice-for-his-younger-self">01:54:27 - Advice for his younger self</a></p><h1>Transcript</h1><h3>01:08 &#8212; Digging into computer science history</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=68">01:08</a>] There&#8217;s this famous quote by Donald Knuth. The big part of it is, &#8220;Premature optimization is the root of all evil.&#8221; And I remember, I&#8217;ve heard that before, but I don&#8217;t think I know the full history behind it. And why is that significant?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=86">01:26</a>] So I guess we can split that into two parts. What&#8217;s the history behind it and why is this significant? The significant part is probably the easiest to answer because, for whatever accident of history, it&#8217;s something that people really remembered. It stuck. And so the reason that it ended up being significant is people keep repeating it, and they keep interpreting it or reinterpreting it. So the effects of that phrase are far beyond what anyone may have done with it in context at the time or what they meant or any of those things, or the history. The effects of it are still felt to this day.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=122">02:02</a>] Some people, I think probably correctly, you might say, interpret the term with a fair bit of nuance and say, oh, well, it can mean a number of different things or it&#8217;s trying to capture something subtle or whatever. Other people are very blunt about it and just think, oh, it just means I don&#8217;t have to think about the performance of my code to the end of the project. &#8220;Premature&#8221; just means any consideration of performance is not worth it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=148">02:28</a>] Eventually, we&#8217;ll see where the code is slow and we&#8217;ll fix it or something. It&#8217;s a significant phrase for that reason, at least to me. You still hear it to this day, and it&#8217;s at least, depending on how you want to account for it, it&#8217;s at least something that&#8217;s 50 years old now. And so that&#8217;s an incredibly lasting impact for a rule of thumb, or however you want to look at it. So on the history side, this was really.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=177">02:57</a>] When I decided to do a talk on this, I knew I wanted to do another computer history talk because I had given one before. This is at something called the Better Software Conference. There&#8217;s been two of them so far. At the first one, I did a history talk, and I wanted to do another history talk. And this was going to be like, I&#8217;m probably only going to do two of these, and this is the second one. I wanted to do it on the history of this phrase because when I started looking into the history of this phrase a while back, I found that there was just a lot of stuff there that I didn&#8217;t know.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=207">03:27</a>] And this is true. I mean, I guess I would say every time I&#8217;ve ever looked at computer history, every time I go, I refer to it as dumpster diving. But it&#8217;s not like that makes it sound derogatory. But I think of that as kind of what I&#8217;m doing. But it&#8217;s not really like, it&#8217;s actually a very enjoyable experience to go back through history and try to pick through what happened. So to do this talk, I basically went and tried to do a pretty detailed analysis of how did this phrase end up being put in print?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=238">03:58</a>] What were the worked examples at the time? What were they thinking about? Why might it have been said?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=242">04:02</a>] What did you learn in doing all that research about these famous computer science phrases?</p><h3>04:08 &#8212; What shocked him</h3><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=248">04:08</a>] The thing that shocks me both times, even though I&#8217;ve now done it twice, is the amount of stuff that you learn about what was going on at that time that you just had no idea about. It&#8217;s just mind boggling. And I would point out that this is for a talk. I&#8217;m not a historian, I don&#8217;t do this for a living. I&#8217;m not at a university somewhere doing computer history research. So what I think of as a pretty thorough read for the history that I did for the talk is actually somewhat shallow.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=279">04:39</a>] If you had time and the resources, you could go access some of these people&#8217;s personal papers, make appointments to see things that are not really in the public record so much, and so on. So even just with publicly available things that you can get, the amount of stuff that I found was just really kind of intriguing to me. And so I&#8217;ll give kind of a brief overview of sort of what I thought was the story of this kind of phrase. In a way, I think there are kind of two.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=310">05:10</a>] I guess I would say there are sort of two threads that are good to understand that come together. And this is what I tried to portray in the talk as well. The first thread is this idea of a thing, and we didn&#8217;t really do an intro to this podcast, but I&#8217;ll say it now. I&#8217;m a big fan of your show. When you said, do you want to be on the podcast? You didn&#8217;t have to tell me what the show was. I&#8217;ve already watched it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=334">05:34</a>] I love it. And one of the people you had on, for example, was Barbara Liskov. And in that interview, she actually refers to something called the software crisis, right? And I think you prompted her a question about this as well, right? So a lot of people don&#8217;t really know about the software crisis. They don&#8217;t know that that was a term. They don&#8217;t know. This is something that happened because it&#8217;s kind of ancient history now.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=358">05:58</a>] It was in the sort of late 60s, early 70s. This was the thing. What this was is originally computer hardware, in a sense that you or I can&#8217;t really conceive of it in practice because it was so far before our time that basically computer hardware originally was very simplistic. And so the things that you were going to do with it could be described very easily. And really all you were trying to do when you programmed a computer was just translate a fairly straightforward definition of some computation into machine instructions.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=396">06:36</a>] The kind of thing that we would think about today as hand optimizing assembly code was really just all computer programming was. It was just like, okay, we&#8217;re going to take this problem description and translate it in. But during the 60s, computing power was advancing to the point where people could consider doing much more significant things with it. They could run payroll systems, they could do analysis of data sets, and they would want to be able to be flexible and change the sorts of computations they were doing and all these other sorts of things.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=426">07:06</a>] They were starting to have graphical displays. Ivan Sutherland&#8217;s Sketchpad is in the early 1960s. What they found was that at that time, programmers were not really ready to make that transition. They didn&#8217;t think about things the way we all now take for granted, where it&#8217;s like, oh, yeah, I built this library for loading JPEG images, and now when I want to load a JPEG, I just call load JPEG and it&#8217;s done.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=452">07:32</a>] All of these things that we take for granted all had to be a cultural shift. The idea that there were going to be these abstractions and these ways of sort of breaking down a much more complex system into parts that we could manage, that actually was a transition that had to happen. And so the software crisis was kind of this period where people were starting to say, the methods that we&#8217;re using for programming will not scale to these larger problems that we want to do.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=479">07:59</a>] We&#8217;ve got to figure out something to do about it. So that&#8217;s one thread of history that&#8217;s coming. Another thread of history that&#8217;s coming. It&#8217;s a little bit more specific to Donald Knuth, who was the first person to really put this into a worked example. He&#8217;s kind of who I would credit with the origin of this phrase, even if there are people who might debate that fact. And to me, it doesn&#8217;t really matter who said the phrase.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=503">08:23</a>] The phrase lives on its own. So I wouldn&#8217;t even argue with someone if they wanted to sort of make other claims about it. But just in general, he was going, like, at Stanford at this time, he was sort of learning about the fact that you could do these profiles, like taking actual execution profiles, figuring out where time was spent in the execution of a program and using that to guide your decision-making about where you&#8217;re going to spend time developing the software.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=535">08:55</a>] Again, something that even today, performance is maybe not as much of an emphasis as it maybe should be in software. But even today, with that in mind, most people understand the concept of an execution profile. They understand that you could go into Chrome and see a network profile of how that&#8217;s working. You could see a flame graph of where time is being spent, or all these sorts of things. Again, you have to remember, if you rewind the clock back far enough, these were all brand new concepts, right?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=564">09:24</a>] A sampling profiler, a block-based profiler, a binary instrumentation profiler. All these sorts of things were not just tools that everyone assumed that they could just go have and had seen at some point. Some people may never have even realized that was a thing. So that&#8217;s another thread that&#8217;s happening. And those two things converge effectively in 1974, right? Roughly. To get us this statement, Knuth is very much into this idea of structured programming, which is sort of a proposed solution to the software crisis.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=602">10:02</a>] Maybe a proposed solution is the wrong way to say it. It&#8217;s a certain set of ideas that are trying to be employed to improve the situation. Let&#8217;s say he&#8217;s very much into this. We can talk about what that is as sort of a separate tangent if we want to later. He&#8217;s very much into that. He&#8217;s thinking about that. He&#8217;s thinking about this idea that programmers can&#8217;t just be about writing all of this very tightly optimized assembly language code.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=631">10:31</a>] And that&#8217;s all we&#8217;re going to spend all our time doing. We have to be thinking bigger picture. We have to be picking and choosing our battles. We have to make concessions to maintainability. If this particular routine is not really accounting for hardly any of our runtime, we should not be obfuscating it with all of these hand-unrolling the loops and all these things. Remember again, underlying all this, they don&#8217;t have massive optimizing compiler passes like we have with LLVM now.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=657">10:57</a>] They don&#8217;t have all of this extra stuff. So if you wanted something unrolled, you were going to unroll it yourself. If you wanted some kind of loop, hoist a loop invariant out, you probably had to do that yourself. A lot of that stuff that we take for granted today, again, not in there. So that sort of. He&#8217;s very unstructured programming. He&#8217;s starting to see these execution time profiles. They do a thing at Stanford in 1970, actually he does it with some students where they take a look at FORTRAN programs and they instrument them to see where the time is being spent.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=686">11:26</a>] Not in exactly the most rigorous way that you might imagine. That&#8217;s another thing from my talk that I kind of go into. But he&#8217;s sort of saying, look, we as programmers need to start taking the decision about what to optimize more seriously. And he kind of has a fairly nuanced take on it. If you actually read the section of his 1974 paper Structured Programming with go to Statements where kind of the first time this was really put to print that I know of.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=714">11:54</a>] He does say it in an ACM Turing Lecture a little earlier, but he would have submitted the manuscript for this publication prior to the lecture. So it was the first time that I know of that he actually committed it to paper. It would have been there. But he may have said this a lot in lectures. He kind of obliquely refers to that in the Turing Lecture that maybe he had said this around town, so to speak.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=740">12:20</a>] So where this was first uttered, I don&#8217;t think anyone knows. But the point being, when he says that phrase, it&#8217;s in this context of saying, look, we need to be very concerned about efficiency. We shouldn&#8217;t be leaving performance on the table when it matters, but at the same time, we need to know that it matters. If we don&#8217;t know that it matters, then we&#8217;re going to apply this optimization. That&#8217;s a premature optimization, and it has real cost.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=768">12:48</a>] It&#8217;s going to make our maintenance and debugging of this program worse for all the same reasons that it would make it worse today. If we go into something and we start applying these optimizations to it that are above and beyond the kind of structural things that we might be doing to that code otherwise, that&#8217;s creating a cognitive burden that other people are now going to have to deal with, and you&#8217;re going to have to deal with in the future.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=790">13:10</a>] So I don&#8217;t want to put too many words in Knuth&#8217;s mouth, certainly, but in general, that&#8217;s sort of the way that this comes together is those two things, this idea that we need to be doing things to solve the software crisis. How have a more structured approach to programming, care about maintainability and debugability, composability, abstraction, all that stuff. But also at the same time, we do care about optimization.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=813">13:33</a>] So how do you dovetail those things? &#8220;Premature optimization is the root of all evil&#8221; is kind of like this phrase he uses, I would say, if I had to imagine his way of sort of reminding himself, think about all this before you make a decision. Right. That would be kind of my take on it. Right. And where it comes from. So hopefully, I don&#8217;t know, that&#8217;s very long, but it&#8217;s a very expansive question. So hopefully I&#8217;ve kind of wrangled most of the stuff into there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=840">14:00</a>] It sounds like this is really an engineering tradeoff, where the software crisis is about the maintainability and the complexity for people to actually write and handle their code, and then the tradeoff between performance and, I guess, code that is maintainable. And it sounds like this Knuth quote is, be mindful of where you pay that cost.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=869">14:29</a>] Yes. And it&#8217;s always important to understand the context in which this was said historically. So I want it to be a historical talk, because who your audience is is a large determiner of what the root of all evil is. If you&#8217;re talking to a bunch of people who don&#8217;t think about optimization at all, then you don&#8217;t really need to say this phrase to them. If you fast forward to today, when a lot of people don&#8217;t really take optimization into account at all during their programming, or very, very little, this phrase doesn&#8217;t make very much sense to someone who does care about performance, because it&#8217;s like why would you ever say that?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=908">15:08</a>] It just seems like it encourages people to not be mindful of the performance of their program. But if you rewind to that time and you understand that, well, a lot of the programs, you know, they had the opposite problem. A lot of the programs at that time, what programming was, was coming up with all these crazy little assembly tricks to shave a cycle off here or there, right? And if you&#8217;re talking to an auditorium full of people who are really into that sort of thing, then it really can be the root of all evil, right?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=941">15:41</a>] It&#8217;s like you guys are spending all your time on this, but we have this really big other problem. And that is not the right tradeoff, right? That is, you&#8217;ve swung, you know, like calling it tradeoffs is a great way to say it. You&#8217;ve swung the pendulum way too far in one of those directions at some point. And so to some degree, while I would highly dispute the description &#8220;root of all evil&#8221; today as a good thing to say to somebody, or that it leaves a good impression, when we rewind to that time period, I think it makes a lot more sense.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=972">16:12</a>] Because when you look at what people were actually spending their time doing, where their priorities were, I think that maybe that wasn&#8217;t so much of an exaggeration. I mean, obviously it&#8217;s hyperbole, but it&#8217;s less of one.</p><h3>16:23 &#8212; Dijkstra was depressed</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=983">16:23</a>] In your slides, there were so many quotes and snippets of really kind of digging into the true history. Was there anything you found that surprised you when you were doing the research?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=993">16:33</a>] I mean, constantly, absolutely constantly. And first of all, I&#8217;ll just say that there&#8217;s, as a <a href="https://en.wikipedia.org/wiki/Meta_Platforms">meta</a> point, I&#8217;ll give you a specific thing that surprised me that&#8217;s kind of funny. That&#8217;s more about a technical thing. But first I want to talk about a broader picture, which is one of the things that really hits home for me whenever I do these, is the personal aspects of it. These people were friends and competitors and all these sorts of things.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1028">17:08</a>] And it&#8217;s all kind of lost to us in a way when we think about computing history in a sterile way. When I read up on what Edsger W. Dijkstra wrote, Dijkstra was the person who kind of kicked off structured programming. No, I&#8217;ll call it a revolution. That&#8217;s a bit too probably grandiose, but he kicked off that train of thought. He wrote a thing called Notes on Structured Programming. It was a manuscript or a monograph, I don&#8217;t remember what they called it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1059">17:39</a>] Got a typewritten thing. It looks exactly like a typewritten thing. Right. When he wrote that, he was very depressed, like he was having a lot of trouble in his life at that time. He had gone through this situation where they had done what people today now recognize as pioneering work in distributed computing. He and these other people at Eindhoven University of Technology.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1087">18:07</a>] I apologize, I can&#8217;t say the proper name for it; it&#8217;s well beyond my pronunciation capabilities. But point being, he had done this really foundational work in distributed computing and published it, and people today now recognize it as that. It is undisputedly a massive contribution to distributed computing. He did it with a part-time team of people at the university, who&#8212;this wasn&#8217;t their job; it was a mathematics department they were in.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1114">18:34</a>] And as a result of that, basically they sort of just got dissolved, like the Department of Mathematics. I guess he wasn&#8217;t really specific about it, but in his notes the Department of Mathematics was just like, they disbanded the team. It was like, this isn&#8217;t math or whatever, I don&#8217;t know, they just had a pessimistic view of it. And he was really depressed about that, and he wasn&#8217;t sure what he should do with his life and, what&#8217;s the deal here, right?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1137">18:57</a>] And he was having this sort of, I don&#8217;t want to call it an existential crisis because I don&#8217;t want to again put words in his mouth. These are historical figures and, you know, I can only. That&#8217;s why I try to use quotes to try to show you like what they said. But that&#8217;s, it&#8217;s just so relatable when you go through the history this way. It&#8217;s not just some random guy who wrote some math down in a paper.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1157">19:17</a>] It&#8217;s like people were really struggling with this, and some of the most important aspects of computer history come out of these amazing human stories. And I found that absolutely fascinating. So I&#8217;ll just put that out there as one of the biggest rewarding things about looking into this history, if you ever do it, is if you can go find the actual writings of the people, like things outside of just their technical papers.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1181">19:41</a>] It&#8217;s just fascinating. And it&#8217;s so much more relatable and fascinating because you feel the story. It&#8217;s not just this abstract computer science thing that happened. So that&#8217;s one thing. But the other thing I was going to say is when you read through this stuff, I just go through piles of documents and I&#8217;m just reading them and seeing, does this fit into this story? Should it be part of what I&#8217;m telling, or is it extraneous?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1207">20:07</a>] Is it something that is interesting, perhaps, but not actually part of it. And one thing that I found, when the VOD of this lecture or talk goes up, I&#8217;m going to include this in the notes because I thought it was so fun, didn&#8217;t make it into the talk. I found a thing, like a thing from Tony Hoare. So C. A. R. Hoare, who again is another massive figure in computer science. Hoare, Dijkstra, and Knuth are like the three amigos.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1238">20:38</a>] And they all write to each other. They&#8217;re very important people in computer science history. So I found a thing by him that was like a thing he did not decide to pursue. It&#8217;s like this note where he&#8217;s talking about&#8212;he&#8217;s talking about this thing that he&#8217;s thinking of doing, and there&#8217;s just a handwritten thing on him from later on in his life, when he was, I guess, categorizing these documents, and he just writes down, like, I decided not to pursue this because I talked about it at this conference that I was at.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1270">21:10</a>] And Peter Naur, who&#8217;s again another famous figure in computer science history, if you&#8217;ve ever looked at context-free grammars and parsing, you&#8217;ve probably heard Backus&#8211;Naur Form. He&#8217;s the Naur in Backus&#8211;Naur Form. If you&#8217;ve ever heard of the language ALGOL, he was a major figure in standardizing that and writing up the standard and all this stuff. Anyway, Hoare. Yeah, I proposed this thing, I went through and said it, and Peter Naur was like, making multipass compilers is easy.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1300">21:40</a>] I just wrote a nine-pass one. So I decided not to pursue this. So what&#8217;s the thing? I read through the thing and maybe I&#8217;m just overreading it with the benefit of hindsight, but he pretty much describes static single-assignment form, like SSA, which is a very standard compiler technique developed in the &#8216;80s. But this note is from the &#8216;60s. So it&#8217;s like, in my head, I&#8217;m like, did Peter Naur accidentally set back compiler science by 20 years by telling Tony Hoare?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1338">22:18</a>] Well, 15, let&#8217;s say, by telling Tony Hoare this is not important, when actually it would have been very, very important. You find stuff like that all the time where you&#8217;re like, whoa, what is this? I saw another one too. Just one more I&#8217;ll mention. I can&#8217;t remember now, I&#8217;m sorry again. Your brain kind of turns to mush when you try to dump this many documents into it for a talk. But there was a pretty interesting thing, I thought, where Margaret Hamilton, so the person who managed the Apollo Guidance Computer, she was one of the core programmers originally on the project and then was the manager of the whole OS, like the whole real-time software that ran the Apollo 11.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1383">23:03</a>] Well, all the Apollo Guidance Computer stuff, right? Which I mean, most people today herald as a very, very significant real-time systems achievement in that era, right? Everyone pretty much agrees with that. She had to write a defense. I think it was in Communications of the ACM, maybe it was in Datamation. They say&#8212;apologize. I can&#8217;t remember the venue because it had been pointed out as a failure because people didn&#8217;t understand the context that the reason that the Apollo computer had to trip those alarms was because that had been used improperly.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1420">23:40</a>] It was used in a configuration that wasn&#8217;t supposed to. And it actually, rather than crashing, went into backup modes and successfully landed. It worked. It actually was a triumph of fault-tolerant engineering. And she had to write a defense of this because people at the time were saying, &#8220;Oh, they screwed up. This is an example of why you wouldn&#8217;t want to do engineering this way, or don&#8217;t write the code this way.&#8221;</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1443">24:03</a>] And so again, stuff like that always makes me think, the more things change, the more they say the same. That&#8217;s exactly what would happen on Twitter now, right? Some very successful thing that people did, right? They would just get slandered and say that it was done wrong, and there would be a flame war, and then they&#8217;d have to come out and have a thing. So anyway, those are just some examples.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1463">24:23</a>] But there&#8217;s so much great stuff in there. I highly recommend it to anyone who likes this kind of thing. You won&#8217;t be disappointed if you go dumpster diving, as they call it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1472">24:32</a>] When you said that you saw the human stories and that Dijkstra was depressed at that time, what did you see that made you realize that?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1482">24:42</a>] One of the nice things about Dijkstra is that he&#8217;s kind of a pretty straight shooter. And so I didn&#8217;t actually have to do any interpretation. He literally talks about being depressed. There&#8217;s a paper, not a paper, like a thing that he did that&#8217;s like a retrospective that&#8217;s handwritten by him later on in life. And he literally says, in this period I was very depressed because of these reasons.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1512">25:12</a>] Now, he could be wrong about the source of his depression. Right. So I don&#8217;t want to claim that just because he said it, it&#8217;s necessarily true or something like that. But one of the amazing things, I think one of the really cool things about academics, especially that era, is they just produced a voluminous amount of material, and so we don&#8217;t have to guess that much. In the private sector, I think it&#8217;d be a lot harder, like if you wanted to know what people at maybe Lockheed or something were thinking or doing at that time.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1547">25:47</a>] I imagine it&#8217;s much more difficult just because things are classified or they&#8217;re not academic, so they&#8217;re not necessarily going to write them all up. So a lot of the internal stuff when you write them, academics don&#8217;t really have that. They don&#8217;t have an incentive or a prohibition on writing up everything they do. So they don&#8217;t just have to have a public-facing statement of what they did or a public-facing paper of what they did.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1567">26:07</a>] And then, oh, there&#8217;s this extra secret stuff. The academics don&#8217;t have that restriction. They&#8217;re sort of... They benefit from being as forthcoming as possible about what they&#8217;ve accomplished, usually.</p><h3>26:17 &#8212; The personal side of goto considered harmful</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1577">26:17</a>] I saw in the slides that you looked at Go To Statement Considered Harmful. That email, do you have any sense of the personal or people side of that?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1588">26:28</a>] So you&#8217;re talking about the original Dijkstra letter to the editor. That&#8217;s like Edsger W. Dijkstra&#8217;s Go To Statement Considered Harmful. The complete story is actually relatively simple according to him. He was at a conference in Tennessee where he was talking to Brian Randell, who&#8217;s another computer science guy. He doesn&#8217;t quite get the same level of name recognition as a Knuth or something like that, but he has a lot of papers at that time.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1624">27:04</a>] You can go find him. He&#8217;s not obscure. He&#8217;s not an obscure figure. I wouldn&#8217;t say to anyone who reads the history. You&#8217;ll see him come up. He&#8217;s talking to Brian Randell and some other people outside. This is exactly like the same thing that happens at modern conferences. There&#8217;s the talks and then there&#8217;s all the stuff that goes down at the bar. Right. It&#8217;s very much that. They&#8217;re outside, they&#8217;re talking.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1645">27:25</a>] And according to him, he&#8217;s basically giving the same sort of example of a problem with the Goto that he gives in the paper. Right. This idea of enumeration, we can talk about that later. But sort of a side note, he&#8217;s sort of saying, like, this is kind of a problem with Goto. And one of the reasons that maybe it&#8217;s not such a good idea. And the people who were listening to him were like, Brian Randell and the others were like, you should publish that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1676">27:56</a>] That would be helpful because there&#8217;s these arguments going on about whether GOTO is good or bad. That was kind of happening at the time already, since around 1959 even, I think there&#8217;d kind of been a little bit of that, had been kind of growing, this idea that maybe Go To Statement Considered Harmful wasn&#8217;t the best idea as a way to structure your programs. And so he does. He goes back to Eindhoven University of Technology and he writes up this same thing he was saying, basically, and he sends it to Communications of the ACM for publication.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1709">28:29</a>] Niklaus Wirth, who is the creator of Pascal, a very prominent figure, also was very heavily involved in ALGOL and all this sort of stuff. Language designer guy, he is the editor who is in charge of getting this thing published. I guess I&#8217;m not exactly sure how things work at Communications of the ACM. He doesn&#8217;t want to wait to have it published. He doesn&#8217;t want to take the time to go through the referee process or whatever.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1738">28:58</a>] I don&#8217;t know at that time what the requirements were for publishing an article in Communications of the ACM, but I&#8217;m sure that it involved a lot of procedure. It wasn&#8217;t just like, oh, hey, I&#8217;m the creator of Pascal at that time. He wouldn&#8217;t have been the creator of Pascal yet, because Pascal comes later, I think. But either way, he&#8217;s like, hey, I&#8217;m this important guy. I&#8217;m just going to put this whatever I want.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1759">29:19</a>] Communications. That&#8217;s not how it worked. And so he decides that in order to get it published more quickly, he&#8217;s just going to publish it as a letter to the editor. Because then there&#8217;s no anything goes there. As long as the editors are fine with the content. I assume it doesn&#8217;t have to be refereed, it doesn&#8217;t have to have any kind of review. It doesn&#8217;t have to be, et cetera, et cetera. So he turns it into a letter to the editor and he takes the name of the thing that Dijkstra sent, which was A Case against the GO TO Statement.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1787">29:47</a>] That was what Dijkstra wrote at the top of the thing as the title. He changes it for letters to the editor to Go To Statement Considered Harmful. Right. So it wasn&#8217;t even Dijkstra&#8217;s plan to have it maybe be that confrontational. But that&#8217;s what happens now. This is not well received, to say the least. According to Knuth, again, I&#8217;m just trying to say who said what here. According to Knuth, Dijkstra said he got letters, angry threatening letters, in much the same way you would today.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1834">30:34</a>] Got people who were very abusive, telling him off. I&#8217;m sure part of that is because by all accounts, Dijkstra himself was also somebody who liked to push people&#8217;s buttons. That&#8217;s well acknowledged. I think Alan Kay is probably the best source for this. He talks about this. He liked Dijkstra, and they apparently got along well. But he&#8217;s often said that Dijkstra was someone who kind of leaned into being sort of brash about stating things about programming and kind of relished that position.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1867">31:07</a>] So it probably didn&#8217;t help that he was already sort of known a little bit in that way. But in general, that&#8217;s how it went down. And, like I said, that wasn&#8217;t the initial foray against Go To Statement Considered Harmful by any stretch of the imagination. But it just kind of maybe, you call it the straw that broke the camel&#8217;s back. It was like a flashpoint, might be the way to say it. And then the title, of course, that Niklaus Wirth picked is quite the doozy and kind of clickbaited everyone, as I call it, into that soap.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1903">31:43</a>] It&#8217;s funny because, I mean, social media, it&#8217;s similar patterns. People have very explosive first lines, and then the comments are full of all this hate and stuff. In their case, I&#8217;m guessing this is all papers. So you submit a letter to that. It&#8217;s a physical thing, and people are reading maybe a newspaper or something. I don&#8217;t know.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1925">32:05</a>] Yeah, it&#8217;s like a periodical. It comes as a bound. I mean, I guess there&#8217;s. All the things around us here are hardbound, but it&#8217;s like usually with soft. It was a maybe a perfect binding kind of bound thing that has you open it up, table of contents and letter to editor. Right. And yeah, like you can go find. There are a few. If you go look at people&#8217;s personal papers, which, like I said, if I did this full time, I&#8217;m sure I could find out way more stuff than I did.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1954">32:34</a>] Right. But some people have gone and looked at the personal papers, and occasionally some of them have been able to digitize some of those and put them online that we can look at. And you can find, for example, online right now, if you search for it, there&#8217;s a back-and-forth letters, personal letters between Edsger W. Dijkstra and a guy who is kind of into functional programming and promoting that. And you can look at the correspondence, and it&#8217;s kind of.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=1982">33:02</a>] It&#8217;s a little, it&#8217;s a little flame war. Like, really, like that&#8217;s what they did. They didn&#8217;t have the ability to do, like, pithy Twitter replies, so it was just on paper. And Knuth did something similar. It wasn&#8217;t really a flame because he had, at least when reading comes through, a tremendous amount of respect for the authors of the book Structured Programming, which is Dijkstra, Hoare, and Dahl.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2012">33:32</a>] One of the things that he published was a series of open letters to them reviewing the book and talking about the things that he didn&#8217;t find compelling in it. Very respectful. So it wasn&#8217;t. That one wasn&#8217;t a flame war. But that&#8217;s what they had to do because they didn&#8217;t have social media. So Edsger W. Dijkstra wrote all these things called EWD with a number. It was his initials, basically, and then a number.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2040">34:00</a>] And that&#8217;s like this serialized list of all the things he wrote. And he labels them this way. It&#8217;s not like some historian characterizes it after the fact. It&#8217;s like, oh, I&#8217;m doing. He&#8217;s like, I&#8217;m doing EWD 937 now or whatever. Right. It&#8217;s hilarious. But anyway, several of those have been put online that are not necessarily technical. For example, his trip reports. You can go read those and you can read about, oh, I went and stayed at, like, such and such&#8217;s house and we went to dinner or whatever, right.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2071">34:31</a>] So you can find some nice personal anecdotes in even the publicly available stuff. But in terms of a real heart-to-heart conversation, I didn&#8217;t have access to anything like that that wouldn&#8217;t have just been in a paper, more or less, normally. So it&#8217;s a shame. One of the problems with this stuff is it&#8217;s so interesting, but it&#8217;s not a job to do. If somehow it was a job, I probably would almost take that job.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2098">34:58</a>] Being the person who crawls through and tries to redocument this whole thing. But, no, computer historian, no one&#8217;s hiring, no one&#8217;s interested in paying for that.</p><h3>35:09 &#8212; The anatomy of a 35 year mistake</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2109">35:09</a>] You have this other talk, and the title was just so catchy. It was &#8220;The Big Oops: Anatomy of a 35-Year Mistake.&#8221; What is that? 35-year mistake.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2121">35:21</a>] So this was a kind of funny thing that happened to me that I then did a historical lecture on. I had worked on systems in the past that were basically like editors. You have 3D graphics editors. You have to multi-select things and move them around. There&#8217;s a bunch of architecture things you have to learn to be able to write that kind of code and certain kinds of problems you have to solve. A very basic one that I would point out that hopefully most people can relate to would be if I have a bunch of things on the screen, some of which have a color.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2160">36:00</a>] So maybe this is a drawing program and I&#8217;ve got some text and I&#8217;ve got some shapes and I&#8217;ve got some strokes, some hand-drawn stuff, whatever, and I want to be able to select a bunch of them, and I want the user interface to present to me which things I could edit on these shapes. So I want to be able to edit the color, and I want it to apply to all of them. Right. This is just a basic architecture problem that you have to solve if you&#8217;re going to write one of these programs and you want it to be any good.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2186">36:26</a>] Now, oftentimes, you will see programs where the person didn&#8217;t solve this problem, and you can&#8217;t do that. It&#8217;s very frustrating, right? So a good program, someone has solved that architectural problem in some way. And I was watching, I just happened to be watching because sometimes I do like to watch old videos of computer science pioneer stuff. And I was watching a demo of Sketchpad. I&#8217;d seen it before, but I was watching a demo of Ivan Sutherland&#8217;s Sketchpad.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2213">36:53</a>] And Sketchpad, for those who don&#8217;t know, is a groundbreaking computer graphics program. It&#8217;s done with a light pen, and it was done in the early 60s. And you could draw things like you do in a modern CAD program, and you could do stuff that even a lot of modern CAD programs don&#8217;t really offer or only offered recently. I think in the talk I said in 2007 was the first time that AutoCAD got some of these features.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2239">37:19</a>] You could do stuff like say these two shapes, this line has to be perpendicular to that line, constrain it, and you could just do whatever you want with the drawing and it would resolve for making that happen. And these are very rare in programs even to this day. That&#8217;s just not a common operation you would see in a drawing package today. I was looking at this and I was like, how the heck did he solve this architectural problem?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2266">37:46</a>] How did he do this stuff in the early 1960s in assembly language, with no editing tools, no debugging tools? I mean, this would have been like either punch cards or one step away from punch cards. It would have been hard to develop this stuff. And when I went back and I went through his thesis and I looked at the code, it&#8217;s actually documented how he did it. And I realized, like, oh crap, this is actually like an entity-component system.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2298">38:18</a>] Basically what we would now consider sort of state of the art of real-time architecture for doing these kinds of operations, meaning cross-cutting, operating on this kind of property across many things at once that are all kind of different. And this was something that object-oriented programming &#8212; I&#8217;m editorializing now, a bunch of people will complain that I&#8217;m saying this, but this is just my opinion.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2327">38:47</a>] Object-oriented programming and the pedagogy around it really struggled with that problem for a long time because they were teaching, despite people who will now claim otherwise. I document it very completely in the talk, so I don&#8217;t think they have any leg to stand on. But the pedagogy at the time was all about build a domain model. The domain model looks like your hierarchy. Basically, like I&#8217;ve got employees and employees, I&#8217;ve got contractors and I&#8217;ve got derived, like is this a full-time employee or a part-time employee?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2355">39:15</a>] That was how they suggested these things should be implemented. That kind of architecture doesn&#8217;t work for these kinds of operations that I&#8217;m talking about. Entity-component systems do. And generally, that kind of rotated architecture, I might call it, does work for these. What I call the 35-year mistake is the fact that I very ironically, in my opinion, Sketchpad sort of showed in its architecture how you could solve this problem and how you could solve it in arguably an object-oriented way.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2392">39:52</a>] I think if you squint at it, a person who does object-oriented programming could easily make an argument that you can do an object-oriented version of this. You don&#8217;t have to violate any particular principles to do it. It&#8217;s just not the hierarchy domain model way that people were conceptualizing it through a lot of the 80s and 90s, let&#8217;s say, and even unfortunately to this day. But I think a lot of object-oriented programs don&#8217;t do that anymore.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2417">40:17</a>] Ironically, that was in Sketchpad. The DNA was there, and we could have had it right away because Ivan Sutherland figured it out. But weirdly enough, the whole idea of this sort of inheritance hierarchy with lots of encapsulation around it grew in part out of Sketchpad because Alan Kay looked at it and said the interesting part of this is the fact that you could take a shape and not know what it was and just have this draw circle call on it or something.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2449">40:49</a>] That was the part they took away from it, which is much less interesting in my opinion and not that architecturally useful in most cases. So in kind of this amusing way, Sketchpad contained, if I&#8217;m editorializing, it contained what I consider to be a good architectural idea, but the takeaway from it was a bad architectural idea, is how I would say it. And saying bad architectural idea is a little bit strong because, as I say in the talk, I think there are times when you do want to think very high level.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2477">41:17</a>] You might want to think about things in kind of the way that Smalltalk thinks about them. So I don&#8217;t want to say bad idea in the absolute sense, but bad in the way it ended up being applied, let&#8217;s say.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2489">41:29</a>] Why does that class hierarchy not work for this problem?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2493">41:33</a>] It&#8217;s pretty easy to understand, actually. So when you think through an inheritance hierarchy that&#8217;s based on sort of what you see in the real world or what you&#8217;re trying to model. So you say, the typical example, and it&#8217;s brought up tons of times by people who wrote the OOP literature, like Bjarne Stroustrup, like the Smalltalk people, the idea is keeping it with a Sketchpad example, I&#8217;m going to have a shape.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2518">41:58</a>] From the shape, I&#8217;m going to derive triangle or circle. And the encapsulation is around that thing. So a circle has a radius, but shapes don&#8217;t have radiuses. Shapes don&#8217;t necessarily know what that is, right? So it tends to get deferred down and the encapsulation boundary is drawn around that sort of derived class. That&#8217;s, like, the most derived version of whatever this thing is.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2545">42:25</a>] The problem is that creates a lot of headaches when now something that wants to work across a lot of shapes needs to actually do something intelligent with modifying, say, the scale of something or wanting to change its color because it doesn&#8217;t really have a way to know what&#8217;s going on in these highly encapsulated things. So typically what you end up having to do in those situations is you then have to take this inheritance hierarchy and kind of, in a sense, throw away the benefits of it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2577">42:57</a>] You end up having to make at the top all of these accessor iterator things that allow you to probe into that lower class and say, tell me all of the colors that you might be using and a name for them. In a sense, what you end up having to do when you want the architecture to work is you have to rebuild the Sketchpad version on top of the inheritance hierarchy version, which doesn&#8217;t really do anything for you.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2604">43:24</a>] And it&#8217;s not to say that there isn&#8217;t some benefit that you can derive from having derived is kind of a pun here, I guess, that you can derive from having this idea of saying this thing is like this other thing. Please give me some of the implementation of it. That&#8217;s not necessarily a bad thing in practice. The problem is when you teach people that this is how it&#8217;s going to work, like you just make this hierarchy and then your architecture is just supposed to work, it really ill-prepares them for the reality that that&#8217;s actually not going to solve most of the problems that you have are not going to be solved that way.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2636">43:56</a>] And again, it just comes from this fact that typically when we&#8217;re working with computer programs, most of the time what we need to do is create cross-cutting operations that work with a lot of the data that&#8217;s inside these derived classes. That&#8217;s actually the thing that we want to do. And it&#8217;s really cumbersome and typically a bad fit for the problem space to put that into some virtual functions that are sitting at the top of this very sort of tall hierarchy.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2663">44:23</a>] It&#8217;s usually much better to do. Again, even if you&#8217;re doing object-oriented programming, you don&#8217;t have to stop doing object-oriented programming. You just have to think about it differently. You have to think about what the objects are differently, where you&#8217;re drawing the encapsulation boundaries. It just helps to think about it differently and say, actually, the more important things to model are what are the things that stuff is built out of?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2684">44:44</a>] How do I build a circle out of objects, right? What are those things? A radius property, a center property. Those maybe should be the objects that I&#8217;m thinking of as more first-class citizens. And this inheritance hierarchy stuff about what the objects are that I see on the screen, like triangle and circle, that&#8217;s a distraction. That&#8217;s not the proper way to model the problem. And again, I don&#8217;t think I would be pissing off any object-oriented programmers by saying that today because I think a lot of them use architectures that do think about objects in that sort of rotated sense and not in the domain model hierarchy sense.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2720">45:20</a>] Not that there aren&#8217;t still people who are doing that. But it&#8217;s not the only way to go nowadays. It&#8217;s not the only way that&#8217;s promulgated.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2728">45:28</a>] because I guess my immediate thought would have been in that case, the base class has the notion of a color or some shared thing, and we just operate. If you&#8217;re a shape, then we assume you have a color. But you&#8217;re saying refactoring and putting up in there is suboptimal and you want all these subclasses to have some handle to some radius property or something like that.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2754">45:54</a>] Is that what you just described is what tends to happen in systems that are architected this way? You end up with your derived class not really being very much of anything because in order to edit stuff, you had to migrate it all up to the top. And there was actually terms for this in the old days. I don&#8217;t still use, like, fat base classes, right? Or things like that. It just ends up your base class has every type of property that any derived class might ever have.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2778">46:18</a>] You have get primary color, get secondary color, get tertiary color, right? Because something needed three different colors or whatever. And most things don&#8217;t use those colors and all these other sorts of things, right? There&#8217;s certainly nothing wrong with running a program that way. It&#8217;s just, again, why did you bother with the hierarchy? You&#8217;re not getting any actual encapsulation benefits from it because you just forced everything in the base class anyway.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2801">46:41</a>] So why didn&#8217;t you just make that be your thing? Just make one class called Shape and be done in terms of the ECS style of implementation. I don&#8217;t necessarily want to advocate for it or not because I don&#8217;t actually personally use ECS or anything like that, but they are very common. Now the idea there is to just turn it on its side and say, look, typically what we want to do is do things like, okay, there&#8217;s stuff that has physics on it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2829">47:09</a>] I&#8217;m doing this thing, and I want to have physical simulation of entities. Well, the physics system is the thing that probably knows how that should be stored. So rather than having that be data in my derived class that&#8217;s encapsulated around the derived class, instead, why don&#8217;t I just have the physics system know what physics is, and then if you want to make something that participates in physics, you just have a handle for whatever this thing is, like a handle for my shape.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2856">47:36</a>] That handle is valid in all the systems. So I can look up the physics properties of the shape and gather it when I want to actually do something with it. But generally speaking, the encapsulation stays around the system. So physics properties are defined by the physics system because, hey, that&#8217;s who knows how physics should work. And I just ask about my physics when I need to do something with it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2879">47:59</a>] And this kind of gets you out of that problem. And then now you can just compose things. Oh, I want to make something that has several physics elements in it. Fine. I can just have multiple physics handles if that&#8217;s what I need, or something like this. I can very flexibly merge these things together. I want something to not have physics. I just don&#8217;t ever insert the. I don&#8217;t ever ask the physics system to have a valid piece of data for this handle in it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2901">48:21</a>] And it won&#8217;t simulate any physics for it. Right. Whether that&#8217;s the world&#8217;s best architecture or not, I&#8217;m not going to argue for it or against it. It&#8217;s just a lot better than domain model hierarchies, that I will say, right?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=2971">49:31</a>] Earlier we talked about the tradeoff between, I guess, performance code and maintainable code. And I saw you had this one video. It said in quotes, &#8220;Clean Code, Horrible Performance.&#8221;</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3003">50:03</a>] I feel like this is on the same topic. What&#8217;s your take on that? Because you also put Clean Code in quotes. What does that mean?</p><h3>50:12 &#8212; Clean code horrible performance</h3><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3012">50:12</a>] So it&#8217;s a fairly subtle topic, right? And that&#8217;s why the quotes are there. Because the first thing that you have to point out is that clean code is clean code. Object-oriented programming. A lot of these phrases, they mean different things to different people. And so if you&#8217;re going to say that this code is clean, it&#8217;s pretty hard to find two programmers who will agree on exactly whether it is or not.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3039">50:39</a>] They all have different ideas about what is clean and what is not. The reason I put that in quotes is I was talking about a very specific thing, which is literally the book Clean Code. The thing where there&#8217;s, here are the things that we would recommend that you do in that particular case. What I was talking about is there&#8217;s a certain set of ideas that is advocated in Clean Code as these are the things you should do.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3064">51:04</a>] And they are like you should have everything should be kind of, I guess I would say dynamic dispatch, or at least code shouldn&#8217;t know the types it&#8217;s operating on, right? Saying it&#8217;s dynamic dispatch is maybe it depends on the language whether that&#8217;s going to be true or not. But when I write code, I don&#8217;t know what types I have. It&#8217;s just, I don&#8217;t know. Some base classes is what I&#8217;m operating on, and they could be anything.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3088">51:28</a>] That&#8217;s thing one. Thing two is functions should be very small. In fact, the numbers are kind of weird. They&#8217;re like four lines or five lines. When you go look at the actual things that are claimed in some of this literature, it&#8217;s really strange. You&#8217;re just like, gosh, that&#8217;s very small, right? And I go through some of those things and I say, look, if we were to actually follow these, we get into a really bad state because a lot of the languages that people are going to be using to implement this stuff, like C, for example, which even would be a fairly good case in a lot of cases.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3121">52:01</a>] At least it&#8217;s a compiled language and so on. In a lot of cases, these things are kind of a recipe for disaster. If you&#8217;re handing out types where you can&#8217;t know the type at compile time, or you can&#8217;t know the type, or it&#8217;s going to be in a different module, so unless you have really aggressive link-time code generation, it&#8217;s not going to be likely that you could figure this out. And you&#8217;re saying all the functions should be very small.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3144">52:24</a>] And of course they&#8217;re going to be virtual because they&#8217;re on these sort of types that you don&#8217;t know what they are. This creates a really toxic combination for the compiler. Normally you can have&#8212; I mean, you can have small functions if you want them only four lines long. If the compiler can know that, it can just inline that, right? If it can see clearly what you&#8217;re doing, it can just merge those things in. We don&#8217;t have a problem if they&#8217;re behind a virtual function.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3170">52:50</a>] So it doesn&#8217;t know, it can&#8217;t guarantee that it is that type. It can&#8217;t do that. And so when you talk about all this stuff, what you&#8217;re giving up... People, I think, also because the video&#8212;I mean, the video was a very short video as part of a series&#8212;I think people sometimes get the wrong idea that I think that virtual function calls cost a lot. They do cost something, but depending on how you want to look at it, they actually don&#8217;t cost that much because it depends on whether it&#8217;s predicted correctly or not.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3199">53:19</a>] But in general, the cost of a virtual function is less than you would think a lot of times. The actual reason that it slows things down is because the compiler cannot merge the code together. It can&#8217;t inline stuff and remove and reduce the waste. That&#8217;s the actual problem. So if you just want to call a virtual function, yeah, if it was a really, really hardcore optimization scenario, then you don&#8217;t want anything, probably calling a virtual function in general, because there is some cost to that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3225">53:45</a>] Just calling a function has cost. But that wasn&#8217;t the really bad part of it, right? So it&#8217;s that. Plus there&#8217;s an additional thing, which is that it can&#8217;t unroll and go wide on things. If the compiler can see everything you&#8217;re doing, it can use SIMD, it can widen loops to operate on multiple things at once. It can unroll those loops. There&#8217;s all these things it can do. If it&#8217;s got these virtual functions, it&#8217;s just like, that&#8217;s it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3248">54:08</a>] I don&#8217;t know. I can&#8217;t make decisions about that. I have no idea what this thing on the other end is doing. That was my point on that. And again, it&#8217;s not meant to be assailing the idea that your code should be easy to read or easy to maintain. It&#8217;s just more like these guidelines seem really bad. And also, I don&#8217;t think we need them for code to be maintainable. I don&#8217;t have trouble maintaining functions that are 30 lines long or 50 lines long.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3273">54:33</a>] I don&#8217;t find that five is a magic number or something, or that it has to be really short. And also I find that it&#8217;s usually pretty easy to write code where I know what the types are, that those can be determined at compile time. I don&#8217;t think that results in code that&#8217;s particularly hard to read.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3288">54:48</a>] When I was in college and very early, I remember people would recommend that book, &#8220;Oh, you should read Clean Code.&#8221; So my understanding is that you wouldn&#8217;t recommend that to people who are software engineers to read that. I guess the way I would categorize it is a mixed bag.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3300">55:00</a>] So I wouldn&#8217;t actually say that I disagree with all of the things in the book, though. I mean, some of the things are things that I definitely do myself. Like giving variables easy-to-understand names is something that I think is good advice, and that&#8217;s in that book. Right. And so kind of more, I wouldn&#8217;t necessarily say do or don&#8217;t read the book. What I would say is, like, I think there&#8217;s some things in this book that don&#8217;t.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3331">55:31</a>] That aren&#8217;t addressed properly. Like, I don&#8217;t think you want to tell people these rules of thumb and not tell them about these other problems. That was exactly what I was pointing out there. And so usually it&#8217;s more that it&#8217;s like I find that oftentimes there is this tendency, I guess I&#8217;ll say, to pretend that we don&#8217;t have to talk about performance and we can just say, in fact, premature optimization is the root of all evil.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3361">56:01</a>] Bringing it back to there, I feel like there&#8217;s this temptation and very prevalent practice of sort of taking that idea to mean we don&#8217;t have to talk about performance when we&#8217;re teaching people things at all. Or like when you write the book Clean Code, you don&#8217;t have to talk about performance at all. You can just include a thing that&#8217;s basically like, hey, there&#8217;s performance issues, but most of the time it&#8217;s not a problem or something like that, which is, or if you look at Refactoring, the book Refactoring, very popular, it literally says exactly that in like the OOP thing.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3394">56:34</a>] It&#8217;s like there might be performance concerns, but they&#8217;re usually okay or something, right? And that is not true. I just think that&#8217;s fundamentally not true. I think these books should include detailed discussions of the actual performance problems that you are very likely to hit if you take some of their advice, because it&#8217;s not that it means you can&#8217;t do those things.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3417">56:57</a>] But I&#8217;ll quote you back to you. It&#8217;s about trade-offs. There is always a trade-off that you&#8217;re making, and sometimes if you understand the trade-off, you will make it. You&#8217;ll say, I know this is going to cost. This may be a serious performance problem for us, but I think that&#8217;s okay. I think it&#8217;s not like our performance won&#8217;t suffer to the point where it&#8217;s a problem for the product. I understand what that is.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3443">57:23</a>] There&#8217;s a huge difference between that and just going like, eh, we don&#8217;t worry about performance till the end. It&#8217;s completely different. Those are two different engineering approaches. And so what I try to do is encourage people to put that analysis back in, understand the performance trade-offs you&#8217;re making, understand how much it might cost in the future to make these fixes. And I don&#8217;t really do any AI stuff, but my sort of feeling is more so now than ever.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3473">57:53</a>] I feel like that&#8217;s got to be pretty important because you&#8217;re instructing these AIs to do what they&#8217;re going to do. And I feel like it would be a bad idea not to know about performance trade-offs because, especially if an AI is going to be doing your bidding and structuring the code the way you want to structure it or doing whatever, it seems like a fairly straightforward part of the process to include performance stuff in those instructions.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3497">58:17</a>] Right? That would be a natural thing that we would want to understand and do. So I feel like really I&#8217;ve been just trying to get that more into the conversation, and I feel like it&#8217;s as relevant now as it ever was and arguably maybe more so. I don&#8217;t know, but I could be wrong about that.</p><h3>58:33 &#8212; How to write high performance code</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3513">58:33</a>] What are some things that are in your mind? They&#8217;re those high-value performance things to consider that don&#8217;t cost a whole lot to think about.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3523">58:43</a>] So to me, awareness is the number one thing, right? Understanding roughly the performance characteristics of the hardware that you&#8217;re working on, which includes things like if there&#8217;s a network, what does that look like? And so on. It&#8217;s really about awareness more than anything else. Because a little bit of awareness can go a very long way if you understand the basic concept that network latency versus network throughput are different things.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3555">59:15</a>] You can make upfront decisions about structuring code such that you batch things properly and so on and so forth. If you don&#8217;t understand those things, you could get very far down a project only to realize that everything you designed and the way it works is all very serial and really just cannot be accelerated over a network at all. It&#8217;s never going to run reasonably or something like that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3576">59:36</a>] Right. And so I think the lowest hanging fruit is just to get some education in performance, get some education in how to think about the way a machine works and what makes it fast or slow, what it struggles with and what it doesn&#8217;t. And just to keep that in the back of your head, because at the end of the day, if you have that knowledge, I think you&#8217;re very unlikely to make the kinds of architectural decisions that will be hard to undo later.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3606">01:00:06</a>] Right. And so that&#8217;s really the majority of it, I think. And that is by far the highest impact, lowest cost thing you can do is just do that training once, because once you have it, it&#8217;s with you forever. Once you understand how to think about performance, it can always be there. And at any time you can sort of have that alarm bell of like, &#8220;I don&#8217;t see...&#8221; You&#8217;re always kind of looking like, &#8220;I don&#8217;t see the path towards this running well,&#8221; then that&#8217;s your cue to stop.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3638">01:00:38</a>] Maybe we need to rethink how we were making some of these decisions. As long as you can see that path, you can delay most optimization work. As long as you can see the path, like here is how we will optimize this, you&#8217;re in pretty good shape because if you&#8217;re making those trade-offs correctly, you&#8217;re not going to paint yourself into a corner where there&#8217;s nothing that you can do other than scrap and rewrite.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3665">01:01:05</a>] You had this one video title. It said, &#8220;Where does bad code come from?&#8221; I guess, yeah. My question to you is: what&#8217;s the answer to that? Where does bad code come from?</p><h3>01:01:18 &#8212; Where bad code comes from</h3><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3678">01:01:18</a>] Yeah, that&#8217;s a good question. My feeling on where bad code comes from is that, for everything that I&#8217;ve seen on projects in the past, the bad code all seems to come from roughly the same source. And that is not dealing with the actual thing that&#8217;s happening. Right. There&#8217;s a lot of ideas in computer science about sort of doing upfront design and making a lot of decisions without ever really implementing anything.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3714">01:01:54</a>] And I have found that universally that leads to the worst kind of code. And the reason for that is, I hate to be pessimistic, and I hate to be pessimistic about my own ability, but I&#8217;ll just say, for me personally, I am not able to correctly hold all of the details for most complex software in my head at once. There&#8217;s just a lot of stuff going on at, like when the actual instructions hit the CPU, there&#8217;s a lot going on down there that&#8217;s very hard to keep all in your head.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3754">01:02:34</a>] And so when you approach an upfront design, typically, unless the problem is very stupidly simple, right? When you approach something with upfront design and you think you&#8217;re going to get all of this architecture worked out ahead of time, you forget some important things. You make decisions without realizing, oh, wait, there&#8217;s this thing that has to happen there that means that this isn&#8217;t really the right way for these pieces of code to interoperate.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3783">01:03:03</a>] Then you end up seeing these hilarious APIs or something where it&#8217;s like you&#8217;re just like, how did this end up being the way that you do this? I have to create all these objects and I have to connect them together with these things and I have to create this weird filter graph thing and then I have to call compile on it. But only you end up with these insane things. And you&#8217;re like, all I wanted to do was call this one thing that said, like low pass filter this buffer.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3806">01:03:26</a>] It could have been one function call, right? It&#8217;s like, how did you get there? And the answer is because you weren&#8217;t actually dealing with the actual problem and looking at what the actual code looks like when you actually want to solve the problem in practice. And so my idea where bad code comes from is usually just that. It&#8217;s that. It&#8217;s this failure to engage with the actual reality of the situation.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3832">01:03:52</a>] Maybe there are people out there. I certainly have never met any, but maybe there are people out there whose brains are so expansive that they can actually hold all that in there and they can therefore just do it upfront. But other than for various sort of constrained problem spaces, I&#8217;ve never seen that work. And there are some times when that&#8217;s the case with, if you&#8217;re just doing, I&#8217;m trying to do this particular mathematical operation, that might be a case where you can work it all out on paper because it&#8217;s very specific what it is.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3861">01:04:21</a>] It takes these N inputs, produces this output. And I can work, right? But when you&#8217;re talking about architecture, like, hey, it&#8217;s a web browser, right? Some complicated thing, the details are all that matter to me, right? It&#8217;s like, it&#8217;s where all the complexity lies. And if you try to do the design without reckoning with them, it just leads to bad results in my experience.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3881">01:04:41</a>] So how do you avoid writing bad code in that case?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3885">01:04:45</a>] I try to start with the actual solutions to the problems and work up from there. So one way to say it would be instead of top-down programming, bottom-up programming as much as possible. And top-down programming isn&#8217;t really the same as designing up front. But I&#8217;m just saying as a contrast there, try to start with the things that you know you need. It&#8217;s like, okay, if I&#8217;m building this thing, it&#8217;s gotta have a rasterizer.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3910">01:05:10</a>] All right, I&#8217;m gonna start making a rasterizer. I&#8217;m gonna see what kinds of stuff that does and how I would like, what&#8217;s the most straightforward way to call this and use it. Okay, let&#8217;s make an API out of that, right? Build it out of steps where the abstraction comes from actual working code that gets abstracted rather than pushing the abstraction down from the top, where I say, this is how I will abstract the rasterizer.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3936">01:05:36</a>] And now I go write the rasterizer and the thing that uses the rasterizer, only to find that that is not a very good way to use a rasterizer. So to me, starting with that and working upwards is the best way to ensure that you will get an architecture that is at least good for one thing. Now, if you want it to be good for multiple things, you probably need a couple different. I want a few different varied things that would use this rasterizer that I will kind of work on together, and I&#8217;ll make abstraction decisions that work across all three of them.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3967">01:06:07</a>] That&#8217;s a good way to make a more reusable API that works for a lot of things. But I never want to just sit down and say, &#8220;How will I design an API for lots of people to use a rasterizer without actually writing any of it?&#8221; It&#8217;s like, no, no, no. And I definitely can&#8217;t do it. And I would question someone who says that they can. Unless one caveat: they&#8217;ve written a lot of them before, right? If you&#8217;ve written tons before, then you&#8217;ve sort of done the thing that I&#8217;m talking about already.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3993">01:06:33</a>] And you might remember, oh, it&#8217;s this way, right?</p><h3>01:06:37 &#8212; Why design docs before code is a bad idea</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=3997">01:06:37</a>] I think there&#8217;s a lot of common advice, especially at these larger companies where there&#8217;s this process to write a design doc and put together everything in advance, present it before you write any code. And it sounds like you&#8217;re saying that that&#8217;s a terrible idea.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4017">01:06:57</a>] I think that&#8217;s an absolutely terrible idea. Now, that doesn&#8217;t mean that I wouldn&#8217;t be okay with that process with a slight modification.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4025">01:07:05</a>] If what you&#8217;re doing during that time is writing test code in the way that I&#8217;m saying. So we&#8217;re going to write a little experimental rasterizer. We&#8217;re going to do those things. We&#8217;re going to start, and then our document that we&#8217;re producing is, here is what we have determined is a good API based on these experiments. I have no problem with that. That&#8217;s following my procedure pretty much to a T. And I don&#8217;t tend to produce documentation because I work on smaller teams.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4052">01:07:32</a>] I don&#8217;t have to have a 1,000-person org know exactly what I&#8217;m doing there or whatever. So I wouldn&#8217;t produce upfront documentation normally in a case like that. But if you wanted to, that seems totally reasonable. Right. And that is communicating our research into how this should be structured out to a wider audience. But we still did the due diligence of determining that it really does work in practice.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4075">01:07:55</a>] It&#8217;s not just our guess about what it will be.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4079">01:07:59</a>] I see. Okay, so you&#8217;re saying kind of this pair design, pair prototype, just build as you go, but it&#8217;s just so you get a sense of the reality of the design.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4090">01:08:10</a>] Yeah. And I think I would be very comfortable with the team that wanted to work that way. If they&#8217;re like, look, we want to produce, we want to document this thing before we start developing. That&#8217;s just how we feel more comfortable with how we&#8217;re going to do it. Maybe it&#8217;s a very large org, maybe we have some very good reasons for that. Then I would say totally fine. It&#8217;s just, I want to hear if I&#8217;m going to be comfortable.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4111">01:08:31</a>] I want to hear that that process is not, we&#8217;re just typing on paper. We are actually testing these API design decisions, and they come from looking at actual usage code that we have made and that we have implemented at least experimental versions of, so that we know they really do account for all the details that we&#8217;re likely to encounter in practice. That&#8217;s what I want to hear. Right. I don&#8217;t want to hear, we thought it through.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4137">01:08:57</a>] I&#8217;m like, did you think it all the way through? Because I&#8217;ve seen a lot of times where people thought they did and they didn&#8217;t, including myself. That is not. That is not. I always think it all the way. So it&#8217;s like, no, this comes from a personal place of I forgot the thing. I forgot this important thing and I made a really stupid design, right?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4154">01:09:14</a>] There&#8217;s another video you had. It was titled The Only Unbreakable Law.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4159">01:09:19</a>] Oh, yeah.</p><h3>01:09:20 &#8212; The only unbreakable law in software engineering</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4160">01:09:20</a>] What is the only unbreakable law in software engineering?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4164">01:09:24</a>] So, yeah, that was a lecture I did where I was trying to think of, if I was going to say one thing about architecture to people, right? If I was going to say one thing about architecture, what could I say that isn&#8217;t probably wrong? Because you think about a lot of our ideas about architecture, it&#8217;s like even the stuff I just said to you, who knows what we&#8217;re going to be thinking 10 years from now.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4187">01:09:47</a>] It&#8217;s like, they&#8217;re not really laws, they&#8217;re just observations that we&#8217;ve had that maybe this is a good way to do things, but it&#8217;s not really a law, right? So the only thing I&#8217;ve seen that feels like a real law to me is the thing I cover in this lecture, which is Conway&#8217;s Law. And this is a. It&#8217;s not a real law in the sense of like the laws of gravity or so, like, it&#8217;s not a physics law.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4214">01:10:14</a>] Like a physicist would laugh at calling it a law. It&#8217;s probably not even a theory; it&#8217;s more like a hypothesis at this point, right? But in terms of things that I would place money on being validated as a law at some point, if we ever had the means to do so, this would be one of the only software engineering things that I think would qualify. And it is that if you would like to organize a series, like a set of people to work on something, or in the modern parlance, let&#8217;s say agents.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4256">01:10:56</a>] So it could be. It doesn&#8217;t have to be human. It&#8217;s any system that has sort of its own internal processing, like a human brain does or like a computer does. If you need to partition a problem into a set of workers in this way, where the communication between those workers is slower than their own internal computation ability, which we would all agree is true about humans, it&#8217;s true about two computers running an AI model as well, right.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4285">01:11:25</a>] The amount of time it takes to send information back and forth to them is significantly slower than what the GPU can be processing directly on its own, right? In systems that look like that, when you start to tackle a problem, in order to get the benefits of that parallelization, you will have to make decisions about who is working on what, at least some decision. For example, if we are making a car and you and I are the two people who are going to make the car, I&#8217;m going to design the body of the car, you&#8217;re going to design the wheels.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4314">01:11:54</a>] Right. Just a decision someone could make. Once you&#8217;ve made that decision, you have now locked in the fact that the iteration on the design will be slower across that boundary than it is interior to either of the two parts. So, for example, your ability to improve the design of the wheels on their own and my ability to design the body of the car and improve that on its own will be much faster than our ability to design the interface between the wheels and the tire.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4352">01:12:32</a>] Or it&#8217;s not even just the interface, the degree of harmony between the designs. So the degree to which your wheels complement my body design and my body design complements your wheels, that they all are considering the trade-offs together. Right. And Melvin Conway&#8217;s paper, which lays this out, this is a very early paper. It&#8217;s in the 60s at least. Was it in the 50s? It&#8217;s way back. When that lays this out, points out the consequence of this is that products or any.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4386">01:13:06</a>] And when I say product, I mean anything that comes out of one of these design processes or development processes. Products will have a similar structure to the organization that produced them. You will be able to see in the product evidence of this because the degree to which things are able to harmonize is restricted across that boundary. So when you look at the object, it will have that boundary visible in some way.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4413">01:13:33</a>] Right. It won&#8217;t be quite as good at the interface between these two things or in the way that they work together as the things are themselves to themselves. It gets sort of flattened, like one way. The rule is pithily stated is products look like the org chart, but that&#8217;s not exactly what it says. But that&#8217;s kind of a downstream consequence, and boy, do you ever see it in the real world.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4443">01:14:03</a>] So, I&#8217;ve heard that in the context of these big tech companies, like, almost like it&#8217;s a bad thing to quote unquote ship your org chart, or basically the product you put out there has these inorganic, yes, things. So it sounds similar to what you&#8217;re saying.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4464">01:14:24</a>] Yeah, it&#8217;s. And I think part of the reason I think it is kind of an unbreakable law is I think it&#8217;s somewhat unavoidable. It&#8217;s more something that you just have to be aware of and mitigate to the degree that you can. Like, you just have to understand, look, that lower communication time, whether it&#8217;s humans or computers or whoever is doing this work, that lower communication bandwidth is whatever.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4494">01:14:54</a>] That just means that where we draw these lines has consequences for what we ship. And so we really want to try, over time, to align those boundary drawings to places where it will have the least bad impact on the result of the product. And that&#8217;s true for org charts, like I said. I suspect it will also become true for AIs and things like that, where it&#8217;s like, you will want to partition these problems in ways that respect that line drawing, because that will be less good than the thing that can be wholly solved by one unit, whatever that unit is.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4536">01:15:36</a>] And so to me, it&#8217;s a very helpful construct for thinking about things. It&#8217;s also a very good explanation for why you see something somewhere. You&#8217;re like, why doesn&#8217;t this just integrate this? It&#8217;s like, because there were two different teams. It&#8217;s like, I&#8217;m sorry, that&#8217;s just the answer. And it costs a lot more for those two teams to have integrated this. That&#8217;s not how we work. Right.</p><h3>01:15:57 &#8212; How he got into programming</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4557">01:15:57</a>] I&#8217;d love to hear more about your career and how you got into programming.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4564">01:16:04</a>] I started programming when I was really little because my father worked at Digital Equipment Corporation, which is actually a very important company in that era. No longer exists. Right. They would be sort of like almost like a case study in how not adapting to a change in technology leads to your demise. They were one of the most important companies in the world for computing. If you ever heard of a PDP-11 or a VAX in computing history, that&#8217;s them.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4595">01:16:35</a>] Right. They made these things that were foundational computers in computing history, gone today. Right. Part of them got absorbed by Intel, part of them got absorbed by Compaq, but they don&#8217;t exist. So he worked for that company. And so we always had computers in the home, and I learned to program when I was little, and that&#8217;s just sort of always what I wanted to do. I really enjoyed it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4619">01:16:59</a>] And so I ended up getting an internship at Microsoft when I was pretty young, and I met some people there. I came out to &#8212; I should say, I grew up on the East Coast, so nowhere near Microsoft. I met some people there. I ended up coming out to the West Coast really early on. This would be in, like, &#8217;95. So it didn&#8217;t feel early at the time. It felt late in computer history. But nowadays, like, oh, wait, &#8217;95.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4646">01:17:26</a>] Oh, my God. Before everything was on the web, right? There was a web, though it was nascent. I ended up coming out and working sort of in the game industry, and that&#8217;s what I did ever since. And I&#8217;ve mostly always worked on game technology. I&#8217;m famously absolutely terrible at understanding game design. People who watch my stuff know this about me. I&#8217;m very, very bad at it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4674">01:17:54</a>] I&#8217;ve tried to make games a couple times and I just cannot do the design side. I&#8217;m awful at it. But I really enjoy the engine stuff, and I feel like I&#8217;m okay at it, and I&#8217;ve been able to contribute to projects. So in general, most of the time, if you&#8217;ve used code that was written by me, you probably used it in the context of a video game that was using technology that I wrote. And an example would be I worked at RAD Game Tools on a character animation system.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4703">01:18:23</a>] I wrote the entire thing myself. That is used. Well, that&#8217;s not entirely true. I wrote the entire thing myself, except for the texture compressor, which Jeff Roberts wrote. There was this texture compressor that you could use as part of the pipeline for the exporting and stuff like that. It was a pretty cool project at the time. I&#8217;m really proud of it. The first version was awful because I was pretty new at the time.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4724">01:18:44</a>] The second version, I thought we did a really nice job. And it was made at a time when people didn&#8217;t think you could do licensable game technology of that kind because it was too hard to integrate into things, like you couldn&#8217;t get the performance or whatever. And we did a lot of things that I think were pretty innovative at the time. And we were able to make something that was very successful. And it&#8217;s.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4744">01:19:04</a>] I mean, we released the first version of that in &#8216;99. It&#8217;s still in use today, largely unchanged from the architecture that I guess was the second version I did in 2001 or something. There are a few changes to the architecture that people have made over time, but it&#8217;s largely unchanged. I found out recently Baldur&#8217;s Gate 3, which was a big game from Larian, came out. I had no idea. It turns out they use it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4770">01:19:30</a>] It&#8217;s in their engine or whatever. So I&#8217;m very proud of that product because I think it was in tons of games and was a very hard problem to solve. And I think we solved it pretty well. It&#8217;s largely irrelevant today. I would say it&#8217;s still in some people&#8217;s engines over time, but it&#8217;s not the kind of product you would make today because nowadays engines are monolithic and you license them as a whole. Typically, you wouldn&#8217;t be getting a character animation library.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4797">01:19:57</a>] It&#8217;s going to be something that&#8217;s built into Unreal Engine or built into Unity. It&#8217;s not the kind of product you would probably consider making today. That was mostly, in terms of things that I&#8217;ve done that people might actually have experienced or used, like folks at home. If you&#8217;ve played games, you may have played something that I had a hand in at some point, but only the technology, not the design.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4820">01:20:20</a>] I also worked a little bit on The Witness, which was a game by Jonathan Blow and team that I thought was absolutely fantastic. And I did some work on the walk system there that I was pretty proud of. I thought it came out pretty well. There&#8217;s some parts of it that are really janky. I don&#8217;t think I did a very good job of the actual implementation of it, but the design was pretty good, let&#8217;s put it that way.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4843">01:20:43</a>] Early in your career. So you worked at Microsoft and then went to go.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4848">01:20:48</a>] I was only ever an intern. I never actually, I mean, I guess that&#8217;s technically working there, but I didn&#8217;t ever work there as an employee. Employee, right.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4855">01:20:55</a>] As someone who&#8217;s writing code, I mean, there&#8217;s a lot of different paths, but there&#8217;s one path I imagine: you go and you work at one of these big tech companies and create their tech products, I guess. And then video gaming seems like another path. And I think there&#8217;s other paths as well, of course.</p><h3>01:21:16 &#8212; Why he didnt work in big tech</h3><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4876">01:21:16</a>] What drew you to going towards video gaming and not continuing down the Microsoft path or something like that?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4885">01:21:25</a>] I think that there&#8217;s a bunch of things I could say about that, but they&#8217;re probably all BS. The truth is probably that I&#8217;m probably more affected by the people around me than I would like to admit. I mean, I guess I don&#8217;t have a problem admitting it now, but at the time I probably wouldn&#8217;t have said that, right. I could envision an alternate past where I did stay at Microsoft and try to get a job there and work there instead of being just an intern and then going off and working at a different place.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4920">01:22:00</a>] The reason that didn&#8217;t happen was because of the people who I worked with there when I was an intern. I fundamentally really like. Nowadays I fundamentally really love doing things like analyzing assembly code or looking at exactly how a microarchitecture is working and these sorts of things. If I had gone to Microsoft and had just happened to be under some people who were doing that kind of work and had taught me how to write device drivers in assembly language or something like that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4953">01:22:33</a>] I might still be there today, right? That&#8217;s not what happened. The capsule summary is the group that I was supposed to be in, the way that they did internships at that time was that a set of interns, say three or four of them, would be underneath a particular manager who was going to be managing those interns. So there were a couple of us who were supposed to be reporting to this guy whose name I won&#8217;t mention, just in case for some reason it&#8217;s ancient history.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=4982">01:23:02</a>] He probably wouldn&#8217;t care at this point, but we were supposed to be reporting to this particular person. And literally the week before we arrived, he has this massive flame-out with upper management. Leaves, just walks out of the building, and has not been heard from since. So we show up, and there was me, a guy named Rudy, a guy named Rajeev, and we&#8217;re just interns. We show up, we&#8217;re like, hey, how&#8217;s it going?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5019">01:23:39</a>] And we report to this test guy, this SDET guy named Scott Latham, really nice guy. And we&#8217;re like, we&#8217;re supposed to be reporting to a program, like a software, or what&#8217;s going on? This is one of the guys in the test org. He&#8217;s like, yeah, that guy&#8217;s gone, he&#8217;s out of here. And me and this other guy, Bruce Johnson, are going to take care of the interns because, eh. Right. We don&#8217;t really know what you&#8217;re going to do.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5053">01:24:13</a>] The first thing I think Bruce had me do was like an ANI cursor loader. There&#8217;s this format, I don&#8217;t know, ancient history now, but there&#8217;s this format for animated cursors on Windows. If you ever see the stupid little walking dinosaur or the, like, no one uses these anymore, I don&#8217;t think. But they were this thing that was in there. He had me write a parser for loading ANI files, right? Like it&#8217;s just meaningless stuff.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5076">01:24:36</a>] So my experience there was pretty lame. I was like, this is kind of dumb. I don&#8217;t really want to work here. But it could have been totally different if I had been on some team that had really inspired me to learn the stuff that I didn&#8217;t know. I didn&#8217;t know assembly language at that time, really, at all. I think the only thing I&#8217;d ever written assembly language was a joystick polling routine for DOS because it was the only way you could poll the joystick in DOS.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5100">01:25:00</a>] I think that&#8217;s really it, if I&#8217;m honest about it. I think most of it is that I think even if you&#8217;re a very brash youngster, which I was, and even if you think that you&#8217;re very independent, I think you definitely gravitate toward people who you see doing things you think are impressive or technologically interesting. And that is just not the experience I had at Microsoft. Now, there were some people who I like, the people I went out to go work with were people in other parts of Microsoft who were leaving Microsoft to do a startup company.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5134">01:25:34</a>] And those were the people that I was really impressed with. Those are the people that I wanted to hang out with. So, if I just look at it in arrears, if those people had just been staying at Microsoft and doing something, maybe I would have gone.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5147">01:25:47</a>] And I think that&#8217;s the truth of it, as best I can determine at this point anyway.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5151">01:25:51</a>] So you moved across the country to go work with those people at a startup?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5156">01:25:56</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5156">01:25:56</a>] That feels pretty risky to do.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5160">01:26:00</a>] It was especially because at that time, a startup is not really what you think of as a startup today. Like, a startup was just some scrappy people in a crappy office doing stuff. Nowadays, you think startup, it&#8217;s like, well, we&#8217;ve got $10 million in VC funding and we have free sodas and whatever, like in the break room, and a masseuse comes in periodically or something like this.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5182">01:26:22</a>] Right. I don&#8217;t know what the now version of a startup is, but it&#8217;s very different, right? So, yeah, it was pretty risky. And I don&#8217;t really know why I did it. It was mostly just because I hadn&#8217;t had really good experiences with education. I didn&#8217;t really want to do college, or I didn&#8217;t want to do any more education. I wanted to actually work, for whatever reason. And this seemed like a.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5212">01:26:52</a>] But this was just the easiest way to do that. And I don&#8217;t really regret it, honestly. I&#8217;ve been able to learn most of the things that I would have needed to, if I was gonna go do a PhD or something. I&#8217;ve ended up doing work that is the sorts of stuff, like I&#8217;ve been lucky enough to be in positions where I could go spend a few months writing this particular linear least squares solver or something like that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5239">01:27:19</a>] The kinds of things you might have done as a thesis or something, or that walks on somebody done The Witness. Those sorts of things are the kinds of things you would have done had you decided to get a master&#8217;s degree or something like that. So I&#8217;m fortunate enough that I&#8217;ve kind of been able to have that part of the education because not everyone gets that opportunity. You may never get, if you don&#8217;t do a master&#8217;s thesis or you don&#8217;t do a PhD thesis, you may never really get the chance to really do deep work on one problem and learn a lot about it, read a lot of papers, maybe make some novel contribution in some very small way usually.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5279">01:27:59</a>] Right. And so I was just lucky enough to do that. So I don&#8217;t really regret not doing something like that, but I think had I not had those opportunities subsequently, I could see being regretful about that because that is something that I do enjoy doing. And if I&#8217;d never had the chance, I would have been sad.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5298">01:28:18</a>] So you went and you traveled to this startup because there&#8217;s no funding, it sounds like.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5305">01:28:25</a>] So, yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5306">01:28:26</a>] What, you? It&#8217;s just a group of guys. Was there pay or...</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5310">01:28:30</a>] It&#8217;s just not really. No, it was. It was pretty rough, and it didn&#8217;t last very long. Right. I ended up going to work at an actual game company shortly after that called Gas Powered Games. And then shortly after that I started working at RAD Game Tools, which I stayed at for quite some time and where I did that character animation system and stuff. So it was pretty quick into not doing that, but it was still a pretty educational experience for me.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5333">01:28:53</a>] And also what I will say is at the time, I worked with a guy named Chris Hecker, who I have to give basically complete credit for teaching me basically about reading technical papers. Before that time, I just thought math was kind of stupid and an annoying thing you had to do in class. Right. And I definitely didn&#8217;t know anything about reading a SIGGRAPH proceedings, really.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5366">01:29:26</a>] And I may have been aware. I mean, I was aware of SIGGRAPH. I knew what it was. But I don&#8217;t think if you&#8217;d handed me that binder, I would have known what that was or what to do with it. I think the biggest takeaway that I got from that, other than startups are hard and maybe don&#8217;t do them unless you have a lot of people and a lot of funding and aren&#8217;t really that much on the line, maybe.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5391">01:29:51</a>] But the biggest takeaway that I got from that that was positive was just a much deeper appreciation for math, a much deeper appreciation for research, how to read it. That&#8217;s where I learned to use CiteSeer, which would be kind of the precursor to Google Scholar or whatever. Crawling references and all that stuff. In a lot of ways you could say if I hadn&#8217;t had that experience, I bet I wouldn&#8217;t have given the two talks that we&#8217;ve talked about for most of the interview.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5417">01:30:17</a>] Because what are those talks? They&#8217;re me crawling every reference. They&#8217;re just me going back and back and back and looking at everything that I can find. And that&#8217;s something that if no one ever conveys to you the importance of reading, the scholarship and being aware of what&#8217;s being done, and also teaches you how to read a tech paper and how to parse through it, I don&#8217;t know that you get that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5442">01:30:42</a>] I don&#8217;t know that you get that.</p><h3>01:30:44 &#8212; Should you work at a startup early on</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5444">01:30:44</a>] I remember early in my career, I think there&#8217;s this common path of these big companies, and then there&#8217;s this thought of, oh, me and a couple buddies, let&#8217;s go build something. Now, with your experience looking back, if someone out there is young and thinking about starting, would you say go that common path or would you say take the chance?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5470">01:31:10</a>] It&#8217;s a really tough question. And I think the advice is that you have to think about what it is that you want to do every day. So in my mind, the things I regret are the times when I&#8217;ve had to do things or chosen to do things that I wasn&#8217;t really that happy doing every day. Right? And so I think you just have to optimize for the experience you want to have. Because a startup might work out, it might not.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5505">01:31:45</a>] It&#8217;s always a risk, right? You may go to the big company and it may be cool. There may be interesting things there, or there might not be. You don&#8217;t really know. You&#8217;re making a kind of blind decision. And so I think you kind of have to optimize for what do you want your day to day to be? What do you want the experience to be? And you want to keep a running talent, you want to be aware of it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5532">01:32:12</a>] If you make a decision and six months in, you are not liking what you&#8217;re doing every day, then you need to get out of that, right? That&#8217;s my opinion. So I don&#8217;t know. I think most people probably, if they sit down and actually clear their mind and aren&#8217;t engaging in too motivated a reasoning just for themselves, go, what do I envision working at the startup will be like? What will I actually be doing every day?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5560">01:32:40</a>] Am I gonna like that? Is this gonna be thrilling? Trying to make this thing work and being kind of on the edge, and do I want the kind of war stories of the things we had to do to pull off the demo or whatever? Right. If all that sounds exciting to you and you wanna have that experience, then you should do that thing. And I would say, really, I really mean that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5585">01:33:05</a>] I don&#8217;t even think you should take into account, even, like, let&#8217;s say, the startup&#8217;s gonna fail. If you want to have that experience, then you need to do it, right? It&#8217;s just like anything else. It&#8217;s like going and backpacking across Europe or playing guitar at a nightclub. It&#8217;s like, you may know that these things are not profitable. Is that the experience you want to have?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5606">01:33:26</a>] Are you going to look back on that six months or a year or five years or whatever, the amount of time you&#8217;re thinking of committing to it? Are you going to look back and say, I&#8217;m glad I did that, or are you gonna be like, that was time that I would have rather spent some other way? Right. And so I think most people can probably, if they&#8217;re honest with themselves, at least make a pretty good guess.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5625">01:33:45</a>] And that guess, if they&#8217;re honest with themselves, is gonna be better than any advice I&#8217;m gonna give them. My guess about what you should do is gonna be worse than yours. So you should just make that determination. Because another way to look at it is, you know, the big company side, maybe you just wanna have some security. Maybe when you&#8217;re starting out, I mean, I&#8217;ll just give some examples.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5644">01:34:04</a>] Maybe you want to spend a lot of time dating. You want to find a partner, and you want to have a relationship, and you want to have a family. Maybe that&#8217;s more important to you. Maybe the startup, yeah, it might be fun. Maybe it would be more money if it worked out or whatever. But maybe the stability is actually going to be something that you would really benefit from because the sorts of things that you imagine being fulfilling in your life are not all around what you&#8217;re going to be doing in computing.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5674">01:34:34</a>] I think you can have some guess about those things if you just let yourself have some space to think about it and don&#8217;t engage in too much motivated reasoning about it. And I think that would be the best way to make a decision, if you&#8217;re going to make one. That&#8217;s my feeling about it. Anyway.</p><h3>01:34:52 &#8212; What video game engineering is like</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5692">01:34:52</a>] We talked a lot about video game engineering, and I have no idea what goes into video games. If you were to just boil it down to the big pieces that you need for a video game, what are those software components?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5707">01:35:07</a>] So the biggest thing that&#8217;s different about a video game is that the sort of original mental model that you&#8217;re taught in programming, which is the standard I O, is not how it works, right? So the biggest mental shift is just like, oh. The way that a video game or a simulation of any kind works is I need to have some world state.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5737">01:35:37</a>] And I&#8217;m constantly updating that world state on a regular interval. Everything is always happening. There&#8217;s no &#8220;I wait for input and then I do something,&#8221; right? Everything is always going because it is real time. The time is going forward. So the components tend to be built around this idea, and it depends on what level of sophistication you end up getting into. But in general, you need a way of storing that world state.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5763">01:36:03</a>] So some kind of. We usually call these entities, like the things that make up a world.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5768">01:36:08</a>] So some way of modeling what is in the world. Where are things in the world, what is their state, what are they doing right now? So you need something that does that, something that is in charge of advancing that state. This is a mixture of several things. It could involve physical simulation, it could involve AI, not necessarily large language model like the modern notion of AI, but pathfinding, making a decision between whether I should attack the player or not, that sort of AI.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5802">01:36:42</a>] So we have the world. Some way of showing the world state, some way of advancing the world state, like an update step that may involve lots of things like that. Then some way of presenting the world state. So a renderer. Right. And this is something that typically has, I guess I would say it&#8217;s traditionally been a very important part of a game because the visuals are often something that is a selling point for games.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5832">01:37:12</a>] People produce trailers, certainly in the AAA space. We&#8217;re trying to show you how cool the new lighting looks and the whatever. So that is a pretty big component of games. Oftentimes you go look at an indie pixel art game, it might be much smaller because that&#8217;s a much smaller part of the problem. Now that&#8217;s generally what the shape looks like. There&#8217;s a lot of little pieces there. Typically nowadays we would have some way of asset streaming. The renderer needs to have things like textures loaded, models loaded, things like that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5865">01:37:45</a>] We need world data, physics, collision models, these sorts of things. They might be too big to fit in memory, or we don&#8217;t want to spend a lot of load time to load them up front. We want to stream them in as they&#8217;re necessary. So there&#8217;s typically this thing that&#8217;s sitting there constantly grabbing things off of disk, caching things, pulling them in and out of a cache that&#8217;s being used by the renderer, by the physics, and so on.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5884">01:38:04</a>] So that&#8217;s another common component that&#8217;s new, didn&#8217;t used to be there. There&#8217;s going to be an audio and music system. Obviously, that&#8217;s in charge of when sounds are triggered in the world, ambient sound effects that are happening in the world, music that&#8217;s playing, continuity. Again, all of that&#8217;s interacting with the world state to know which ones of those things are happening. The update step, which will be triggering things in that sound system.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5908">01:38:28</a>] And of course, the asset streaming to load what sounds are being played or load the music. So we typically have that. And then I&#8217;m trying to think, I&#8217;m trying not to leave out any major components. In a modern context, there&#8217;s often networking. So we want to have a way for multiple of these game clients to communicate with a server. So typically what that means is that world, that state of the world might be provisional. It might be sort of a predicted model of the world that&#8217;s not the real model of the world.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5941">01:39:01</a>] The simulator is actually running. The authoritative simulator is actually running on a server somewhere. And I am merely communicating with it to find out what the world state is. And then, because I don&#8217;t want to wait, I don&#8217;t want the latency of going all the way around, I am predicting the motion of things forward in time based on the last information I have. So that&#8217;s another kind of way that things tie in.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5967">01:39:27</a>] If you want to start doing things like competitive multiplayer and all these sorts of things. Right.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5971">01:39:31</a>] That makes a lot of sense because I used to play video games, and then these MMO online games, when I start to lag, everyone continues forward.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5983">01:39:43</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5984">01:39:44</a>] And then my Internet catches up, and then everyone jumps to where they&#8217;re actually supposed to be. Okay, that makes a lot of sense. I pulled some of your top tweets. I thought that might be interesting to discuss.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5997">01:39:57</a>] I&#8217;m sorry.</p><h3>01:39:57 &#8212; Why preventing recursion is reasonable</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=5997">01:39:57</a>] So someone said, NASA does not allow recursion in their code. How crazy is that? And then you said, not even slightly crazy, and it went really viral. Can you explain why that&#8217;s not even slightly crazy?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6012">01:40:12</a>] When you&#8217;re writing functions in a procedural programming language like we have, and that NASA is probably using, you are able to use the program stack for storage. I mean, that&#8217;s what it&#8217;s there for. It&#8217;s what local variables are. I call a function, I get some storage space on the stack. The compiler did that for me. It also saves the return address when I make a function call. That program stack is keeping track of where in my code I was, so that when the thing that I called is finished, I get back there, not some other point.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6050">01:40:50</a>] Neither of those two things are things you couldn&#8217;t implement yourself. They&#8217;re both things you could do. You could break the thing up into pieces. You could have ways of remembering what they were, like a state machine. You can have your own stack that you put data on and access it. Right? So the only thing recursion really does is it allows you to leverage the fact that someone already wrote that code for you and maybe more convenient to use because it&#8217;s built into the language.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6078">01:41:18</a>] Now, there are some things, if you really want to get super technical about it, there are some things that happen at a CPU level when functions are called, but we can put those aside for now if you want to talk about them after we wrote the code. So the reason that I don&#8217;t think it&#8217;s crazy to go, we&#8217;re not going to use recursion to implement an algorithm, is because if you&#8217;re doing that, you&#8217;re kind of just YOLO swagging that there&#8217;s room on the stack for whatever it was that you were doing, right?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6104">01:41:44</a>] And you can&#8217;t even really check unless you&#8217;re going to do some kind of weird thing. We could sort of do a thing where we go like, okay, let&#8217;s try to determine how many more iterations, how many more recursion depths we have before we hit the end of our stack. We can do those calculations, but it&#8217;s like, eh. And also if the compiler changed something about the layout, it wouldn&#8217;t be true anymore, and so on and so forth.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6127">01:42:07</a>] So it&#8217;s like, so if I&#8217;m NASA and I&#8217;m like, hey, I don&#8217;t want my astronauts to crash into the moon, it seems much more logical to say, don&#8217;t use recursion. Just figure out whatever this thing was you&#8217;re going to do, turn it into a loop, keep a stack, and make a state machine for it. We can reason about that much more clearly. We know exactly how far it&#8217;s going to take. It doesn&#8217;t matter what compiler, it&#8217;s going to work the same way every time, and it will know the bounds.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6151">01:42:31</a>] Precisely.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6153">01:42:33</a>] Very sensible to me.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6155">01:42:35</a>] And again, you assume at NASA that they&#8217;re not doing it for their health. They&#8217;re doing it because they have hard constraints on the problem domain where someone&#8217;s life is at risk, or at a minimum, many millions of dollars in equipment is at risk. If you screw up, if your stack falls, if you get something where you&#8217;d recurse too many times and hit the end of the stack, that is not just a, oh, I rebooted the computer or restarted the software.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6188">01:43:08</a>] So I don&#8217;t know. Hopefully that makes sense.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6191">01:43:11</a>] Andrej Karpathy had this famous tweet that kind of coined the phrase vibe coding.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6196">01:43:16</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6197">01:43:17</a>] And then you said, if you thought software was bad today, buckle up, because it&#8217;s about to get a whole lot worse.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6203">01:43:23</a>] Yes.</p><h3>01:43:24 &#8212; Is vibe coding bad for the industry</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6204">01:43:24</a>] It&#8217;s been about a year and a half since that tweet came out. The tweet came out in February of 2025. Would you say that that was an accurate prediction?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6215">01:43:35</a>] To be clear, I think the whole lot worse part is predicated on something that may not happen. And that is that the idea that we just kind of type some stuff into a computer and ship it to prod, right? Basically like, hey, could you make me a thing and then publish it, becomes a common way of doing things, right? For example, also done by people who maybe don&#8217;t have a computer science background. I think it&#8217;s fair to say, and that I wouldn&#8217;t be being sort of overly dismissive of AI at this point, to say that a person who is well trained in computer science using an AI to make code right now can make substantially better code than someone who doesn&#8217;t know anything about computer science, who is just given Fable and types some stuff in.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6267">01:44:27</a>] The difference is rather dramatic, I would say, from everything that I&#8217;ve seen. So part of my concern that I was trying to express in that tweet was like, if the idea is like, we&#8217;re just going to type stuff in and we&#8217;re not going to really be checking the code, someone who knows computer science is not really going to be looking at it. Worst-case scenario, it&#8217;s literally just like some random person in marketing somewhere who has no idea what programming is, just types of stuff in and crosses their fingers.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6297">01:44:57</a>] Right. I think we&#8217;re in for a world of hurt right now. It&#8217;s a race, so it&#8217;s hard to say because it&#8217;s basically a race of how good can you make the AI versus how much adoption does it get? It&#8217;s a curve. If you can make the AI good enough that the people who are adopting it at a particular rate are always using an AI that&#8217;s good enough for what they&#8217;re adopting it for, we wouldn&#8217;t expect software to get significantly worse.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6324">01:45:24</a>] If those curves go the other way, we&#8217;re in a lot of trouble. Right. So I feel like right now we&#8217;re almost kind of teetering on this knife&#8217;s edge. It&#8217;s really, to me, it feels like a footrace of improving AI so it can be more autonomous and make better decisions without you needing to make them for it, versus the capability level of people who are using it and the degree to which they&#8217;re paying attention to its output.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6349">01:45:49</a>] It&#8217;s like these two curves that are just, what&#8217;s going to happen? I don&#8217;t have a prediction. I don&#8217;t know where we&#8217;ll be in a year. Obviously, for all of our sake, I&#8217;m hoping that the AI curve wins because I agree. I understand certainly the perspective of people who maybe just don&#8217;t like AI and don&#8217;t want there to be AI. I can understand wanting it to fail. I understand that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6381">01:46:21</a>] Right. And I&#8217;m no fan of AI myself, so it&#8217;s not like I&#8217;m going to criticize someone for taking that position, but at the end of the day, if you&#8217;re talking about something that tons of people are using, you&#8217;re going to kind of want it to be good. I think at this point, given the level of adoption of AI, I really don&#8217;t think it would be great if it stopped getting any better right now.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6406">01:46:46</a>] Like, if this was as good as it was going to get, I think that might be bad. Certainly six months ago, I think that was true, and I think it&#8217;s probably still true today. So I think ideally, if you want software to not be terrible, you have to kind of still be hoping that six months from now the AIs are again significantly better than they were. Right. Like, that is the only way out of the current situation, as I see it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6435">01:47:15</a>] Right. I don&#8217;t know if that&#8217;s fair, but that&#8217;s my sort of feeling on that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6441">01:47:21</a>] One of your other top tweets, it was Shopify put out this internal memo and you just replied, Slopify.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6450">01:47:30</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6450">01:47:30</a>] So I guess it&#8217;s because in this tweet it&#8217;s leadership pushing the adoption curve, maybe harder than the capabilities of the AI in this case.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6459">01:47:39</a>] Yeah. Although I also just like the pun. One of the things that I think is most unfortunate about the AI adoption, as I&#8217;ve seen it, is just the, because people think that it&#8217;s going to be this major. I guess if I had to categorize the way it appears that companies are reasoning about it, they&#8217;re assuming that if they don&#8217;t get in early it will be a big disaster for them, right? Like there&#8217;s a tremendous.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6487">01:48:07</a>] They don&#8217;t just think, &#8220;Oh, well, we can just wait until the AI does what we need it to do and then start using it.&#8221; They&#8217;re like, &#8220;No, we have to do it now. Even before we know whether it can really do the thing that we want it to do or whether we know whether the outcomes are good, everyone has to do it right now. Let&#8217;s do this, right?&#8221; And I understand why they want, why they&#8217;re going about it that way, because they think that that&#8217;s critical, right?</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6507">01:48:27</a>] They obviously believe that&#8217;s very important. And to me, that&#8217;s just kind of terrifying because, as with any technology, the sane way to do it is to measure its capabilities, see how well it is able to solve problems that you have, see if it solves them faster than the way that you were doing it, and put it into a workflow at such a time as you&#8217;ve determined that it is a net positive. That&#8217;s just the same; that&#8217;s what you would do with any technology.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6537">01:48:57</a>] And you&#8217;d probably have your team of people whose job it is to assess this thing, and they&#8217;re out there YOLO swagging it. They&#8217;ve got 3,000 agents working on this cluster, talking to each other and doing God knows what, right? So there&#8217;s going to be that, and someone&#8217;s going to do that. But that should not be every org. You wouldn&#8217;t just be like, everyone needs to use a ton of tokens, right? So yeah, I do have concerns about that.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6564">01:49:24</a>] I don&#8217;t think that the way AI adoption was done was the best way for quality in software. But I would temper that statement with just the obvious fact that we were not exactly a five nines industry to start out with. Software quality was really pretty low rolling into the AI era. So I always try to just also caveat most of the things that I have to say that might be critical of a particular thing happening with AI with just the fact that, look, it wasn&#8217;t particularly great beforehand either.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6601">01:50:01</a>] A lot of this software was pretty low quality, and so you can&#8217;t. Some AI things may make things worse, but it&#8217;s not like software was amazing and the AI showed up and ruined everything. That is a completely ridiculous narrative. That is not true at all.</p><h3>01:50:17 &#8212; Technical reading recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6617">01:50:17</a>] Do you have any top technical book recommendations for engineers that might be listening?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6625">01:50:25</a>] No. I would probably use this opportunity to try to pitch reading technical papers. I think something that really the industry could use more of is people being more aware of what&#8217;s happening, both the historical papers, but also just current papers. It&#8217;s daunting at first, to be sure. If you don&#8217;t tend to read technical papers in your field, you will probably find it confusing. You won&#8217;t know where to find them, you won&#8217;t know how to approach them.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6660">01:51:00</a>] It will seem to take too long to read them because you don&#8217;t know how to skim them properly and determine whether it&#8217;s worth your time to investigate a particular section of a paper and all that stuff. But if you&#8217;re willing to spend a few months just, at night I look at a paper, or on my lunch break I look at a paper, if you&#8217;re willing to spend a few months just doing that, you will get your bearings and you will start to know where the good papers are in your field.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6685">01:51:25</a>] You&#8217;ll know how to find them, you&#8217;ll know where they tend to be published. You&#8217;ll be able to read them much more effectively, you&#8217;ll be able to know how to spend your time on them. And I think in all but probably a few small fields that probably exist somewhere that maybe people don&#8217;t tend to write many papers in or something like that, it&#8217;s tremendously valuable, and I think it&#8217;s so valuable that I read papers in disciplines I don&#8217;t even do.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6710">01:51:50</a>] And I find it extremely rewarding. I read security research papers all the time. I don&#8217;t even work in a field where there are security research implications. Games don&#8217;t really do much of that. They tend to be run very sandboxed, and they don&#8217;t. Sometimes there are. But they&#8217;re not that kind of thing. They&#8217;re not like a web server authentication protocol or something like this. And I just find it incredibly rewarding because there&#8217;s just so much good stuff out there that you can learn.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6738">01:52:18</a>] So I would say that would be it. Instead of a book rec, I would say find a paper, try reading a paper.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6744">01:52:24</a>] How do I go and find that first paper or someplace to get started?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6749">01:52:29</a>] That will be valuable for most people who are watching because they might be generalists or working in web dev or in just a general tech org at a big tech company or something like that. I would say pull up the proceedings of USENIX, look through it for a paper that sounds interesting to you and try reading that paper. Or oftentimes there&#8217;s awards. I think USENIX has them where it&#8217;s like best paper of the conference.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6780">01:53:00</a>] Try reading the best paper of the conference, see what you think, because that&#8217;s a collection of papers that&#8217;s usually about operating system stuff and systems design stuff. So it&#8217;s going to be something that most people can relate to. It&#8217;s not going to be really esoteric. If you were to open up the proceedings of SIGGRAPH, for example, it&#8217;d be like, oh, neural networks for cloth simulation or something, and you&#8217;re like, okay, this is not.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6808">01:53:28</a>] I can&#8217;t relate to this because I don&#8217;t do graphics or whatever. Right. So USENIX might be a good place to start. But in general, yeah, if you had another option would be to go to scholar.google.com, which is their search that just specializes in papers, and type in a topic description that you find interesting. So if you wanted to learn, maybe you were very interested in consensus algorithm Paxos or something, and I don&#8217;t know what that is, or it came up at work and I haven&#8217;t really ever looked at it.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6842">01:54:02</a>] You can just type that in, Paxos, and they&#8217;ll just be a list of papers, and they&#8217;ll say a thing like cited by. You can see how many citations they have. That&#8217;s usually how influential that paper was, to a certain extent. So those are some ways you could get started and find something that might interest you. It won&#8217;t be for everyone. You may bounce off it, but that&#8217;d be my recommendation. It&#8217;s something to try. You might find it interesting.</p><h3>01:54:27 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6867">01:54:27</a>] And then, yeah, last question for you is if you could go back to the beginning of your career, when you just entered the industry, and give yourself some advice, what would you say?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6878">01:54:38</a>] I think the advice I would have given to myself was to get into low-level programming earlier. I didn&#8217;t really learn how to properly analyze and even really write assembly language code until probably 2010 or 2015 or something. I mean, it&#8217;s recent because I&#8217;m pretty old, and it&#8217;s like the last 10 years or so. Right. And I feel like I always wanted to know how to do it and just never seemed to.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6921">01:55:21</a>] And the advice I would have given to myself was, just go down the hall. What I was interested in, go down the hall and be like, grab somebody. Show me. How the heck do you write this? Just show it to me. I think one of the problems is a lot of people, especially when they&#8217;re young, they are afraid of appearing like they don&#8217;t know things. And I think that can be a real impediment to learning.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6948">01:55:48</a>] I&#8217;m sure it was for me. And I probably just didn&#8217;t want to literally say, &#8220;I don&#8217;t understand any of this stuff. Can you explain it to me?&#8221;</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6955">01:55:55</a>] To me?</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6958">01:55:58</a>] And I think that&#8217;s one of the most useful things you can do. I don&#8217;t understand this. Please explain it to me. It&#8217;s very useful, and most people will be very happy to do that, if they&#8217;re not a dick. Most people will be like, oh, sure. Now, they might not be good at explaining it. There are plenty of engineers who are really good at something and suck at telling you how they do what they do. So you won&#8217;t always get a great explanation, but sometimes you will.</p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=6986">01:56:26</a>] Sometimes you&#8217;ll find the type of person who can explain very clearly how it is they do what they do. So if you just keep asking, you&#8217;ll get the good explanation. Nowadays, I wouldn&#8217;t really need to give myself exactly that advice because the Internet has so much great information on it. I could have taught myself that did not exist at the time.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=7006">01:56:46</a>] Awesome. Well, thank you so much for your time, Casey. I really appreciate it.</p><p><strong>Casey:</strong></p><p>[<a href="https://youtu.be/jHLbL1Eg4gM?t=7009">01:56:49</a>] Thank you so much for having me. I love the show. It was an honor to be invited on, so thank you very much.</p>]]></content:encoded></item><item><title><![CDATA[How Anthropic Builds And How Engineering Will Change Soon | Thariq Shihipar]]></title><description><![CDATA[When I talk with my friends who work at Anthropic it always surprises me how far ahead they are in adopting AI within their processes and workflows.]]></description><link>https://www.developing.dev/p/how-anthropic-builds-and-how-engineering</link><guid isPermaLink="false">https://www.developing.dev/p/how-anthropic-builds-and-how-engineering</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 07 Sep 2026 13:17:19 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214514209/8354e12b84e1ff9eba08e771121af1c4.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>When I talk with my friends who work at Anthropic it always surprises me how far ahead they are in adopting AI within their processes and workflows. Generally what engineers in the industry start doing, these top AI labs have already been doing for a while. </p><p>In this conversation, I asked <a href="https://www.linkedin.com/in/thariqshihipar/">Thariq Shihipar</a> who is an engineer on Anthropic&#8217;s Claude Code team how Anthropic makes the most out of the models for engineering and how the industry will change soon.</p><p>For instance, we discussed:</p><ul><li><p>How Anthropic maintains their code given so much more code is landing now</p></li><li><p>What has worked in preventing AI-written breakages</p></li><li><p>How internal software engineering practices have changed (e.g. onboarding, glue work, etc)</p></li><li><p>How engineers within the company typically leverage the models</p></li><li><p>What percent of Anthropic&#8217;s work is fully autonomous</p></li></ul><p>Hope the conversation is helpful for giving you some insight in how the models are and will change software engineering more in the near future.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/2Kch3tWMnw8">YouTube</a>, <a href="https://open.spotify.com/episode/5dT5nBPv2JXPKz14M5zlE1?si=sQsWgOiZS1SzsEjWRJ0ruA">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-2Kch3tWMnw8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;2Kch3tWMnw8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/2Kch3tWMnw8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/214514209/0029-onboarding-at-anthropic">00:29 - Onboarding at Anthropic</a></p><p><a href="https://www.developing.dev/i/214514209/0253-internal-capabilities-vs-external-perception">02:53 - Internal capabilities vs external perception</a></p><p><a href="https://www.developing.dev/i/214514209/0616-model-vs-harness">06:16 - Model vs Harness</a></p><p><a href="https://www.developing.dev/i/214514209/0855-what-percent-of-anthropics-changes-are-fully-autonomous">08:55 - What percent of Anthropics changes are fully autonomous</a></p><p><a href="https://www.developing.dev/i/214514209/1451-computer-use">14:51 - Computer use</a></p><p><a href="https://www.developing.dev/i/214514209/1742-how-to-make-the-most-out-of-your-compute">17:42 - How to make the most out of your compute</a></p><p><a href="https://www.developing.dev/i/214514209/2045-loop-engineering">20:45 - Loop engineering</a></p><p><a href="https://www.developing.dev/i/214514209/2247-where-the-industry-will-go-soon">22:47 - Where the industry will go soon</a></p><p><a href="https://www.developing.dev/i/214514209/2602-which-model-do-anthropic-engineers-use">26:02 - Which model do Anthropic engineers use</a></p><p><a href="https://www.developing.dev/i/214514209/2738-is-learning-a-particular-model-worth-it">27:38 - Is learning a particular model worth it</a></p><p><a href="https://www.developing.dev/i/214514209/3056-prompting-tips-for-todays-models">30:56 - Prompting tips for todays models</a></p><p><a href="https://www.developing.dev/i/214514209/3504-how-to-get-the-models-to-do-tasteful-work">35:04 - How to get the models to do tasteful work</a></p><p><a href="https://www.developing.dev/i/214514209/3900-how-much-of-writing-is-done-by-ai-at-anthropic">39:00 - How much of writing is done by AI at Anthropic</a></p><p><a href="https://www.developing.dev/i/214514209/4536-code-ownership-and-maintenance-at-anthropic">45:36 - Code ownership and maintenance at Anthropic</a></p><p><a href="https://www.developing.dev/i/214514209/5204-how-anthropic-prevents-breakages">52:04 - How Anthropic prevents breakages</a></p><p><a href="https://www.developing.dev/i/214514209/5524-visibility-and-sharing-your-work">55:24 - Visibility and sharing your work</a></p><p><a href="https://www.developing.dev/i/214514209/5842-luck-surface-area-example">58:42 - Luck surface area example</a></p><p><a href="https://www.developing.dev/i/214514209/010057-should-people-still-learn-to-code">01:00:57 - Should people still learn to code</a></p><p><a href="https://www.developing.dev/i/214514209/010742-advice-for-his-younger-self">01:07:42 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:29 &#8212; Onboarding at Anthropic</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=29">00:29</a>] My goal with this conversation is to ask you as much as I can about how you can get the most out of the models specifically for software engineering, so that people in the industry can learn from the best practices where <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>&#8216;s having success. And so to start the conversation off, I&#8217;d like to ask you, let&#8217;s say I was coming from a company that&#8217;s maybe less AI-pilled, and I was onboarding onto your team.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=61">01:01</a>] What are the most impactful things that you tell me to start doing to successfully onboard?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=68">01:08</a>] I think the number one tip we have for both people inside and outside <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> is that if you treat <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> like a thought partner and give it the context that you need, then you can usually figure out the next steps. And so I think that starting with, okay, can <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> do it? If not, why not? And using that as the starting place, I think is really important. And I think just always thinking about, okay, reflecting on, can you move up a higher abstraction level?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=102">01:42</a>] Can you set up a system? Can you build the system that builds a system instead of just building the product itself?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=112">01:52</a>] I remember when we used to onboard people, you would get assigned an onboarding buddy, someone who really knows the code base. And you can ask all of your trivial setup questions too. It sounds like <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> fills in a lot of those gaps. And in practice, is it 100% you don&#8217;t need an onboarding buddy and you can just ask the model for everything?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=134">02:14</a>] You do need an onboarding buddy, but more from a social and cultural perspective than a technical perspective. I think, from a technical perspective, you can basically just work with <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> and, if you&#8217;re good at it, you can get what you need to onboard. But I think just having the context of how to work with the team, and also how to get buy-in on what you&#8217;re building, or understand how the team works together, or just even have a friend to work with, is really important.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=166">02:46</a>] So, yeah, we still have onboarding buddies, but it&#8217;s not nearly as much technical lift as it used to be.</p><h3>02:53 &#8212; Internal capabilities vs external perception</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=173">02:53</a>] There&#8217;s a big difference between the external perception of the AI capabilities and the internal <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> usage. Because I&#8217;ll talk to friends and say, yeah, what Boris is saying is the reality. And so my question is, why is there such a big difference between that internal perception and that external perception?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=195">03:15</a>] I think we have a good track record when we make these claims. You&#8217;re like, oh, it turns out to be true on a large enough timescale. I think a lot of engineers are not really in the IDE anymore. The code being written is like, I talked to large enterprise customers where they haven&#8217;t typed a line of code themselves in six months. We think of it part of our jobs to figure out how to work at a higher abstraction level.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=225">03:45</a>] And if I spend a day trying to figure out how to get <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to do this autonomously and I fail, that&#8217;s fine, maybe even good. Because now I can be like, oh, hey, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> is not good at this. How do we make it better? But I think if you&#8217;re working an average job, your job is to produce the output. And I think automation always has a cost because it&#8217;s an investment. And if it works, then great, you&#8217;ll make a return on your investment over time.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=254">04:14</a>] But if it doesn&#8217;t, now you&#8217;ve wasted all this time. And so I think the nice thing about the models is that generally the chance of the investment working out is higher and higher because the models are getting smarter and smarter. But you still have to take that leap as an individual. Your boss is not going to be super happy with you if you&#8217;re like, you know, if you&#8217;ve been working on your harness setup the entire time and not shipping code.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=280">04:40</a>] But yeah, I think it&#8217;s mostly just culture and thinking of it as our job to sort of live in the future. And then we try and build it into the product and our harness as well so that you don&#8217;t have to do as much manual step. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=293">04:53</a>] Are there examples that come to mind when you think of things you don&#8217;t see people doing at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> that are really making a big difference and you&#8217;d kind of recommend?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=306">05:06</a>] Using <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> for knowledge work I think is really valuable. The way a technical person does knowledge work, I think, is very different than the way a non-technical person does knowledge work these days because you can get the models to do it. And so I think increasingly, if you can figure out, like, hey, this is a task, the task is made up of code-like things, how do I tell the model the code steps to take in order to do it right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=336">05:36</a>] You can do a lot more, right. So I think I do a lot of accounting with <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> on the personal side, using scripts and Python instead of <a href="https://en.wikipedia.org/wiki/Microsoft_Excel">Excel</a>. I do video editing using FFmpeg and these libraries to render visuals and things like that. And so I think most knowledge work is reducible to code and coding agents, if you think about it well.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=367">06:07</a>] And I think that that&#8217;s one big difference that I see, other people doing less of.</p><h3>06:16 &#8212; Model vs Harness</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=376">06:16</a>] I think we talked a little bit about the model and the harness, and I want to ask you, what&#8217;s the relationship between the model and the harness in terms of getting the best end results? Obviously the model is important, so how important is the harness, and are there any examples you could share?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=396">06:36</a>] Yeah, I mean, the harness is super important. I think that sometimes there&#8217;s this idea that the harness doesn&#8217;t matter because the models will get better and better. And if the model just does everything perfectly, then why do you need a harness at all? I think that in practice, what we see is the models get better and better, and so the harness needs to become more and more complicated to allow the model to do more things.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=424">07:04</a>] An example of this is Auto mode. And so Auto mode is a classifier that runs after every task that you normally have to ask a permission prompt for <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>. And back when we were on <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 4</a> or even <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 4.5</a>, maybe it was not so bad to hit Enter on the permission prompts because the turns would only last a few minutes anyway. Now <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can run for hours, and so you really need that ability for it to do work safely.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=458">07:38</a>] Stick to your instructions. And Auto mode is really complicated software. Sandboxing is really complicated software. And then you also think about things like, okay, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can do so much more work. Now, how does it represent the work that it&#8217;s done? It&#8217;s worked for eight hours and you want to know what&#8217;s done in eight hours. This is what we use Artifacts for. And Artifacts themselves are a form of prompting, because how does <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> represent that in a useful way?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=486">08:06</a>] There&#8217;s a lot of different ways that it could represent that information. And so I think roughly what we see is the harness. Harness engineering is definitely this mix of science and art. I think it&#8217;s very unintuitive in a lot of different ways, but I think it has just big abilities to unlock new parts of model behavior. And I just see the harnesses get more and more complicated, and it&#8217;s kind of harder and harder to actually vibe code to your own, which is a little bit unintuitive to me because of how good the models have gotten.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=525">08:45</a>] But yeah, things like Auto mode and workflows and things like that, which are actually quite complicated pieces of software, become really load-bearing.</p><h3>08:55 &#8212; What percent of Anthropics changes are fully autonomous</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=535">08:55</a>] I think as <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> have advanced, more and more of the human part can be taken out of the loop. At <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, what percent of your or your team&#8217;s changes are fully autonomously made versus actually pairing with a model? People were doing more of that around a year ago.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=558">09:18</a>] I think it&#8217;s very dependent on what you consider autonomously made or what team or what function you&#8217;re working on as well. For example, a designer might give you a Figma file, and then you pass the Figma file to <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>. <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is great at using the Figma MCP, but the designers put a lot of work in there. I think roughly our goal is to sort of get <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> being able to do all of the glue work that ties everything together, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=595">09:55</a>] So, okay, you&#8217;ve got a Figma file. Do I really need to rewrite that in React or something? Probably not.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=600">10:00</a>] I think that work has been done once, right? And so the way I think about it is, what&#8217;s the unique work that I need to do every day, right? And the more I&#8217;m doing that unique work, the better. And I think there&#8217;s just a ton of demand for unique work and thinking. And the more I&#8217;m doing something I&#8217;ve done before, I&#8217;m like, okay, can <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> do this?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=624">10:24</a>] Or if I&#8217;m just translating what someone else has done, I&#8217;m like, oh, can <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> do this as well? It&#8217;s really blurring the lines of what does it mean? Even if <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> has fully generated his PR, you&#8217;ve probably made a lot of decisions and added a lot of context along the way.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=642">10:42</a>] How close are we to a world where someone just says, &#8220;Hey, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>, here&#8217;s the ticket. Just don&#8217;t tell me until you&#8217;re done&#8221;?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=650">10:50</a>] Yeah, so I mean, it depends on how good the ticket is. If someone has perfectly written out essentially a software spec for this ticket, then yeah, actually <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can probably do that, right? But I think that to start, let&#8217;s say that we get a GitHub issue. Is this worth fixing? Oftentimes issues are also feature requests or something, or there might be multiple things happening that maybe you combine together in a different way.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=678">11:18</a>] And so I think that a lot of that work is, okay, what&#8217;s the vision for the product? Where do we want to go?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=687">11:27</a>] How do we make sure what we&#8217;re building is cohesive? And so I think for a specific spec, a specific enough spec, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can do it. In practice, people haven&#8217;t figured out what they really want and haven&#8217;t figured out the unknowns or the shape of the problem and things like that. And maybe the way you would have done it before is you start writing the code and you&#8217;re like, oh, what am I supposed to do now?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=715">11:55</a>] Or you figure it out that way. Now I think you need new ways of figuring out what you don&#8217;t know yet. And I think you can still chat to <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>, really. But I think that&#8217;s the hard part. And I think definitely a failure mode is like, you tag <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>, tag. And you&#8217;re like, hey, please do this one sentence or less description, no previous context or memory. And then it does it. And you&#8217;re like, oh, no, I don&#8217;t like it.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=742">12:22</a>] And then you&#8217;re like, just iterating forever on that, right? Versus figuring out what you actually want, really quickly.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=750">12:30</a>] When I was working at Meta, the ambiguity of the task really ranges, oftentimes proportional to people&#8217;s&#8212;I mean, levels aren&#8217;t everything in software engineering, but roughly, senior engineers kind of take that ambiguous business need and convert it into something that&#8217;s a lot more concrete. And then people who are maybe newer in their careers kind of do the implementation work, and intern projects were kind of almost line by line specified.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=782">13:02</a>] If I was still working at Meta and I had an intern, I was like, just kind of give the intern back. Give <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> my intern&#8217;s work back, and it would kind of do it. What kind of work does an intern do then at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, if <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can kind of handle it? What we see in practice is that there are just so many new types of work that no one has ever done before, right?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=835">13:55</a>] And I think that we all have to figure that out. And so, for example, how do you do an eval against these new coding behaviors, right? How do you make sure that <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>&#8216;s performance across millions and millions of users doing all sorts of tasks? Sometimes these tasks might have tradeoffs, you know. So I think there&#8217;s a lot of new work to be done, especially on the research side of things.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=799">13:19</a>] And I think that it&#8217;s harder and harder to write those specs, but there&#8217;s also less and less experience that is relevant, you know what I mean? And so I think that, yes, having some experience means that you can sort of, you know, you know how to get work done and you know how to learn. But I think you also get those skills out of college. And now I think as an intern trying to figure out, okay, what are the things that people have not done before?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=865">14:25</a>] And doing that, I think, is really exciting. And I think that there&#8217;s, yeah, I think internships are less &#8220;write the React code,&#8221; you know what I mean? And there&#8217;s so many new problems to solve. We need people to solve them. It&#8217;s helpful if you have a fresh set of eyes on it. But it is definitely more proactive and opportunistic, maybe than before, where it was maybe more of like a pipeline.</p><h3>14:51 &#8212; Computer use</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=891">14:51</a>] We talked a little bit about knowledge work, and that makes me think of computer use and browser use. And I wanted to ask you, how far away is that from being widely adopted in the industry and impactful? Maybe you can talk about how <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> uses it, since it feels like you&#8217;re always far ahead.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=911">15:11</a>] I think computer and browser use, the models have gotten a lot better. Opus 5, I think, is a really good computer use model. But there are these weird edge cases where, for example, it can&#8217;t type a password on my behalf because my passwords are in <a href="https://en.wikipedia.org/wiki/1Password">1Password</a> or something, and it can&#8217;t access that. And so it gets stuck. I think there are still those edge cases, which are kind of UX edge cases. And then I think also, obviously, computer use is also the more and more people turn APIs and MCPs into new ways to use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>. I think that can also take a lot of the use cases that you&#8217;re using computers for.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=951">15:51</a>] And so going back to, oh yeah, knowledge work, everything is code. I think there are some things where there&#8217;s no <a href="https://en.wikipedia.org/wiki/API">API</a>, there is no way to execute in code other than just opening up your browser. But I think increasingly there are more and more ways. And I think, yeah, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is a great example of we&#8217;ve just sort of tried to roll up all these things into APIs that it can use. And so it feels like it can do a lot of work on your behalf, even though it&#8217;s not literally running a virtual computer and clicking things.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=984">16:24</a>] I have noticed that whenever I use any kind of computer use tooling, it feels painfully slow. I look at the <a href="https://en.wikipedia.org/wiki/Cursor_(company)">cursor</a>, and it&#8217;s sitting there for 10 seconds, moves over, goes there for 10 seconds. Where is all that latency coming from? Do you have a sense?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1001">16:41</a>] I think it&#8217;s hard to make a small model that&#8217;s really, really good at it. I think just because there&#8217;s a lot of knowledge that you need to have about this task and how these things work together and stuff. And so you want a smart model, and smart models take a little bit longer time. I think a lot of people thought we&#8217;d get here sooner, but I think it&#8217;s just been a harder task than we expected.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1023">17:03</a>] I think one thing someone&#8217;s told me about computers before is that it&#8217;s a state machine where you don&#8217;t control the entire state. If you are on the <a href="https://en.wikipedia.org/wiki/DoorDash">DoorDash</a> website or something, you want to add something to the cart, and you add it incorrectly. Now you have this new flow to undo it. Now you need to go click, and now you need to go delete it, and you can make a mistake along that side as well.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1045">17:25</a>] Whereas in code, you can sort of undo, Git, whatever. You control all of the state. But for computer use, each action is, if not irreversible, it&#8217;s much harder to reverse than others.</p><h3>17:42 &#8212; How to make the most out of your compute</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1062">17:42</a>] At <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, I get the sense that employees have a lot of budget in terms of the compute to speed up whatever it is they need to do. And so if we were to spur the imagination of people who use the models to make them more productive, assuming they had infinite compute, what type of workflows would you start telling someone to do if they had infinite compute?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1090">18:10</a>] I think this difference is slightly more exaggerated than you think. I think that, for example, I use my <a href="https://en.wikipedia.org/wiki/List_of_HBO_Max_original_programming">Max subscription</a> on the weekends, and I have almost never hit a five-hour limit. I think the models are really smart, and I think that a lot of times when we&#8217;re spending a lot of compute, we&#8217;re just trying to find capabilities or we&#8217;re trying a bunch of different things, and it&#8217;s more about us figuring out model possibilities than getting a lot of work done.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1123">18:43</a>] When we&#8217;re testing for math, for example, we&#8217;re trying to understand how smart the model is. And that is useful to us, that&#8217;s useful output to work. And it can also sometimes solve the Riemann Hypothesis or something, or not make progress, you know what I mean? People can replicate what we do at home just by thinking at a higher level of abstraction. For example, I&#8217;m trying to get <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to draft feedback for you.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1147">19:07</a>] And so I want the funnel for, like, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> has drafted some feedback to the user, has submitted feedback, to be really good, right? And so I monitor that funnel. And then I asked <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>. I had some ideas, but then I was also like, oh, what if I ask <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to try and improve the funnel? And it&#8217;d be like, here are some ideas, can you come up with some as well? Let&#8217;s figure it out. All of those things you can kind of do yourself right now.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1173">19:33</a>] If you&#8217;re like, okay, when I&#8217;m making a feature, let me annotate it with events. Let me make sure that <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> has access to those events. Let me run a loop in the morning every day to check what events have fired and what changes were made. Maybe even let me proactively suggest some ideas. This can all happen, I think, within a fairly reasonable amount of compute, actually. What I see more often is people running into limits where they&#8217;ve actually done the opposite.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1204">20:04</a>] They&#8217;ve started with a small mid-scope task. It&#8217;s like, oh, hey, refactor this function in this way, and then <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> does it, and maybe it has some follow-on effects because refactoring this means you have to do some other work too. And that&#8217;s not exactly correct, or you&#8217;re iterating there, and it&#8217;s like you&#8217;ve just spent a lot of time. Whereas if you had sort of stepped up a level, told <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> your goals, then figure out, okay, what are the details, do some exploration, maybe write out the schema, and then work with it, then let it run.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1241">20:41</a>] You could probably get the same output.</p><h3>20:45 &#8212; Loop engineering</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1245">20:45</a>] I remember it was going kind of viral, this idea of creating loops. And I mean that does feel like a pretty expensive sort of thing to set up if you&#8217;re just asking it to keep hammering away. Maybe first could you define this loop engineering thing, and then I&#8217;m curious, how often do you use it, and would you hit limits if you were on a <a href="https://en.wikipedia.org/wiki/List_of_HBO_Max_original_programming">Max subscription</a> doing that kind of stuff?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1270">21:10</a>] Yeah. Okay, so loop engineering, roughly, it&#8217;s like instead of prompting <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> directly, you&#8217;re setting up a system that prompts <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>. You don&#8217;t have to think of it, if you use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>, a lot of this comes naturally. You just ask it, like, hey, every day do this thing. And that will, that&#8217;s a loop. You can definitely do that sort of work right now. But I think if you&#8217;re trying to set up, let&#8217;s say, like 10 loops or something that are triaging your feedback and implementing and things like that, we do that sort of work, but we also spend a lot of time making sure our skills and things like that are useful and are good at triaging, that we can have, we&#8217;re good at verification and so that we can, like, make sure that the changes land, you know, and then I think at that level, you know, what it means is like, now we have someone monitoring issues that we just could never have kept on top of before, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1336">22:16</a>] And it just increases our software development velocity, and it&#8217;s worth a lot of value to us. And so, yeah, I think that, if you&#8217;re kind of at the scale of like, okay, you want it to essentially be autonomously running, doing a software engineering job for certain cases, I think it can do that if you set up the verification well, if you set up your skills well, if you give it the right data sources.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1358">22:38</a>] But it is a lot of work, right? And I think that you have to make sure that that work is valuable to you in the same way that if you hire someone, you have to make sure that work is valuable.</p><h3>22:47 &#8212; Where the industry will go soon</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1367">22:47</a>] We&#8217;ve talked a little bit about <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> being kind of just ahead because that&#8217;s the job of the company, on adopting AI. Is there something that you feel is kind of the next big shift that maybe <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>&#8216;s already felt that you feel is going to diffuse into the industry, and if so, what might that be?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1386">23:06</a>] Yeah, I think there are a few different ones. Obviously, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is the way we do a lot of our work right now. I think the way I think about it is that <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is really good at the implementation of code, and <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is for the rest of the software development lifecycle, so getting feedback, doing code review, babysitting, building CI/CD incidents, all of these other things. <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is really good at at a high level.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1418">23:38</a>] Turning every part of your software development lifecycle into kind of like a routine or a loop using <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> or something like that, I think is probably where things are headed. I think generative interfaces are still to come, and I think that Artifacts and, yeah, basically Artifacts, I think, are going to be an increasingly large way of how you interact and read with <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1446">24:06</a>] And so I think that I use Artifacts for almost everything. And so I think that we&#8217;re going to see that also become big. And I think that also ties into <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>. Right? So, like, let&#8217;s say that you are on your phone and <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> has done a bunch of work for you. It creates a report. The report is readable on your phone because it&#8217;s made the, like, the Artifact look really great.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1468">24:28</a>] Could you explain artifacts a little bit, or kind of what are they and how do people use them?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1474">24:34</a>] Artifacts are basically <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> can upload, essentially a web app for you to use. And I think that you can use it in a really wide range of capabilities. They can now, for example, call your MCP. So you can make an artifact, for example, to read your inbox and display it or sort it or tag it in one way. You can also use an artifact to show you a plan for coding. And that might show diagrams and file snippets and code and schemas.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1509">25:09</a>] And so it shows you sort of, it&#8217;s like an interactive display for the job you&#8217;re doing right now. And as <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> does more and more jobs, we&#8217;re realizing that just text in, text out is probably not useful for everything, right? And <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> is better and better at creating essentially the exact interface for you at the right time. I think that there&#8217;s a lot more to do there. And I think it&#8217;s another one of those things where you have to think about, like, oh, could I be interacting with <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> in a different way?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1544">25:44</a>] Could I interact with it through an artifact? Could I use it to learn more or to understand it better or stay in the loop or improve some of my own knowledge work or something? Yeah, I think it&#8217;s pretty exciting. We&#8217;re still kind of early to it.</p><h3>26:02 &#8212; Which model do Anthropic engineers use</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1562">26:02</a>] Does the daily driver model vary among engineers, or do people typically just pick the most intelligent one that&#8217;s available at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1571">26:11</a>] I do think similarly with models. If you use the smart models and you use them well and you give them tasks where they are well shaped and you&#8217;ve spent some time setting up a good verification harness and things like that, you can get a lot out of them. I don&#8217;t think it&#8217;s like choosing which model at the right time. I think it&#8217;s like, oh, how do you get the most out of the frontier models? Because I think technology works the way it does, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1597">26:37</a>] Everything gets more abundant, more available. And so I think we&#8217;re in this current weird spot where we don&#8217;t quite have enough compute for everyone to have <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> at 100% of the rate limits. But I don&#8217;t think this is a durable skill to build, figuring out, like, oh, when do you use <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> and when do you use <a href="https://en.wikipedia.org/wiki/Sonnet">Sonnet</a>? I think it&#8217;s unintuitive there. But I think in practice, if you&#8217;re today, what I would do is I use <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> for planning, for brainstorming, for finding unknowns, for coming up with a detailed spec, and I&#8217;d use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 5</a> to implement it, and I would probably use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 5</a> with workflows using a verification aid like Schema or Harness that&#8217;s built with <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1641">27:21</a>] So I&#8217;d use <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> for those high-leverage tasks and, yeah, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 5</a> for execution and implementation. But yeah, I think increasingly, probably next year, I think you&#8217;re just not going to be thinking that much about which model.</p><h3>27:38 &#8212; Is learning a particular model worth it</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1658">27:38</a>] You mentioned that skill of prompting, and I&#8217;ve heard some people say it&#8217;s not too durable of a skill because it&#8217;s really specific to a model. These models, they almost have their own unique, I guess, spiky intelligence. So if you knew everything that was very specific to, let&#8217;s say, today&#8217;s <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>, and you&#8217;re a master of today&#8217;s <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>, maybe you don&#8217;t need to know any of that stuff a year from now.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1688">28:08</a>] And <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> 3 or 4 or 5 or whatever comes out. Let&#8217;s say you were talking to a software engineer who&#8217;s looking for career advice, and they&#8217;re thinking, hey, should I really become a master of engineering?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1703">28:23</a>] My prompt, I do want to say prompting is a little bit more than just a prompt you put in. It&#8217;s also, you know, you might have made a skill or you might have added some data or something. It&#8217;s not just a prompt you write, but it&#8217;s like everything you&#8217;ve done before that builds up into your context. Sometimes people see us write small prompts and they&#8217;re like, oh, what does that mean?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1724">28:44</a>] But we just spent so much time on the harness and the verification and the skills. So I think it will be really valuable to just get better at prompting. I think, to what you&#8217;re saying about each model is different, you&#8217;re right. I think each model is kind of its own kind of almost organic digital thing. And so there are quirks you have to learn, and you do have to unlearn them. We recently wrote about how we removed 80% of the system prompt from <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1758">29:18</a>] One of the learnings we had was we needed to remove examples from the tool descriptions. This used to be the only way you could get good output from the models was through tools or through examples. You&#8217;d have to be like, hey, this is the right tool. Use this here. Here&#8217;s an example of writing a file well, and here&#8217;s an example of not doing it well. Now, we found that examples are mostly negative, I think, unless you really see <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> doing something you don&#8217;t like because it&#8217;s just quite imaginative.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1792">29:52</a>] It&#8217;s good at sticking to your intention and working with you. And so we removed a lot of examples. So in that case, yes, you do have to sort of adapt, but the skill you&#8217;re building is the skill to adapt. You learned how to use <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>, and now you know a lot about <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>, but you also know how to learn. You learned how to work with a model, and so <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Opus 4.5</a> comes out. You need to learn how to use it again as well, but you&#8217;ll be much faster because you&#8217;re better at learning.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1822">30:22</a>] You&#8217;re better at working with <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> 5. And I think, for me, I think the first model I worked with was GPT-2. And I remember it was so hard to get a JSON output out of GPT-2. If you could just get it to choose one of the categories that you gave it, that would be incredible. And so I think that building that skill&#8212;GPT-2 is such a different model than <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> 5&#8212;but I feel like the time I spent doing that has helped me be better at prompting <a href="https://en.wikipedia.org/wiki/Fable">Fable</a> 5.</p><h3>30:56 &#8212; Prompting tips for todays models</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1856">30:56</a>] Is there any tribal knowledge or quirky tips in today&#8217;s models where you&#8217;d say someone should know that to get more when they prompt? I actually need to counter one.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1863">31:03</a>] I think that people have been saying where it&#8217;s like, oh, just believe in yourself or something. I know that Jared&#8217;s post about the Riemann Hypothesis. He was just like, keep going, just believe in yourself. I think in this case it was mostly just Jared saying, it&#8217;s okay to use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> to solve this problem, and I&#8217;m giving you permission to do it. And I think that that&#8217;s not exactly the same as I believe in you.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1893">31:33</a>] You know what I mean? It&#8217;s really just letting the model use compute. So I think that is something I like to tell the models right now is like, okay, hey, I think this is a hard problem. Use subagents, use workflows if you need it, right? So I always tell it to use its own judgment, but I&#8217;m giving you permission to do this stuff. I think that you have to remember that the models by default do what maybe the average user wants, which is they want it to respond and start doing work as fast as possible.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1931">32:11</a>] Roughly complete the task but not spend a crazy amount of compute on it. I think that you have to sort of, if you want the model to do it differently, you have to nudge it slightly, right? So you might have to be like, okay, hey, I don&#8217;t want you to do any work yet. I want you to brainstorm. I want you to think with me, right? And if you prompt it that way, it will start doing that if you want it to spend a lot of compute.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1956">32:36</a>] If you&#8217;re like, hey. Sometimes I&#8217;ll say, yeah, hey, I think this is a hard problem. Feel free to use workflows. If I&#8217;m running overnight, I might just be like, hey, I&#8217;m going to sleep. Set a /goal or something and then let it run. So yeah, I think there is just giving it permission to do the thing you want.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1977">32:57</a>] When you recently removed so much of the system prompt, how did you prove that the end result was better?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=1986">33:06</a>] We have a bunch of user metrics, just like how much do people like the output of <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>? Sometimes you get that survey and you see it. We run evals against our internal eval and external evals to see how it performs at these different tasks. But I think it is hard sometimes. You don&#8217;t realize that if <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> is telling the user, if it does all its work and then it&#8217;s like, hey, maybe you should go to sleep.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2015">33:35</a>] There&#8217;s no eval for <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> telling you to go to sleep. You know what I mean? And now we have to catch this new behavior. So it is hard. I think we spent a lot of time basically just removing lines in system prompt, running evals, seeing how it works, seeing how people reported it internally, and then adjusting. But it was a full-time job for several people over long periods of time. And so I don&#8217;t think I necessarily recommend everyone do this.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2043">34:03</a>] I think that&#8217;s kind of why we wrote that post about what we learned from adjusting the system prompt. And we think that&#8217;s pretty general. So hopefully you don&#8217;t have to now go through this crazy iteration process.</p><h3>35:04 &#8212; How to get the models to do tasteful work</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2104">35:04</a>] I think a lot of people, when they use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>, they get excellent results when it&#8217;s kind of getting very objective work done.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2113">35:13</a>] But when it&#8217;s prompting models to do beautiful or tasteful work, it&#8217;s not always. It&#8217;s pretty much more hit and miss. And so, how do you best instill a very particular style you&#8217;re going for, or particular taste in the model&#8217;s outputs when it&#8217;s a much more subjective domain, maybe like front end?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2135">35:35</a>] The way you do it is sort of like you stay in the loop. I think you give it references, and I think the more references you give it with data, the better. So it&#8217;s better to give an <a href="https://en.wikipedia.org/wiki/HTML">HTML</a> file than a screenshot. It&#8217;s better to give a Figma file than a raster image or something. Because now if <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> wants to know the border radius of this thing, it&#8217;s like, oh, what&#8217;s the Figma component? Border radius?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2160">36:00</a>] And let me just copy it over. I think giving it a bunch of references, ideally in code, is a really good way of getting it to stick to this. I think if not, you can then ask it to. If you don&#8217;t have, let&#8217;s say you&#8217;re not a designer, there&#8217;s probably step one is being like, okay, I&#8217;m not a designer. There&#8217;s a lot I don&#8217;t know about design, and there&#8217;s a lot I don&#8217;t know about iteration. I don&#8217;t even know what good looks like.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2189">36:29</a>] And I think this is part of the art of working with the designers. They will just know what good looks like, and they&#8217;ll be like, this isn&#8217;t good enough in this way. And they prompt it. And that, I think, is also a skill that will keep getting valuable and even more valuable over time: what is good output, what is worth doing? I think, for example, with the Riemann Hypothesis, Jared prompted it, but he had no idea if it was correct.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2216">36:56</a>] Until Lev, who is one of the world&#8217;s best mathematicians, was like, okay, is this correct? And he worked with it, and he asked it tons of follow-up questions, and we could not follow that at all. We had no idea what he was saying, but he was really intrigued. I think more and more, being that high-taste user and knowing a lot about a problem in a domain space is how you get good outputs. That&#8217;s how you solve these problems.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2247">37:27</a>] Otherwise, maybe <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> did solve physics or something, but you just wouldn&#8217;t know it. You don&#8217;t know enough about it. And so I think when you&#8217;re talking about design, the first thing you do is how do you become more tasteful with design? And so you can ask <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> that as well. You can be like, hey, I&#8217;m not a designer, I want to be better at design. I don&#8217;t even have the language first. Maybe let&#8217;s find some reference sites.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2270">37:50</a>] And then maybe you pull some reference sites, and then you&#8217;re like, this is what I like, or this is what I don&#8217;t like. And then you sort of build up those references, and you give <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>, and now maybe you tell it, hey, let&#8217;s do some exploration. This is sort of my taste. And then it&#8217;ll do a few different mockups. I like to do these mockups all in <a href="https://en.wikipedia.org/wiki/HTML">HTML</a> because it&#8217;s all self-contained, easy to edit.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2293">38:13</a>] And then once you have that reference, now it&#8217;s a reference. So now you can make a new session and be like, hey, this is a mockup of a design. I want to start implementing this, and it will have all of it in code, and you can start getting there. But I think that the really hard part is just knowing when something is good enough versus when can you push harder. And you see this with a lot of the math proofs too.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2321">38:41</a>] When Terence Tao is like, okay, enough with the half-complete theorems, just do the full theorem. And so I think it sounds simple, but it takes a lot of domain knowledge and expertise to get to the point to be like, okay, it seems like you&#8217;re smart enough to have done that. Now do this.</p><h3>39:00 &#8212; How much of writing is done by AI at Anthropic</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2340">39:00</a>] When I was a software engineer, there was a lot of, I guess it was like glue writing work kind of where you finish some work and you write a launch post, or you are part of some work stream and you got to post an update every two to four weeks. Or maybe you have a direction doc or design doc and all this writing around the software that you actually write. And curious your thoughts on if that&#8217;s changed at all at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, how much of that is written by <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> and how much of that is human written?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2379">39:39</a>] Still, my rule of thumb is if I would be happy to show someone the prompt, I would send them the output, right? And so a lot of times the prompt is just collecting context. Context is the most common one. For example, before every one-on-one with my manager, I asked <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to read every Slack message and GitHub message or PR and compile a report of what I did. That&#8217;s context that my manager doesn&#8217;t have because they haven&#8217;t been literally reading every PR or something.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2415">40:15</a>] And so I don&#8217;t feel bad if my manager could also run this command, but it doesn&#8217;t have my context. And I don&#8217;t feel bad with them knowing that this is the prompt I used to generate this command. I think likewise, if you&#8217;re sharing updates or something like that, you might want to prove that you&#8217;ve read it. I think this is important. And sometimes little edits are ways of proving that you have also understood this work.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2445">40:45</a>] And so if it&#8217;s just like a data readout, maybe you&#8217;re sanity-checking the numbers make sense and these are all the numbers you intended to include. And as part of that, maybe you format it differently or you add a little sentence on your behalf, but I think ultimately it&#8217;s a data readout. Everyone knows the prompt you wrote is like, hey, generate a readout on this feature based on this, and you&#8217;re fine, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2470">41:10</a>] But I think if you&#8217;re pitching a new concept or a new idea, and your prompt was like, help me come up with a new concept for this product, right? People probably at least want to know that you&#8217;ve believed in it, right? And so even if you think <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>&#8216;s idea is incredible and just verbatim, you wouldn&#8217;t change anything. What I would say is, I&#8217;d be like, hey, <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> generated this.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2494">41:34</a>] But I think it&#8217;s great, you know, like I did 100 different generations. And I think this was really good. And here is something, you know, that I want to send you, right? And so I think that it&#8217;s good to be upfront about it, I think, right? Because I do think what people don&#8217;t like is when they feel like they&#8217;ve been misled a little bit. Like, oh, we thought you were doing this work, but, you know, it&#8217;s really <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2521">42:01</a>] Right. And I think that writing sometimes can have a lot of... There are some parts of writing where individual words matter, funnily. Tweets are, I think, an example of this, where the individual tweet matters more and more. You can&#8217;t use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to do that because it&#8217;s like every word has some thought that you&#8217;ve put into it and some intention that you&#8217;ve put into it. So yeah, a pitch or an essay or something like that.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2549">42:29</a>] We do a lot of internal essays at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, being like, hey, this is why I think we should do this. And that&#8217;s generally all human written. It&#8217;s very looked down upon, I think, to have <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> write your essay for you, you know what I mean? Because every word is something that you are intentional about.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2568">42:48</a>] This sounds like the proportion of the writing. That&#8217;s all that boring rote writing, like the data readouts, the one-on-one updates, the work stream updates, that&#8217;s increasingly becoming AI but always reviewed by a human. And then the novel thoughts, novel direction, is still very human written and feels like it should remain that way, even if the models were a little bit better too.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2596">43:16</a>] Yeah, I think it&#8217;s like if your writing is also a way of thinking, and so maybe if you need to think about the data stream or data readout more than maybe you need to summarize it. And so, but I think especially gathering context is one of those things where no one will ever hold it against you, kind of. Right. Like, oh, you know, you did a bunch of research and <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> did this.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2621">43:41</a>] But yeah, like I don&#8217;t want to ask my agent to do the same research. Like, you know, it&#8217;s a shortcut, but just acknowledging it is good.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2628">43:48</a>] Someone else I was talking to, they had this thought that it would be valuable to have almost like a Git blame, but it&#8217;s like a prompt blame because it would be nice to reverse look up what was the prompt that generated this change, to kind of debug things. Do you have any sort of meta version control on the prompts that generated the software, or is it still very vanilla Git history?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2657">44:17</a>] Yeah, that is a little bit tough because, again, what goes into a prompt is not just the prompt, but also context and skills and things like that. So maybe a prompt might seem basic, but isn&#8217;t. I think it&#8217;s kind of hard to judge, but I do, going back to when I&#8217;m giving a PR, I don&#8217;t think a PR at this point is any different than an artifact or something. <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> is doing basically all the code writing.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2684">44:44</a>] So if I send someone a PR, I usually also attach an artifact of every prompt I sent to <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>, including failed approaches and things like that, so that they can see, okay, I&#8217;ve considered a lot of other things. And if I haven&#8217;t, if this is just a one-shot, I just tell people, I&#8217;m like, hey, this is the one-shot example. This is the prompt I used. And I think that&#8217;s usually impressive and interesting to them as well because they&#8217;re like, oh, it&#8217;s cool that <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> could one-shot this.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2718">45:18</a>] But the worst is when you get, like, I&#8217;m just trying to avoid cases where I send in this 10,000-line PR and they&#8217;re like, how much have you read this? Or how much do you work with <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> on it? And if I&#8217;m doing a large PR, I am going to show my work.</p><h3>45:36 &#8212; Code ownership and maintenance at Anthropic</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2736">45:36</a>] as much as possible on maintaining code, because AI can generate such high volumes of code at this point. Do you have any tips on what&#8217;s worked well at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> for code ownership and maintenance?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2752">45:52</a>] I do think you have to revisit what is important with code maintenance. I think that there are some things where naming used to be really important because it was how you as a team thought about this abstraction and feature. But I think naming is becoming less and less important. A lot of stylistic things in code are becoming less and less important. And I think that this is not the same to me as maintenance.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2779">46:19</a>] I think that you might want to sit down and be like, okay, what things really matter now? What don&#8217;t? What opinions do we have that don&#8217;t matter? That&#8217;s one. And then I think on the second, on the maintenance side, is having good scaffolding. So I think that it&#8217;s like having a good verification harness, having a good sort of skills. We use the simplify skill a lot. I think that sometimes even from a model perspective, it might do a lot of work.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2811">46:51</a>] And then even just giving it permission to be like, &#8220;Hey, I think this is the right idea. Let&#8217;s simplify.&#8221; It gives it that permission to do it. But you actually don&#8217;t want a model to, by default, do work and then simplify because maybe it&#8217;s not correct. You don&#8217;t want it to simplify work. That&#8217;s not correct. It&#8217;s like you&#8217;re wasting tokens. And so I think this is also how humans think. You&#8217;re like, you think generatively, you try and approach, and then maybe you simplify and abstract a little.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2843">47:23</a>] So I think there&#8217;s some work inside your own code base or setting up those skills and that process for those practices where you&#8217;re not just submitting the first take, but you&#8217;ve simplified it and tried it. And then having a really good verification harness where you feel like you&#8217;re catching it. You have a good belief that <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> is testing every part of it. When you submit a PR to <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a>, you get a recording back of it using the feature and testing it as an example.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2876">47:56</a>] And you can just get really, really creative with different ways to test stuff. I think you should basically have on the order of, I&#8217;d say more like 100 times more testing code than you&#8217;ve ever had before. You know what I mean? So you should have fixtures for everything. You can just pull production code and create fixtures and mockups on the fly for databases. You can have Storybook for front end and things like that.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2902">48:22</a>] And you can have all these different ways of testing and verifying your code. And that, I think, is really valuable for maintainability. It&#8217;s like just having all these ways of verifying it. And then, of course, there&#8217;s just the human element of where do you want your codebase to go? If you know that&#8217;s the case, probably you start thinking about how you&#8217;d replay, and undo and redo becomes really important, just that sort of stuff.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2932">48:52</a>] But if it&#8217;s single-player only, that actually matters a little bit less. The average user is not going to undo 100 times or something. And so many single-player applications&#8217; undo stop working pretty quickly after four or five times, just because it&#8217;s not that important. But in multiplayer, the ability to compose different things together is really, really important, or compose different operations together is really important.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2959">49:19</a>] That&#8217;s just an example where, if you know the direction of your codebase, there are things you care about and you want the models to know, and you can include that in your skills. You can also just. That&#8217;s sort of what maintainability means to me, is having a vision for what your codebase is good at, where you&#8217;re going, keeping it in line. I do think the models are getting better and better and better.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=2979">49:39</a>] And almost every code base probably, I think, will have this moment where you&#8217;re like, do you just ask the model to rewrite all of it? Because now you&#8217;re like, oh, I can do it in the most performant language, I can mix and match. There&#8217;s not a reason that every software in the world shouldn&#8217;t run in WebAssembly in the browser. You know what I mean? There&#8217;s not a reason why I don&#8217;t know your Xbox game can&#8217;t run in your browser.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3009">50:09</a>] Actually, the consoles are much weaker than your average MacBook right now. But it&#8217;s just the code base. But we could. And so I think probably over the next year or two, everyone&#8217;s going to need to think about that. And so I don&#8217;t think you want to spend too much time worrying about maintainability in this way. That might not matter anymore. I&#8217;m not saying maintainability doesn&#8217;t matter.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3037">50:37</a>] You just have to update your mental model of what does it mean for code to be maintainable.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3043">50:43</a>] I was talking to this friend who was saying that he leaves tons and tons of tech debt lying around because the next iteration, why waste time fixing that now? In the next iteration of <a href="https://en.wikipedia.org/wiki/Fable">Fable</a>, it&#8217;s going to one-shot all this tech debt and refactor this for me. So kind of funny.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3065">51:05</a>] Yeah, I don&#8217;t think that&#8217;s completely incorrect. It depends on the circumstance, right? I think this is also just a classic thing in startups where you&#8217;re like, even in normal human engineering, you always have this problem of like, hey, do I refactor? Do I do tech debt or do I deliver more customer value? Would refactoring help me deliver more customer value right now? I do think that, generally, you should think on your projects on shorter timescales, so you&#8217;re like, okay, how do I deliver value over the next month or two?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3096">51:36</a>] And if the project you&#8217;re talking about is delivering value over six months or 12 months, then maybe wait for the next model a little bit and just deliver value on the short timescale. Make sure it&#8217;s really good. And if the models are not quite good enough, they might get there soon. It depends very much on the specific case.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3119">51:59</a>] Right? But I&#8217;m not saying this is true of everything, but just something you should keep in mind.</p><h3>52:04 &#8212; How Anthropic prevents breakages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3124">52:04</a>] I&#8217;ve talked to some friends who work at big tech companies like Google, Facebook, those types of places. And then they, as they&#8217;ve become more and more AI-pilled, one thing that people have noticed is there&#8217;s a lot more incidents or sevs in their usage, and it&#8217;s natural in those organizations because they read the code less and there&#8217;s more code flying out. What countermeasures have worked really well for <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> to prevent breakages, given that the code velocity is so much higher?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3158">52:38</a>] Yeah, I think this is something that is a byproduct of moving faster sometimes, and we have to figure it out. I don&#8217;t think our uptime is exactly where we want it to be either. But also, as a company, we&#8217;re almost six years old, I think, around. So it&#8217;s like no company has grown this fast before. And a lot of that is because we&#8217;ve been able to create more products faster than ever before.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3184">53:04</a>] And so you can use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to make your uptime better. And I think the way we think about this is just really good. What&#8217;s the dream testing environment? The dream deployment environment? Can you take requests and replay them across mock databases and fixtures? Across everything. Can you Chaos Monkey everything?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3206">53:26</a>] In my personal workflows, I have so many more custom random tools and scripts that just make everything faster. And just curious, on your team at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, for instance, does everyone have a set of miscellaneous tools that help them get random things done?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3228">53:48</a>] Yeah, everyone does for sure. I think part of this is the job. I think that sometimes what we try and do is we play around with harnesses. Sometimes some people on the team build their own harnesses for a little bit to figure out, oh, is this useful or not? And then they figure out if it works and if not, they integrate it. I think that&#8217;s part of the job in some ways. I think other examples of sort of misc stuff people will do.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3258">54:18</a>] A lot of people will have unique <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> setups. I think <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> is one of these things where, you know, you can have it like I have it scheduled, my calendar advice, right? And so if someone wants to schedule something with me, I&#8217;m just like, hey, you can just tag <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> here in this channel and I&#8217;ll accept whatever it puts on my calendar. You know what I mean? So I think there&#8217;s some stuff like that.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3281">54:41</a>] I know a lot of people use it for email. I think that I&#8217;ve seen some people do interesting multi-closing, sort of tmuxing-like setups. What&#8217;s the ideal scenario for you to display 50 different <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> Codes, and what&#8217;s the best way for you to figure out what&#8217;s going on at any one time? I think there&#8217;s a lot of different ways, and we&#8217;re trying to make <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude Code</a> more hackable as well so that more people can even make their own version of <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3318">55:18</a>] Code more different, and everyone might have their own little twist on it.</p><h3>55:24 &#8212; Visibility and sharing your work</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3324">55:24</a>] Your role at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> is really interesting because of the external visibility that you have. And I think a lot of people, when they give career advice, visibility is a good thing, but they don&#8217;t have this level of external visibility. So do you recommend to software engineers that they should be posting on <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> and X? And if so, what advice would you give in that sense?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3349">55:49</a>] Generally, the thing I say to people is that you should share your work externally as much as you can. Especially, I think, within certain companies. You might not be able to, but maybe you have side projects or something like that. I think just before I joined <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, what I did was I spent a bunch of time working with different companies, building stuff, and writing about it and talking about it.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3373">56:13</a>] And this was really valuable because it increased my surface area of luck. And so I think that the bar is much lower than you think. Basically, whenever someone asks me for advice, I&#8217;m like, okay, I&#8217;ll sit them down. I&#8217;ll be like, I know statistically I tell this advice to a lot of people, and almost no one does it. And everyone who&#8217;s done it is either in a job that they&#8217;re pretty excited about or running their own company.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3405">56:45</a>] And there are reasons why you&#8217;re not going to want to do it, but this is it. I just can tell you one: I don&#8217;t care about your resume. Don&#8217;t do that. You have to choose an interesting project to work on. You have to work really hard on it. You have to lock in, and then you have to ship it and write about it. You have to do it. You can&#8217;t. There&#8217;s going to be so many reasons why you don&#8217;t want to.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3426">57:06</a>] It&#8217;s like not good enough yet. Or like you haven&#8217;t, like you think the writeup is not very interesting or something, but you have to do it and you can&#8217;t get discouraged. Maybe the first one might not work, but you have to do it again. And maybe <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> is not the right place. Maybe it&#8217;s <a href="https://en.wikipedia.org/wiki/Reddit">Reddit</a>, maybe there&#8217;s a specific <a href="https://en.wikipedia.org/wiki/Reddit">Reddit</a> or Hacker News or something like that. Part of what you&#8217;re trying to do is you&#8217;re just trying to find people who like what you&#8217;re doing, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3449">57:29</a>] But I think that people in general want to really, you know, we consume more content than ever, right? And you know this, right? I think that people want really high-quality content. High-quality content, as you know, is a lot of work. And so I think going back to what we said before about would you show someone the prompt? You did, right? I think a lot of times people are like, oh, what&#8217;s the shortcut?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3475">57:55</a>] &#8220;Oh, should I just ask <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to manage my <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> account? And is that how I grow?&#8221; And I&#8217;m like, no. Like, you should not do that.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3483">58:03</a>] You have to sort of engage authentically. You have to post, do good work, and talk about it. And I think build up networks and things like that. But I think that as long as that&#8217;s one of your goals and you try hard at it, I haven&#8217;t seen anyone not succeed at it, but it is really hard. It&#8217;s kind of like saying, oh, I want to go to the gym every day, or I want to lose weight or something.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3506">58:26</a>] There are simple things that you can do that take a lot of discipline and are easy to get discouraged with. But they&#8217;re worth doing if you do it. And I think posting, or more specifically writing about your work, is really, really valuable.</p><h3>58:42 &#8212; Luck surface area example</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3522">58:42</a>] You mentioned LLM surface area. Do you have an example that kind of illustrates the value of expanding your LLM surface area?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3532">58:52</a>] Yeah. I mean, okay, how did I get my job at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>? I did a fellowship with a company called Goodfire where I did some applied research. And so they&#8217;re an interpretability AI company. And I was trying to figure out how do I use interpretability to make better products? And so I learned about their research, and I felt I want to work on a project, and it was really important for me to share it.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3561">59:21</a>] And so, when I agreed to work with them, I was like, that&#8217;s my goal: I want to share what I&#8217;m building. I think that&#8217;s good for you as well, because Goodfire is a startup, and people want people to learn about them. And so this is something I&#8217;m going to do. So I built this interpretability visualization, and I shared it. This was my first big post on <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a>, but it was only 500 likes or something.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3582">59:42</a>] Now, you know, that&#8217;s not very much to me, but it&#8217;s like, back then it was like a huge post, and several people saw it, and some of them DMed me about it. And then I got an intro into a role here. And so I think that I spent a month on that project in particular, so it wasn&#8217;t crazy. And, yeah, it just showed people that I could do interesting work.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3617">01:00:17</a>] And I got paid for it too, so it wasn&#8217;t like I was doing it for free. And there&#8217;s plenty of examples of, I think people are happy to do that sort of thing for you, if you sort of show that kind of initiative. But yeah, you have to. It is like, you really have to do interesting and novel work and work hard and be proud of your work as well. There&#8217;s not a shortcut to that, I think.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3643">01:00:43</a>] But if you do, I think everyone is always looking for interesting work and wants to support you and wants to hire you or, yeah, give you money, like, lots of good things.</p><h3>01:00:57 &#8212; Should people still learn to code</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3657">01:00:57</a>] I see this tagline going out a lot with some of the stuff that <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>&#8216;s been maybe, like, I think it was Boris went on some podcasts, and the tagline was, &#8220;Coding is largely solved.&#8221; Should people still learn to code if coding is largely solved?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3674">01:01:14</a>] I think being technical is really, really important. Knowing how computers work, how computer programs work, how languages work, what are the hard things? And what&#8217;s a backend service, what&#8217;s a cache, what is memory allocation? All of these things are actually kind of really important to learn. I do think it&#8217;s hard to motivate yourself sometimes to do it the same way that doing math by hand was not that motivating to me, but some people just love math and did it.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3709">01:01:49</a>] I think that that&#8217;s probably something that people have to figure out, but I think it is really worth it. Being technical is really, really important. We talked about at the start, or earlier, where the only way you can tell you&#8217;ve solved the Riemann Hypothesis or made progress or whatever is if you&#8217;re a great mathematician. And in the same way, the only way you can tell if you&#8217;ve built great software is if you&#8217;re a great software engineer.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3739">01:02:19</a>] I think that how do you do that is hard. But we&#8217;ve talked about some of this stuff before, just staying in the loop, reflecting on your process and getting better. And you can use <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to learn as well and sort of explain things to you. So treating it like a thought partner, but learning truly. I think Karpathy says learning should feel like effort. And I think that&#8217;s one of the hard things is a lot of times, even if you ask <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to explain something to you, you might just nod along and you&#8217;re like, oh, yeah, I learned it.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3772">01:02:52</a>] But you didn&#8217;t really, because you didn&#8217;t put any effort in. So I think it is really technical. You should learn. It&#8217;s hard to learn. And sometimes what school forces you to do is learn that. I don&#8217;t know, truly. I&#8217;m not learning programming from scratch, so I don&#8217;t know exactly how to do it. Now, I think if there are better ways, I would still type out and build programs and run them and learn them, you know what I mean?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3800">01:03:20</a>] There might be better ways that are more <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>-informed as well. But I think it&#8217;s really important. I think when Boris says coding is solved, I think it just means we don&#8217;t get stuck in the same ways that we used to before. I think coding used to be this very high-variability thing where you&#8217;re like, oh, could this bug take a day or could it take two weeks? You have no idea sometimes. And I think on the whole, coding used to be in real terms something that was very rare for something to go well.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3836">01:03:56</a>] Very few people in the entire world could write software, and they were very, very rare. And even if you got them all together, there were so many other reasons why it wouldn&#8217;t work. And coding was one of the rarest things in the world where the chance of a software project going well was, on absolute terms, very low. And if you were hiring someone to make software for you for something that&#8217;s not a huge product, like if you&#8217;re hiring someone to make software for your car dealership, you are almost certainly not going to get the software you wanted.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3871">01:04:31</a>] You&#8217;re going to get essentially scammed, you know what I mean? Not because anyone was trying to scam you. It was just like, software is really, really hard and it could only be spent on the most important scalable things in the world. And now that coding is solved, I think what we mean is that you can use coding to do all these other things that we&#8217;ve not done before. But that&#8217;s not to say that that&#8217;s not a lot of work still.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3896">01:04:56</a>] It&#8217;s just not this incredibly rare, difficult thing that mostly just doesn&#8217;t work, and you have to spend, like, eight hours a day locked in to do well.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3907">01:05:07</a>] It&#8217;s really insane. That shows the difference in expectation is when I used the right software. I would be shocked if it worked on the first try. I was like, whoa, wait, why is this working? And you expect to kind of bash your head against the wall a little bit and then it works, even if it&#8217;s just like you missed a semicolon or something like that. And now I almost have the flip expectation where when it doesn&#8217;t work, I&#8217;m like, wait, what, why, why did <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a>.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3934">01:05:34</a>] What happened here? Usually I expect it to work almost the opposite, like on the first try.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3941">01:05:41</a>] It&#8217;s crazy immediately. Yeah, I think humans get really used to abundance, right? There was this article someone shared recently about going through a modern apartment and talking about all the wonderful things that we have now that people could not have imagined before. Something that can play music on demand that&#8217;s suited for your mood. Before, you&#8217;d have to hire a musician to go write incredible music, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3969">01:06:09</a>] But I&#8217;d say, on the whole, probably more people are getting paid to make music now than ever before, right? And I think that music is reaching more people than ever before. And I think probably the same thing will happen with software, will happen with math, will happen with all of these things where, when you get abundance, people are like, great, I want more abundance.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3990">01:06:30</a>] I want my software and everything, you know. Like, I want my music everywhere, you know.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=3994">01:06:34</a>] So, yeah, one potential other data point from this conversation saying that maybe people should still be technical or learn how to code is, I think earlier in the conversation we talked about automating knowledge work, and you mentioned that people who are technical had kind of a leg up because they kind of understood how to coordinate <a href="https://en.wikipedia.org/wiki/Claude_(AI)">Claude</a> to do a variety of things, like you&#8217;re calling FFmpeg to automate some video editing.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4021">01:07:01</a>] And I don&#8217;t think someone who is not technical would have that thought. So it does seem like even in this case where models are doing a lot, there&#8217;s still so much value in</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4032">01:07:12</a>] Being technical, knowing how computers work, in the same way that probably the most technical CEOs are, like, the best CEOs are technical, like Zuckerberg, Elon, but they probably haven&#8217;t written a line of code truly in a long time. But they understand how systems work, they understand constraints and things like that. And that&#8217;s really, really important. And so even if all that work has changed, I think being technical is really, really important.</p><h3>01:07:42 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4062">01:07:42</a>] And then last question for you is, if you could go back to when you just entered the industry and give yourself some advice, knowing what you know now, what would you say?</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4072">01:07:52</a>] They&#8217;re like kind of like two wolves, sort of. I think you have to believe in yourself. I think this is really important. And I think that at least when I joined the industry, it was rarer to believe in people when they were kind of younger. I think that was the whole point of Y Combinator or something was believing in young people early on to do really great things.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4095">01:08:15</a>] And so I think that, you know, even going back to the internship discussion we had earlier, right, I think that thinking of yourself not as someone who&#8217;s being trained to do something, but someone who can do something incredible right away, I think is really important. And I think I wish I had done bigger, bolder things that I&#8217;d written and shared about.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4118">01:08:38</a>] I think there were lots of ideas where I was like, wow. I think I had thought I had done original work that I wish maybe I&#8217;d even shared in some form, but maybe not in a form that survived the Internet, you know, and it wasn&#8217;t a goal of mine. And so I wish I&#8217;d sort of been bolder and done and shared some of this work. But at the same time, you also.</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4143">01:09:03</a>] There is a lot to learn from people. And I think that when you&#8217;re young, you tend to fall. Either you&#8217;re not bold enough or you&#8217;re too bold and you don&#8217;t figure out what people have done before and figure out why this is not working, right?</p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4169">01:09:29</a>] And there&#8217;s this balance. And I think everyone has different failure modes, but I think I probably didn&#8217;t take enough advantage of mentors and people who had learned a lot. And I think I probably had to relearn a bunch of things as a result. So there&#8217;s probably not universal good advice, but just things to think about as you&#8217;re entering.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4193">01:09:53</a>] Awesome. Well, thanks so much for your time, Thariq. Really appreciate it.</p><p><strong>Thariq:</strong></p><p>[<a href="https://youtu.be/2Kch3tWMnw8?t=4196">01:09:56</a>] Yeah, of course. Thanks, <a href="https://en.wikipedia.org/wiki/Nathan_Peterman">Ryan Peterman</a>. It was fun.</p>]]></content:encoded></item><item><title><![CDATA[Creator of Scala: Comparing Languages And How AI Will Impact Them | Martin Odersky]]></title><description><![CDATA[There are so many choices every programming language creator makes while navigating engineering tradeoffs in building one.]]></description><link>https://www.developing.dev/p/creator-of-scala-comparing-languages</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-scala-comparing-languages</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 31 Aug 2026 13:05:35 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211791435/2857a28d4760be8b26ba11ba481ad003.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>There are so many choices every programming language creator makes while navigating engineering tradeoffs in building one. </p><p>This made me curious about <a href="https://en.wikipedia.org/wiki/Martin_Odersky">Martin Odersky&#8217;s (the creator of Scala)</a> thoughts comparing different languages. His deep domain expertise made his responses comparing Rust, Zig, Python and Scala interesting.</p><p>He also had a lot to say about how AI will impact programming language ecosystems in the future. For instance, he talked about <a href="https://arxiv.org/abs/2603.00991">some research he was working on</a> to strengthen the static guarantees programming languages give to provide more safeguards for LLMs. </p><p>Hope you enjoy the conversation!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/LdN4sPWM-WY">YouTube</a>, <a href="https://open.spotify.com/episode/0EHxXlNhBNbZuXc876W52m?si=4HTiVaY4S-Cn5-JlycKS8w">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-LdN4sPWM-WY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;LdN4sPWM-WY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/LdN4sPWM-WY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/211791435/0044-why-care-about-functional-programming">00:44 - Why care about functional programming</a></p><p><a href="https://www.developing.dev/i/211791435/0635-why-should-people-learn-scala">06:35 - Why should people learn Scala</a></p><p><a href="https://www.developing.dev/i/211791435/0926-rust-vs-scala">09:26 - Rust vs Scala</a></p><p><a href="https://www.developing.dev/i/211791435/1242-rust-vs-zig">12:42 - Rust vs Zig</a></p><p><a href="https://www.developing.dev/i/211791435/1545-scala-vs-python">15:45 - Scala vs Python</a></p><p><a href="https://www.developing.dev/i/211791435/1831-the-programming-languages-that-influenced-him">18:31 - The programming languages that influenced him</a></p><p><a href="https://www.developing.dev/i/211791435/2216-how-running-on-the-jvm-works">22:16 - How running on the JVM works</a></p><p><a href="https://www.developing.dev/i/211791435/2633-why-writing-a-compiler-is-hard">26:33 - Why writing a compiler is hard</a></p><p><a href="https://www.developing.dev/i/211791435/2919-why-twitter-adopted-scala-early-on">29:19 - Why Twitter adopted Scala early on</a></p><p><a href="https://www.developing.dev/i/211791435/3100-how-he-believes-ai-will-impact-programming-languages">31:00 - How he believes AI will impact programming languages</a></p><p><a href="https://www.developing.dev/i/211791435/4340-will-there-be-less-engineers-in-ten-years">43:40 - Will there be less engineers in ten years</a></p><p><a href="https://www.developing.dev/i/211791435/4434-top-programming-languages-to-learn-to-grow">44:34 - Top programming languages to learn to grow</a></p><p><a href="https://www.developing.dev/i/211791435/4618-top-technical-book-recommendation">46:18 - Top technical book recommendation</a></p><p><a href="https://www.developing.dev/i/211791435/4651-why-he-chose-academia-instead-of-industry">46:51 - Why he chose academia instead of industry</a></p><p><a href="https://www.developing.dev/i/211791435/4828-reflecting-on-scala">48:28 - Reflecting on Scala</a></p><p><a href="https://www.developing.dev/i/211791435/5542-advice-for-his-younger-self">55:42 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:44 &#8212; Why care about functional programming</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=44">00:44</a>] What is functional programming? And why should an imperative programmer care about the functional way of programming?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=52">00:52</a>] So functional programming is programming with values, so you don&#8217;t have state that changes, like mutable variables that change or arrays that change. You have values and you have functions, and the functions transform these values into other values. So it&#8217;s a very restricted form of programming. And most real functional programming languages are impure in the sense that, yes, of course, at some point you have to touch state and change things and write to standard out and things like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=84">01:24</a>] But the idea is to do that sort of very late, so to have the largest part of your program represented as values and functions that transform values. And the benefit you get from that is essentially much better predictability and fewer bugs, because these side effects to, let&#8217;s say, a global variable or a field in some object graph or things like that, they can really trip you up. They can essentially, they&#8217;re an undocumented effect of your function, so it&#8217;s not reflected anywhere typically.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=117">01:57</a>] But maybe you have a comment if things are good. And these things are very hard to keep in your head and generally reason about. So the advice from functional programming is essentially to keep these things to the bare minimum and possibly to nothing if your function maps into these things. The other reason why functional programming is nice is that it&#8217;s very directly linked to mathematics. So in mathematics you have theories of, let&#8217;s say, polynomials or strings or lists or things like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=153">02:33</a>] And in none of these theories will you find the concept of mutation; that just doesn&#8217;t exist. So you can take a polynomial and you can transform it into a new polynomial, but you can&#8217;t change a coefficient at 0.3 at the third coefficient and pretend it&#8217;s the same polynomial. It&#8217;s not in mathematics, right? It&#8217;s clearly something different. So functional programming, you could say, follows quite closely the way mathematics sees things.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=178">02:58</a>] What would you say to someone that doesn&#8217;t do functional programming because it&#8217;s not convenient?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=184">03:04</a>] I guess two things. One is it&#8217;s a learning curve. So initially your brain is wired to do imperative programming from when you were very young. I imagine most people learned an imperative language first, and then it&#8217;s hard to see how you would express things differently. But if you persist in that a little bit, and it doesn&#8217;t really take much, then you will essentially reap the benefits very quickly, that you say, well, actually my program got a lot clearer.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=216">03:36</a>] I can understand it better. I don&#8217;t really have to track some steps, go step by step through my program with a debugger or things like that to figure out what it does. Typically I don&#8217;t need a debugger when I write functional code. That&#8217;s really not necessary. The other thing is to essentially know when to stop. So there is a strand of functional programming called pure function programming, which tries to express as much as possible in pure functions and to delay the side effects and wrap them up in monads or whatnot.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=255">04:15</a>] And that can get very inconvenient very quickly. So my answer to that would be, it really depends. And in most cases, it&#8217;s completely okay to have your side effects as long as these are well documented and you use them. So basically, use these things in moderation. I would say functional programming is great for 95% of your program, and if you need some side effect, some imperative feature, for the last 5%, that&#8217;s also okay.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=285">04:45</a>] Is there any quantitative measure of the, I guess, the correctness benefits that functional programming gives you over imperative?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=296">04:56</a>] I think there are some studies. It&#8217;s not quite clear how significant they are. I think people criticize them a lot. The studies tend to say that a language like <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> would have fewer bugs than a language like <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a>. But I don&#8217;t really want to insist on that because these studies are very hard to do, and the methodology is very hard to get right. I believe the effects become more important in large programs than in small ones.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=328">05:28</a>] For small ones, you can do anything, really. You can write a loop, you can write a recursive function. It doesn&#8217;t really matter. And it&#8217;s very much a matter of taste which one you prefer. But in a large system, I believe static types have certainly proven their worth, even if we don&#8217;t really have a conclusive study that shows it. I mean, I can point you to a study, and if you&#8217;re a dynamic typing enthusiast, then you will point to another that essentially claims the opposite.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=359">05:59</a>] So it&#8217;s very, very hard to do empirical software engineering on that scale. And so I&#8217;m afraid I can&#8217;t really give a good answer. I just have a feeling that mostly functional statically typed languages is sort of a sweet spot for getting programs right. Maybe the other thing is when you really need to get them absolutely right, like with proofs and things like that. If you do <a href="https://en.wikipedia.org/wiki/Rocq">Rocq</a> or Lean or any of these things, these are all functional languages.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=389">06:29</a>] So people wouldn&#8217;t even attempt to do something like <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a> here.</p><h3>06:35 &#8212; Why should people learn Scala</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=395">06:35</a>] When you think about <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> in the environment of all programming languages, what sets <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> apart? Or maybe put it another way, why should someone learn <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> in 2026?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=407">06:47</a>] Well, <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> is the only functional language that is also a very capable object-oriented language. In fact, it was born by the idea that we can actually make a fusion of the two in a way which is not side by side, but which really combines the features in a nice synthesis. So that was what <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was about. That&#8217;s what I wanted to show, and that&#8217;s been largely successful. So you can write really beautiful programs in this combination.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=439">07:19</a>] Object-oriented programming comes in when you talk about components and modules and essentially encapsulation, these sort of things where functional programming typically doesn&#8217;t really have a very strong story. I mean, there are languages like Standard ML or <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> that do have very capable module systems, but other functional languages don&#8217;t. And people, when they think of functional programming, they don&#8217;t really think much about components and interfaces and these sort of things, which are things that I believe also matter very much.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=470">07:50</a>] Could you give an example? Maybe an object-oriented thing that you could do in <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, but you couldn&#8217;t do in Haskell, which is pure functional.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=480">08:00</a>] It&#8217;s sort of ingrained in the whole fabric of the thing that everything is essentially objects, what you do. So one thing that you get is what I think Simon Peyton Jones called the power of the dot, that you say it&#8217;s super convenient. You have an object and then you do dot, and then you have the environment that immediately tells you what are the methods and fields of that object, that you can use them.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=504">08:24</a>] So it focuses your mind. Whereas in functional programming, it&#8217;s typically you have a sea of functions that can be applied to arguments, and you have to figure out what they are.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=514">08:34</a>] When you think about the systems programming languages like <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a>, Go, <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a>, in your opinion, which one would you say is kind of the best one? And then maybe we can compare it to <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=526">08:46</a>] In this day and age, systems programming languages need to be memory safe, guaranteed memory safe. So that would already exclude quite a few of them, but it would leave <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> and Go, I think. And <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> and Go are at different levels. Go is really not a sort of nuts-and-bolts systems programming language. This is more a language in which you would write, let&#8217;s say, an application server or some middleware or some cloud infrastructure, or things like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=553">09:13</a>] It&#8217;s not something you would use for embedded, say, which you would use <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> for. So I think in the domain, both of them are probably the ones that are the leading ones that I would take most seriously.</p><h3>09:26 &#8212; Rust vs Scala</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=566">09:26</a>] And then when you compare <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> or Go with <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, what are the pros and cons of the different language designs? What&#8217;s something that <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> does better than <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>? Something <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> does better than <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=580">09:40</a>] So <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> is closer to the metal, you have better performance guarantees. I guess the fact that <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> is a garbage-collected language means that you always have some pauses. I mean, garbage collectors have become quite, quite capable. I mean, they&#8217;re brilliant and the pauses are really very, very small. But it&#8217;s a fact that you do need a big chunk of memory to run fast, and I guess <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> could run in much smaller memory.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=608">10:08</a>] So I believe that&#8217;s better for embedded systems and things like that. I believe right now <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> is actually overused because a lot of people push <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> for things higher up in the stack where your garbage collector is fine. But essentially, you still want to write code without, and that is for me a bit an exercise in, I don&#8217;t know, just intellectual that I can do it.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=637">10:37</a>] I don&#8217;t really think there&#8217;s a big sense in it to write memory management for a garbage collector. You should absolutely use one because it makes a lot of things simpler. So yes, of course, if given enough brains, I can write code around it, and I can do that. But why should you? I mean, it can be much simpler. So that was for <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, for Go. Go is sort of a language that lags behind. I mean, it was intentionally designed to be very, very small and essentially to be the standards of the &#8216;90s.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=675">11:15</a>] So now they have gotten generics, which is a big step. So I believe that sort of is a step forward for that. But it&#8217;s still a fairly limited language. What you can do, which has an advantage that essentially it forces a very uniform style because there&#8217;s not much different things you can do. And there&#8217;s probably a culture fostered to do things in a certain way, which makes it easier to essentially read one another&#8217;s programs and essentially jump in a new code base and things like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=707">11:47</a>] So I think that would be my main advantage of Go that I see there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=711">11:51</a>] Is it possible to turn off the garbage collector in <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=715">11:55</a>] Not in production. So we have some research that would let you use essentially your own memory allocators and things like that, the way let&#8217;s say <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> does it. The problem with that is always that you can leak references into memory that you reclaim, and then all hell breaks loose. Essentially, your pointers, you point to memory that&#8217;s undefined. But in <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> we now have a very good way to actually track these references, so we can actually prevent that statically by the type system, that this will never happen, and you can still have your own allocator.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=756">12:36</a>] That said, we haven&#8217;t really shipped that in production yet. So right now I would say you use the garbage collector. <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>.</p><h3>12:42 &#8212; Rust vs Zig</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=762">12:42</a>] Yeah, I see a lot on the Internet, people comparing <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> and <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a>, or kind of a fierce debate on which one is better. When you think about <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> versus <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a>, what are the merits of the two compared to each other?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=778">12:58</a>] So what I can see is <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> has a really nifty compile-time construct, essentially inlining, where the compiler does smart inlining, and that is quite clean and quite powerful. And <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> has macros, but I think they&#8217;re more clunky than the <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> version. In <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, we have something quite close to <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a>. We also have essentially a thing that&#8217;s based on inlining and optimizations by the compiler. But we have a restriction, which I believe <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> doesn&#8217;t have, and that is that there cannot be additional type errors after inlining.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=821">13:41</a>] So the thing is, you inline, and then the question is: Is the inline program guaranteed correct, so type-correct, or might you have type errors? The typical example where you might have type errors is <a href="https://en.wikipedia.org/wiki/Template_(C%2B%2B)">C++ templates</a>. In fact, that&#8217;s quite scary in C++ that you can expand a template, and then you get very, very complex type errors and very, very hard-to-debug things. So <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a>, I believe, its inlining mechanism is much saner.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=850">14:10</a>] When you say inlining, can you explain the concept?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=853">14:13</a>] Inlining just means that in <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> we can write `inline` in front of a function, and that means that the compiler will, before it starts code generating, during the time when it looks at the type code, which are trees, take the function body. When it sees a function call to that function, it will take the call and replace it by the body, and then it will do some optimizations. It can say, okay, so here we have essentially an application to, let&#8217;s say, a value which is a lambda, but I know where the lambda points to, so let me forward the call right to the function.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=894">14:54</a>] And with that you can already do quite a bit of essentially optimizations, which are guaranteed because the inliner must inline. That&#8217;s not a thing. An optimizer has essentially discretion whether they want to inline things or not, and so you can never rely on that. But an inliner, which is essentially a compile-time based on the typer, must inline, so you can rely on it, and constexpr in <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> and C is essentially very similar.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=925">15:25</a>] Okay, so it&#8217;s a way to get rid of the function call, or the overhead of a function call, and tell the compiler you can do more because it&#8217;s all, I guess, inline.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=937">15:37</a>] Yeah, you reveal the implementation, and that means the compiler can do something with that.</p><h3>15:45 &#8212; Scala vs Python</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=945">15:45</a>] When you think about all the dynamically typed languages, which one stands out as one that you think is kind of the best among them?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=953">15:53</a>] Python is ubiquitous, and it has a nice syntax. A lot of Python programs look like they&#8217;re very easy to read, so that&#8217;s definitely an advantage. And the other one would probably be something like Scheme, which is sort of very grounded in computer science theory and lambda calculus and things like that. So both of them are sort of interesting in their own way. But of course Python is 100 times more popular.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=983">16:23</a>] If you compare <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> to Python, for instance, I guess what are the trade-offs that the two languages take?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=990">16:30</a>] I think the gap is closing because Python now actually has an optional type syntax and a number of type checkers that check that syntax. And Python is getting some of the features that <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> had since the beginning, like pattern matching is in one of the recent Pythons. So I think actually the gap is closing not just between Python and <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, but between a lot of programming languages in general. They&#8217;re sort of all drifting to a standard set of features, which mostly come from functional programming.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1021">17:01</a>] Actually, pattern matching, for instance, strong type systems, generics, polymorphism, all these things came from closures, all these things came from functional programming. And so the gap is closing in that sense. Python is, I think the main advantage of <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> over Python is that it has a strong type system that is always on, and that gives you essentially guarantees that certain bad states can&#8217;t happen.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1054">17:34</a>] So really you can rely on it. Whereas I think the Python type system, well, it&#8217;s just syntax. It has a number of type checkers, but in general you have fewer guarantees. And I believe it&#8217;s also the ecosystem and culture that doesn&#8217;t value types as much in Python. So I think that&#8217;s probably the main difference syntactically. The two languages, I think with <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> 3, are actually also quite close.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1079">17:59</a>] So <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> 3 looks a lot like Python in that sense. You could say, well, you can see it as a language that has strong types and runs on different runtimes than Python. Or I should say the other thing with Python, which is really great, is that Python is a fantastic glue language because I can essentially have very, very efficient linkages to high-performance C libraries, <a href="https://en.wikipedia.org/wiki/Pandas_(software)">pandas</a> or <a href="https://en.wikipedia.org/wiki/NumPy">NumPy</a> or things like that.</p><h3>18:31 &#8212; The programming languages that influenced him</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1111">18:31</a>] You mentioned that a lot of the programming languages, they&#8217;re kind of drifting or kind of being inspired by the other ones and adopting new features. In the design of <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, is there a programming language that you admire most and influences <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> the most?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1128">18:48</a>] Historically, <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was essentially, you could say it was a blend of <a href="https://en.wikipedia.org/wiki/Java">Java</a>, <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, and Standard ML, which is close, an ML-like language, and Haskell. So I&#8217;m by trade an imperative programmer. PhD is from Niklaus Wirth. I know Pascal, <a href="https://en.wikipedia.org/wiki/Modula">Modula</a>. That was sort of my first generation of languages, and I like them a lot. And then I became sort of a converted functional programmer. So I looked very closely at ML, <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, Haskell, and these things.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1163">19:23</a>] And then <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> came about. There was a predecessor language called <a href="https://en.wikipedia.org/wiki/Pizza">Pizza</a>, and I was working on that with Phil Wadler, who&#8217;s one of the original Haskell designers. And the idea was to have essentially an accessible functional language on a very widespread platform, which was the JVM at that point. And in order to prepare for that, I wrote a <a href="https://en.wikipedia.org/wiki/Java">Java</a> compiler. So that&#8217;s how I learned a lot about <a href="https://en.wikipedia.org/wiki/Java">Java</a> and what <a href="https://en.wikipedia.org/wiki/Java">Java</a> was.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1193">19:53</a>] And in the end I had to say, well, it&#8217;s actually quite useful. I was quite dismissive at first, but afterward I found it actually quite useful. I mean, this was very early <a href="https://en.wikipedia.org/wiki/Java">Java</a>, this was <a href="https://en.wikipedia.org/wiki/Java">Java</a>, even pre-1.0. So now <a href="https://en.wikipedia.org/wiki/Java">Java</a> is, of course, even more useful, but also a lot bigger than what it was at the time. So when we came up with <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, which was sort of the second iteration after <a href="https://en.wikipedia.org/wiki/Pizza">Pizza</a>, <a href="https://en.wikipedia.org/wiki/Pizza">Pizza</a> was the language I did with Phil Wadler.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1224">20:24</a>] The ideas came mostly from <a href="https://en.wikipedia.org/wiki/Java">Java</a>, <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> for the modules and the component model, and Haskell a lot for the standard libraries. So if you look at the function names in the <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> standard library, then there&#8217;s sort of a mixture of <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> and Haskell, you could say, but you recognize a lot coming from both of these languages.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1248">20:48</a>] If I recall correctly, <a href="https://en.wikipedia.org/wiki/Java">Java</a> had some sort of licensing with it, or it was owned by a company. How were you able to build on top of <a href="https://en.wikipedia.org/wiki/Java">Java</a>? Or did you have to pay for licenses, or how&#8217;d that ecosystem work?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1265">21:05</a>] The language standard was open source. There was a lot of experimentation with <a href="https://en.wikipedia.org/wiki/Java">Java</a> at the time. I think the only thing that Sun, that was the company at the time, they&#8217;re very particular about, is that if you did something that was using <a href="https://en.wikipedia.org/wiki/Java">Java</a> and had <a href="https://en.wikipedia.org/wiki/Java">Java</a> in it, you had to call it <a href="https://en.wikipedia.org/wiki/Java">Java</a>, and you couldn&#8217;t have a different name. So that&#8217;s why initially, for instance, <a href="https://en.wikipedia.org/wiki/Microsoft">Microsoft</a> clashed with them, and then <a href="https://en.wikipedia.org/wiki/Microsoft">Microsoft</a> went away and designed C#, which was sort of a.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1300">21:40</a>] <a href="https://en.wikipedia.org/wiki/Java">Java</a>, some steroids. They threw in a bit more than what <a href="https://en.wikipedia.org/wiki/Java">Java</a> had. At the time, I think the nastiness came, Sun was actually quite an open company. So at the time, we were not worried about that. So that was in the time frame from 1998 to 2004, 2005, something like that. And at the time, there was no issue. I think the issue came later when Sun was acquired by <a href="https://en.wikipedia.org/wiki/Oracle">Oracle</a> and <a href="https://en.wikipedia.org/wiki/Google">Google</a> forked <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> on Android.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1331">22:11</a>] <a href="https://en.wikipedia.org/wiki/Java">Java</a> on Android, and that&#8217;s when the fight started, but that was much later.</p><h3>22:16 &#8212; How running on the JVM works</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1336">22:16</a>] So when you say that <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was using the JVM or built on top of it, like concretely in the tech stack, what does it mean that <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> uses the JVM? Like, how does <a href="https://en.wikipedia.org/wiki/Java">Java</a> source code eventually execute on a machine?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1352">22:32</a>] So the JVM is essentially defined by its bytecode. The bytecode is essentially an intermediate format which you can translate your program into. And then the bytecode is run first by an interpreter and then essentially by an optimized compiler, just in time. Compiler <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a>. So that&#8217;s called <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a>. They do that. And so if you know how to output bytecode, and that&#8217;s not very hard, then you can run on the JVM.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1386">23:06</a>] So that&#8217;s essentially the idea. The harder part then is interop, that you say, okay, I run on the JVM. I have to make sense of all these <a href="https://en.wikipedia.org/wiki/Java">Java</a> libraries out there. So I have to somehow map a <a href="https://en.wikipedia.org/wiki/Java">Java</a> concept into a <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> concept so that the compiler can understand what it is and I can document what it is for <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> programmers that maybe don&#8217;t understand much <a href="https://en.wikipedia.org/wiki/Java">Java</a>. So that is sort of work that sort of goes more into the details.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1415">23:35</a>] That said, I mean, that&#8217;s not the only <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> platform, so it was born on the JVM, but now it also exists on <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and <a href="https://en.wikipedia.org/wiki/Node.js">Node.js</a>, on WASM and on <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> Native. So it&#8217;s really a multi-platform language.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1505">25:05</a>] Let&#8217;s say I write a <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> program. I have the source code, I have to compile it, and then I have some <a href="https://en.wikipedia.org/wiki/JVM_bytecode">Java bytecode</a>, and then I call the JVM and I pass in this bytecode, and it just works.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1519">25:19</a>] Yeah, exactly.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1521">25:21</a>] What would the advantage be of compiling <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> to <a href="https://en.wikipedia.org/wiki/JVM_bytecode">Java bytecode</a> instead of doing something like <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a> where you compile it all the way to machine code from source?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1533">25:33</a>] We have that too. So <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> Native compiles directly using LLVM to native code. The advantages of compiling to bytecode is essentially, again, the interop can use all the <a href="https://en.wikipedia.org/wiki/Java">Java</a> libraries, then the garbage collectors. So the JVM has actually very good garbage collectors, high-performance ones, and the runtime loading. So what I can do on the JVM is I can compile, let&#8217;s say, a one-line thing into a little <a href="https://en.wikipedia.org/wiki/JVM_bytecode">Java bytecode</a>, and I can immediately load that bytecode and execute that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1573">26:13</a>] And that makes it very easy, for instance, to have a REPL, because that&#8217;s how a REPL would work, a read-eval-print loop. So I would just generate a snippet of <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> code to bytecode, load it into the running JVM process, and it gets executed. And that&#8217;s much harder if you go to binaries; basically, binaries don&#8217;t really have a concept for that.</p><h3>26:33 &#8212; Why writing a compiler is hard</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1593">26:33</a>] You mentioned that you wrote a compiler for <a href="https://en.wikipedia.org/wiki/Java">Java</a> before <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, and I&#8217;ve heard that compilers are notoriously hard to build. Can you explain the hard parts and why they&#8217;re hard?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1606">26:46</a>] Compilers are very intricate because I think there are a lot of requirements on a compiler. So you have languages which are already complex artifacts. Then you have a thing called type inference, so you have a type system, but essentially you want the compiler to infer a lot of types that make sense. Because you as a programmer think, well, the compiler should know that, and it should, but actually for the compiler to do that is quite hard.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1641">27:21</a>] And then, of course, you also demand that it would generate very efficient code for your program, because again, the compiler should know that, right? So there are a lot of demands on that. Plus, there&#8217;s a demand that it should be very, very fast. And to square all these demands requires quite a bit of work, basically. There are also things that make it easier for compilers because they&#8217;re fundamentally deterministic programs.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1667">27:47</a>] So essentially, you run a compiler on a source, and you always get the same output, so it means they&#8217;re easier to debug. You can just replay things and repeat things. Whereas if I would have some cloud service or distributed application that has other nightmares, right? Every run is different, and how do you even figure out what goes wrong?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1687">28:07</a>] So when you wrote that compiler, I think Espresso for <a href="https://en.wikipedia.org/wiki/Java">Java</a>, how long did it take to write that by hand?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1696">28:16</a>] At the time, it took me about three months, I think. Not full time, maybe half time. Three months, half time, something like that. But that was a very simple compiler. So that&#8217;s the other thing with compilers: they typically start simple, and when you&#8217;re at the 10-year mark or 20-year mark, then there&#8217;s lots and lots of essentially other requirements to a compiler that make them more complex.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1717">28:37</a>] Were there libraries that you could rely on to kind of piece together components of it, or did you have to write everything from scratch?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1727">28:47</a>] There was a library to generate bytecode, I think. I think I used that in that version. Yeah, I did, yeah. So there was a library, essentially a high-level library, to just assemble bytecode and put them into the right format and things like that. But that was essentially it. The rest I wrote by hand.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1747">29:07</a>] That&#8217;s crazy because, well, I mean, these days a lot of people, even more so now, we&#8217;re thinking less about using AI and more about kind of writing the code. So that&#8217;s pretty impressive.</p><h3>29:19 &#8212; Why Twitter adopted Scala early on</h3><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1759">29:19</a>] I saw one of the big, I guess, moments for <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was that <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> adopted it, and I wanted to know the story behind that. How did they choose such an obscure language at that time?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1770">29:30</a>] It was very obscure at the time, that&#8217;s true. So I believe the story was <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> originally was written in <a href="https://en.wikipedia.org/wiki/Ruby">Ruby</a>, and it wasn&#8217;t very reliable because I believe it&#8217;s again a problem with garbage collector and memory management and things like that. And it was a small company at the time, so 25 people, including the ops people, and essentially the board and VC investors said you can&#8217;t go on like this, being unreliable like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1803">30:03</a>] You should do <a href="https://en.wikipedia.org/wiki/Java">Java</a>, because <a href="https://en.wikipedia.org/wiki/Java">Java</a> was sort of the solid choice at the time. But some of the engineers at <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a>, they wanted essentially something more fancy as a programming language. And there are some people who knew <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> and liked <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, but of course <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> didn&#8217;t run on the JVM. And then they found <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> and said, well, <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> is actually quite a lot like <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> and we can tell our investors that we do <a href="https://en.wikipedia.org/wiki/Java">Java</a> because it&#8217;s not a lie, we do JVM bytecode.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1833">30:33</a>] And that&#8217;s essentially what it comes down to. So that&#8217;s why they picked it. And they were quite the first. But once they picked it, sort of the floodgates opened because at that time they were a very interesting company and a lot of people admired them. So a lot of people followed and did the same thing. Mostly they came from dynamic languages, so <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> came from <a href="https://en.wikipedia.org/wiki/Ruby">Ruby</a>, others came from <a href="https://en.wikipedia.org/wiki/PHP">PHP</a> or <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, things like that.</p><h3>31:00 &#8212; How he believes AI will impact programming languages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1860">31:00</a>] I mentioned a little bit that AI is kind of generating a lot of code, and I wanted to ask you what you thought maybe the future of programming languages might look like if you speculated or drew it out further into the future. If AI is generating more of the code, how do you think that might affect the programming language ecosystem?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1882">31:22</a>] Yeah, I think right now we&#8217;re sort of in an existential crisis, right? So we have AI generating the code, but humans being asked impossible tasks, like to review all these mountains of code and things like that, which will never work. And at the same time, AI has also gotten extremely good at exploiting vulnerabilities in code, like we all heard of Fable and things like that, that you can&#8217;t use it anymore because it&#8217;s too dangerous. It will exploit things.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1913">31:53</a>] So we are at a moment where it&#8217;s essentially very dangerous that we lose control as humans of what actually happens here. And that&#8217;s a challenge that I think programming languages can help meet. And probably definitely not only programming languages, that&#8217;s not a silver bullet, but they definitely can help things. So I think one of the things is that the focus, if the code is AI-generated, then the focus has to go elsewhere.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1946">32:26</a>] And I think the focus will go to the interfaces and to the types. So I expect types will become a lot stronger and more precise than what we had. Because types are essentially the handle that we can make a contract between the human and the AI that the human can understand and that&#8217;s concise enough to be reviewed and that essentially the AI can keep to. We have to level our game quite a lot because right now, I mean, let&#8217;s face it, type systems are mostly recommendations. They&#8217;re mostly things that mostly hold, but not always. There are no guarantees because you can always have a cast or value, you use some dirty memory or.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=1993">33:13</a>] I mean there are a number of techniques to undermine the type systems, and we have to close all these holes from the beginning because once there is a hole, somebody can exploit it. So I think strong types, strong high-level types, will help. And then I think the other part is generally the programmer has to think much more about what the requirements are and what are essentially the high-level specifications and be able to leave the code to be generated by somebody else in confidence.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2030">33:50</a>] And I think we&#8217;re not quite there yet, but we have some ideas about how we could get there. So one technique that I believe we can use, and it has been around for a long time, but maybe its time has come now, is capabilities. Capabilities essentially were used in operating systems to give very fine-grained permissions to entities, users, programs, and things like that. And I believe that can be used also for agents and agentic AI to say, well, once we have agents, we have to give agents very precise and fine-grained capabilities for what they can do, and that lets us essentially be confident about what they will not be able to do.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2074">34:34</a>] They will not be able to leak my API keys or my email or do other things. So I think that&#8217;s an important part. And the existing languages are not there yet. I think <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> is halfway there. At least it&#8217;s there in essentially stuff we&#8217;re working on, which we have in the lab and we have released as an experimental feature. So I&#8217;m quite excited about that. The first thing that you have to do is definitely be memory safe.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2108">35:08</a>] So a language that essentially is not memory safe, that lets you essentially access undefined memory, is immediately out because you can&#8217;t guarantee anything. So that&#8217;s. In that sense, it&#8217;s good that there is a drive to use, let&#8217;s say, <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> as a memory-safe language that was even promoted by the American government, I believe. So that&#8217;s definitely a very useful drive. But I think you need a lot more, because you need much.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2136">35:36</a>] <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> talks essentially mostly or only about memory. You need to talk about a lot more things than memory. You need about essentially read permissions, write permissions, access to secrets, all these things that go beyond that. And you could say, okay, Martin, you&#8217;re totally unrealistic because all our software is written in <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a> and we will not be able to rewrite that. But I believe AIs can help there, right?</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2169">36:09</a>] So AIs are great in rewriting software. So if we know what to rewrite, too, I think we might be able to get there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2177">36:17</a>] You mentioned some of those experimental features in <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> that might have some sort of safety guarantees or signal capabilities. Can you explain what that might look like, or maybe give an example?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2190">36:30</a>] So a simple example would be, let&#8217;s say somebody gives me a file and I have access to the file, let&#8217;s say a log file or something like that. I have access to a file for a limited time, and then I need to close it. So typically I have an operation that essentially somebody passes a file to me, to an operation that my program provides, and the program does something with a file, and then the environment will close it.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2219">36:59</a>] But how do we make sure that I don&#8217;t hold on to the file after I gave it back to the environment, or after I pretended I&#8217;m finished with it? Because, hey, I have a file, I could have stored it in a variable, I could have stored it on the side, I could have gone back to it and done something with it. So capabilities help me prevent that because essentially I can say, okay, so this file is a capability.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2243">37:23</a>] And then I can further say, well, this capability can be used only in a limited scope, and the type system will make sure that the capability doesn&#8217;t escape. And the way we do that is that if a type refers to capabilities. So if I have a thing that I say, I give you back a lambda or a stream and it holds onto the file. So the stream holds onto the file in secret. In our language, that won&#8217;t be a secret anymore because the type has to declare that the thing I return does hold on to the file.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2276">37:56</a>] The file is a capability, and I can&#8217;t essentially hide capabilities I have access to in my type. I have to declare them. And that gives me essentially this control that then I can also enforce, to say, well, at this point you&#8217;re not allowed to have any capability because the type that I enforce you to have is a type that doesn&#8217;t hold capabilities. And that way I enforce with the type system something which previously hasn&#8217;t really been enforceable for memory safety.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2306">38:26</a>] It&#8217;s essentially the same thing with arenas, that I have an area where I allocate memory and then I want to get rid of it. I have to make sure I don&#8217;t have pointers pointing into it. And that&#8217;s exactly the same situation, and let&#8217;s say for accessing secrets again. So it&#8217;s a very common pattern that I say in certain situations I want to make sure that you don&#8217;t have, or that you only have a set of defined capabilities that I give you, and nothing else.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2337">38:57</a>] You mentioned memory safety is an absolute table stakes. What are the programming languages that you think of that are not memory safe? I know there&#8217;s <a href="https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B">C/C++</a>, but what are the other ones?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2348">39:08</a>] The big ones are C. I don&#8217;t know about. I think <a href="https://en.wikipedia.org/wiki/Zig_(programming_language)">Zig</a> or Nim or other low-level systems languages are not memory safe. So that was sort of in <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, the big achievement, that you can be a low-level systems language and be memory safe. Nobody sort of thought that was possible before <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> came. So that&#8217;s why I would think, I don&#8217;t want to say anything wrong, but I would think that essentially most other low-level systems languages would not be memory safe.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2380">39:40</a>] But you really need more than memory safety. You really need capability safety, that you say, when I essentially hang on to something, I can&#8217;t sort of forget capabilities to say I hang on to something and I just conveniently forget that I have access to that. And I can&#8217;t forge capabilities to say, well, if I need a capability, I just make one up. These two things need to be prevented. And that goes beyond memory safety.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2407">40:07</a>] But memory safety without memory safety, essentially you have nothing because you can fake everything.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2412">40:12</a>] A lot of programming language design and how we write code in the past is writing the source code so that it&#8217;s nice for humans to read. But if humans are no longer interacting with the code, what kind of things come to mind that we might not care as much about but are good for machines to read?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2434">40:34</a>] For instance, the first thing is maybe sometimes you want to read it, but you probably wouldn&#8217;t have written it. So easy to write is definitely not a big criterion anymore. Easy to read to some degree, yes, but that means you don&#8217;t need any of these sort of syntactic hacks, like to write plus plus in C or things like that. That&#8217;s easy to write, right? But I don&#8217;t, I mean I&#8217;m sure we will still have that, but it doesn&#8217;t really matter anymore.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2463">41:03</a>] I mean, whether I write an assignment in long form or with X, whatever. So I think these things won&#8217;t matter much less. The things that matter much more are, that continue to matter, and are essentially high-level ways to constrain and specify what my program should do. So constrain what it should not do, that&#8217;s one of the things. And also specify what it should do. And I mean some people say a golden age for formal verification because we can be very precise in our specifications, and our AI can actually not just furnish the program, but also the proof that the program actually meets the specification.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2508">41:48</a>] And I think that&#8217;s true. That&#8217;s really very exciting in a lot of areas. But in the large, the problem is you often don&#8217;t really have the formal specification, or it&#8217;s just as hard to write a formal specification than to write a program, or sometimes even harder. So that means that we will still live in a world where we specify things by natural language, by prompt to the agent, and the agent then will do the code.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2536">42:16</a>] But I imagine that also that will be a lot more formal in a way. So in a sense right now the prompts, I mean, it&#8217;s fantastic what they can do with the prompts, but then we throw away the prompt or it&#8217;s hidden in the chat history with the agent. So that&#8217;s really a shame. So I really should have the prompts as first-class values in my program, that I can say, well, essentially that&#8217;s what the program is about, and if I change the prompt, then the AI will know essentially what the incremental change was and change the program incrementally, these sort of things.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2571">42:51</a>] That&#8217;s another thing that I think we&#8217;ll see in future programming languages for agentic programming.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2577">42:57</a>] Maybe some kind of meta information about the program, almost like a Git blame, but like a prompt, I guess, blame of what generated that part of the code.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2591">43:11</a>] You should be able to keep that around and come back to it. Yeah. Also because you might want to change, right? You might want to say, well, now it&#8217;s exactly the same, but I want to change this little detail. But I don&#8217;t want the LLM nondeterministically to generate a new program because that way it might get a lot of other things wrong that I reviewed already. So there really should then be an incremental small change to the code, and that means I need to keep the prompt as a part of my program.</p><h3>43:40 &#8212; Will there be less engineers in ten years</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2620">43:40</a>] Ten years from now, do you think there will be more software engineers than today or less?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2626">43:46</a>] I think there will be less. And it will be a higher profession that essentially has higher standards. So it will be harder to become one. You will need to know quite a lot of logic and math to be competent at essentially keeping AI on the right track and things like that. It&#8217;s sort of, I can say, sort of like a control engineer for a factory where you don&#8217;t understand many things initially. It means that there must be more&#8212;you must be higher skilled than a factory worker of 50 years ago or something like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2667">44:27</a>] And I think the same will happen for software, where you will need fewer but more qualified people.</p><h3>44:34 &#8212; Top programming languages to learn to grow</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2674">44:34</a>] One thing that a lot of people recommend is to become better at programming: you should learn multiple languages. And I wanted to know, aside from <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, what are the top programming languages you would recommend people learn to expand their mind?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2689">44:49</a>] So definitely a systems language. I think to know how hardware works and how software links with hardware, I would learn a systems language. I&#8217;m sort of torn between C and <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> there because C has going for it that it&#8217;s very simple and very close to the metal. In <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, you have to learn a lot of abstractions, but on the other hand, it is memory safe. So I would say probably initially C, to figure out what these things are.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2717">45:17</a>] And then if you decide to become a systems programmer as a career, then you should switch to <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>. I guess that would be one thing. The other thing would be something more with a verification, improving background because that will be a lot more, will become a lot more important. So I would actually do a course in Lean or <a href="https://en.wikipedia.org/wiki/Rocq">Rocq</a>, one of these languages, to say we. We should be able to use AI or to.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2748">45:48</a>] To have to develop first an intuition what it means for a program to be correct, because I guess for most people have only a very, very fuzzy intuition for these things. And learning one of these languages would sharpen the mind. Yeah, so I think those. And then, of course, <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> too as a language that essentially sits right in the middle where fairly provable, strong types, high expressivity, these sort of things.</p><h3>46:18 &#8212; Top technical book recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2778">46:18</a>] When you think of a top technical book that you might recommend to people, does anything come to mind?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2783">46:23</a>] I got a lot out of Structure and Interpretation of Computer Programs. That was a book from the 90s. It was an intro text at MIT at the time, or 80s even, I think. And in fact, the courses I taught on Coursera and at EPFL are based quite a lot. Well, to some degree they&#8217;re based on that material. So, of course, not the same language and strong types instead of dynamically typed.</p><h3>46:51 &#8212; Why he chose academia instead of industry</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2811">46:51</a>] But still, when you look back on your career, why did you choose to work in academia instead of industry?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2817">46:57</a>] When I came to the end of my studies, I was asked to do a research project, and the project I picked, or the professor then asked me to modulate that a little bit. In the end, I found it super interesting to say I work on something that I don&#8217;t know whether it has a solution or not. So it might be that. The question is yes, it might be. The question is no, we have to do research, and that&#8217;s sort of how it started.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2844">47:24</a>] And then I believe the main advantage of being in academia is really the long term and being independent. So in the long term, I mean, in industry, I&#8217;m sure there are many periods, including now, where I could be paid 10 times what I&#8217;m being paid at university if I joined <a href="https://en.wikipedia.org/wiki/Google">Google</a> or any of the other companies in Silicon Valley. But then times also change, and sometimes essentially what you do is no longer relevant, and then it&#8217;s very easy to get fired. And then you say, no big deal, I can do something else.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2880">48:00</a>] But at a university you have essentially long-term tenure to do exactly what you want. So no manager, you can define your own research agenda. Of course you have to get some money. That is hard. You have to get grants and things like that. You have to convince students that what you do is the right thing. But that&#8217;s also very rewarding, working with students. So I think in the end I&#8217;m very happy that I took the career I took.</p><h3>48:28 &#8212; Reflecting on Scala</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2908">48:28</a>] What about looking back on <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>? You have so much experience there. What are some things that went well or things that didn&#8217;t go well, and maybe some learnings you could share?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2919">48:39</a>] <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was initially this experiment that we could combine object-oriented and functional programming. And technically that experiment was a big success. In terms of the ecosystem, it had a lot of challenges. I don&#8217;t know whether there was A Brief, Incomplete, and Mostly Wrong History of Programming Languages by James Iry, which is quite hilarious. And there&#8217;s a <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> entry there that says I discovered Reese&#8217;s Peanut Butter Cups.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2947">49:07</a>] And I had an idea to essentially have a language where you put in both object-oriented and functional, and then it ends with these pieces of both communities. And they promptly declared jihad. And at the beginning, it was very funny. But in retrospect, that&#8217;s to a large degree what happened. I mean, there were a lot of fights and cultural things like that, and it was a challenge. And I don&#8217;t know, in retrospect, I think it might have been more prudent to be more careful introducing functional features because that sort of was.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=2991">49:51</a>] We went in essentially quite complete. So essentially you can port most Haskell programs to <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, maybe that you find them a bit less attractive looking, but you can do it. And that brought in essentially different cultures. And it caused a big clash of cultures. So, for instance, in Go, Go was a lot maligned because they didn&#8217;t even have generics. Right. And it&#8217;s true, it was a ridiculous language, but by not having it initially you sort of form a culture that when you add it, nothing much will go wrong.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3030">50:30</a>] And people want to over-abstract or things like that. And by essentially going full hog into it with <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> from the start, we were not protected against that. And people did reach abstraction peaks and over-abstract it and things like that, and they still do it to this very day. So essentially having fancy abstractions is great, but you have to use them responsibly. And then it sort of clashes with the natural urge to just try out all these fancy things and do something that in the end maybe neither you nor nobody else understands very well.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3065">51:05</a>] And it could have just been a simple map or things like that. But yeah, why do I use that when it can be complicated?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3073">51:13</a>] Why would that upset someone?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3075">51:15</a>] Because of the library ecosystem. I think that the libraries are important because they sort of set the agenda, how you express your programs. And in particular the libraries that come from the Haskell side, they&#8217;re essentially monadic frameworks, and they require you to express your whole program as a monad, which is a concept that works well in functional programming. I have my reservations. I wouldn&#8217;t actually do that in my <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> programs.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3108">51:48</a>] But by having prominent libraries out there, it&#8217;s quite normative that people feel like being a <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> programmer, they&#8217;re required to program this way. And then, of course, there are opinion leaders and conferences and all these things that tell you how you should be writing your programs. But I would say I think most of the <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> community is actually very, very level-headed, and I think most of the advice is great.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3138">52:18</a>] But then, yeah, it&#8217;s still a challenge because it&#8217;s just too easy to. Mostly it&#8217;s not really the opinion leaders in the community, but let&#8217;s say it&#8217;s your boss. You have a small team in a company, your boss came from Haskell and says, &#8220;Hey, this is great. We do exactly like Haskell, the whole thing.&#8221; And then two years later the boss&#8217;s boss says, &#8220;No, nobody can understand this code. This code in the project is cancelled.&#8221;</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3166">52:46</a>] So these things happened in the industry. I&#8217;m not saying that all projects are like that, not at all. I mean, there are lots of really great success stories of <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a>, but that&#8217;s essentially a challenge in terms of techniques. I think, in the end, I wish we had been a bit less dependent on the JVM in the sense that we took a lot from <a href="https://en.wikipedia.org/wiki/Java">Java</a> that intuitively makes sense but, in the end, and was very important for interop, but in the end could have been done better.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3205">53:25</a>] So for instance, in <a href="https://en.wikipedia.org/wiki/Java">Java</a>, being an object-oriented language, you have universal methods like toString and equals and hashCode, and they&#8217;re defined for everything. And languages like <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> or Haskell, they&#8217;re no more discriminating. They have a thing called <a href="https://en.wikipedia.org/wiki/Type_class">type classes</a> where, essentially, at compile time, you tell exactly where you have equality and hashCode and these things, and it&#8217;s a bit more tedious to set these things up, but in the end it&#8217;s also safer.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3235">53:55</a>] So I wish we had that, and we couldn&#8217;t because we sort of adapted <a href="https://en.wikipedia.org/wiki/Java">Java</a>&#8216;s notion of what an object is, and it already came with all these things, which is sort of very convenient, but in the end caused friction.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3250">54:10</a>] If I went back to you at that time, when you were starting it, and I said, what do you think is going to happen 20 years from now? What would you have said?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3257">54:17</a>] I would probably have said that while it was a footnote in memory, because, I mean, let&#8217;s face it, you do a thing and initially we had maybe five users or something like that outside our group, something like that. Right. And you don&#8217;t really expect that it would change a lot. And so no, I think it was quite by surprise that this actually happened. And I think the reason why it happened was that at the time <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> was a good bridge between dynamic languages that are slow and sometimes they crash.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3302">55:02</a>] And solid languages, statically typed languages like <a href="https://en.wikipedia.org/wiki/Java">Java</a>, which were at the time quite cumbersome to write code in and there was a lot of ceremony and things like that. And <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> had inferred types, so you had types but you didn&#8217;t see them much. So it felt like a dynamic language, but it had also the solidity of essentially a good platform, and I think that&#8217;s what made it. And we&#8217;ve been copied a lot by a lot of other languages that put in these features then five or ten years later or things like that.</p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3333">55:33</a>] But at the time, it was sort of <a href="https://en.wikipedia.org/wiki/La_Scala">Scala</a> that was the first one that had it, and that&#8217;s sort of why it happened. But I wouldn&#8217;t have foreseen that by</p><h3>55:42 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3342">55:42</a>] No means looking back on your career, when you just graduated college and kind of started your career, knowing what you know today, what advice would you give your younger self?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3355">55:55</a>] I think to take risks, be adventurous, pay it out. So essentially don&#8217;t follow the mainstream. If you take a fancy to do something wild and crazy technically, take the time to do it and do it. But I mean, yeah, so that&#8217;s what I would. Of course, you still need to sort of stay on track somewhat, but I don&#8217;t want to exaggerate that either. But yeah, so basically be a bit nonconformist.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3388">56:28</a>] Awesome. Well, thank you so much for your time, Professor. I really appreciate it.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/LdN4sPWM-WY?t=3392">56:32</a>] Thank you, Ryan.</p>]]></content:encoded></item><item><title><![CDATA[Sergey Levine: Current State of Humanoid Robotics, China & Future Predictions]]></title><description><![CDATA[Sergey Levine is one of the world&#8217;s top robotics researchers and co-founder of Physical Intelligence.]]></description><link>https://www.developing.dev/p/sergey-levine-current-state-of-humanoid</link><guid isPermaLink="false">https://www.developing.dev/p/sergey-levine-current-state-of-humanoid</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 24 Aug 2026 13:05:33 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212043203/2d07767bb4579d653af29fa9f397aa7e.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/sergey-levine-5a31a24/">Sergey Levine</a> is one of the world&#8217;s top robotics researchers and co-founder of Physical Intelligence. Recently I talked to him to dig into where we are today with humanoid robotics.</p><p>He also had a lot to say about the Chinese robotics ecosystem. I asked him questions about that and about competitors such as &#8220;what if Anthropic and OpenAI started their own robotics labs?&#8221;</p><p>We went over a bunch of questions about predictions on future timelines. I asked him what a concrete roadmap might look like so we could get a sense of how deployment will happen. </p><p>Hope this episode is helpful in learning where we are today and where we will be when it comes to humanoid robotics.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/9OSbaPjv0Rc">YouTube</a>, <a href="https://open.spotify.com/episode/19amTrz0o0NECMeEtblMq1?si=rnFMXeqXQ9S-ZA8Yku6Zqg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-9OSbaPjv0Rc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;9OSbaPjv0Rc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/9OSbaPjv0Rc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/212043203/0037-where-are-we-today">00:37 - Where are we today</a></p><p><a href="https://www.developing.dev/i/212043203/0420-most-surprising-capabilities-so-far">04:20 - Most surprising capabilities so far</a></p><p><a href="https://www.developing.dev/i/212043203/0703-the-most-inspiring-real-world-robotics">07:03 - The most inspiring real world robotics</a></p><p><a href="https://www.developing.dev/i/212043203/0836-if-openai-or-anthropic-got-into-robotics">08:36 - If OpenAI or Anthropic got into robotics</a></p><p><a href="https://www.developing.dev/i/212043203/1022-chinese-robotics">10:22 - Chinese robotics</a></p><p><a href="https://www.developing.dev/i/212043203/1315-will-one-lab-breakout-from-the-rest">13:15 - Will one lab breakout from the rest</a></p><p><a href="https://www.developing.dev/i/212043203/1659-thoughts-on-a-concrete-roadmap">16:59 - Thoughts on a concrete roadmap</a></p><p><a href="https://www.developing.dev/i/212043203/2103-generalization-and-demonstrating-it">21:03 - Generalization and demonstrating it</a></p><p><a href="https://www.developing.dev/i/212043203/2604-types-of-data-and-which-is-best-for-robotics">26:04 - Types of data and which is best for robotics</a></p><p><a href="https://www.developing.dev/i/212043203/3434-why-humanoid-robotics-differs-from-waymo">34:34 - Why humanoid robotics differs from Waymo</a></p><p><a href="https://www.developing.dev/i/212043203/3710-if-humanoid-robotics-failed-here-is-why">37:10 - If humanoid robotics failed here is why</a></p><p><a href="https://www.developing.dev/i/212043203/3955-are-there-hot-take-modeling-architectures-in-robotics">39:55 - Are there hot take modeling architectures in robotics</a></p><p><a href="https://www.developing.dev/i/212043203/4205-thoughts-on-ai-safety-in-robotics">42:05 - Thoughts on AI safety in robotics</a></p><p><a href="https://www.developing.dev/i/212043203/4644-top-robotics-research-paper-recommendation">46:44 - Top robotics research paper recommendation</a></p><p><a href="https://www.developing.dev/i/212043203/4935-why-is-boston-dynamics-less-top-of-mind">49:35 - Why is Boston Dynamics less top of mind</a></p><p><a href="https://www.developing.dev/i/212043203/5347-advice-for-his-younger-self">53:47 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:37 &#8212; Where are we today</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=37">00:37</a>] <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> have caused unprecedented impact and investment in that area. <a href="https://en.wikipedia.org/wiki/Physical_artificial_intelligence">Physical AI</a> and humanoid robots could potentially be even bigger. And so I wanted to ask you today about where are we today with humanoid robotics, and how do you foresee this technology actually being deployed into the world?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=61">01:01</a>] I guess with machine learning, what we&#8217;ve learned over the last few years, or the last decade rather, is that it works when you do it at scale. And this is very obvious now, but it wasn&#8217;t always obvious. But there&#8217;s a caveat, which is you have to scale the right thing. And initially when people started working on models for language, for example, the dominant design was LSTMs. Some people remember what those are.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=89">01:29</a>] They were kind of okay; they were a lot better than what came before that, but they didn&#8217;t really scale as well. And then the big thing with transformers was not that transformers were somehow particularly mathematically elegant or anything like that. It&#8217;s just that they scaled better. So they were easier to train on very large amounts of data with lots of parameters. So the technology proceeds in phases.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=108">01:48</a>] First you figure out what you can scale, basically what is the scalable technology. And then you pour on a lot more of kind of an industrial-scale effort and adding lots of data, adding to model size. And that&#8217;s when kind of the magic happens. So when we&#8217;re doing kind of more fundamental technology development, the key is to understand what are those scalable levers. So figure out the design, figure out roughly the mixtures, and that by itself does something pretty cool.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=134">02:14</a>] But that&#8217;s not the thing that actually changes the world. It&#8217;s like when you start pulling that lever that things actually change. So with <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a>, when the first GPT models came out with <a href="https://en.wikipedia.org/wiki/GPT-2">GPT-2</a>, it did some stuff, but it was sort of like a parlor trick, so you could get it to synthesize a story about unicorns in Peru or something. And it was coherent English, but it wasn&#8217;t a thing that would solve lots of real-world problems.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=159">02:39</a>] But the folks that worked on this kind of recognized that, hey, there&#8217;s something kind of magical that happens because as you add more data and you make the model bigger, this stuff gets more coherent and more effective. So they could see that if we do a lot more of that, then it&#8217;ll become a lot more powerful. So to come back to your question, what I would say about robotics is that it&#8217;s not in the <a href="https://en.wikipedia.org/wiki/GPT-4">GPT-4</a> to <a href="https://en.wikipedia.org/wiki/GPT-5">GPT-5</a> stage, where it&#8217;s an industrial-scale effort to kind of make the model bigger and get more capability out of it.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=190">03:10</a>] It&#8217;s in that stage where we&#8217;re establishing the fundamental technologies. And because of that, what one should expect to see right now is not necessarily that each month the model gets bigger and more powerful by some predictable kind of scaling curve. It&#8217;s that the scaling properties themselves are evolving as we develop the right technologies. So to bring this back to something closer to reality.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=216">03:36</a>] I&#8217;m very happy with the demos that we&#8217;re doing here at <a href="https://en.wikipedia.org/wiki/Theory_of_multiple_intelligences">Physical Intelligence</a>. And I think that a lot of the results that other people are coming out with are really cool. But these are, to put them in context, we should not expect these to be the things that are actually illustrating the power of scale. We should expect them to be developing the fundamental technologies that will be scaled up after that.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=235">03:55</a>] So where I think we&#8217;re at now is that we&#8217;re actually kind of getting all those puzzle pieces in place. And I think it&#8217;s actually very close. I think a lot of the puzzle pieces are falling in place, but it&#8217;s not like. What makes it so hard to prognosticate about where the technology is going to go is that it&#8217;s not yet at that predictable scaling stage. It&#8217;s at the stage where we&#8217;re figuring out the puzzle pieces, which I think is really exciting.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=255">04:15</a>] But it means that it&#8217;s also very hard to foresee sort of what the coefficients on that will be.</p><h3>04:20 &#8212; Most surprising capabilities so far</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=260">04:20</a>] What are the most astonishing emergent capabilities you&#8217;ve seen so far?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=265">04:25</a>] That is the thing that is the most fun. And certainly we&#8217;ve seen a lot more of that happening as we progress. In the very beginning, it was kind of little things, but they were kind of magical because in robotics, basically prior to 2024, the stuff never happened. So the little things that happen, and this was maybe at this point about two years back, we would see things like, okay, we train our policy for folding laundry and it takes individual shirts out of the hamper and tries to fold them.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=295">04:55</a>] And then one very vivid memory I have in late 2024 is we were watching one of the evals, and it takes out two shirts at the same time. And I&#8217;m watching this, I&#8217;m like, okay, it&#8217;s done for. There&#8217;s no way you can possibly do this. And then it puts the two shirts on the table, disentangles them, puts one of them back, and then starts folding the other one. It&#8217;s like, wow, okay. In retrospect, you can do some detective work and figure out where it got that from, some piece of training data.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=322">05:22</a>] But that was one of those moments where I don&#8217;t think anybody watching that eval thought that the robot was going to do this. It&#8217;s like a little thing. It&#8217;s exhibiting the common sense you expect people to have. But actually, the thing that I find more interesting recently is some of the mistakes, because one of the things that&#8217;s pretty remarkable about <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> is that once they got good enough, even the mistakes kind of made sense in the sense that they weren&#8217;t like crazy mistakes where just outputs, like, zzz all the time.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=350">05:50</a>] But their mistakes are sort of semantically sensible. We had an evaluation last year for &#960;0.5 where the robot was cleaning up a kitchen, and it&#8217;s told, like, put away all the utensils. There are some spoons, spatulas, et cetera. And it tries to open the drawer where it thinks the silverware goes, and it can&#8217;t really get the drawer open, so it slides over and opens the oven, which is right next to it, and then starts putting this stuff in the oven.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=376">06:16</a>] It&#8217;s like, you can sort of imagine that if you ask a child to clean stuff up and put it away, they might decide to do that because, okay, it&#8217;s like a container, and you can put stuff there, and nobody sees it. Another experiment we had is washing all the plates. So it would pick up the plates, wash them with a sponge, and put them on the drying rack. And this was an experiment on memory, because it has to keep track of everything that it&#8217;s doing and has a scratch pad of memory where it&#8217;s writing down, hey, I had three plates.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=400">06:40</a>] I cleaned the gray one, I cleaned the green one. And then it drops one of them on the floor and drives the base over so you can&#8217;t see. And I was like, okay, I&#8217;ve cleaned the gray plate. It&#8217;s done. So, I mean, obviously these are not the things that we want to see. But it&#8217;s kind of interesting that some of the mistakes, they&#8217;re almost like what you would associate with a child trying to do the task.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=420">07:00</a>] So now it just needs to grow up.</p><h3>07:03 &#8212; The most inspiring real world robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=423">07:03</a>] You mentioned the advancements. They&#8217;re kind of happening all over the industry. And I know there are a lot of people building humanoid robots. There&#8217;s Figure AI, there&#8217;s Tesla, and there&#8217;s many other competitors, and you have the expertise of what is hard and what isn&#8217;t. When you look at the competitors, has there been any advancement or achievement where you think, oh, that&#8217;s really admirable and that&#8217;s impressive?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=449">07:29</a>] I actually think that one of the most inspiring things to me in the industry is to see the kind of takeoff that autonomous driving systems have had. Because one of the criticisms that is sometimes leveled against robotics researchers is, it&#8217;s like nuclear fusion. It&#8217;s the technology of the future, but it&#8217;s always in the future. But that&#8217;s what people said about autonomous driving, too. And now we&#8217;re in San Francisco.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=473">07:53</a>] You can go outside and you can take a Waymo, and it will actually take you to your destination. And there was no driver sitting there. So, without getting too much into the technical details, I think what&#8217;s really inspiring about that is just this case in point that, yes, you can actually have one of these technologies of the future, and it actually does land. And I think it&#8217;s not an accident that it&#8217;s landing now in the mid-2020s, because a lot of the puzzle pieces for large-scale ML are getting to the level where we can put them together with actual physical systems.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=504">08:24</a>] And there&#8217;s a lot of differences between driving and robotic manipulation, of course. But I think that the illustration that, yeah, we can actually land learning-based technologies in the real physical world, I think that&#8217;s really inspiring.</p><h3>08:36 &#8212; If OpenAI or Anthropic got into robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=516">08:36</a>] If <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a> or <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> started investing more heavily into robotics, how do you think that would impact the industry? Do you think competitors would be worried about that?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=529">08:49</a>] Robotics is an area where, to be fair, the ecosystem hasn&#8217;t been as healthy as it has in other areas of machine learning. And what I mean by that is that computer vision and NLP are things that sort of lend themselves naturally to a machine-learning-based ecosystem, because there&#8217;s freely available data. People kind of have a general acceptance that they&#8217;re going to be using learning. There aren&#8217;t concerns as severe concerns about safety, at least physical safety.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=560">09:20</a>] Right. You know, people rightfully are concerned about AI safety, but it&#8217;s not the same as like a physical device causing some physical harm. And because of that, I think it&#8217;s a bit easier to spin up a very serious large-scale ML effort in those areas. Robotics is not like that. Robotics traditionally is not a discipline that really embraces sharing of data, for example. So I think that the more activity there is around learning and robotics, the more I think it&#8217;ll shift people&#8217;s thinking toward this kind of future where we accept that robots will be controlled by learned models, not by hand-designed controllers, that there will be data, that data will need to be shared because there&#8217;s no way that somebody can build out a true foundation model in a single vertical and basically kind of shift the entire thinking around robotics to look more like how we think about vision and NLP as opposed to traditional factory automation.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=612">10:12</a>] So I think in that sense, much as I&#8217;m proud of the work that we&#8217;re doing at <a href="https://en.wikipedia.org/wiki/Theory_of_multiple_intelligences">Physical Intelligence</a>, I think it&#8217;ll take more than one company to kind of shift everyone&#8217;s thinking in that direction.</p><h3>10:22 &#8212; Chinese robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=622">10:22</a>] As a bystander, I see on <a href="https://en.wikipedia.org/wiki/X_(social_network)">Twitter</a> these really impressive demonstrations of Chinese robotics. It almost feels like they&#8217;re ahead in some sense, but I don&#8217;t have that deep domain expertise. So I was curious about your thoughts on what you think of China&#8217;s robotics, and are they further along?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=645">10:45</a>] I think that one thing that is very useful and constructive for us to do, those of us that work on these things in the United States and in Europe, is to ask what is the lesson to learn. And to me, one lesson is that it&#8217;s important to have a healthy ecosystem. And ecosystem means that there should be obviously good researchers, good engineers working on these things. There should be healthy open source.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=678">11:18</a>] But it also means that the different industries that contribute to robotics need to individually be very healthy. And those industries are not just the computer science, ML, and model building stuff. It&#8217;s also supply chains, manufacturing hardware, R&amp;D, these are all very important things. And aspects of those things are things that the United States does quite well. Other aspects of them are things where the United States has sort of let things go a little bit.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=712">11:52</a>] And I think that what we should do is we should look at what&#8217;s going on in the world, look at some of the excellent results that Chinese labs are doing, that labs in other countries are doing, and we should take away that lesson that we should strive to build a healthier ecosystem. And that means investing in all the different facets that contribute to this. So I&#8217;m not much of a business person, I&#8217;m not much of an investment person, so I can&#8217;t claim to know how to do this, but I think that it&#8217;s important to sort of fully embrace that this is a holistic thing and not something where we can do just one piece of it and outsource everything else.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=745">12:25</a>] Basically.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=747">12:27</a>] Of all the pieces of the ecosystem, are there parts of it? As a robotics researcher in the US, if it was better, that would have the biggest impact on the advancement of robotics.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=759">12:39</a>] Yeah, I think certainly availability of reliable, low-cost hardware is a big deal. And right now, I mean, certainly for hardware used for robotics research, a lot of that does come from China. And it&#8217;s good, it&#8217;s relatively inexpensive, it&#8217;s of high quality, and meets the standards that people generally need. But it would be awfully nice to be able to source all that domestically as well. And I don&#8217;t think there&#8217;s anything impossible about that.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=787">13:07</a>] I think it&#8217;s just a matter of embracing the fact that that entire ecosystem needs to be supported rather than just one piece of it.</p><h3>13:15 &#8212; Will one lab breakout from the rest</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=795">13:15</a>] You can imagine that the lab that is first to get to scale will be the first to break out. There may be this exponential growth effect where once you deploy, deployment helps you grow faster, and growing faster helps you deploy, and this flywheel. And so do you believe that that will happen in this industry, where one lab, whether in the US or China, will hit some breakout point and they&#8217;ll kind of jump everyone else?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=827">13:47</a>] There is a lot of truth to that. I think that there&#8217;s an important detail to keep in mind. The detail is that you kind of have to scale the right thing. So I think that basically the statement that the way that I would phrase this is having an effective positive feedback loop where more deployed robots translates to more model capability. That&#8217;s kind of the key, and that makes total sense. The trick is that there are lots of ways that it can be done wrong.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=863">14:23</a>] So I&#8217;ll tell you a few obvious ones. One obvious one is let&#8217;s say that I&#8217;m a car company and I have a robotic arm that is welding cars, and it&#8217;s there on the assembly line, and it welds cars every day, and it gets a million welds every month, and so on. If I just use that as my data flywheel, I&#8217;m unlikely to get something more capable than a robot that welds cars. So since the robot is already there and it&#8217;s already welding cars, presumably the marginal improvement for that is not all.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=894">14:54</a>] That valuable. So that&#8217;s an example of how you kind of have to scale the right thing, because data is not quite as fungible. It&#8217;s not like electricity or oil. You can&#8217;t just buy more of it. It has to be heterogeneous. So I think that basically it&#8217;s right, but you have to scale the right thing with the right technology and the right kind of source of diverse learning, like data is more like an education program for your robot than it is a fungible commodity.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=919">15:19</a>] If you think about that data flywheel and, I guess, putting out the proper platform that would generate data that matters, that would then improve that platform, when do you foresee this kind of happening in the world?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=934">15:34</a>] I think, on the technology side, things are advancing very rapidly toward that. And I think that maybe a particular balancing act to strike that would significantly determine that timeline is how much structure somebody&#8217;s willing to admit. So on the one extreme, you could imagine jumping straight to fully unstructured deployment domains like home robots. And that could be really exciting because then you have a lot of diversity right off the bat, but the bar is a lot higher to be effective enough in that domain to be safe enough.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=964">16:04</a>] Safety is a much bigger issue there because it&#8217;s around people in their home. On the other extreme, you could imagine much more structured tasks, maybe not quite the welding robot, but sort of like maybe the robot down the hall that does something more unstructured. And that could be a much easier domain in the sense that there&#8217;s much less safety concerns because it might be around trained humans, the task might be more predictable, and so on, but the marginal value of each bit of data you get in that domain is lower because there&#8217;s less variety.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=996">16:36</a>] So it&#8217;s like you could take off earlier, but with a smaller slope, or later, but with a larger slope. And you kind of have to calibrate that. But my sense is that whichever end of that extreme we&#8217;re talking about, it&#8217;s in the single-digit years rather than the double-digit years at this point. So it could be that the more structured ones might be happening now or next year. The less structured ones might be a few more years out, but probably not like a decade.</p><h3>16:59 &#8212; Thoughts on a concrete roadmap</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1019">16:59</a>] A lot of people talk about timelines, and it&#8217;s kind of more nebulous. And I&#8217;m wondering if you were to put down, I guess, somewhat of a roadmap, but not so concrete, but just milestones toward that north star of a robot in my home doing.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1036">17:16</a>] Yeah, that&#8217;s a really good question. So maybe to preface this answer, I think that there&#8217;s one thing that is very important to say about robotics that is very easy to miss from the hype cycle and the demos that people put out, which is that the hard thing in robotics was always generalization. But when somebody shows a demonstration of their system, the demonstration alone usually doesn&#8217;t make it clear what level of generalization is being shown.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1068">17:48</a>] The highly acrobatic robot demos, for example, are really exciting to look at. But typically, if it&#8217;s something that is a little bit more staged, sometimes literally on stage, if it&#8217;s a show that is obviously rehearsed, and that&#8217;s okay, because it&#8217;s literally just a show, but it&#8217;s not the same as doing a task every time reliably in any home. And the generalization piece often doesn&#8217;t look that impressive when viewed in isolation, because generalization is sort of a property of many trials, not of one trial.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1100">18:20</a>] So you might see the robot doing something fairly mundane and unimpressive. But what&#8217;s exciting about it is that it&#8217;s doing it with an object that it&#8217;s never seen before in an environment that has never been tested before. And that is actually harder than doing an acrobatic backflip that it&#8217;s practiced millions of times. So with that said, my answer to your question is that the roadmap is all about both achieving better generalization and kind of that second-order effect of having a mechanism to get more generalization as you generalize.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1133">18:53</a>] So one of the steps on that roadmap is to have a very concrete demonstration of a robotic system that gets better with autonomous experience that is collected in a setting that it wasn&#8217;t originally trained for. So I do whatever I do in the backend in the lab and whatever. I get my model, I get my adaptation algorithm, I put it in a new setting. Maybe it&#8217;s a home, maybe it&#8217;s a factory, whatever it is, something where it&#8217;s doing something real, and it does okay.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1164">19:24</a>] But then over time, it gets better and better. And it gets better and better to the point where it reaches sort of practically relevant levels of robustness without sort of capping out at like 50%. I think that would be a major milestone, because now that says, okay, if this is truly an automated process, it&#8217;s improving, it&#8217;s getting better, even as it&#8217;s collecting useful experience. Now, I can take it and I can put it in lots of different domains, collect useful experience, do something that people actually want, and it&#8217;ll improve the model.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1191">19:51</a>] So that, I think, is a really major step on that road. I think another really major step is to demonstrate a very concrete and practically useful way to transfer knowledge, to transfer common sense, to achieve robustness. So that&#8217;s like kind of the other scenario we don&#8217;t get to practice. If you&#8217;re driving your car on the road and you see a fire truck and you see a bunch of traffic cones, even if you&#8217;ve never been in that situation before, your common sense tells you, like, hey, I should slow down.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1220">20:20</a>] Maybe I don&#8217;t know how to react optimally, but I shouldn&#8217;t just barrel on through the cones and upset the firefighters and so on. So that&#8217;s common sense. And if you can apply that common sense to effectively recover from unexpected situations, like when the robot put the spatula in the oven, it should probably open it up, take it out. It knows that this is not the thing you do semantically.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1240">20:40</a>] That, I think, is another important step because that tells us that we can use common sense to fix mistakes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1246">20:46</a>] I could imagine a more narrowly scoped robot. Let&#8217;s say it&#8217;s a humanoid robot. That&#8217;s one step in assembly line, in manufacturing. And my understanding, you&#8217;re less interested in that because that&#8217;s not really a step toward that general intelligence.</p><h3>21:03 &#8212; Generalization and demonstrating it</h3><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1263">21:03</a>] What I would say, it&#8217;s not that I&#8217;m less interested in it, it&#8217;s that I think that the real world has these leaky abstractions that make that kind of stuff a lot more complex than it seems. Let me try to explain this with an analogy. So in the &#8216;90s, when people really started working kind of full steam autonomous driving, there was this idea that we could avoid a lot of the hard problems by kind of instrumenting the environment a little bit.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1291">21:31</a>] Like, we&#8217;ll have magnetic sensors along the highway and so on, and cars will have a little sensor on them and a little transmitter so they can tell where each of the cars are. Kind of like the way you do it with aircraft, basically. And then people thought, well, we don&#8217;t really need fancy AI. We&#8217;ll just have these sensors and it&#8217;ll just work. And that just basically didn&#8217;t go anywhere because the real world has so many messy exceptions and special cases that even if 99% of the time, the magnetic sensors and all that stuff just allow the car to drive the 1%, when someone steps in the middle of the road or there&#8217;s a piece of trash or whatever, it just messes everything.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1325">22:05</a>] Up. So the thing that actually worked was when Waymo said, hey, we&#8217;re going to not try to avoid the hard problem. We&#8217;re going to actually try to deploy our cars not in the middle of nowhere, but in San Francisco, very messy, and let&#8217;s just deal with it head-on. And that allowed that community to make progress. And I think robotic manipulation is going to be the same way, that past the fully structured world of the factory, if you want to go even a little bit outside of that, even if 99% of the time it&#8217;s all pretty straightforward, that 1% when something weird happens means that you really need the full scope of the problem basically to be addressed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1361">22:41</a>] This kind of reminds me, I remember Figure AI had this demo where they live-streamed the robot sorting packages. When you watch that demo, does it demonstrate generalization in your opinion?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1374">22:54</a>] Yeah, I think it does. And by the way, this is something that I find very encouraging is that, like I mentioned, it&#8217;s hard to show generalization in a video, and it&#8217;s clear that lots of people are thinking about that and are thinking about how do you present something that somebody can watch and take in kind of at a glance what generalization is. I think you can see a lot of creative steps towards that.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1396">23:16</a>] Live demos, these kind of really long time lapses. I think it&#8217;s a great idea, and I think that&#8217;s a really nice way to move towards elevating the importance of generalization in people&#8217;s consciousness. When we were working on the &#960;*0.6 project, the RL project that we did late last year, we wanted to do some longer horizon experiments. In some cases it&#8217;s obvious. We had our robot assembling boxes at Dandelion Chocolate Factory.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1426">23:46</a>] So they&#8217;re like, it&#8217;s an actual chocolate factory, so they need the boxes. So we ran it for several days. But we had this coffee task, which was the robot using an espresso machine to make espresso. So what we did is we ran it for 13 hours making espresso drinks. And we are, we try to be very, I guess, environmentally conscious about this. So we didn&#8217;t want to throw out the coffee.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1445">24:05</a>] So after 13 hours, everyone in the office was a little wiry because someone has to drink the coffee. But it ran for 13 hours, and it was pretty cool. It screwed up a few times. It&#8217;ll spill all the coffee grounds, and then it needs to go and get a cloth and wipe it down, but it does it and nothing exploded. Thirteen hours went by. Probably the most negative consequence was loss of sleep from too much caffeination.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1469">24:29</a>] But when it spills the coffee grounds and cleans it up, it did that by itself.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1474">24:34</a>] Well, so the way that that experiment was done is that there is a high-level prompting. So roughly the prompt is updated maybe every five minutes or so in between semantically coherent tasks. So you tell it, like, make espresso, clean up the machine, et cetera. So those steps, the actual clean up the machine, was commanded by a person. In principle, we could automate that. In fact, one of the things we&#8217;re spending a lot of effort now on is improving our high-level policy that does those commands.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1499">24:59</a>] But for that experiment, every five minutes, somebody basically updates what it&#8217;s being asked to do, the way we intended. It was like the commands would be like, if you go to an actual coffee shop, you say, &#8220;Oh, I want a latte, I want an espresso.&#8221; That was supposed to be the prompt. Except then you also have to tell it, &#8220;I want you to clean it up before you do the next one.&#8221;</p><h3>26:04 &#8212; Types of data and which is best for robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1564">26:04</a>] It sounds like on the way to generalizing data is a very important part of that.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1570">26:10</a>] And I was reading there are different types of data. There&#8217;s simulated data you could collect, like physical interactive data. And I want to hear your take on what&#8217;s the best data to get, what&#8217;s the worst, and what are the pros and cons.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1585">26:25</a>] This is, by the way, a question that is, I guess, quite. There&#8217;s a lot of discussion in the robotics community about this question. And some people have very opposite opinions on it. My own take on this is that a lot of different data sources are easier for the model to internalize if it can ground them in a thorough physical understanding of the world. So let me try to explain what I mean with a few examples.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1613">26:53</a>] If you want to learn to fly an airplane, you will probably use a simulator, at least during part of your training. But the simulator makes a lot of sense to you, because when you start using the simulator, you have a lot of world knowledge that you can use to ground what&#8217;s going on. You know that when you are using the flight simulator to learn how to fly the airplane, you&#8217;re not just playing a video game.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1635">27:15</a>] You&#8217;re trying to acquire knowledge that you will then use with a real airplane. And you understand that there&#8217;s sort of an abstraction there. Same thing if you&#8217;re playing a really cartoony Atari game or something. You know that all the symbols on the screen, you can sort of connect them to things that you&#8217;ve experienced in your life, and you can make an analogy there. So a lot of that, even though it kind of seems like these simulated environments reflect aspects of the real world, to us, they make a lot of sense because we kind of bring to bear a lot of our own prior experience, and we fill in the blanks that the simulation has.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1672">27:52</a>] And also, if you want to use data yourself as a person of somebody else doing something, if you watch someone, let&#8217;s say, cooking a meal, even though you don&#8217;t experience every movement they&#8217;re experiencing, you have a lot of that knowledge that you bring to bear. And you&#8217;re like, okay, I see they&#8217;re picking up the salt shaker. I&#8217;ve put salt on things before, so I kind of roughly know what&#8217;s going on there.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1693">28:13</a>] And I can file it away at this level of abstraction of add salt without having to figure out all their muscle movements. So my point with this is that once you have that understanding of how you do things physically with your own body and how you experience the physical world now, all these other sources of knowledge can be connected up to it, because that foundation you get from your experience helps you ground everything.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1714">28:34</a>] So where I&#8217;m going with this is that if we have a robotic foundation model that is trained on lots of real embodied data, that provides that grounding, and it might actually be much better able to absorb other sources of knowledge. And this is actually a little bit upside down relative to how some people think about it, because it&#8217;s very tempting looking at the success of Internet data for <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> to say, well, maybe we should do the opposite.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1736">28:56</a>] Maybe we should start with <a href="https://en.wikipedia.org/wiki/YouTube">YouTube</a> videos and then put robot data on top of that. But I think it&#8217;s actually the other way around, and I even have a little bit of evidence for this. So my colleague Suraj Nair, together with <a href="https://en.wikipedia.org/wiki/Simar">Simar</a> from Georgia Tech, they had a project together a while back where they took our robot foundation model and they added human video data. But they didn&#8217;t start with human video data.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1762">29:22</a>] They actually started with a model trained on robot data and then added video data on top of it. And what they did is they looked at the representations inside the model. Basically, how does the model represent human experience versus robot experience? And they found that if you use a small model with a small amount of robot data, predictably the human experience and the robot experience are fully separated, meaning that the feature representations are different.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1782">29:42</a>] But if you train on lots of robot data from lots of different robots, then the features are grouped much more by what task is being done rather than by whether it&#8217;s a human or a robot. And when we looked at the feature plots, it was just mind-boggling because literally when you crank up the amount of robot data to 100%, they just line up perfectly. You do this t-SNE embedding, you see the shapes of the embeddings, and it&#8217;s just all task identity and minimal sensitivity to embodiment.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1810">30:10</a>] And to me, that&#8217;s kind of mind-blowing because the base model wasn&#8217;t trained on any human data at all. But once you start adding human data, it represents it exactly the same way. And I think that&#8217;s really exciting. And I think that to me is one of the strongest indicators that if you have that good foundation of robot experience, you can put everything else on top of it. It&#8217;s actually better at absorbing that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1829">30:29</a>] Is it important that that base model has data that was collected using that specific set of motors, specific set of joints?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1839">30:39</a>] So far we&#8217;ve obviously put a lot of effort into cross-embodiment models that can handle many different robot types. But generally, you do need data of the robot you&#8217;re going to be deploying on to get good performance. So kind of the metric of generalization there is not can you zero-shot a new robot, but it&#8217;s mostly can you get away with less experience from the new robot and transfer skills from other robots?</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1862">31:02</a>] Okay, so that&#8217;s the current state of things. Now there is a little bit of a surprisingly positive read on that, which is even though you need data from these robots, the amount of special stuff that the model is doing is kind of minimal. When we started doing all this, I had a big long list of all the cool research I wanted to do to better accommodate different morphologies. Can you factorize the model&#8217;s representation in some way so that there&#8217;s a six degree freedom arm head and a seven degree head and a gripper head, et cetera?</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1892">31:32</a>] We didn&#8217;t do any of that. The model just outputs a big vector of numbers. If the robot has fewer degrees of freedom than the number it outputs, it just zero-pads it. There&#8217;s just nothing fancy, and that&#8217;s it. And then it just trains on all the robots and outputs the correct actions based on what is seen through the camera. But now, to your point about whether you can handle new robots.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1915">31:55</a>] So far, the thing that we focused on, and I think this is showing some promise, is being able to transfer skills between robots. And this is actually where the particular choices in how the model works seem to matter. For example, you can have a model that does some intermediate thinking, and that thinking can be done in different modalities. So you can think in text, and thinking in text is really good for transferring high-level behavioral structure.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1947">32:27</a>] So that&#8217;s basically how you understand that, hey, if I want to clean the kitchen and put away the silverware, first open the drawer. That&#8217;s kind of a semantic inference. And you can transfer that very well because obviously that&#8217;s largely agnostic to any embodiment or anything like that. But even lower level things can be transferred if you use the right representation. So one experiment we did is we had a thinking stage that is expressed in images, where you basically dream up an image of the next milestone in the task.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=1975">32:55</a>] And with that, we can actually get a robot, the UR5 robot, to fold a T-shirt, even though we didn&#8217;t have any T-shirt folding data on the UR5. Because while getting the arm motions correct is very hard, because the robot basically requires very different joint angles to do the task, cooking up an image of what it looks like for it to fold a shirt is not that hard because you&#8217;ve seen the robot arm in all sorts of different poses, you&#8217;ve seen the shirt in all different stages of being folded and unfolded, roughly where it should hold it.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2001">33:21</a>] So getting a good generative model to cook up that image is pretty straightforward. And once you have the image, then, from that backing out the correct actions is easy too, because you can just look at the synthesized arm angle and just back out what the angle should be. So it&#8217;s not changing the problem, but it&#8217;s just introducing this intermediate step that makes it easier to solve. Just like if you&#8217;re solving a math problem, if you figure out the right intermediate step, kind of the answer is obvious from that intermediate step.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2027">33:47</a>] And I think that&#8217;s really exciting because now that shows that this level of generalization across robots, and I&#8217;m sure other generalization too, can be facilitated with thinking, just like in large language models, but with a twist that you have to think in the right modality.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2040">34:00</a>] Interesting. So it outputs a, I guess that image is what its video sensor is seeing, and it&#8217;s like the next step.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2049">34:09</a>] Yeah, you can almost think of it like image editing. You can do the same thing with video. You can do it with video prediction. But the key is to imagine what it would look like to progress on this task. That seems like a very human thing to do.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2062">34:22</a>] Some things you plan semantically and some things you plan spatially. If you&#8217;re doing rock climbing, you&#8217;re probably not thinking like, &#8220;Hey, left arm to rock, 37 centimeters to the left.&#8221; You&#8217;re probably more imagining your arm reaching for the rock.</p><h3>34:34 &#8212; Why humanoid robotics differs from Waymo</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2074">34:34</a>] Earlier in the conversation, you mentioned that Waymo was very inspiring, and their kind of path to productionization is proof that you can do real-world generalized robotics. And if I recall correctly, when I was a lot younger, it was kind of this early promise of, &#8220;This is going to happen,&#8221; and then in reality, it took a lot longer. So I guess my question is, in the case of humanoid robotics, what would make you say single-digit years, it&#8217;s coming, versus a long tail and policy challenges as well?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2117">35:17</a>] I think one big difference between how robotic foundation models address the problem and how more traditional engineered systems address the problem is that the stack is really thin. So it&#8217;s not easy to train a foundation model. Obviously you need to get the right data. There&#8217;s a lot of work that goes into curating, labeling, all that other kind of stuff. But the actual software that runs on the robot is very, very simple.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2142">35:42</a>] So you might have some kind of thinking or reasoning stage. You might have the model produce actions; it needs to be fast enough. But if you think about it in terms of raw lines of code, it&#8217;s much, much lower than a more traditional AV stack. And partly that&#8217;s because modern autonomous vehicles, the work on that started a lot earlier with very different technologies and evolved over time. Partly it&#8217;s also because the problem is more safety-critical.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2173">36:13</a>] Like, yes, you don&#8217;t want a robotic manipulator to drop a fragile object, but at the end of the day, that&#8217;s a lot less bad than having a car hit somebody. So that is not to say that the safety challenges with robots are not real. They&#8217;re very real and it&#8217;s very important to tackle them. In fact, it&#8217;s probably one of the harder ends of the problem. But they are not as much of a hard stop to practical deployments because you can come up with tasks and environments and domains and also physical hardware where those problems are a lot less severe.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2207">36:47</a>] So I think that combination, radically simpler software stack, plus less drastic software challenges, actually make it a lot easier. And because, to your earlier point, there is this kind of flywheel effect, that there is a positive feedback loop, that starting to get things out in the world, even under some constraints, will actually facilitate getting them out more and more.</p><h3>37:10 &#8212; If humanoid robotics failed here is why</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2230">37:10</a>] I think a lot of people are familiar with this idea of postmortem, looking back on why something failed. But in this case, I&#8217;m curious, what would you say to a pre-mortem in the sense of if humanoid robotics did not succeed in single-digit years, what do you think would be the most likely reason why humanoid robotics failed?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2254">37:34</a>] Ultimately, for these things to be truly useful, they do need to reach a level of reliability and robustness and generalization that is higher than what we typically expect from <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a>, for example, or generative AI for images and video. Because typically these tools, they are very much human-interactive tools. You get an LLM to do something, and it doesn&#8217;t do quite what you want, so you sort of revise your prompt, and you basically iterate with it.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2281">38:01</a>] And that&#8217;s why even the earlier LLM tools, like the first version of <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a>, even though they were much more primitive than what we have now, they were still already useful because somebody could just keep hammering at it until it basically solves their problem. Just like if you&#8217;re using a search engine, you type something in the search engine, you don&#8217;t get quite what you want, you revise your query and then you get what you want.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2300">38:20</a>] Whereas with a robot, the full value of it is unlocked when it&#8217;s actually doing the thing autonomously. So to have somebody constantly iterate for every single task is almost antithetical to the benefit that you&#8217;re getting. So I think a lot of the risk has to do with how easy is it to get that level of reliability and robustness. And that&#8217;s again where some of the demos might be a little bit misleading, because if someone shows a demo of their robot doing something cool, I mean, obviously if everything is presented in a forthright way, that could still be a very good indicator of progress.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2332">38:52</a>] But it doesn&#8217;t make it obvious how far or how close it is to reaching that practically relevant level of robustness. So I&#8217;m personally a big believer in using techniques like reinforcement learning that can actually benefit from autonomous experience to kind of fine-tune those last few percentage points to make it go from 95 to actually 100%. But that&#8217;s really important, and it&#8217;s not yet a solved problem.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2358">39:18</a>] If it did take longer than expected, it&#8217;s because the bar is higher.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2363">39:23</a>] Because the bar is higher. And in particular, those last few, kind of the last inch, so to speak, is something that requires not just really good models, but also new innovations in technology. I mean, I don&#8217;t think that I&#8217;m not the kind of person that would say, oh, we should throw out everything that we know about foundation models to start over. I don&#8217;t think it&#8217;s that at all. I think that roughly the puzzle pieces that we have are actually very good puzzle pieces.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2386">39:46</a>] But still, we should acknowledge that right now, the methods and the models need more work to cross that level of robustness.</p><h3>39:55 &#8212; Are there hot take modeling architectures in robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2395">39:55</a>] In <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a>, it feels like everyone is doing kind of the same thing, but different flavors in the robotics industry. Is everyone doing kind of the same thing? Are there any hot take architectures that are different direction?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2409">40:09</a>] I actually think that there&#8217;s a lot more heterogeneity than it might seem. One big dividing line that I think is maybe not as obvious from just looking at the results is the distinction between fully embracing the foundation model ethos, so to speak, versus focusing on specific vertical areas. And I think this is hard to tease out sometimes because obviously everyone&#8217;s going to say, &#8220;Oh, I&#8217;m doing the thing that <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> did,&#8221; because that&#8217;s the cool thing.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2440">40:40</a>] But the foundation model ethos fundamentally is something like this: if you have a particular problem you want to solve, it is better to train a more general model that can use data from a breadth of problems. And if you do it right, it&#8217;ll actually be better at the specialized problem you want to solve than a narrow specialist. So again, to come back to the LLM analogy, if you want to do machine translation, don&#8217;t build a machine translation system.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2465">41:05</a>] Build a language model that understands all language tasks and throw it at machine translation. And in robotics, I think that is actually very deeply uncomfortable to people. Because if someone is actually working on an application, they&#8217;re doing warehouse automation. It is very awkward to think, oh, if I want to do warehouse automation, let me collect data of putting away silverware in kitchens.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2487">41:27</a>] It just sounds bizarre. But that is the foundation model lesson: if you have enough breadth, if you collect data from a wide range of different tasks, then you will acquire those generalizable skills. And if your model is built correctly, it will repurpose those skills for whatever situation it encounters. So I think it is actually true that even if you want to build a warehousing robot, you are better off collecting a breadth of data and will be better at handling all the weird edge cases you might encounter, even in that warehouse domain.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2515">41:55</a>] But this is not something that&#8217;s easy for people to accept because it&#8217;s just so antithetical to the principle of building kind of a traditional, vertically integrated robotic system.</p><h3>42:05 &#8212; Thoughts on AI safety in robotics</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2525">42:05</a>] I noticed this new, interesting phenomenon with these AI companies, which is if they&#8217;re wildly successful, it creates this, I guess, worry or new set of things. So, for instance, when <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> had a very powerful model, then the government comes in and there&#8217;s these worries about safety and risk and all that. And I&#8217;m curious how you think about that. Like, if <a href="https://en.wikipedia.org/wiki/Theory_of_multiple_intelligences">Physical Intelligence</a> this year had a phenomenal, incredibly capable, generalized model, how do you think about those kinds of topics that might come up?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2568">42:48</a>] Working on AI safety is not a new thing. My colleague at UC Berkeley, Stuart Russell, was talking about this stuff over a decade ago, and lots of people spent a lot of time working on it. It&#8217;s just that the trouble is when the technology moves so fast, the important problems are not just a function of the core principles. It&#8217;s also a function of how society reacts to it, what kind of tools are adopted, and so on.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2591">43:11</a>] And I think that&#8217;s very, very hard to anticipate. So I don&#8217;t have a very satisfying answer here. In terms of how we are approaching it, our philosophy around all this stuff is basically one of empirical experimentation. Let&#8217;s get stuff out there. Let&#8217;s see what happens in the real world. Let&#8217;s see what goes right and what goes wrong, so that we have as much of a preview for what the technology can do, what are its weaknesses, what are its strengths, and so on.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2622">43:42</a>] But at the end of the day, you kind of have to just keep your eyes open, see what happens, and adjust as you go. It&#8217;s very hard to anticipate. And I think your question, though, is very spot on. Even though I don&#8217;t have a great answer for you because, yeah, if we&#8217;re having this much concern and issues with AI systems that are basically limited to using computers, we&#8217;re presumably going to have strictly more concerns and issues with AI systems that can do everything in the physical world that we can do.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2654">44:14</a>] Right. So the issues are real. It&#8217;s just you kind of have to see what happens and then adjust.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2659">44:19</a>] And that&#8217;s kind of in the scary path, but in the happy path. If everything goes well and we have incredibly capable models and robots, in 10 years, is the North Star, that&#8217;s the end of human labor.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2677">44:37</a>] I believe it&#8217;s a mistake to think of robots as mechanical people. Computers at some level are kind of like mechanical brains. But when personal computers really took off in the &#8216;90s, early 2000s, et cetera, it&#8217;s not like the first thing that happened is that people replaced their brains with computers. Rather, what we saw is actually a proliferation of very different kinds of computers. We saw kind of ubiquitous computing.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2703">45:03</a>] So you would have a computer on your desk, but you might also have one in your pocket. You might have one in your refrigerator and in your car. Because computing became so accessible, you could have a little bit of computing in everything. And I don&#8217;t think that&#8217;s what the people that first started thinking about this stuff in the 40s and 50s would have imagined. They would have imagined room-sized computers whose job is to control the policy of an entire country or something, rather than a little bit of computer in everybody&#8217;s refrigerator.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2732">45:32</a>] So I think by analogy, we might imagine there might be a little bit of <a href="https://en.wikipedia.org/wiki/Physical_artificial_intelligence">Physical AI</a> in everything, and it might just be lots of everyday things that you have to do yourself. Now you can get a little bit of help with it. I think the other example that&#8217;s worth thinking about is modern coding agents. So I think that this is something where, of course, the jury is still out as to what the end game of coding agents is.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2757">45:57</a>] But certainly from the experience of software engineers today, it kind of seems like it&#8217;s fair to say that most would consider coding agents to be more empowering than somehow causing them to panic. I mean, obviously some people might panic, but in general, at least from the software engineers that I&#8217;ve talked to and from my own experience, it&#8217;s more empowering to be able to amplify how much work you can do with AI tools.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2788">46:28</a>] So I think from that, and that maybe is a pretty direct analogy, because that is straight-up an example of an actual real job where AI has entered into it and has actually provided more leverage to the people doing that job. So I think that&#8217;s another example that we can look to. But the truth is that I think it remains to be seen.</p><h3>46:44 &#8212; Top robotics research paper recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2804">46:44</a>] So in <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a>, there are a few seminal papers that, if you&#8217;ve read those papers, you kind of get a sense of the lineage of the breakthroughs that mattered in understanding where we are today in the robotics industry. Are there a set of top papers that you really think kind of show the breakthroughs that people should know about if they&#8217;re curious about the state of the art in terms of humanoid robotics?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2831">47:11</a>] One thing I would point out, and this is partly a shameless plug, because I am a coauthor on that paper, though candidly, 99.9% of the work on this was done by Tony, who was the lead author, is the original ACT paper, the ALOHA paper. It&#8217;s an interesting example because in some ways the ideas weren&#8217;t really that new, but they were illustrated in a really nice way. And the idea was that, hey, if you set up the right kind of low-cost robot setup, in his case it was based on these robot arms from Trossen Robotics that are, they&#8217;re like $7,000 hobbyist arms.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2868">47:48</a>] He set them up in a bimanual setup with a leader-follower teleoperation device. And he showed that actually if you do it right, without really any particularly fancy tricks, you could easily collect teleoperation data of extremely dexterous tasks that people had previously thought would require very sophisticated hardware and all sorts of really expensive stuff, and then set up a fairly straightforward transformer-based model, and it could actually do a lot of those tasks.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2893">48:13</a>] And it&#8217;s kind of like an interesting thing because usually in academic research we put a big premium on do you have some sophisticated new mathematical thing or some sophisticated technical insight. And in that paper, which I think at this point has been hugely influential, the insight is really just, yeah, just put together the right pieces and have a little bit more faith in what a simple robot could do, so to speak, equipped with a good intent learning system.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2919">48:39</a>] And he showed things like replacing batteries in a remote control. He even got a mannequin foot, and he showed that he could put a shoe on it for an assistive task, sort of. Some people need help getting their shoes on, so that&#8217;s good. But what people found, I think, so interesting about that paper is just how far you could get with relatively simple building blocks. And at this point, he open-sourced the code for it.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2946">49:06</a>] And the ACT code has been used by lots of people. Sort of like if someone wants a very basic starter kit for robotic learning, that&#8217;s usually what they grab. And I think it&#8217;s worth for somebody who wants to get into the field to go through that paper and really understand what&#8217;s going on there. Because even though in some ways it&#8217;s not that sophisticated, I think it provides a bit of calibration on what matters.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2971">49:31</a>] The details matter, but the details don&#8217;t have to be complicated.</p><h3>49:35 &#8212; Why is Boston Dynamics less top of mind</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=2975">49:35</a>] Before all this large-model robotics kind of wave, prior to that, Boston Dynamics had these really impressive demonstrations and tons of mind share. I guess I wasn&#8217;t even in the field, but I think, wow, they&#8217;re really doing incredible robotics. And then in the last, I don&#8217;t know how many years, I don&#8217;t really hear about them much anymore. Is there some shift in the industry that made that so, or is that something you could explain?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3012">50:12</a>] So the way I would explain it is this: robotics at some level is about building complex systems. So even though it&#8217;s very tempting to say, oh, there&#8217;s different areas of AI, there&#8217;s <a href="https://en.wikipedia.org/wiki/Large_language_model">LLMs</a> and vision and robotics, one of those is not like the others because for robots you actually need all the parts, everything from how you wire up the robot, what the power source is, what does the actuator look like, all the way to how does it do high-level planning to determine what tasks to do next.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3046">50:46</a>] And even though we could look at these things and say all of these different videos and different companies and different demos, they&#8217;re all robotics, they&#8217;re really different parts of the stack. A lot of what the classic Boston Dynamics results show is very sophisticated hardware, very carefully designed hardware with a traditional control approach, with very smart controls engineers setting everything up, but with comparatively less emphasis on the kind of decision-making aspect.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3080">51:20</a>] And I think that at a particular point in time that actually made a lot of sense. Because if we can&#8217;t build the physical body, it doesn&#8217;t matter what kind of decision-making system is running on it. But to our earlier discussion about generalization, at this point we&#8217;re at a stage in the development of these things that even though we can do more on hardware, in many ways it&#8217;s good enough. And the big challenge is how to have the decision-making loop that actually works and that reacts intelligently to everything in the environment.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3112">51:52</a>] And the place where I would draw the dividing line between those is decision-making loop doesn&#8217;t mean symbolic decisions. It could mean low-level decisions. The question is, do you need to take the rest of the environment into account, or are you just dealing with a robot? So if you want to do a backflip on flat ground, you most have to deal with a robot. But if you want to pick up a coffee cup off of a table, even though that&#8217;s maybe in some ways simpler than doing a backflip, you really have to understand what&#8217;s going on in the rest of the world rather than just your own body.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3140">52:20</a>] And that dividing line, I think the way that technology has panned out, I think it&#8217;s fair to say that that is the dividing line between AI and controls. Controls is when you have to control the robot body. AI is when you have to take into account what goes on outside of the robot. And I think that that&#8217;s why you see this divide, because I think a lot of the demos where you mostly needed to deal with the robot itself and not the rest of the world, really good controls, could allow you to admit a very good solution there.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3171">52:51</a>] And another thing I would say here is, okay, if there&#8217;s a lot of controls work that goes into doing some particular skill, well, there is actually something to learn from that. Because if you can hand-design a controller that performs a sophisticated behavior, very likely you can also learn that controller. So just that proof of existence that the thing is possible, and not only possible, but also simple enough that a person could build it.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3195">53:15</a>] Because remember, people are, at the end of the day, even with code these days, the kind of complexity that people can handle is not as high as the kind of complexity the AI can handle. So if a person can hand design something to do a backflip or do some acrobatics, that&#8217;s a really great proof of existence that there exists some relatively parsimonious control law for doing that skill. And parsimonious does mean generalizable.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3216">53:36</a>] So if it&#8217;s simple enough for a person to design, probably there&#8217;s something fairly general in there. And if you can learn it and automate it without having to have the human controls engineers in the loop, that&#8217;s good news.</p><h3>53:47 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3227">53:47</a>] And then last question for you is, if you could go back to when you just entered the industry and give yourself some advice, knowing everything you know now, what would you say?</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3237">53:57</a>] One thing that I&#8217;ve learned over the last few years, which I think is a little different than my original mindset, is that addressing robotics effectively requires using very broad prior knowledge. And I think there&#8217;s this idea that a lot of people in robotic learning have, which I think I shared initially, that since people learn things from scratch, maybe robots should learn things from scratch too.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3266">54:26</a>] So, for example, in some of our early work on large-scale robotic learning at <a href="https://en.wikipedia.org/wiki/Google">Google</a>, we had this, what we call the ARM Farm project. We set up a bunch of robot arms in a conference room, actually, because we didn&#8217;t have a proper lab. It was a conference room. And we had them all grasping objects. And the idea was, well, if they grasp millions of objects, they&#8217;ll learn very general grasping strategies.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3291">54:51</a>] And it basically worked. They could learn to grasp objects, but it was very hard to take it further than that to the next level. Now I can pick up anything, but so what? It didn&#8217;t serve as a very good stepping stone for more complex skills. And I think part of that was that we were approaching this from a very blank slate. Let&#8217;s start from zero and see if, knowing nothing in advance, the robot could start picking up behaviors.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3312">55:12</a>] But I think that it&#8217;s much, much more practical to get all this to work if you can combine robot experience with knowledge that you can pull in from other sources. For example, I was very skeptical initially about the utility of language. And I think scientifically this is defensible, which is that, hey, animals can do some pretty impressive things. Monkeys can do really cool stuff. But monkeys, as far as I know, can&#8217;t speak, at least not very eloquently.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3339">55:39</a>] So maybe our robots should also be able to do stuff. And they don&#8217;t necessarily need to understand language. But I think the subtlety there is what&#8217;s important is not language, it&#8217;s prior knowledge that you can put in as a scaffold on your learning process. And you can pull in that knowledge in all sorts of ways. Humans don&#8217;t necessarily pull it in entirely through language. Humans and monkeys certainly don&#8217;t.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3361">56:01</a>] They do it from observation, from observing other people, other creatures, and so on. So there&#8217;s lots of sources of prior knowledge, but the point is that you got to get that prior knowledge in there. Otherwise, you&#8217;re actually faced with a harder problem than what humans and animals have to solve. Because if a person had to figure out how to assemble IKEA furniture, but they&#8217;ve never actually encountered any article of furniture in their entire life, okay, that would be pretty difficult because they don&#8217;t even know what the point of this is or what the end game looks like.</p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3388">56:28</a>] So, yeah, prior knowledge is important. And while I&#8217;m still a big fan of learning things through experience, I think that my advice to myself would have been, take prior knowledge more seriously.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3398">56:38</a>] Awesome. Well, thank you so much for your time, Sergey. I really appreciate it.</p><p><strong>Sergey:</strong></p><p>[<a href="https://youtu.be/9OSbaPjv0Rc?t=3401">56:41</a>] Yeah, thank you for your questions.</p>]]></content:encoded></item><item><title><![CDATA[Creator of TypeScript: 10x Faster Typescript, Why AI Won't Replace SWEs | Anders Hejlsberg]]></title><description><![CDATA[It was interesting talking to Anders Hejlsberg, the creator of TypeScript and C#, about all the technical details behind rewriting the TypeScript compiler in Go.]]></description><link>https://www.developing.dev/p/creator-of-typescript-10x-faster</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-typescript-10x-faster</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 17 Aug 2026 13:03:31 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211501947/742daed19902300f9b81ff80035cd341.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>It was interesting talking to <a href="https://www.linkedin.com/in/ahejlsberg/">Anders Hejlsberg</a>, the creator of TypeScript and C#, about all the technical details behind rewriting the TypeScript compiler in Go. </p><p>I had a lot of questions about why they chose Go (vs Rust) and if they used LLMs in the rewrite. Turns out they used LLMs much less than expected in the rewrite. He explained why in the conversation.</p><p>Anders has strong takes on how AI will impact software engineering so I made sure to run all the famous takes I&#8217;ve seen on Twitter to see what he thought like:</p><ul><li><p>Will AI write all the code in ~1-2 years?</p></li><li><p>Will AI replace junior engineers in a few years?</p></li><li><p>Will you still need to use an IDE by the end of the year?</p></li></ul><p>And also I asked him why his opinion seemed to be so different from the more AI-pilled companies like the labs. Hope you enjoy the episode, full transcript below as usual if you want to skim.</p><p>Check out the episode wherever you get your podcasts: <a href="https://www.youtube.com/watch?v=cywK3XYYJ2o">YouTube</a>, <a href="https://open.spotify.com/episode/1thvWd3ClYQHvaUX9qKckk?si=pmojE4jTSPmHVnm4fydbFg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-cywK3XYYJ2o" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;cywK3XYYJ2o&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/cywK3XYYJ2o?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/211501947/0048-why-write-a-compiler-in-javascript">00:48 - Why write a compiler in JavaScript</a></p><p><a href="https://www.developing.dev/i/211501947/0729-why-rewrite-the-compiler-in-go">07:29 - Why rewrite the compiler in Go</a></p><p><a href="https://www.developing.dev/i/211501947/1449-llms-for-large-migrations">14:49 - LLMs for large migrations</a></p><p><a href="https://www.developing.dev/i/211501947/2012-why-javascript-is-so-popular">20:12 - Why Javascript is so popular</a></p><p><a href="https://www.developing.dev/i/211501947/2632-why-ever-use-javascript-on-the-backend">26:32 - Why ever use Javascript on the backend</a></p><p><a href="https://www.developing.dev/i/211501947/3259-what-it-takes-to-build-a-programming-language">32:59 - What it takes to build a programming language</a></p><p><a href="https://www.developing.dev/i/211501947/3706-will-there-be-fewer-languages-in-10-years">37:06 - Will there be fewer languages in 10 years</a></p><p><a href="https://www.developing.dev/i/211501947/4257-hands-on-engineering-vs-delegation">42:57 - Hands on engineering vs delegation</a></p><p><a href="https://www.developing.dev/i/211501947/4914-why-fast-tooling-matters-more-now">49:14 - Why fast tooling matters more now</a></p><p><a href="https://www.developing.dev/i/211501947/5116-ai-software-engineering-predictions">51:16 - AI software engineering predictions</a></p><p><a href="https://www.developing.dev/i/211501947/5852-the-most-technically-challenging-work">58:52 - The most technically challenging work</a></p><p><a href="https://www.developing.dev/i/211501947/010204-top-book-recommendation">01:02:04 - Top book recommendation</a></p><p><a href="https://www.developing.dev/i/211501947/010350-advice-for-his-younger-self">01:03:50 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:48 &#8212; Why write a compiler in JavaScript</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=48">00:48</a>] The big thing that we&#8217;re going to talk about here is obviously <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> 7 with the native rewrite. And the first thing that I think of when I see this was I didn&#8217;t even realize that the compiler was written in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> to begin with because I don&#8217;t usually think of compilers being written in high-level dynamic languages.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=69">01:09</a>] Yeah, no.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=71">01:11</a>] So why was <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> written originally, the compiler written in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=77">01:17</a>] No, it&#8217;s a good question. And honestly, if you had told me before the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> project started that, &#8220;Anders, you&#8217;re going to be writing compilers in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>,&#8221; I&#8217;d be like, &#8220;No, no, I&#8217;m not.&#8221; But I mean, the original prototypes for <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, the <a href="https://en.wikipedia.org/wiki/Strada">Strada</a> project, as it was called in the very early days, were actually written in C because it was like an adaptation of the parser, the <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> parser that we had in <a href="https://en.wikipedia.org/wiki/Internet_Explorer">Internet Explorer</a> and whatever.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=109">01:49</a>] And then we&#8217;ve moved to <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, but we wrote it sort of <a href="https://en.wikipedia.org/wiki/C">C#</a> style. But why did we go with <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>? Well, I mean, if you can self-host in the ecosystem that you want to be a part of, then that is just dramatically better than putting yourself outside the ecosystem and trying to target that ecosystem. Do you know what I mean? Because by writing it in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, we were also daily users of <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> and daily users of the tooling that we were building.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=141">02:21</a>] And whenever something didn&#8217;t work right, we knew it immediately. And whenever something didn&#8217;t perform right, we knew it, and so in the early years, it was a huge boon, I think. And then, of course, <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> runs everywhere, right? And that meant automatically our compiler could run on every platform, anywhere, even in the browser. Had we gone with a native code solution at that time, it wouldn&#8217;t have been possible to run it in the browser.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=169">02:49</a>] Now we have <a href="https://en.wikipedia.org/wiki/WebAssembly">WebAssembly</a> or <a href="https://en.wikipedia.org/wiki/WebAssembly">WASM</a>, but that didn&#8217;t exist at the time. Right. So it was a meaningful choice. And honestly, at that time also, people didn&#8217;t really realize how fast <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> had gotten. I mean, it used to be very slow, and then Google did all of their excellent work with <a href="https://en.wikipedia.org/wiki/V8_engine">V8</a> and got it to within 2 or 3x of native code, which is pretty darn impressive. And it was like full well possible to write compilers in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, and you could even get them performant because often it&#8217;s the algorithms that determine how performant you are.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=213">03:33</a>] It isn&#8217;t necessarily the runtime environment. Right. I mean, it&#8217;s a combo. But yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=218">03:38</a>] Is there anything unique about the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> compiler? I think of a compiler, I think of <a href="https://en.wikipedia.org/wiki/Clang">Clang</a> or one of those ones that is used for maybe C or those static languages. What&#8217;s unique about the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> compiler when it was written in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> compared to these native compilers?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=240">04:00</a>] I think there are a couple of things that are pretty unique about <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>&#8216;s compiler. First of all, it doesn&#8217;t target machine code, it targets <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, and some people call that a transpiler. And in a sense, the compilation phase is mostly about removing the type annotations and turning your <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> back into <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. Right now, in the early days, also an important part of the compilation process was to downlevel your <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, meaning all of the runtime environments at the time were not evergreen.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=276">04:36</a>] And sometimes people were lagging multiple years behind on which level of the standard was implemented. And for example, when classes got standardized, most <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> runtimes didn&#8217;t implement classes, but it turns out that you can downlevel to constructor functions, transform the code. And so part of what we did was this transformation, downlevel transpiling of the code. But then, of course, the type checking is the big thing that we do.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=306">05:06</a>] And the type checker in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> is also quite unlike any other type checker in other compilers. Typically, type checkers exist in compilers to guide the code generator because you want to know: are we dealing with a float or an int or a string here so we can generate the correct machine instructions for that data type, et cetera, et cetera. That&#8217;s not the case in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>. In fact, since we erased the types, the types have no impact whatsoever on the runtime behavior of the code.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=340">05:40</a>] They purely exist for tooling sake and for the developer&#8217;s sake and to guide things like statement completion and refactoring and code navigation, which is a whole different thing. Also, they don&#8217;t necessarily have to be there. And so <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> has a gradual type system. You can sort of have half of your code have types and the other half is just anys that we don&#8217;t check. Now very few languages have anything like that.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=373">06:13</a>] So it&#8217;s a very different compiler in that sense and a very different language in that sense. That actually made it fascinating to work on because we were solving problems that no one had solved before.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=385">06:25</a>] I know a lot of compilers, they take the higher-level code and they typically apply optimizations as well. Maybe they unroll loops or they get rid of duplicate instructions, simplify things. Does anything like that happen in the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> to Java compilation process?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=404">06:44</a>] Very little. Like I said, when we do the downleveling, there are some transformations that go on that are complex. But it is less and less the thing that&#8217;s important about <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>. In fact, a lot of people use <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> only for the type checker, and then they use some bundler like esbuild or SWC or whatever to package their app, erase the types, do whatever needs to be done in order to make it runnable in the browser.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=435">07:15</a>] And so we&#8217;re not really there to optimize your code for runtime. We&#8217;re there to make you more productive as a developer or make AI more productive as a developer.</p><h3>07:29 &#8212; Why rewrite the compiler in Go</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=449">07:29</a>] So with the native rewrite of this originally fully <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> compiler, what was the problem you were trying to solve with the rewrite, and why did you end up rewriting it in native code?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=463">07:43</a>] Well, the problem we were trying to solve is real simple: performance and scalability. <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> was never really optimized for compute-intensive workloads like compilers, right? <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is more about creating UI that runs in a browser. It was originally intended for just maybe 100 lines of code, but now, of course, it&#8217;s gotten to be a lot more. But it&#8217;s not a place that you would optimize for the kind of workload that we are.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=496">08:16</a>] So in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, first of all, you pay a 2 to 3x perf penalty compared to native code. I mean, it depends on the workload, but if you&#8217;re doing compute, that&#8217;s about where it&#8217;s at. And then secondly, in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> there are a lot of restrictions around use of concurrency. <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> was always engineered, in fact, to be a single-threaded language. That&#8217;s why we have callbacks and async and whatever.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=524">08:44</a>] It&#8217;s because you can&#8217;t spin up threads, and that&#8217;s a good thing. Mostly because concurrency in a language with mutable data is very, very hard, because you can have races and deadlocks and all of this good stuff, right? That&#8217;s why functional programming languages are easier with concurrency, because all the data is immutable, and it&#8217;s much easier to reason about. So <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> doesn&#8217;t give you access to concurrency other than web workers.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=555">09:15</a>] But web workers can&#8217;t share data between each other other than by remoting it. So you can have one web worker compute something, but if it wants to give it to some other web worker, it has to turn it into JSON, or it can&#8217;t just hand it objects. And that means you can&#8217;t have one worker compute some data structure and then have someone else use it. In other words, you can&#8217;t have shared memory concurrency.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=581">09:41</a>] And really that&#8217;s what we wanted in order to gain all of the perf that we were leaving on the table by not utilizing multicore CPUs that everyone has today, right? Because <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> has stopped giving us faster CPUs, it&#8217;s giving us more CPUs, and we got to find ways to use them in our compute-intensive workloads, or else we&#8217;re leaving money on the table. And so we were leaving money on the table in multiple ways, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=608">10:08</a>] And so native code, and that was why we looked for a language that would get us out of both of those binds. It had to be native and it had to have access to shared memory. Concurrency.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=621">10:21</a>] I saw for the language choice you went with Go. And when I think of all the systems languages, the ones that are often hot are maybe Rust or Zig, why did you choose Go for the native rewrite?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=635">10:35</a>] I think the decision process was actually pretty structured. The first decision we made was we&#8217;re not going to rewrite, we&#8217;re going to port. Because only by porting can we preserve the semantics and the algorithms and the exact behavior of our existing compiler, which everyone depends on for backwards compatibility. Now, we could have cleaned the slate and started completely from scratch, but we would have come up with a different language in the sense that it would give you different errors or behave differently in certain situations where it has to make choices between multiple possibilities and what have you.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=675">11:15</a>] So we wanted to port that meant our code makes certain assumptions. Our code, for example, assumes the existence of garbage collection. It assumes the existence of first-class treatment of functions. You can have functions within functions, and you can close over outer state and so forth. As we evaluated all of these, Go was the one that checked the most boxes. It gives us very robust and mature native code generation on all major platforms.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=713">11:53</a>] It has garbage collection and it has excellent access to shared memory, concurrency. And those were the high-order bits that we wanted to check. And every other language had something that kind of worked against those objectives. So for this particular workload, Go was the right choice for us. And it&#8217;s worked out well.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=736">12:16</a>] If you think about Rust, what didn&#8217;t it have that if it had it, you would have picked it?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=742">12:22</a>] I mean, I think there are two things. It doesn&#8217;t have garbage collection. It&#8217;s done manually, or it&#8217;s done through the borrow checker. But the borrow checker doesn&#8217;t allow circular data structures. And our compiler is chock full of circular data structures. We have trees with parent pointers, we have types that are recursive, we have symbols that refer. I mean, it&#8217;s everywhere. And like I said, if we had chosen to rewrite, we could have probably solved all of that.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=774">12:54</a>] But it wouldn&#8217;t have been just porting the code. It would also have been solving a whole bunch of new problems that we had bought ourselves by going that route. We tried, and it just wasn&#8217;t feasible. And so we were just trying to optimize for doing as little work as possible to get us to the payout. And honestly, Rust is no better when it comes to the quality of the generated code or the concurrency gains that we could have gotten out of it.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=807">13:27</a>] Go is just great. I mean, so we got all the bennies for less work by going that route.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=815">13:35</a>] I know some languages, when you say they&#8217;re garbage collected, there&#8217;s something in the language itself that is doing garbage collection. But there&#8217;s also independent libraries that can provide garbage collection, right?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=832">13:52</a>] There is, but it often comes with a whole bunch of restrictions, and it doesn&#8217;t necessarily come with the safety guarantees that we were looking for. I mean, Go is type-safe and memory-safe, meaning that you don&#8217;t get to just have stray pointers or whatever. With Rust, you can do ref counting or you can do all sorts of different strategies, but there&#8217;s always that little bit of unsafe where you have to do this in order to make it all work.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=866">14:26</a>] Where if garbage collection is engineered into the language, it really is a separate concern that is you don&#8217;t have to do anything special. I mean, of course, you have to not hold on to data that you no longer need and whatever. But that&#8217;s true of ref counting as well, right? But it is inherent in the language.</p><h3>14:49 &#8212; LLMs for large migrations</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=889">14:49</a>] When you ported it from <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> to Go, I&#8217;m curious how much LLMs were used in that process, because I know famously there&#8217;s some major projects that have been doing rewrites of their entire codebase. I know there&#8217;s this <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> project called Bun that rewrote everything from Zig into Rust, so it seems like a pretty useful tactic. And I&#8217;m curious if you leveraged that in this process.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=916">15:16</a>] Not a whole lot, but it also has to do with timing. We started this project two years ago, and two years ago, LLMs were nowhere near as good as they are now. I mean, that&#8217;s how crazy fast it&#8217;s all moved, right? So what we did, though, is we certainly had a bunch of tools that aided us in the port. I&#8217;d say the prototypes we wrote, the scanner and the parser we wrote manually. And at that point we could convince ourselves, wow, we can actually get 10x out of this.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=951">15:51</a>] Now what do we do with the rest of this code base? What we did was we wrote a tool that syntactically translates <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> into Go, this code that comes out of it. It has no syntax errors in it, but it doesn&#8217;t compile either because it&#8217;s full of... It uses types as if they were <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, which they&#8217;re not in Go. And so all of that had to be refactored. We had to redo all the data structures or whatever, but we got the same code base out of it, and then we could sort of bang on that, and then in a more localized fashion, maybe use AI occasionally to help us do the transformation.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=989">16:29</a>] So it was sort of manual and automated and a little bit of AI. Now, if we were to start the project today, would we do it differently? Possibly. Although I will say, if we just let AI loose on what is it, like half a million lines of code that we have in the old compiler, or somewhere between half a million and a million, and asked it to translate it all, well, I don&#8217;t know that that would absolve us from then having to go in and carefully examining every line that came out of it to make sure that there were no hallucinations.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1024">17:04</a>] Right. Unless you have 100% perfect test coverage, you probably still got to go check all of that. Now, I think a better approach, quite honestly, would be to enlist AI to write a program that helps you translate from <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> to Go, because at that point you can park the stochasticness in that program and then you can get deterministic behavior whenever you run the program, which means you get the same transformation every time you run it.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1055">17:35</a>] That&#8217;s always the thing about AI that people forget: it&#8217;s not deterministic. So you can&#8217;t really trust that it&#8217;s going to do the same thing twice. We might have done it that way, but AI was not ready to do it at the time. Now, that doesn&#8217;t mean that we don&#8217;t use AI now down the line. We use AI a lot in our internal processes around the compiler, as everyone does, right? Help it to write tests and to move pull requests and what have you.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1085">18:05</a>] And whenever an issue is logged, first thing we do is put <a href="https://en.wikipedia.org/wiki/Microsoft_Copilot">Copilot</a> on and see if it can fix it for us, right? I mean, so, yeah, in that approach.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1093">18:13</a>] where you use the LLM to generate tooling and then the tooling is deterministic, but you&#8217;d still have to validate that thing.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1103">18:23</a>] Yes, but that&#8217;s a much smaller surface area that you have to validate, right? And honestly, I&#8217;ve seen that even with friends. We did some trip together, and there was some accounting of it. Well, he paid for the hotel rooms and we paid for this and whatever, and we needed to sort it all out, right? So, okay, well, here&#8217;s what all the costs were, and here are the pieces, and have AI fix it for us, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1126">18:46</a>] Or like tell who owes who what. And it very authoritatively came up with the wrong answer, right? Because it just forgot about some of the settlements, right? But if we had asked it to write a spreadsheet, which is really just a program, right, then we would have gotten the right answer because it&#8217;s very good at writing programs. So it&#8217;s kind of funny how sometimes you don&#8217;t ask AI for the answer, you ask it for a program that computes the answer.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1157">19:17</a>] What are the potential benefits that you left on the table by doing a port instead of a full rewrite?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1163">19:23</a>] To be honest, I don&#8217;t think there are any, and I don&#8217;t think we would consider a rewrite. We were quite happy. We&#8217;re quite happy with the way this codebase works. I don&#8217;t think we&#8217;re leaving any money on the table currently. And as I said earlier, rewrites are toxic often for ecosystems, because they sacrifice compatibility, because you never rewrite and you never get back to exactly what you had before, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1194">19:54</a>] Oh, but in the name of making it better or prettier or whatever, we change this and that and the other, right? And then while all the users get to suffer through that, right? So backwards compatibility is super important if you want to keep your ecosystem happy. And we do want that.</p><h3>20:12 &#8212; Why Javascript is so popular</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1212">20:12</a>] In studying for this interview, I just keep reading about <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and the strange semantics of the language. And I think throughout my career too, people have typically not been super satisfied or happy with <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> for some reason or another. My question is, why is it so popular despite that?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1236">20:36</a>] First of all, I don&#8217;t think <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is as bad a language as some people make it out to be. I will say in the three, what, three, four weeks that <a href="https://en.wikipedia.org/wiki/Brendan_Eich">Brendan Eich</a> had to create this language back in the mid-90s, he did a lot of things right, and thank God that he had been educated on functional programming and got first-class functions done, right? Where you can have functions within functions and the inner functions can access the outer function state and you can pass functions around as first-class values.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1269">21:09</a>] That is super powerful. <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> actually got that very right. Now there are some quirks in the language too. All of the automatic conversions and all of the bizarre differences between double equals and triple equals just drive everyone mad. But that&#8217;s exactly what type checkers are happy to do. They can ferret out all the little differences because they understand all the deep semantics of the language that no one seems to be able to keep in their mind.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1299">21:39</a>] So I think that&#8217;s part of why <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> has become so popular, right? It can actually truly bring out the best parts of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and leave the bad stuff behind. And of course you can&#8217;t. The thing about <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is it is the cross-platform language, like even Java, that was engineered and intended to be the way to write cross-platform applications. It&#8217;s not really cross-platform.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1326">22:06</a>] It doesn&#8217;t run in the browser, I mean, it doesn&#8217;t run everywhere, right? <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> runs everywhere. So there are a lot of reasons why you should like <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. And with <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, I think we&#8217;ve managed to sort of capture all the badness and park it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1347">22:27</a>] I saw there were some projects. I don&#8217;t know if you&#8217;re familiar with <a href="https://en.wikipedia.org/wiki/Jaxson_Dart">Dart</a> or <a href="https://en.wikipedia.org/wiki/CoffeeScript">CoffeeScript</a>, and it seems like these were some efforts to replace <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> or improve it, but they&#8217;re not popular today. So I&#8217;m curious what happened.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1364">22:44</a>] I mean, I think <a href="https://en.wikipedia.org/wiki/CoffeeScript">CoffeeScript</a> was mostly a different syntax for <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, so it doesn&#8217;t really change the semantics, and it didn&#8217;t provide any additional tooling or whatever. And <a href="https://en.wikipedia.org/wiki/Jaxson_Dart">Dart</a> was also, in a sense, about fixing <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. Right, but if you can fix something by fixing it instead of trying to replace it with something else, then you&#8217;re doing the ecosystem a much bigger favor, I think.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1396">23:16</a>] Right. And that&#8217;s what we set out to do with <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>. We weren&#8217;t trying to change <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>; we were trying to improve <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. And I think that&#8217;s ultimately why we succeeded.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1408">23:28</a>] I mean, in this new age where it&#8217;s very feasible that maybe AI could rewrite code bases, I imagine people will start to go toward the desired languages rather than the ones that are the incumbent languages. And I&#8217;m curious if you think that <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> might become less used in the future.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1431">23:51</a>] See, I don&#8217;t think so. What I&#8217;ve observed is that AI is best at the languages it&#8217;s seen the most of in its training set. AI gets trained on all the code that exists in the world, which means there&#8217;s an awful lot of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, an awful lot of <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, an awful lot of <a href="https://en.wikipedia.org/wiki/Python_(programming_language)">Python</a>, and therefore AI is good at <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> and <a href="https://en.wikipedia.org/wiki/Python_(programming_language)">Python</a>. It is less good at my favorite little language that I just invented.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1462">24:22</a>] In fact, there&#8217;s none of it in the training set, which means you have to sacrifice God knows how many tokens in the prompt to teach AI the language before you can even ask it a question about the language. Do you know what I mean? So in that sense, I think incumbent languages are actually, if anything, going to flourish because of AI. And we&#8217;re sort of seeing that right now with <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> too. Right.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1487">24:47</a>] There&#8217;s a noticeable knee in the curve of <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> adoption that coincides with the advent of AI. I do think the incumbent languages are actually, if anything, getting stronger.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1498">24:58</a>] What percent of all <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is written in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1504">25:04</a>] Well, so in the last year, <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> became the number one used language on <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a>. So there is <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> more than <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. So more than 50% on <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> at least is written in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>. So we&#8217;re actually larger now than the language that we set out to improve.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1528">25:28</a>] Wow.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1529">25:29</a>] Yeah. And larger than <a href="https://en.wikipedia.org/wiki/Python_(programming_language)">Python</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1531">25:31</a>] If you were to just draw that out in the future, do you think that eventually <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> would become the default and almost 100% of all <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> would be typed?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1541">25:41</a>] Well, certainly. I mean, if you&#8217;re having AI write it, I almost challenge you to find a tool that writes <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. They all write <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> because it helps AI write better code, because the type annotations guide the LLM toward writing, or to write, making fewer mistakes, and the ability for us to statically validate the code before you run it. I mean, that&#8217;s the distinction, right? The only way you can validate <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is by running it, which means you have to set up a runtime environment, which might not be possible for AI to do because AI lives in a sandbox, but AI can run through tools, the compiler, and check the code that it just generated and be told, oh, you made a mistake.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1587">26:27</a>] Oh, well, let me fix it then before I even tell anyone that I made the mistake. Right?</p><h3>26:32 &#8212; Why ever use Javascript on the backend</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1592">26:32</a>] After learning about this 10x improvement and performance from rewriting in native, it makes me wonder why anyone would use <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> on the back end. Why not just always use native languages on the back end?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1608">26:48</a>] Well, it depends on your problem, I think, because the thing that overlooks is what kind of thing am I going to write on the back end? There are an awful lot of very, very good web server frameworks, for example, that use <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> or <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, right? So if there are frameworks in your ecosystem that get you to a solution quicker than if you had to write that because your native language wasn&#8217;t really intended to target that particular community.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1641">27:21</a>] So I think it very much depends on what your problem is affinitized with. And often performance of your code isn&#8217;t the bottleneck. I mean, look at <a href="https://en.wikipedia.org/wiki/Python_(programming_language)">Python</a>, right? It&#8217;s used to basically script all of the LLM training in the world, right? I mean, a lot of that code is ultra time critical, but that&#8217;s written in native code, right? And then you use <a href="https://en.wikipedia.org/wiki/Python_(programming_language)">Python</a> as the orchestration language, and in a sense, in a web server, <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> is just an orchestrator of talking to a database and doing this and that. And the things that ultimately control the performance of applications like that aren&#8217;t necessarily whether that inner for loop is running as native code or not.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1692">28:12</a>] Which actually it is because of the JIT anyway, right? I mean, so measure first and then make decisions, right? Because you always get surprised when you measure, when you talk about how much faster it was to.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1704">28:24</a>] One of my immediate thoughts is, why wasn&#8217;t it done sooner?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1713">28:33</a>] At the time we started <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, I don&#8217;t know that we had ever envisioned that people would be writing projects that are the size of the projects that people are now writing in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, right? <a href="https://en.wikipedia.org/wiki/Visual_Studio_Code">Visual Studio Code</a> is like 2.3 million lines of code. I mean, we have in-house projects that are over 10 million lines of code, which is insanity. We would have just at the time gone, well, that&#8217;s clearly no one&#8217;s ever going to do that, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1741">29:01</a>] But that was a decade ago, more than a decade ago, right? We started in 2012 with the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> project, and the world looked very different. And so gradually, people have started writing larger and larger and larger projects, right? And at the same time, our compiler has gotten smarter and smarter. We&#8217;ve added all sorts of cool features like union types and discriminated unions and control flow analysis and blah, blah, blah.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1766">29:26</a>] And all of these things make our type checking even better. But it also makes it go a little slower every time we add a new feature, right? In fact, our <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> 6 runs only about 50% of the speed of <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> 1.5, for example. But <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> 1.5 also did a whole lot less. So projects have gotten bigger and the compiler does more. All of which detract from your effective output. And as I mentioned, <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> also, unfortunately, stopped delivering faster CPUs, and <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> keeps you from taking advantage of the additional cores.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1802">30:02</a>] And it&#8217;s not clear that we could have foreseen all of that, right? But it&#8217;s just the way it panned out, right?</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1810">30:10</a>] OpenAI, <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>, Cursor, and Vercel all use this product to make their lives better. And the problem it solves is when you&#8217;re building SaaS or an AI product and you want to sell to other companies, there&#8217;s all these requirements you need to meet. There&#8217;s SSO, there&#8217;s SCIM, there&#8217;s RBAC, there&#8217;s audit logs. These are all things that take time to integrate but aren&#8217;t the main focus of your app. WorkOS is an <a href="https://en.wikipedia.org/wiki/API">API</a> layer that lets you meet all of these requirements in just a few lines of code.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1840">30:40</a>] So let&#8217;s say you have a new SaaS product and you want to sell to other companies. WorkOS will solve all of these critical feature gaps for you. You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. Jira by <a href="https://en.wikipedia.org/wiki/Atlassian">Atlassian</a> isn&#8217;t just for tracking work anymore. Now you can pick your favorite AI agent to assign tasks to, and they&#8217;ll get access to the rich context that&#8217;s already in Jira.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1867">31:07</a>] When the agent is done, it surfaces a pull request. That way you can get more done with your favorite agents, all in one place. Learn more at jira.dev. That&#8217;s J I R A. D E V. Appreciate them for sponsoring the podcast. And back to the show. I first became a software engineer and I went to start working at Facebook. At the time they had this popular framework called Flow, which was doing type inference on top of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1897">31:37</a>] It was this thing that would take a long time, and then I&#8217;d never hear about it these days. So it sounds like <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> won, if there was a competition to type in.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1906">31:46</a>] There was a bit of it, yeah. In the early days, I think the thing that worked for us was actually the fact that we were self-hosted, which made it a lot easier for the community to contribute to the project. Flow was written, as I recall, in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. And so in order to contribute to Flow, you had to go learn a whole different way of programming in a different language. Right. And Flow also did not focus all that much on IDE-based tooling, which we knew full well going in.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1942">32:22</a>] We&#8217;re writing a compiler, but we&#8217;re not writing a compiler in order to generate machine code. We&#8217;re writing it to make better tooling. That was why we wrote it right there, hand in hand with our compiler. We had a language service. It was deeply integrated into Visual Studio, <a href="https://en.wikipedia.org/wiki/Visual_Studio_Code">Visual Studio Code</a>, all sorts of other editors through an open protocol, and that&#8217;s where the problem that needed solving was.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1967">32:47</a>] I mean, it&#8217;s nice to have a type checker, but boy, you want that. You also want statement completion. You want red squigglies, you want it interactive, right?</p><h3>32:59 &#8212; What it takes to build a programming language</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1979">32:59</a>] So I know in your career you&#8217;ve, you&#8217;ve worked on multiple programming languages. What does it take to build a programming language?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=1987">33:07</a>] Well, first I&#8217;ll say that the world needs another programming language. Like it needs another hole in the head. I mean, it&#8217;s like you kind of got to be a certain kind of person to even embark on it to begin with, right? It&#8217;s a super fascinating area of computer science and it&#8217;s one that&#8217;s been around ever since the beginning of computers. Now every compiler, every language has a whole bunch of things you need to understand, like parsers and scanners and lexers and code generators and whatever, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2023">33:43</a>] But that&#8217;s sort of the mechanics of implementing the language and then there&#8217;s the design of the language, sort of the art of making it feel right, whatever, right? And I think you sort of have to master both. And then you also got to appreciate that like every new language is actually only 10% new. And then it&#8217;s 90% the same drudgery that every other language has to go do. And so be prepared for a lot of work that maybe isn&#8217;t as interesting, you know, and then finally maybe I&#8217;ll say that, you know, like, never forget there that you stand on the shoulders of giants.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2064">34:24</a>] I mean, in order to make a great programming language, you got to understand a bunch of other programming languages first and understand all the, all the distinctions between the different styles of programming like procedural and object oriented or functional or what have you. And then last but not least, it&#8217;s like it&#8217;s a long game. I mean, it is probably a longer game than anything else. I mean, like every language project I&#8217;ve worked on, I&#8217;ve worked on each of them for at least 10 years and shipped multiple versions.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2098">34:58</a>] And it&#8217;s really never until version three that it truly starts to get okay, do you know what I mean? And it just takes an incredible amount of devotion to that. So you gotta really not be the type that tires of a problem quickly and then moves along that you&#8217;re not meant to do language design. And you see that too when you talk to language designers in the industry. They&#8217;ve been doing it for a long time.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2123">35:23</a>] You mentioned those two parts. There&#8217;s the objective pieces that you need to implement and then there&#8217;s the art of the design and making the language, I guess, ergonomic for developers. When you think of that second part, that art, and you look at other programming languages, are there any that you admire outside of the ones that you&#8217;ve worked on?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2145">35:45</a>] Of course, learning and fully appreciating the beauty of functional programming has been very educational. I mean, because it really is a different way of thinking about programming, a way of thinking of programming that&#8217;s much closer to math than it is to machines. And I think there&#8217;s a lot of incredible goodness that has come from that. And honestly, Even in building <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> project, large portions of the <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> compiler are written in a highly functional style and work on immutable data structures that are then.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2185">36:25</a>] Now, because of shared memory, concurrency can be shared between different threads or processes that don&#8217;t mutate the data and therefore they can share one data structure that&#8217;s incredibly powerful. But mastering this, like writing islands of pure functional programming inside an imperative program and putting it together with concurrence, I mean, it&#8217;s complex, but I think functional programming has brought a lot to the world.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2215">36:55</a>] I know object oriented programming as well, but I wouldn&#8217;t single out any particular language. Every one language you come in contact with, you learn something.</p><h3>37:06 &#8212; Will there be fewer languages in 10 years</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2226">37:06</a>] You mentioned the world doesn&#8217;t really need more programming languages. And if you think 10 years from now, do you think there will be less programming languages than there are today?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2239">37:19</a>] It&#8217;s hard to say. But like I said, AI tends to favor the incumbents. Right? And I do think, and I think this has been true even before AI, the bar keeps going up for what it takes to successfully implement a programming language and create an ecosystem around it. I mean, it used to be, oh, you just need a compiler. Well, you also kind of need tooling now. I mean, you need a language service, you need debuggers, you need profilers, you need frameworks and libraries, you need code generators that can target all sorts of.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2279">37:59</a>] I mean, it&#8217;s like it just keeps getting harder and harder to get all the way there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2288">38:08</a>] When you think about all those pieces, really, they&#8217;re all for the people. I mean, the IDE and the types. I mean, the computer looks at that too, but it&#8217;s so that we can read source. Maybe one day the LLMs will generate machine code directly.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2308">38:28</a>] I wonder. I think LLMs are better at what they do when they don&#8217;t have to repeat themselves a lot. I think, like when a program is expressed in text, it is closest to its sort of ultimate condensed meaning and representation. Right. If you translate it into machine code, there&#8217;s an awful lot of noise in that machine code. Like all of the instructions, like, half of the instructions are memory addresses that have absolutely no bearing on what&#8217;s going on here.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2349">39:09</a>] So half of all the data is noise already. Right. And then whichever register you pick, well, that doesn&#8217;t matter either in the code generator. And honestly, you could change your mind at any point in time. So teasing the truth out of that noise is a lot harder than if you just have a program where the variable is called I and whatever. And by the way, that you can relate to the written instruction that the user just gave you.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2374">39:34</a>] I mean, keep in mind that, like, AI are, they&#8217;re just emulators of humans in a sense. Right? I mean, it&#8217;s like neural networks. Right. There&#8217;s a reason we have programming languages because we&#8217;re terrible at writing machine code straight off out of our head. Right? AI is not that different from us, so the same would probably be true there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2396">39:56</a>] When you look back on working on C and <a href="https://en.wikipedia.org/wiki/TypeScript">Typescript</a> and these language projects, they were long journeys and I think, yeah, one thing people might want to know is what did you Learn through those journeys that if you knew at the beginning of when you embarked on working on them, you might have been better off.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2419">40:19</a>] Gosh, well, each journey is different. I mean like the first project I worked on, Turbo Pascal, I think</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2430">40:30</a>] that</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2430">40:30</a>] project started back in the old days when it was 8bit micros and 64k of memory and one person could do everything themselves and have ultimate control of everything. And so I was very much a one man shop. But that quickly was not scalable. Right. And capacities of machines. We went from 64 to 640k and then kaboom, the lid went off. Right. And you could have as much memory as you wanted. And so not one person couldn&#8217;t do it.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2459">40:59</a>] And I had to learn to become a team player. Right. And that, that was a big journey for me. I mean to learn to let go. Right. If you&#8217;re a perfectionist, that could be hard, but that&#8217;s one thing that I would say I learned there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2474">41:14</a>] Well, another way to word this question is, you know, what&#8217;s the most common mistake that you see people make when they embark on building a programming language?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2484">41:24</a>] Well, I think people tend to over index on one idea that they have that they think, oh, wouldn&#8217;t it be cool if my language could do blah. And then they under index on all of the mundane stuff that every programming language has to do and then it ends up doing that one thing they love, maybe better, but then it does everything else not as good. And the net is a detractor. Right. And so it&#8217;s so hard to convey how many things you have to do that really doesn&#8217;t have anything to do with the problem you&#8217;re trying to solve per se.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2523">42:03</a>] When you&#8217;re implementing programming languages, when you</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2526">42:06</a>] were working on C, do you keep tabs on the other languages and what they&#8217;re doing better, what you could do better with C?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2535">42:15</a>] Yeah, no, you got to keep yourself informed because like I said, we all stand on the shoulders of giants. No programming language was ever created in perfect isolation. Everyone learns from everyone and then someone comes up with a good idea and then you see it adopted elsewhere. We&#8217;ve adopted many ideas from functional programming in C, for example. Then other languages have adopted some of the good ideas we had in C, like Async now is in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and in some of the other programming languages out there.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2566">42:46</a>] And that was pioneered in C. You know, it&#8217;s, it&#8217;s, we all learn from each other and that&#8217;s how the industry moves forward. Really.</p><h3>42:57 &#8212; Hands on engineering vs delegation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2577">42:57</a>] You mentioned the, I guess, learning to let go. I think that&#8217;s something that a lot of people, as their career advances, they start with, you know, individual contributor work and then later they start delegating things. How did you learn that? Or what was the. Is there like a, maybe a story or particular project where you let go and you develop that skill?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2601">43:21</a>] Well, it&#8217;s almost like I got a counter story to the letting go there because in the C Sharp project, my main function was not implementing the C compiler, it was being the language designer and writing the spec and designing all the new language features. Before that, when I, when I had worked on Turbo Pascal, I was, I was doing a bunch of active coding, you know, on, on the thing itself, right? But then I sort of let go of that, but I let go of it too much and I sort of drifted up into the stratosphere and became more of an architecture astronaut, do you know what I mean?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2639">43:59</a>] And I wasn&#8217;t getting my hands dirty. I was writing stuff in the framework and in the libraries, but it wasn&#8217;t, it didn&#8217;t quite. These are not like the deep, really super hard problems to solve. And I didn&#8217;t really know it at the time, but I was gradually getting less happy with the work that I was doing. I still enjoyed, but there was something missing. And so when this <a href="https://en.wikipedia.org/wiki/TypeScript">typescript</a> opportunity came along, I decided to jump in and actually actively participating writing compiler.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2670">44:30</a>] And it brought so much happiness to my life. I finally realized I&#8217;m just a happy coder. I mean, I like the spec stuff too, but coding is like, man, that&#8217;s what gets me up in the morning, make coffee and write some code. You know, that&#8217;s when I&#8217;m in my happy space.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2688">44:48</a>] You know, when people are at these really, really high levels, they kind of have this expectation of a certain level of impact. And so how do you continue to write code as such a high level engineer?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2702">45:02</a>] Well, the specific code I&#8217;m writing isn&#8217;t necessarily having all that impact, but it&#8217;s the fact that I also do the architectural guidance as part of a group of people in a project, there&#8217;s plenty of work for everyone to go around, but you stay involved in a corner of the project because then you see what happens in the code. When you&#8217;re checking in your code, you go, what was, what was that over there?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2731">45:31</a>] Let me just, what&#8217;s happening over here? Or you realize, oh, we need to refactor this. This is not feeling right anymore. But if you let go of that, you can&#8217;t really talk about the thing that people use day to day. You can talk about the Syntax of the thing that people use day to day, but you can&#8217;t talk about that thing. So both are important, I think, and I&#8217;m happy that I get to do both.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2757">45:57</a>] Now, I know a lot of companies have this idea of engineering archetypes or the way that high level engineers get some work done. And you mentioned architect, for instance. A lot of companies have different words for architect. I&#8217;m not sure if you&#8217;re saying it&#8217;s just your working style or are you saying that all high level engineers should be hands on at some level?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2782">46:22</a>] No, I&#8217;m not saying that because I don&#8217;t think it&#8217;s like I made a conscious choice to remain an independent contributor as opposed to become a manager. I mean, I have plenty of opportunities, but that is not where I do my best work. Right. So I think you have to decide for yourself what is it that motivates you, makes you happy, keeps you doing this for a lifetime. Do you know what I mean? And you only get one run at it in life here.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2814">46:54</a>] And if you&#8217;re going to do your best work, do the thing that makes you the most happy. I think that&#8217;s what I learned along the way, that coding is part of that, a big part of that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2830">47:10</a>] So then when you were working on C, but before <a href="https://en.wikipedia.org/wiki/TypeScript">Typescript</a>, how did you know something was missing?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2836">47:16</a>] I just knew something. I don&#8217;t know how. I mean, I just wasn&#8217;t feeling as fulfilled with only doing the design part. Right. I missed the coding and whenever I had a chance to write a little bit of code, I would feel a lot happier. And I go, that&#8217;s kind of interesting. Do you know what I mean? But now I got to go write some more specs.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2861">47:41</a>] So then was <a href="https://en.wikipedia.org/wiki/TypeScript">Typescript</a> something you kind of went towards because you wanted to write more code?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2866">47:46</a>] Well, I went towards it for many reasons. I mean, it wasn&#8217;t specifically that to begin with. I just found the problem fascinating. Right? I mean, because the way <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> came about was actually indirectly through C. Because the Outlook. com team, I think it was, approached the C group and asked whether we would pretty please productize something called Script Sharp. Script Sharp was this thing that took C and transpiled it into <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> so you could run it in a browser.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2899">48:19</a>] I&#8217;m like, what? Why would you want to do that? Why don&#8217;t you just write the job? Well, because that way we can get great tooling, then we can use Visual Studio and we can have projects and we can have type checking and we can have code refactoring and navigation and we can write interfaces and people can understand how this all works. And I&#8217;m like, wow, is <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> really that busted? And you&#8217;re saying, like, the way you want to fix that problem is by abandoning <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, treating it as an instruction language.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2931">48:51</a>] Right. And then having a code generator targeted. I&#8217;m like, couldn&#8217;t we fix <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>? Wouldn&#8217;t that be better? Do you know what I mean? And that&#8217;s how <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> kind of got off the ground, Right? It was like, gosh, there&#8217;s something here that&#8217;s really broken that needs to be fixed. Right. When it comes to <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> tooling, and that&#8217;s what we set out to do.</p><h3>49:14 &#8212; Why fast tooling matters more now</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2954">49:14</a>] One thing I thought would be interesting is in one of the talks that you gave, you said that you believe that the speed of your tooling is more important now because of AI, And I want to hear thoughts on why is that the case.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=2970">49:30</a>] Well, more and more, our workflows now are dominated by agents that are, like, running away in the background. I mean, like, writing code for you. And they make mistakes just like we do. But you can have as many of them as you want. So there are like, many, many more little workers in the ecosystem writing much more code. Right. And that whole workflow is like, LLMs already have performance challenging.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3000">50:00</a>] And if you add, like, if you&#8217;re writing code in a big project and you have to type, check it, and it takes two minutes every time the LLM writes something. Do you know what I mean? That&#8217;s horrible. Right? So making that go faster is a big benefit.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3014">50:14</a>] Yeah, absolutely. I guess. Because it&#8217;s hammering away.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3017">50:17</a>] Yeah, Because, I mean, well, we all know LLMs, they like to use GREP and whatever. Right? But then they&#8217;re also smart enough now to realize that if you&#8217;re trying to change this property called version, and like every other object has a version property, you can&#8217;t just grep for version and then go rename it to something else because, like, you&#8217;re breaking all these other interfaces that have a version in them, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3045">50:45</a>] So you got to do semantic search, and the only thing that can do semantic search is a compiler. And so you got to use the compiler through, typically through lsp, through language Services, or through our command line tool. That would give you the equivalent. Right. And so more and more AI is doing that. And that means, you know, compilers run at. You may not even see that they&#8217;re running, but they&#8217;re running and they&#8217;re doing work on your behalf.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3075">51:15</a>] Yeah.</p><h3>51:16 &#8212; AI software engineering predictions</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3076">51:16</a>] When you talk about AI, it&#8217;s very realistic and grounded in the software engineering. And so I thought it might be interesting to run some common AI takes that I hear from the Internet and get your sense of how accurate is this? What are your thoughts on when or if it&#8217;ll be true? Okay, so the first one that I&#8217;ve heard this was more common at the end of last year, is that AI is going to write 90 plus percent of the code within a year.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3109">51:49</a>] Curious what you think on that.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3111">51:51</a>] Well, for certain classes of apps that might be like all of I coded apps are writing 100% of the code. Right. And there&#8217;s an awful lot of that being written. Are they going to write 90% of the high quality code on the Internet? I don&#8217;t know about that. So the thing is like 90% of what the volume of written code is just going like that right now. Right. Because AI writes so much code. So in a sense that&#8217;s a self fulfilling prophecy.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3142">52:22</a>] Right. But when it comes to the super high quality code that no one has ever written before, it&#8217;s not clear to me that AI is going to write 90% of that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3155">52:35</a>] So when you think about cases that AI is not writing all the code, maybe some examples of the super high quality stuff, what is that?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3164">52:44</a>] Well, I mean it&#8217;s algorithms that no one has ever seen before, for example, or particular solutions. Like, I mean, I know that AI could not write our compiler, the <a href="https://en.wikipedia.org/wiki/TypeScript">Typescript</a> compiler. It just can&#8217;t. I mean we&#8217;ve tried it can&#8217;t do it. And I&#8217;m not asking it to either. I mean I&#8217;m not expecting it to either. But our compiler is also not a typical piece of code. That&#8217;s why it&#8217;s bad at writing it, because it&#8217;s seen none of it in the training set.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3196">53:16</a>] Well, I&#8217;ve seen some compilers, but not like this one, not with these concepts, not with done the way that we&#8217;re doing it. Right. And so it&#8217;s just not a good fit there. AI, we can&#8217;t forget that AI is at heart a big stochastic machine that has memorized the entire Internet. Right. And has an ability to extrapolate somewhat over what it&#8217;s memorized. Right. But if it hasn&#8217;t seen it before, it&#8217;s not necessarily going to be all that good at doing it.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3229">53:49</a>] It&#8217;s getting better. I mean, I&#8217;m not saying it&#8217;s not impressive, it&#8217;s wildly impressive what it can do, but I just think there&#8217;s a long way still to it&#8217;s doing all of it. And I don&#8217;t know that we ever get there. I hope we don&#8217;t. Because then what&#8217;s the point of humans anymore then? I mean, it&#8217;s.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3249">54:09</a>] What I see this take often too is you won&#8217;t need to use an IDE within the year or you don&#8217;t need to read the code anymore.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3259">54:19</a>] Well, good luck to you. I mean, if you&#8217;re going to hand AI the keys, I mean, then all bets are off, right? I mean, the minute AI you can&#8217;t convince AI to fix this issue or this new feature that you want built, what are you going to do? I mean, if you don&#8217;t understand what&#8217;s below? Right. That&#8217;s one problem. The other problem is, let&#8217;s say a user comes to you and goes, your app just stole my bank account.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3286">54:46</a>] Well, it&#8217;s going to sue you. They&#8217;re going to sue you. Not the AI. I mean, it&#8217;s like at the end of the day, someone has to take responsibility for what AI is doing. And if you&#8217;re giving it the keys, well, it&#8217;s on you. But I don&#8217;t know, I personally would not feel comfortable. I wouldn&#8217;t be able to go to sleep at night, I mean, if I didn&#8217;t understand what it is that I am vouching for.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3313">55:13</a>] Why is there such a big difference between this opinion and what I see from maybe people who work at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3322">55:22</a>] Well, the difference is, all I&#8217;m saying here is not that AI can&#8217;t do amazing things. It does, it can. It does that every day. But people love these, like, take it all the way to 100%. Right. And that&#8217;s where I get off the bus a little bit. Right. I think, oh, yeah, okay, fine, we can get higher and higher percentage. But there still has to be someone understanding what&#8217;s going on and how it&#8217;s relevant to our business problem and how it relates to what our organization is doing.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3353">55:53</a>] And like, I mean, there&#8217;s so many other things that have to be connected that it&#8217;s. I think it&#8217;s wrong to just view it as AI did it all and we were completely unnecessary, you know?</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3364">56:04</a>] Yeah. The next one I was going to ask you was, you know, three years from now, will AI replace junior software engineers or what junior software engineers do today?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3373">56:13</a>] Well, if it does, then how do we ever get senior software engineers? Right? I mean, it&#8217;s just like, I don&#8217;t think so. Well, I&#8217;m not. Unless you also believe that you will be able to do your business three years later than that without any senior software engineers.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3390">56:30</a>] Right.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3390">56:30</a>] Because where are they going to come from you got to train them, I mean. But to me, I think maybe one thing that to some extent has happened is that for certain classes of programmers, the pyramid has narrowed at the bottom. There are fewer people coming in, and we are wanting them to advance higher up faster so that they can get to this more supervisory role. I think the craft of being a software engineer is definitely changing from your typing in lines of code to you&#8217;re having agents type in lines of code and you are reviewing the work of the agents.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3434">57:14</a>] So you do more reviewing, less typing. Some people love that and really are fantastic at that workflow. But it&#8217;s a change in nature of what the job looks like as a programmer and this tool that you now have in a toolbox is having that impact. But it is but a tool and there are still programmers involved in the process that I firmly believe.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3462">57:42</a>] It sounds like you enjoy typing out the code. Would you be less happy if five years from now, for 40, 50 years I have enjoyed typing out code?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3468">57:48</a>] I&#8217;m actually. And so I will continue to type certain parts of the code, but there are other parts of the code that I don&#8217;t particularly care to type. I don&#8217;t care to type in tests and follow all the rigor of the particular testing framework. Heck no. If I can farm that out to AI happy me, I do think reviewing has always been harder for me. It&#8217;s harder for me to spend time on reviewing other people&#8217;s code than it is writing the code.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3504">58:24</a>] But I also think that over time we can make that process of reviewing code much more ergonomic than it is today. Right. I mean, often today you just get, here&#8217;s like a list of the files that were changed and here are the deltas. You figure it out. Right? I mean, AI could help us more there by explaining what the changes are and whatever. And so I think we&#8217;re going to get better at this. It&#8217;s going to get more interesting, but it will change the nature of, of the craft.</p><h3>58:52 &#8212; The most technically challenging work</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3532">58:52</a>] You said that you&#8217;ve been writing code for 40, maybe 50 years. What is the hardest piece of code you&#8217;ve ever written? Or what&#8217;s the most technically challenging thing you&#8217;ve ever done?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3546">59:06</a>] There are so many things that in their time were challenging. I mean, this last project we worked on, we knew that the easy part was getting the native code that&#8217;s just like just move it to this other language. But then taming this beast called shared memory concurrency and getting the compiler to use as many CPUs as there are available on your box that is not a simple problem to solve in a compiler because we&#8217;re all taught to write sequential stuff, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3588">59:48</a>] I mean, first you do this, then you this. But how do you do all of this at the same time and then share the data structures afterwards without anyone like accidentally modifying the other guy&#8217;s data structure? And you know, how do you, in the cases where you have to do that, how do you then synchronize it and avoid deadlocks and race conditions and whatever, Right? This, it was tricky. There were some good technical problems that we licked that I think we can be proud of, you know.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3613">01:00:13</a>] But what about when hardware was more constrained earlier in your career? Is there any wacky thing you had to do to kind of, you know, make it work?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3623">01:00:23</a>] Oh, totally. I mean, when I started, it was a. Like being a programmer was a very, very different thing, right. I mean, I started out with, with 8 bit micros that had <a href="https://en.wikipedia.org/wiki/Microsoft_BASIC">Microsoft Basic</a> in ROM, right? And then the BASIC interpreter and you could buy these books and type in these <a href="https://en.wikipedia.org/wiki/Star_Wars">Star wars</a> games or whatever. They were all text, you know, there&#8217;s just scrolling and telling you there&#8217;s <a href="https://en.wikipedia.org/wiki/Klingon">Klingons</a> over in that quadrant, you know, and K4, you got whatever, right.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3655">01:00:55</a>] But I got fascinated with understanding how it was all put together and quickly learned how to write assembly code. And Turbo Pascal, the first product I worked on was all written in <a href="https://en.wikipedia.org/wiki/Zilog_Z80">Z80</a> assembly code. The compiler, the editor, the runtime library, everything was assembly and writing structured. A compiler in assembly code is like that. A lot of fiddling. And you were counting like, okay, I think we need to squeeze this into this <a href="https://en.wikipedia.org/wiki/EEPROM">EEPROM</a>.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3690">01:01:30</a>] And then there&#8217;s only 12k and I&#8217;m 20 bytes over budget here. Oh, but I can make this jump, this long jump instead. I could jump to this short jump and then jumps to that jump and then I could save one. That&#8217;s one bite saved. Okay. You know, moving right along and you would just like sitting there fiddling, you know, it was a craft. It was like woodworking almost. I mean, it was just so different from what we do now, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3717">01:01:57</a>] Everything is now it&#8217;s bottomless pits of capacity and bottomless pits of expectation from users.</p><h3>01:02:04 &#8212; Top book recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3724">01:02:04</a>] Right. Is there a top technical book recommendation you have for software engineers?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3729">01:02:09</a>] I usually recommend the same, but the one book that was like the IT book for me when I was younger, and it&#8217;s called Algorithms plus Data Structures Equals Programs by Niklaus Wirth, the inventor of Pascal and then later Modular and <a href="https://en.wikipedia.org/wiki/Oberon">Oberon</a> and whatever. And it&#8217;s a book that is light on math and Symbolism and rich in instructive examples and good explanations about how data structures work. I mean, that&#8217;s how I learned about hash tables, right?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3763">01:02:43</a>] I mean, like, when you&#8217;re self taught, you don&#8217;t know about hash tables, you sort of know about linked lists, right? And fine. So in the original version of Turbo Pascal, all the symbol tables were just, well, they were just linked lists, you know, or whatever. And then of course, as you had many local variables, well, you&#8217;re going exponential in search time. And then I read about these hash tables and like, wow, what you mean I can just do this?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3791">01:03:11</a>] And then it&#8217;s basically like linear lookup time and implemented and the compile went twice as fast. Right? Like, holy cow, that was useful reading. Do you know what I mean? So I was always like the engineer, I mean, understanding how to do error recovery. He explains that in how you construct a simple compiler. And then like, you know, and it was, for me, a great book. And now it&#8217;s actually like, I mean, it&#8217;s just a big PDF, you can download it out there.</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3826">01:03:46</a>] It&#8217;s out of print since 40, 30 years. Yeah.</p><h3>01:03:50 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3830">01:03:50</a>] Last question for you is, with all the experience that you have now, if you go back to yourself when you just entered the industry and give yourself some advice, what would you say?</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3841">01:04:01</a>] Well, a couple of things. Maybe like we talked about being a team player, you know, I would have probably tried to educate myself a bit more about that to begin with. I think also one thing in retrospect is like, don&#8217;t let people tell you that it can&#8217;t be done. When they tell you it can&#8217;t be done, it&#8217;s because they can&#8217;t do it. That doesn&#8217;t mean that you couldn&#8217;t possibly do it. Do you know what I mean?</p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3869">01:04:29</a>] So, so don&#8217;t let that discourage you.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3871">01:04:31</a>] Is there a project where you. Someone said you couldn&#8217;t do it and then you did it and that&#8217;s how you.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3878">01:04:38</a>] Well, I think the first thing I worked on, like Turbo Pascal, people were telling me, whenever we told them, here&#8217;s what we have, they said, that can&#8217;t be done. That&#8217;s like, nah, you guys are beautiful. That&#8217;s like not possible, you know? Well, it was and I didn&#8217;t know that it wasn&#8217;t possible.</p><p><strong>Ryan:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3895">01:04:55</a>] Right, awesome. Well, thank you for your time. I really appreciate it. Anders.</p><p><strong>Anders:</strong></p><p>[<a href="https://www.youtube.com/watch?v=cywK3XYYJ2o&amp;t=3898">01:04:58</a>] You&#8217;re welcome. No, this is a lot of fun.</p>]]></content:encoded></item><item><title><![CDATA[Creator of Lean: Handwritten Math Will Change Dramatically | Leonardo de Moura]]></title><description><![CDATA[In 2024, AlphaProof from Google Deepmind broke through in competition math achieving a silver-medal in Interational Mathematical Olympiad (IMO).]]></description><link>https://www.developing.dev/p/creator-of-lean-the-end-of-handwritten</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-lean-the-end-of-handwritten</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 10 Aug 2026 13:03:04 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/209970277/4b08d5b419149374be2933e522f2097d.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>In 2024, <a href="https://deepmind.google/blog/ai-solves-imo-problems-at-silver-medal-level/">AlphaProof</a> from Google Deepmind broke through in competition math achieving a silver-medal in Interational Mathematical Olympiad (IMO). At the time, that was an impressive breakthrough by formalizing math problems in Lean and then having an LLM try to advance the proof. Lean was essentially the scaffolding for a &#8220;game&#8221; that AI could play.</p><p>More recently, frontier labs like OpenAI have been generating massive candidate proofs to unsolved problems in frontier math <a href="https://openai.com/index/ten-advances-in-mathematics/">(e.g. OpenAI recently released these 10 advances)</a>. Lean plays a critical role in formally verifying these proofs.</p><p>I wanted to learn more about Lean so I talked with the creator, <a href="https://www.linkedin.com/in/leonardo-de-moura-26a27b5/">Leonardo de Moura</a> recently. He&#8217;s also a <a href="http://amazon.science/news/amazon-is-investing-in-the-lean-focused-research-organization">key scientist at AWS</a> for their neurosymbolic AI initiatives, bringing Lean and formal verification together to prove agentic AI systems correct. He explained how Lean works and why LLMs plus Lean will fundamentally change how we write software and do math. Hope you enjoy this episode!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/KzdYKeAqWhY">YouTube</a>, <a href="https://open.spotify.com/episode/34dbrI4zjw94K2R4S43BI3?si=ZiaycTbmTG6xgr3Ofq51Dw">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-KzdYKeAqWhY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KzdYKeAqWhY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KzdYKeAqWhY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/209970277/0028-how-formal-verification-works">00:28 - How formal verification works</a></p><p><a href="https://www.developing.dev/i/209970277/0521-a-new-way-of-writing-software">05:21 - A new way of writing software</a></p><p><a href="https://www.developing.dev/i/209970277/1315-proof-assistants-vs-programming-languages">13:15 - Proof assistants vs programming languages</a></p><p><a href="https://www.developing.dev/i/209970277/2106-how-lean-has-assisted-in-mathematical-breakthroughs">21:06 - How Lean has assisted in mathematical breakthroughs</a></p><p><a href="https://www.developing.dev/i/209970277/3203-when-is-it-worth-formalizing-software">32:03 - When is it worth formalizing software</a></p><p><a href="https://www.developing.dev/i/209970277/3329-how-lean-will-impact-handwritten-math">33:29 - How Lean will impact handwritten math</a></p><p><a href="https://www.developing.dev/i/209970277/3855-the-z3-theorem-prover-project-he-started">38:55 - The Z3 theorem prover project he started</a></p><p><a href="https://www.developing.dev/i/209970277/4544-the-most-technically-challenging-work-of-his-career">45:44 - The most technically challenging work of his career</a></p><p><a href="https://www.developing.dev/i/209970277/5110-lean-vs-its-competitors">51:10 - Lean vs its competitors</a></p><p><a href="https://www.developing.dev/i/209970277/010037-the-future-of-lean">01:00:37 - The future of Lean</a></p><p><a href="https://www.developing.dev/i/209970277/010410-technical-book-recommendations">01:04:10 - Technical book recommendations</a></p><p><a href="https://www.developing.dev/i/209970277/010615-advice-for-his-younger-self">01:06:15 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:28 &#8212; How formal verification works</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=28">00:28</a>] There&#8217;s this famous <a href="https://en.wikipedia.org/wiki/Dijkstra">Dijkstra</a> quote I wanted to start the conversation with. It&#8217;s that program testing can be used to show the presence of bugs, but never to show their absence. And my understanding is that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> and formalizing proofs can be used to show the absence of bugs. And so in your words, what is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> and how do people use it to show bugs can&#8217;t occur in programs?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=59">00:59</a>] <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a programming language. You can write code, but you can also write proofs. You can reason about your code. You can write state properties about your code and prove them. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> gives you machine-checkable proofs. You can check your proofs and get absolute assurance they are correct. You have many checkers, independent checkers, but you should view <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> as a platform. You can write code, you can write properties about your code, and you can prove them.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=90">01:30</a>] So could you give a concrete example? Because when I think of a proof, I think of what I learned in math. But then how do you couple that with the software that we write?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=102">01:42</a>] Yeah, it&#8217;s a great question. It&#8217;s not that different from proof. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is actually very popular for math, but for software verification, there are two very different use cases. You can reason about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> programs. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a programming language. You can write programs in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> itself. Then a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> program is not that different from a definition you have in math. The techniques are very similar. But if you want to verify programs written in a different programming language, there are basically two different approaches.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=140">02:20</a>] One of them, they translate, called shallow embedding. They translate for <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>. This happens today. We have a tool called Aeneas that maps <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> into <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and you can verify the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> translation. There&#8217;s another technique called deep embedding, where you have the semantics. You write a semantics of the programming language of C in. And now you have a data structure that represents a C program, and you can state properties about it, almost like your programs become <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> objects that you can reason about.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=178">02:58</a>] To make it really concrete, to give someone a sense of, here&#8217;s this thing I want to prove about a simple C program. Maybe no buffer overrun or something like that. What&#8217;s the step by step where we could use <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to prove that?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=193">03:13</a>] Yeah, yeah. Let&#8217;s get an array. You&#8217;re trying to access an array in C. You want to make sure the index is in bounds. You&#8217;re not accessing elements that your array, for example, has 10 elements. You&#8217;re not trying to access element 11, right? Basically, you can write that. You can write something that says the value of i at this point in your program is going to be greater than or equal to 0 and less than 10.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=224">03:44</a>] You can write that as a mathematical statement. Another way to view it is that if you can express in math what you care about in your program, you can verify using <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. That&#8217;s another way to view it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=242">04:02</a>] So I have my C source files, and then somehow there&#8217;s an equivalent <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> proof that&#8217;s almost like metadata on top of the C program.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=251">04:11</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=252">04:12</a>] And that&#8217;s a line-by-line proof. And <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> will go and check.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=256">04:16</a>] Exactly. People will build automation for automating the process. They&#8217;re going to use techniques like Hoare triples, that says like we have a precondition, some mathematical facts that should be true before executing that statement, the statement and what is true after. Right. And they will have a lot of automation to process makes your proof modular. Right. I mean, complexity is a challenge in software verification.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=288">04:48</a>] Handling the complexity is a big deal. And all these frameworks for verifying programs, they are trying to manage the complexity, make your proofs modular, even if AI. Now AI can prove things automatically for us, but they have to be modular, the proofs, if you want them to scale.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=308">05:08</a>] And what you&#8217;re saying sounds similar to software, where you have to write it cleanly. It needs to be easy to edit and reason about. So it&#8217;s almost like a second software layer on top of the software.</p><h3>05:21 &#8212; A new way of writing software</h3><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=321">05:21</a>] Yes, yes, you can view it this way. You can also flip it and imagine a future where you&#8217;re writing what you want mathematically, precisely, and AI synthesizing the code and a proof that the code that was synthesized meets your specification. Right.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=342">05:42</a>] Oh, interesting. So you could start by writing what you want to be true and then ask the AI, please.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=350">05:50</a>] And AI is going to just go and hammer at it until the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> proof says you&#8217;re good.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=355">05:55</a>] Yes. It feels like science fiction. Six months ago, I would say this is science fiction. But for example, now a colleague of mine, Kim Morrison, a few weeks ago, she started a project I thought was, six months ago, I would say it&#8217;s not possible at all. She said, we have zlib, this compression library written in C. She says she creates a very complicated prompt for AI saying I want you to translate to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=388">06:28</a>] Ensure the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> version passes the test suites for zlib. Then I want you to prove that if you compress data and you decompress, you get the original data back. It&#8217;s a really strong property and, believe it or not, after one week, she succeeded doing the whole thing, and now it&#8217;s just asking to optimize the code. But you cannot break the proofs. You have to keep still proved all the properties you care about.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=422">07:02</a>] I mean compressing and decompressing, getting the data back is a really important property for a compression engine. Right? Yeah, this is rich now. I mean, surreal.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=436">07:16</a>] I mean, it&#8217;s crazy. And I hear in the industry, a lot of people use a really comprehensive test suite coupled with AI to do some really amazing rewrites because they have some more confidence that the rewrite is accurate and the AI can check itself. And it sounds like a specification is even better than a comprehensive test suite.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=465">07:45</a>] Yes, because your quote from <a href="https://en.wikipedia.org/wiki/Dijkstra">Dijkstra</a> captures it perfectly, right? With a test suite, you can show the presence of bugs, but not their absence. It&#8217;s almost like with a good test suite, you may say, well, probably there are no bugs here, but you may have a really corner case that&#8217;s not covered by your test suite. But with a proof, you&#8217;re covering all possible cases. So property-based testing is really popular now.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=499">08:19</a>] People are writing properties. They want to ensure they are true, but they are checking with testing. But now we can prove them. And you say, look, there&#8217;s no point testing anymore.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=512">08:32</a>] It seems like having a well-written specification is a superset of a test suite. But in terms of the human labor required to create a reasonable test suite versus a reasonable specification, how much more work is it to come up with a great specification? It varies a lot by the program.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=544">09:04</a>] In many cases, it&#8217;s not uncommon for someone to start developing a piece of software, and they don&#8217;t know exactly what the spec is. But properties are usually vaguely in your mind. Another thing that I tell people to keep in mind is that an inefficient program can be viewed as a specification. Usually, writing an inefficient program is way easier than writing the super efficient one that has many clever tricks.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=575">09:35</a>] You can write a very, this is what I want, in a very naive way. And you ask the AI, look, generate the efficient version and optimize and prove that equivalence to my inefficient one. There are many scenarios. I mean, I will not say this. Specifications are always easy to come up with. But properties, usually the developers have good ideas about properties they care about. An inefficient program is a spec, right?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=608">10:08</a>] You can view it as a specification, and the technology form of verification complements testing, right? I think your quote from <a href="https://en.wikipedia.org/wiki/Dijkstra">Dijkstra</a> captures...</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=619">10:19</a>] Perfectly, and I think I saw this on Twitter because Jane Street was more heavily investing in formal verification, and I read their post, and they talked about this. I think it was seL4 or some software that was entirely formally verified.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=637">10:37</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=638">10:38</a>] And the main drawback to why it wasn&#8217;t. You know how much more time would it take writing a program versus verifying it?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=648">10:48</a>] Your example, seL4, is a great example. This was a major milestone. It&#8217;s a microkernel they verified manually before AI. This project was done before AI was a big deal, and the cost is super expensive, right? At AWS, we have been using formal verification for decades, but only for the super-safe, critical components because it&#8217;s expensive. Until now, right? With AI, it changes the game.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=681">11:21</a>] You have to come up with a spec. But this is not the most painful part. The most painful part is developing the proofs manually. If you have tools before AI and maintain the proofs as you change the code, it&#8217;s messing you up. I&#8217;ve seen people complaining that, &#8220;Oh, I changed the program. Right now I have a bunch of failures in my test suites, and I have to patch them one by one.&#8221; Imagine with proofs, it&#8217;s the same process.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=710">11:50</a>] You have to fix the proofs. Sometimes you don&#8217;t remember anymore why. What&#8217;s the story behind this proof? It&#8217;s a lot. But AI is extremely good at proving, writing formal proofs, maintaining formal proofs. For example, yesterday I was changing something. I want to modify some proofs for technical reasons. I didn&#8217;t even know what the proofs were about. Someone else wrote. Then I said, look, I want you to&#8212;I asked AI to rewrite these proofs without using this feature because I&#8217;m going to change it.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=751">12:31</a>] I don&#8217;t want to break the libraries instantly. It came up with the new proofs for me. I mean, it&#8217;s really good. And this is crucial for making formal verification mainstream because otherwise maintaining the proofs was almost like if your program took X amount of time to do the formal verification in the past, it would take 10x. That would be normal. But imagine if your program is changing. That&#8217;s something that&#8217;s really common now.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=785">13:05</a>] You have to keep maintaining the proofs too. It&#8217;s a lot of work. But AI eliminates this pain for us.</p><h3>13:15 &#8212; Proof assistants vs programming languages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=795">13:15</a>] You mentioned that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a proof assistant, and then you also said it&#8217;s a programming language. Is that typical for proof assistants to be both a programming language and a proof assistant?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=810">13:30</a>] Some of them, especially the ones that are based on <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a>, they are like <a href="https://en.wikipedia.org/wiki/Rocq">Rocq</a> and <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> &#8212; programming languages and proof assistants, right? You can write definitions like when you&#8217;re defining concepts in math. But some of these definitions can be programs, and you have types, you have structures. It&#8217;s a programming language, but it&#8217;s in the family of what&#8217;s called functional programming languages.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=841">14:01</a>] I mean, is it a specific kind of programming language for people that are familiar with programming languages like <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>? <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is close to <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, but with the support for proofs. That&#8217;s a way to view <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=856">14:16</a>] Are there major use cases where people use it as a programming language but not as a proof assistant?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=861">14:21</a>] Well, the first big use case is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. We have many of our tooling implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. Like the documentation authoring system called <a href="https://en.wikipedia.org/wiki/Recto_and_verso">Verso</a> is implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. The build system that&#8217;s called <a href="https://en.wikipedia.org/wiki/Lake">Lake</a>, like <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> make, is implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. At AWS, we have a compiler for AI accelerators. It&#8217;s half a million lines of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> and using <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> as a programming language. They&#8217;re proving some properties about the program using <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=894">14:54</a>] But the main goal is to use <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> as a programming language. In this project, the proofs are like a bonus that you can get the proofs and find problems in the design.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=906">15:06</a>] I think most people are familiar with programming languages and the toolchains they have. But what are all the major components that you would need for a proof assistant?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=919">15:19</a>] It is not that different. I mean, if you&#8217;re used to modern programming languages like <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, the tooling, for example, <a href="https://en.wikipedia.org/wiki/Lake">Lake</a> is our <a href="https://en.wikipedia.org/wiki/Cargo">Cargo</a>, right? I mean, you&#8217;re going to open <a href="https://en.wikipedia.org/wiki/Visual_Studio_Code">Visual Studio Code</a> the same way. And I want to get all the IntelliSense. One big difference is that we have something called the Infoview in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. Your screen usually is going to be split in two. You have your file on the right-hand side, you have the Infoview that tells you information about your proofs, about your code.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=955">15:55</a>] It&#8217;s giving you constant feedback about your development. That&#8217;s basically the main difference. But the tooling, VS Code, everything works the same way.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=966">16:06</a>] I mean, it seems like the programming language is the fundamental layer, and then there&#8217;s some additional layer on top that keeps track of that stuff.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=974">16:14</a>] Good question. In <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> you have definitions that show your program, but you have theorem statements. You&#8217;re going to say, for example, factorial is always greater than or equal to zero, or something like that. Or if you add two even numbers, you get an even number. You can write statements like that, and immediately you can write actually a program that&#8217;s the proof. But most people don&#8217;t do that. They go into something called tactic mode.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1007">16:47</a>] You can view it as a domain-specific language for writing proofs. When you write &#8220;by,&#8221; it switches to this domain-specific language, and you have steps like &#8220;simplify my goal,&#8221; the state of my proof. You can say, &#8220;Oh, apply this rewriting step.&#8221; Apply, for example, we know that x + 0 equals x. You can ask <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to apply this rewriting, and you&#8217;re going to see in the Infoview the state of the proof changing. You get immediate feedback, and you get this feeling that you said.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1041">17:21</a>] You keep telling it, applying transformations to your proof step by step. You can see what&#8217;s happening until you get no goals left. You are done; the proof is complete. And some people view this process as a game. I have users that told me, &#8220;You built my favorite computer game.&#8221; That&#8217;s funny.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1066">17:46</a>] Is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> itself verified in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1069">17:49</a>] <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a massive program. You only have to trust the kernel, where the proofs are checked. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> itself is the kind of program that, because there are so many new things we are adding, it&#8217;s not even clear what the specification is. For example, what is the specification of a simplifier? You can write general ideas, but users want to be able to customize the behavior. They keep asking; they keep changing the specification all the time.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1098">18:18</a>] I want this, I want that. No, at this knob here. So it&#8217;s really hard to have a full specification of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, but the kernel is possible to have a specification. We have a kernel. The kernel that comes with <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is not verified. But there are other kernels that you can use. One of them is implemented by Mario Carneiro. It&#8217;s called <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> for <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and it&#8217;s implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. And he&#8217;s verifying, it&#8217;s proving that this kernel has been verified with respect to the semantics of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1132">18:52</a>] This is a cool project, but for us, having multiple kernels is the best way to ensure that your results are correct. Some users implement their own kernels. We have kernels implemented in <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> in different programming languages.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1150">19:10</a>] And when you say kernel, what&#8217;s the responsibility, or what&#8217;s the input and output of that portion?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1156">19:16</a>] Yes, in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, proof checking is type checking. These kernels, they are type checking your programs. Basically, you can export your <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> developments. You get a big blob, and you read this blob, and they&#8217;re going to check it. When you say we have a proofing link, basically you have a term that has a type, and you&#8217;re checking if the type of this term matches the type that you claim it has. For example, the example for the even numbers, this is a typing link saying that the sum of two even numbers is an even number.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1197">19:57</a>] You can view that it is exactly a typing link. And the proof, what the kernel is checking is whether the type of the proof matches the type you claim this term has. The kernel will check, you do this type checking. The kernels vary between high-performance kernel. It&#8217;s like 5,000 lines of code. I mean, the goal of the kernel should be something you can write yourself. Of course, sometimes people put bells and whistles. One thing concerns people have is like, okay, but how do I know?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1239">20:39</a>] Suppose that I wrote Fermat&#8217;s Last Theorem in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. How do you know that when you export it, it&#8217;s really Fermat&#8217;s Last Theorem? It&#8217;s not 2 plus 3 equals 4. I mean, and people, some external kernels, they write print printers. I mean, it will print the statements that have been proved, all the dependencies. You can have all these fancy tools to make sure you&#8217;re not being misled.</p><h3>21:06 &#8212; How Lean has assisted in mathematical breakthroughs</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1266">21:06</a>] <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, obviously, it&#8217;s so powerful, and I see on social media all these amazing results from formal verification. What are the top ones that you think of, or top examples more recently, that have impressed you that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> was able to accomplish?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1285">21:25</a>] Well, there are so many. I mean, I thought it was impossible. For example, getting a gold medal in the International Mathematical Olympiad a few years ago, everybody thought it was impossible. Now everybody use it as a benchmark of an easy problem. They say, oh, this is like an IMO problem. I mean, this is easy, but it&#8217;s not. I mean, these are really challenging problems. The other conjectures that people close using AI with <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> also impresses me.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1319">21:59</a>] There is this unit distance conjecture for others. First, <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a> proved using formal was not formal. And we have a system now in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. It&#8217;s a website where we call <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Eval, where we collect challenges. The same Kim Morrison saw this proof <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a>, and she put on <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Eval a challenge saying, okay, I want to see someone formalizing. We knew it was huge to formalize. It&#8217;s serious math. It depends on.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1358">22:38</a>] And Boris Alexeev from <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a>, he did, he made a 1 million line proof for formal proofing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> for this conjecture. We estimate something though that background math that is needed for the proof. We knew it would take months, I mean for experts to do by hand.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1378">22:58</a>] I mean, how long did it take in that case? Like the time from when the proof came out to when the challenge was solved on <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1386">23:06</a>] I think it was one month, yeah. And after Kim was a few, I think less than two weeks. After Kim puts as a challenge only involved, it took I think two weeks to get the, or less. I mean, that&#8217;s insane. Yes. Yeah, yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1404">23:24</a>] We talked about verifying programs, but also there&#8217;s using it just directly for mathematics, like in this case. How popular is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> in terms of when you look at the users of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>? What percent are using it for software engineering? What percent are using it more for just direct math?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1424">23:44</a>] Well, historically <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> got popular first with math, right? I mean, with the beginning of the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> mathematical library in 2017, we had a big project called Liquid Tensor Experiment. It started beginning of 2020. It was a big deal because it was to verify a result from a Fields Medalist, Peter Scholze. It was a result he was unsure about. He has not published it. He felt like this was one of the most important results in his career.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1462">24:22</a>] He wanted to be sure it was correct. It was done manually. The verification was a big deal because the team that formalized it, led by Johan Commelin, not only formalized the results without fully understanding it, they did not fully understand the proof. But having this Infoview helped guide them. Step by step, they formalized and simplified the proof without fully understanding the proof. It&#8217;s mind boggling.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1494">24:54</a>] How can you simplify a proof from one of the greatest living mathematicians without fully understanding it? I mean, but they managed to simplify. It&#8217;s almost like when people do refactoring code, you start changing the code. It&#8217;s faster now, but the programmer doesn&#8217;t really know why it felt like that. It&#8217;s like you have a gut feeling, I mean, that you&#8217;re going in the right direction. And this, for us at the time, lots of people got excited about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, the math community, because of this project. It&#8217;s shown that it&#8217;s not about verifying, but enabling people to work together in large numbers.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1536">25:36</a>] Because you can trust, you don&#8217;t need to trust someone else&#8217;s proof. Right. They can fill holes for you. This was, I mean, what attracts a lot of attention from the math community. Then came Terence Tao. He starts using <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> after this project. He has really cool projects with and without AI. He got addicted, actually. I think the first time he used it, he said, &#8220;I don&#8217;t think I&#8217;m going to do it again.&#8221; The formal proof, one week later he had another project using <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and he did a new result he had.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1573">26:13</a>] When you say addicted, you mean because of that game, that completion engine.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1577">26:17</a>] Yeah. Some people, I feel, we feel like when I talk to professors that use <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> for teaching, they tell me that the class is more or less split. Some people love it, some people don&#8217;t like it. But people that like problem solving, people that get medals in the IMO, they love <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> because you get this excitement of solving. It&#8217;s like solving millions of really hard Sudoku problems, right? And you can keep solving one, and it&#8217;s easy to get.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1609">26:49</a>] I mean, I&#8217;m a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> developer, not a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> user. But when you&#8217;re developing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I&#8217;m coding in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and sometimes I have proved things. It&#8217;s really easy to get addicted. You get lost proving things. You get this buzz every time you prove something.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1627">27:07</a>] You mentioned the IMO gold medals. How is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> used in that kind of... I think maybe you&#8217;re referring to AlphaProof by DeepMind, or maybe something else.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1637">27:17</a>] Yes, DeepMind was. They got in 2024 was a big surprise. In 2024 they got a silver medal. And now we have gold medals from startups like <a href="https://en.wikipedia.org/wiki/Harmonic">Harmonic</a>. <a href="https://en.wikipedia.org/wiki/ByteDance">ByteDance</a> got a medal. I never imagined. <a href="https://en.wikipedia.org/wiki/ByteDance">ByteDance</a>. They&#8217;re behind <a href="https://en.wikipedia.org/wiki/TikTok">TikTok</a>. Right. I didn&#8217;t even know they care about IMO math, but they have a 400 math team. They got medals. Also their prover is really good.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1668">27:48</a>] So how&#8217;s it work? Let&#8217;s say, I mean, there&#8217;s a series of math problems. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is just used for verification, right? So I imagine there&#8217;s other components there.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1677">27:57</a>] Yeah, there&#8217;s the AI. You can view it like we have Google. That was very popular with AI. The AI is playing the game. I told you that. Many people see <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I guess. It&#8217;s the same. The AI is viewing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I guess. They have the statements of what you want to prove. You have the by keywords. Now you have a blank field. Make sure that you have no goals left. It keeps applying steps in this game and seeing the state of the board.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1710">28:30</a>] That&#8217;s the Infoview changing. In the AI, they use reinforcement learning to try to get to no goals left. It&#8217;s a single-player game. Some people play together these days, but the AI is learning to play the game.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1729">28:49</a>] So before we were talking about that game, it kind of starts with the specification. So in that case, then the problem is kind of the specification, and then it hammers over the problem.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1740">29:00</a>] Yeah. You have basically, for each problem in the International Mathematical Olympiad, someone translates the statements to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. It&#8217;s super important to have the mathematical library because you want to be able to talk about the problems, right? For example, suppose the problem uses the real numbers. You need a definition in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and we have it inside of mathlib, the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> mathematical library. The first thing you have to do is be able to write the problems in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1767">29:27</a>] And this now because of mathlib, it&#8217;s the easy part. And after that you have to provide the proofs. I mean, sometimes some problems you have to come up with a definition to some objects you have to create that has some property. But yeah, that&#8217;s how it is.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1787">29:47</a>] I remember you said one of the first use cases of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> was this Fields Medalist had this novel mathematics, and then we used <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to verify it. But can <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> be used with LLMs to generate novel mathematics, or just verify existing?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1808">30:08</a>] It&#8217;s a good question. I mean, we don&#8217;t see lots of evidence. We see no evidence it can find novel proofs, right? But coming up with new mathematical concepts is still at the limits, right? I mean, right now in this <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Eval, we have challenges where the AI has to come up with the objects themselves, right? I mean, but this is. People will keep investing in this area. I would not bet against AI here.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1842">30:42</a>] I mean, but right now we don&#8217;t have evidence they can come up with new math.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1894">31:34</a>] When I see on Twitter this major conjecture, they made headway on it. It&#8217;s that they made headway on confirming something or disproving.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1903">31:43</a>] They also have these 1 million lines. They showed the conjecture was false. They have a proof showing that it&#8217;s false, but it&#8217;s a formal proof. They did not come up with a new theory or anything like that. They are coming up with a proof, right?</p><h3>32:03 &#8212; When is it worth formalizing software</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1923">32:03</a>] And at what point would you say someone should evaluate formalizing something in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1932">32:12</a>] I think if it&#8217;s safety-critical or if you don&#8217;t understand it really well, I mean this is an important part. I don&#8217;t really understand. Anybody that went through the process of formalizing something understands the subject way better after that. I almost feel like I remember when I was in college, people would say, wow, after I implemented this algorithm, now I understand it much better. The next level is that you implement the algorithm, you prove the properties you expect, your level of understanding grows.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=1972">32:52</a>] Another cool thing is that it enables you to be much more bold on your optimizations, because sometimes I&#8217;ve seen that all the time. People fear implementing optimization because they don&#8217;t really understand why the piece of software works. They feel like if I do, that still works and it&#8217;s faster. But they are not confident with proofs. You eliminate this discomfort, right? You can prove it again, or the AI can prove it for you or find a counterexample.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2006">33:26</a>] They are good at both things.</p><h3>33:29 &#8212; How Lean will impact handwritten math</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2009">33:29</a>] But if you were to speculate or draw into the future, maybe three to five years from now, if the cost of formalizing things goes down, how does that change software? How does that change handwritten math?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2026">33:46</a>] Oh, I think it would change dramatically. We have to keep in mind the big labs, they only started training for formal verification and <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> very recently. Seriously, before that, it was like, oh, is in the data sets. I mean, you don&#8217;t really have the reinforcement learning pipelines to optimize the behavior we see today. That is already amazing. Will get way better in the future. The costs will reduce. Programming languages like <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> and <a href="https://en.wikipedia.org/wiki/Rocq">Rocq</a> will become more mainstream.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2063">34:23</a>] Because of that, many people are not so, in the past, people say, &#8220;Oh, functional programming, I don&#8217;t like it.&#8221; But if I&#8217;m not the one that&#8217;s writing most of the code anyway, it doesn&#8217;t really matter. What matters? The specification level, right? It doesn&#8217;t really matter how the code has been written. Yeah, I think you change a lot because of that. At least that&#8217;s the direction we are pushing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2093">34:53</a>] What about, let&#8217;s say, 10 years from now? <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is, everything&#8217;s going really well with <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. Could that be the end of handwritten math or handwritten proofs?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2107">35:07</a>] I think there will always be aspects that are handwritten. Some people like to make a proof look. They want to use the proof as an artifact. You communicate ideas to others. I can&#8217;t imagine there will always be people polishing, making them super easy to understand for another human to communicate ideas to other people. There will always be people like that. The same way today we have people that, we have machines that build furniture, but people, they like to create them by hand and polish them, make them perfect.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2151">35:51</a>] This will always exist, but it will be a mixture. I will be surprised if there is someone who&#8217;s completely AI is not in their workflow somehow. Right. I mean, there&#8217;ll be hybrids, many hybrids. Some people feel uncomfortable about this future, but for me, it&#8217;s super exciting because I view developing software as a super painful process. And with AI, it&#8217;s crazy how it brings you awareness of how many steps are just repetitive and there&#8217;s no creativity, or just patching things.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2197">36:37</a>] And AI automates and removes lots of this pain, right? I mean, I cannot go back to that. I&#8217;m looking forward to this future you describe.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2210">36:50</a>] Yeah, I guess the thing that gives people. I mean, I&#8217;m guessing the discomfort is the worry that if it kept going, then we need less mathematicians or less computer scientists or something like that.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2225">37:05</a>] People don&#8217;t see that AI can bring more people. There&#8217;s also the specification. I mean, I see people saying, oh, we are going to leave AI, we come up with new math. But if there&#8217;s no connection to our world, this is some alien thing that&#8217;s going by itself. You need an interface, for example. We want to build programs because we want to accomplish something. This something, whatever it is, has a specification.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2259">37:39</a>] There will always be humans in the loop saying this is what we know, this is what we want, writing this interface, interacting with the AI. The AI will have math libraries and everything to prove things about these programs we are writing, these artifacts, whatever we are trying to build. But we have to be able to interact with these libraries. We have to understand the abstractions that are there.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2292">38:12</a>] I don&#8217;t see humans being eliminated. We are always going to be there. We are in the interface, we are telling them what we want. Specifications will be there. I can see people always writing code. Another thing is, a lot of people like to write this, and love coding. My interpretation is they love to write prototypes. This is fun. This is the fun part, to try a new idea. But to transform it into a product is never fun.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2325">38:45</a>] I can tell you it&#8217;s never fun. And I can take over these parts. Nobody really likes doing outside of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>.</p><h3>38:55 &#8212; The Z3 theorem prover project he started</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2335">38:55</a>] I know you worked on the Z3 SMT solver, and that sounds like a really difficult thing to build. So first, what is that solver in your words? And yeah, how does it differ from a SAT solver?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2350">39:10</a>] Yeah, I started this really long time ago. It&#8217;s going to be, yeah, 20 years ago I started Z3, I mean when I joined Microsoft Research. Yeah, Z3 is an SMT solver. Like I said, solver. But you have background theories, like you have support for arithmetic, for arrays. These are not random choices, right? This is because we use it for doing test case generation, software verification. Z3 is fully automated. It&#8217;s a push button.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2385">39:45</a>] Although <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> and Z3 are called theorem provers, they are completely different beasts. Z3 is fully automatic. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is interactive. It has automation, but it&#8217;s interactive. Z3 is not a programming language, it&#8217;s more like a constraint solver. It turns out you can prove simple things about it. You cannot do advanced math with abstract math. You can solve constraints with Z3.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2416">40:16</a>] What&#8217;s an example of the inputs to this SMT solver, and what do you get out of it?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2422">40:22</a>] Oh, I can give you one. I mean that&#8217;s even in the Z3 manual. You can encode a Sudoku problem as a set of constraints, and you can ask Z3 to solve it. It will give you back the answer instantaneously. For real applications, Z3 was used very successfully for finding bugs in software. People would convert. For example, suppose that you have a path in your code. You know that has a security vulnerability, but you don&#8217;t know which inputs to the program allow you to execute this path.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2458">40:58</a>] You can convert that into a set of constraints that you send to Z3, and it will say unsatisfiable. It means it&#8217;s impossible to execute this path and you are happy. Or it gives you back an example saying, look, with these inputs you&#8217;re going to be able to do it. And people use it for doing software verification. But because the problem becomes undecidable at that level, you have many universal quantifiers for stating properties about your program, your pre- and postconditions.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2495">41:35</a>] The solver has heuristics, and this was all before AI. The heuristics were hand coded. They would always fail. And for simple things, people would be very happy with the fact that Z3 is push button. But when the property is not trivial, they would come back saying, &#8220;Come on, I know the proof. I mean, why can&#8217;t Z3 find it?&#8221; That&#8217;s why <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> started. I started <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to make sure that we would have a system that&#8217;s really good for software verification, right?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2527">42:07</a>] Z3 was successful for finding bugs, but not so much for software, for proving the absence of bugs. It was never super successful there. But <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> was born to fill this gap.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2545">42:25</a>] You said undecidable, but in practice, in the real world, if you run it, does it typically terminate?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2552">42:32</a>] Yeah, great question. Z3 goes the whole complexity ladder, right? You have SAT solvers, you have NP-complete, PSPACE-complete, EXPTIME-complete, you have the whole two undecidable rights. I mean, surprisingly, even for SAT solvers, you can write really tiny SAT problems that are really hard to solve. No SAT solver will solve them. But in practice, the problems we get for hardware verification, especially if you bound everything, say, oh, I&#8217;m trying to look for a bug in the first 10 steps, right?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2593">43:13</a>] Everything&#8217;s bounded. You&#8217;re not trying to prove, but you&#8217;re trying to capture a class, like a space of scenarios, right? They are very effective there. I mean, I think the lesson there is that programs and hardware, they are not correct for esoteric reasons, right? They are correct because for very simple reasons. I mean, that&#8217;s why these tools are super effective there. Yeah, but when you get too undecidable and you&#8217;re trying to prove, even if the property is not trivial there, I mean, it runs out of steam.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2636">43:56</a>] I mean, and they time out frequently. And sometimes there is a colleague of mine here at <a href="https://en.wikipedia.org/wiki/Amazon_(company)">Amazon</a>, Iminator chose the name proof instability because sometimes if you change the problem, you just flip. You have A and B, you write B and A, where B and A are complicated formulas, you may fail to prove. And when she was just trying to maintain things, proofs would break if using this kind of technology.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2668">44:28</a>] But with <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, she switches to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and it&#8217;s super smooth, right? Because you&#8217;re controlling the proof in the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> case.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2680">44:40</a>] Why is it so much more efficient?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2683">44:43</a>] Your proof is basically, you can view the sequence of steps for solving the problem. In Z3, you can view that you have only one proof step, solve, right? I mean, you&#8217;re saying you have options to solve goals, but you have very little. You cannot influence what Z3 is going to do. It&#8217;s much harder to influence this kind of system. In <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, if you want to give a step, super detailed, step-by-step proof, you can. You can use proof automation like is available in Z3, but you can also break it down step by step.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2725">45:25</a>] And the fact you can do that, humans can do it. But the happy surprise is that AI can do it because now you can say step by step why something is true. The AI can convince <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> that it can provide a proof.</p><h3>45:44 &#8212; The most technically challenging work of his career</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2744">45:44</a>] When you were working on Z3 and <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, what was the most technically challenging part that you had to build for either project?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2755">45:55</a>] I underestimated how much harder <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is in comparison to Z3. It&#8217;s always a magnitude harder. I mean, I talked to many colleagues about that, why I felt like this <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> was so much harder. I think it&#8217;s the surface. The interface for humans is way fuzzier. You have a defined language for it. It&#8217;s called SMT-LIB. It&#8217;s a very simple language. It&#8217;s not meant for humans; it&#8217;s meant for tools. I mean, Z3 is used as the backend of many different tools.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2792">46:32</a>] Someone generates some program generating input for Z3, and people expect a counterexample or saying it&#8217;s impossible, it&#8217;s unsatisfiable to come up with a counterexample. The interface is really, really simple, right? You can view Z3 as a command-line tool that you pass this file on this very low-level language that&#8217;s super easy to parse, and you come back with yes or no. And <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a programming language.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2824">47:04</a>] You have libraries, you have mechanisms, you have interactivity, you have user interface, you have LSP, you have build system, you have GitHub. You have, it is so vast. That&#8217;s another challenging part for me. As I mentioned, Z3 was a backend. The Z3 users are very sophisticated software developers, people that speak the same language I speak. It&#8217;s way easier to talk to people that speak the same language.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2861">47:41</a>] With <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, it&#8217;s completely different. The first users are all math people. They have a completely different background, different expectations, different everything, and different community. And then you have people that want to choose <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> as a programming language. I&#8217;m not a programming language person. My background is automated reasoning. And yeah, it&#8217;s a different language, different expectations.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2891">48:11</a>] Was there a singular component that was just really technically challenging?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2898">48:18</a>] I told you that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>&#8216;s implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. Of course, it was not always like that, right? I mean, it had to be implemented in something else at the beginning. The switch from originally was C++. The switch from C++ to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> was extremely painful. Really? Really. I remember I literally wanted to cry when I managed to compile <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean 4</a>. I was just... Sebastian Ulrich and I at the time were building <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean 4</a> together.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2936">48:56</a>] But this was before we had a nonprofit for <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. I remember calling him and I said, wow, man, insane. Say, are you not excited? He said, yes, I am. I mean, super excited.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2953">49:13</a>] What made that switch hard?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2956">49:16</a>] The first thing is, imagine you&#8217;re going to implement the language in itself. The first thing you want is to minimize the number of features as much as possible, is that you want to implement <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> using bare-bones features. Because you&#8217;re going to have to be able to compile it with itself then. Now you have like 100,000 lines, more or less. I don&#8217;t know the exact number, but it was around 100,000 lines.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=2988">49:48</a>] And you start trying to compile, you fail the first. You can&#8217;t even compile the first file in the pipeline, the one that more than 1,000, and the first one fails, then you fix the bug, you can compile the first one, then you can compile the second, and you keep moving, and you are always finding discrepancies between the new and the old one. And you&#8217;re trying to reconcile, make things easier for the new one because you want to replace the old one.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3022">50:22</a>] This process and <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is a complicated language because of these <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a> and so on. The proofs, for example, when you&#8217;re implementing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, you still need some proofs there. There are some basic proofs you need, but you have to construct these proofs with no interactivity, nothing. Bare bones. You have to provide the proof too. It&#8217;s almost like programming in assembly. The proof. This was also super painful.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3055">50:55</a>] Thank God. Yeah. But it took. Yeah. Many people thought we were going to fail. Sebastian, I would not be able to do it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3065">51:05</a>] 100,000 lines is a lot, like all human written.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3067">51:07</a>] Yes, all human written. Yeah.</p><h3>51:10 &#8212; Lean vs its competitors</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3070">51:10</a>] We talked a lot about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and I know there are competitors to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. What are the pros and cons of the different proof assistants? In what scenarios is one preferred over the others?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3083">51:23</a>] For instance, the first disclaimer: I&#8217;m a completely biased person here, right? But I can tell you what users tell me about. For example, one thing users love is the fact that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is super extensible. Because <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is implemented in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, you can add extensions. So imagine you&#8217;re doing your math proof. In the middle of this math proof, you say, oh, I want this fancy automation here. You can write in the same file, or the AI can write for you, the extension for automating a proof.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3117">51:57</a>] And it will do it. I mean, even the AIs, they know about the fact that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is extensible. If I ask the AI to isolate an issue in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I give the AI a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> file. It will start writing a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> metaprogram, a program about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> tunnels to validate the conjecture it has about why it doesn&#8217;t work. It&#8217;s crazy. It keeps writing <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> meta extension there. The fact that <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>&#8216;s extension is popular, really popular with so many people.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3153">52:33</a>] For example, there is Patrick Massot. He&#8217;s a French mathematician. He wrote something called <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Verbose. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Verbose you can use for teaching. We have this language for writing the proofs, but he made the language look like English, and he has the Infoview now. You can click there. It gives you suggestions about the next move that&#8217;s written in structured English, like you find in a textbook.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3187">53:07</a>] You see, the student has a really good idea on how to write informal math proofs. And he did that without asking me any questions. I mean, all this stuff, the point and click, the new language, the new interactivity, he did all by himself. He&#8217;s not a computer scientist. He has a math degree and he does all this stuff. And it&#8217;s for English and French. I mean, you can choose, you can write the proofs in French and it looks like textbook proof.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3221">53:41</a>] I mean, there are people that write visualizations. You&#8217;re trying to prove something about a math object. You can write extensions that visualize these objects in your Infoview. The people that write new domain-specific languages embedded in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> for different purposes. For example, for protocol verification there is a language called Veil. It&#8217;s a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> file. You open Veil, you feel like it&#8217;s a different system for protocol verification, but it&#8217;s just a <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> file with these extensions for protocol verification, the language for writing protocols in a very convenient way.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3265">54:25</a>] These folks wrote the whole thing without ever talking to us. They all talked to us after they had done it, said, &#8220;Look, I want this part of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to be faster.&#8221; That was the only interaction we had. And so interactivity is a big deal. Another big deal now is the mathematical library. It&#8217;s vast. I mean, for stating problems, open conjectures, you need a library with the concepts to even state the problem.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3296">54:56</a>] So <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> has a massive library and a massive community. The community also plays a big role. Before AI, I think now most people ask questions about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> to AI. But in the past, people would go to the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Zulip channel, ask a question about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. They would get an answer in five minutes. People would say human beats AI. I mean, people would be writing answers instantaneously to your problems. The community played a big role.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3329">55:29</a>] Another one was we listen to our users. I mean the math. I mean if you talk, for example, <a href="https://en.wikipedia.org/wiki/Jeremy_Avigad">Jeremy Avigad</a> was the first user. I mean he has math backgrounds. You ask him, look, he said, look, I could ask anything, any new feature, I would get back the same day. I mean fix the new feature same day. And this attracts people, right? I mean, because you&#8217;re making people improvements, making sure the system does what they want, they come back for more.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3370">56:10</a>] I mean, this also has a huge impact in growing the community.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3376">56:16</a>] When I was doing some research, there was this idea I think you mentioned in this conversation too. There&#8217;s this <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a> proof assistant, and then there&#8217;s <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a>. What is that difference there?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3389">56:29</a>] At the beginning when I started <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I won&#8217;t choose <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a> because it&#8217;s much easier to implement. I mean, <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a> is way harder. But the math community, I mean, Jeremy is the one that convinced me that I would never be able to attract serious mathematicians like Fields Medal-level math people with <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a>. Right? His point is that <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a> is good for concrete math.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3421">57:01</a>] But if you want to talk about abstract objects, <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a> is way more powerful, and it&#8217;s beautiful. It&#8217;s easy to explain. Why is it called dependent? For example, you can have a structure in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> when you have several fields, like X and Y are natural numbers or integers, let&#8217;s say integers. You can have another field. The type of the field is a proof that X is greater than Y. The type of these fields is X.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3452">57:32</a>] Let&#8217;s call it greater colon. You say X greater than Y. The type depends on the value of the previous fields. That&#8217;s why it&#8217;s called <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a>. You can have types that depend on the values of other parameters, other fields, and so on. But the beautiful thing about that is that you have this very small language that is so expressive. For example, this field, now that&#8217;s a proof. You have to provide a proof.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3485">58:05</a>] You can view it as an invariant. I can only build elements of this type if I give the X and Y, like in other programming languages. But I have to give a proof that X is greater than Y. It&#8217;s impossible to construct elements without providing this evidence, right? You can view this as an invariant. You don&#8217;t have to invent invariants, right? The language is just the fact you have these dependencies; you can express them in the functions.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3517">58:37</a>] You can have a function, for example, that says it takes X, a Y, and a proof that Y is different from zero, right? I mean, it&#8217;s impossible to call the function if you do not provide evidence that Y is different from zero. This was always cool, but in the past people would say, wow, providing these proofs is really annoying. But with AI now, the AI can synthesize the proofs for you, and it&#8217;s really cool.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3547">59:07</a>] In <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a>, though, could you express the same?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3550">59:10</a>] No, no, that&#8217;s... You cannot. You don&#8217;t have that. You lose the dependencies. For example, one thing that you cannot do in <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a> in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>. In serious math, people have a bunch of structures they manipulate. You have something, a field, I mean a ring, a group. You can write a function in <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> that takes a group and turns it into a new group, a new structure. You&#8217;re not manu.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3583">59:43</a>] You don&#8217;t really care about the elements of the structure. You are viewing the structure as a first-class citizen. That&#8217;s something that <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a> can do easily; in <a href="https://en.wikipedia.org/wiki/Higher-order_logic">higher-order logic</a>, you have to play encoding tricks. It&#8217;s a mess. I mean, some people say, &#8220;Oh, it works for me.&#8221; None of the mathematicians agree with this statement. None. I mean, you talk to Terence Tao, to Alex Kontorovich, <a href="https://en.wikipedia.org/wiki/Jeremy_Avigad">Jeremy Avigad</a>, Kevin Buzzard, Patrick Massot, they would say, no, no, you have to do <a href="https://en.wikipedia.org/wiki/Dependent_type">dependent type theory</a>.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3618">01:00:18</a>] I mean, that&#8217;s another example for me that listening to your users is important. If you want to appeal to this community, it&#8217;s totally okay to say I don&#8217;t care about this community. But if you care, listening to what they really want is important.</p><h3>01:00:37 &#8212; The future of Lean</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3637">01:00:37</a>] So when we think about the future of <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I&#8217;m curious to hear your thoughts on where you think <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is going, things you&#8217;re excited about in the future. What might it look like in a few years?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3651">01:00:51</a>] Yeah, I think we have these nonprofits behind <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> since 2023. I mean, <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is 13 years old. The first 10 years was a such project, right? I mean, only when we got the nonprofits behind <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> that it became, you can view it as a product. You have a team of engineers, and we managed to do it because of the impact on math. But Sebastian Ullrich and I, we co-founded this nonprofit. What we are really excited about is <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> as a programming language.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3689">01:01:29</a>] A programming language where you can prove things about your programs, right? That&#8217;s a direction we are pushing really hard. <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean 4</a>. We are super grateful for AWS, <a href="https://en.wikipedia.org/wiki/Amazon_(company)">Amazon</a>. They&#8217;re making the largest donations so far to these nonprofits, where the goal is to accelerate this path. I mean, <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> is doing super well in the math path. But let&#8217;s make <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> a programming language, <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> a system for software verification, hardware verification.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3721">01:02:01</a>] Let&#8217;s push to the extreme. Let&#8217;s give some love to these people, to this path that&#8217;s right. Now we do not really have funding to push seriously this path. That&#8217;s the nonprofits, right? For me to be in a world where you can reason about your code is part of my life. I&#8217;m not writing unit tests anymore. I&#8217;m writing properties and proving them. The AI is proving most of them. For me, this is the direction we are pushing hard.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3757">01:02:37</a>] And one thing that people don&#8217;t realize is that when you have proofs, it enables optimizations for free. You can ask the, for example, today, if you ask the AI to optimize, you have to inspect the code to make sure no bugs were introduced in the process. But if the AI is telling you, look, I optimized it, it still computes the same thing. Here&#8217;s the proof. Game changer in my point of view.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3790">01:03:10</a>] Yeah, I&#8217;ve heard multiple people say this decade will be the decade of formal verification of software. And yeah, maybe <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> will be a huge part of that.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3803">01:03:23</a>] Yeah, for sure. We&#8217;re super excited to make it happen. Scalability is super important, right? Because there is a big difference between math and software verification. In math, the statements are usually really tiny or small. I mean, Fermat&#8217;s Last Theorem is an example. I mean, super small. But the proof&#8217;s insanely cheap. For software verification, it&#8217;s the opposite, right? The statements are big.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3839">01:03:59</a>] I mean, but the proofs are shallow. The reason why this is shallow, but you have to manipulate these big objects, right?</p><h3>01:04:10 &#8212; Technical book recommendations</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3850">01:04:10</a>] So for someone who wants to learn more about formal verification or learn more about <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, do you have a top technical book recommendation in the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> websites?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3858">01:04:18</a>] I mean, we have several books there that introduce <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Functional Programming, <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Theorem Proving, <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> Mathematics, and <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>: The Mechanics of Proof is great for educational purposes. We have a collection of proofs. If people go to <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">lean</a>-lang.org, you find all these books there. But one thing I tell people these days is that learning <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> using AI is super efficient. I mean, you keep talking to AI in natural language, asking what do you want?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3903">01:05:03</a>] Asking it to write examples. Many people split the screen in three now, right? We had the <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a> code, the Infoview, and on the bottom. Now many people use an AI agent there that&#8217;s writing the code and explaining in natural language what&#8217;s going on there. It&#8217;s a super effective way. I mean, Terence Tao, he told me when he learned <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, he uses the old version. It was before agents. He would have ChatGPT on one window and <a href="https://en.wikipedia.org/wiki/Visual_Studio_Code">Visual Studio Code</a> in the other window, and he would copy and paste between them.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3941">01:05:41</a>] And that&#8217;s how he learned. But now it&#8217;s even more effective with the AI agents. It&#8217;s easy to pick up. I mean, just talk to the agent. Sometimes people say, how do I start? Start talking to the agents. It will help you, it will customize. Right? You can explain what you know already, right? I mean, what&#8217;s your background? And for example, if you tell, oh, I know <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, it&#8217;s so much easier, right?</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3972">01:06:12</a>] You can customize the process.</p><h3>01:06:15 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3975">01:06:15</a>] And then last question for you is, if you could go back to the beginning of building C3, building <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, and give yourself some advice, knowing what you know now, what would you say?</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=3988">01:06:28</a>] I&#8217;ll keep it a secret. I mean, I think eagerness is bliss. I mean, you don&#8217;t know how hard things are, and when you start the adventure, maybe I&#8217;ll keep secrets. What I would tell, I think one thing, especially before starting <a href="https://en.wikipedia.org/wiki/Lean_(proof_assistant)">Lean</a>, I&#8217;m super introverted. And I would tell, look, you should work on your people skills because it helps a lot. I mean, when you have to interact with a community, with people.</p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=4018">01:06:58</a>] For me, it was hard to learn that, and I would tell myself, I think it&#8217;s really important to have people skills, too.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=4026">01:07:06</a>] Well, thank you for your time today.</p><p><strong>Leo:</strong></p><p>[<a href="https://youtu.be/KzdYKeAqWhY?t=4028">01:07:08</a>] Thank you. Thank you.</p>]]></content:encoded></item><item><title><![CDATA[Creator of Lua: Scripting, Programming Languages, Predictions | Roberto Ierusalimschy]]></title><description><![CDATA[When I worked at Instagram, most of the codebase was in Python but occasionally I&#8217;d see these bindings into C++.]]></description><link>https://www.developing.dev/p/creator-of-lua-scripting-programming</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-lua-scripting-programming</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 03 Aug 2026 13:05:15 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207184056/e7e481befa19d616fb4655496dfdc9b3.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>When I worked at Instagram, most of the codebase was in Python but occasionally I&#8217;d see these bindings into C++. I never really understood how that worked and just took it for granted since some other infra team had set that up.</p><p>In this conversation with <a href="https://en.wikipedia.org/wiki/Roberto_Ierusalimschy">Roberto Ierusalimschy</a> (the creator of Lua), he finally clarified how this cross-language mechanism works for me. It was relevant since Lua is a scripting language purpose built to be embedded in other languages.</p><p>The programming language is famously used in World of Warcraft and Roblox where you have all the high performance rendering code written in something like C/C++. The parts of the game that change frequently though are written in Lua and called from that lower level language.</p><p>He also had some interesting recommendations on programming languages (aside from Lua) that he&#8217;d recommend people learn to expand their minds.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/jCZnFKk6M9A">YouTube</a>, <a href="https://open.spotify.com/episode/3SWEjGkoNCdgSEvOA1sU1A?si=8VrcgJWHQkm8suGWO-Ud8Q">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-jCZnFKk6M9A" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;jCZnFKk6M9A&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/jCZnFKk6M9A?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/207184056/0043-what-sets-lua-apart">00:43 - What sets Lua apart</a></p><p><a href="https://www.developing.dev/i/207184056/0835-comparing-lua-with-python">08:35 - Comparing Lua with Python</a></p><p><a href="https://www.developing.dev/i/207184056/1304-top-book-recommendation-on-language-design">13:04 - Top book recommendation on language design</a></p><p><a href="https://www.developing.dev/i/207184056/1420-how-jit-works-and-why-it-is-hard">14:20 - How JIT works and why it is hard</a></p><p><a href="https://www.developing.dev/i/207184056/2321-compiling-python-and-interpreting-c">23:21 - Compiling Python and interpreting C</a></p><p><a href="https://www.developing.dev/i/207184056/3031-how-cross-language-calls-work">30:31 - How cross language calls work</a></p><p><a href="https://www.developing.dev/i/207184056/3658-lua-unique-design-decisions">36:58 - Lua unique design decisions</a></p><p><a href="https://www.developing.dev/i/207184056/5117-predictions-for-ais-impact-on-languages">51:17 - Predictions for AIs impact on languages</a></p><p><a href="https://www.developing.dev/i/207184056/010010-top-3-languages-to-learn-to-become-a-better-engineer">01:00:10 - Top 3 languages to learn to become a better engineer</a></p><p><a href="https://www.developing.dev/i/207184056/010321-advice-for-his-younger-self">01:03:21 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:43 &#8212; What sets Lua apart</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=43">00:43</a>] In 2023, <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> was the second highest programming language by growth according to contributors on GitHub. And I saw that it was also used in famous games like World of Warcraft and Roblox. And so I wanted to ask you what kind of programming language is <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> and what sets it apart?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=67">01:07</a>] <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is a language that usually is used together with another language like <a href="https://en.wikipedia.org/wiki/C">C</a> or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and then you have an architecture where <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> takes care of the more dynamic part of your program, things that change more frequently, things that are not so resource intensive, the hard parts you have written in <a href="https://en.wikipedia.org/wiki/C">C</a> or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. So it&#8217;s typically what is called a scripting language. There is some confusion, I think, between scripting languages and dynamic languages.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=105">01:45</a>] And people just, I think, considered a dynamic language to be more or less the same thing as a scripting language. And they are not exactly the same thing. So, for instance, a big difference between <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> and several other scripting languages is that for scripting you have two typical uses. You can have the app that I call who has the main loop. So you can have the program written in <a href="https://en.wikipedia.org/wiki/C">C</a> and calling <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, or you have your program written in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> calling <a href="https://en.wikipedia.org/wiki/C">C</a>.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=143">02:23</a>] And for several scripting languages, actually, you don&#8217;t have to, but it&#8217;s much easier to use if you have your program written in the scripting language, and the <a href="https://en.wikipedia.org/wiki/C">C</a> language is only for your libraries. <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is a very good fit. When you want the reverse, you want the program written in <a href="https://en.wikipedia.org/wiki/C">C</a> and then it calls <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> from time to time. So the point is that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is very good at this, that you have the main loop, the program written in <a href="https://en.wikipedia.org/wiki/C">C</a>, and then it calls <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> for some tasks or for whatever you want to do.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=189">03:09</a>] And for instance, in games, this is very important to keep the frame ratio of the game. So, for instance, you have the loop. The loop keeps the rhythm of the game and keeps everything. And then at every frame it calls <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> updates all characters, all images, everything in the game, and then it returns to <a href="https://en.wikipedia.org/wiki/C">C</a> to do the renderization, etc. So it is a very good fit. But the main point is that people sometimes call this the embedding versus extending.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=226">03:46</a>] So you can extend the scripting language with <a href="https://en.wikipedia.org/wiki/C">C</a>, or you can embed the scripting language into <a href="https://en.wikipedia.org/wiki/C">C</a>. And so several languages are very good only for extending, while <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, it&#8217;s really good for both embedding and extending.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=245">04:05</a>] Why is <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> good for embedding and Python? It&#8217;s almost the opposite, because <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, since the beginning, we thought about <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> as a library.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=251">04:11</a>] Since the beginning, <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> was written as a library. So I think that&#8217;s the main difference. <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> has a standalone program that, I mean, you can call <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> in your console, in your command line, but that program is just a client of the official client of the library, following the same <a href="https://en.wikipedia.org/wiki/API">API</a> that any other program can use. So this idea of language as a library, I think it&#8217;s not very common. It&#8217;s something that is very particular in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=289">04:49</a>] When you design a language as a library, what are the unique design decisions?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=296">04:56</a>] I think that goes a lot of small and big details. For instance, your idea of what is your global space, or about scopes of variables, about exception handling. For instance, it&#8217;s very important that it&#8217;s very common in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> that you raise an exception in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> but you catch the exception in <a href="https://en.wikipedia.org/wiki/C">C</a>. So all these things about how you do exception handling in the language. Another, it&#8217;s not a big deal, but it&#8217;s.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=331">05:31</a>] For instance, most dynamic languages, that&#8217;s almost a definition of a dynamic language. They have an evaluation function that you can give a piece of code and then it executes that code. In <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, we don&#8217;t, instead of <a href="https://en.wikipedia.org/wiki/Eval">eval</a>, we have a function that we call load, that you give a piece of code and it returns a function, an internal function. Like it was a <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> function that when you call that function, then you execute the associated code. It doesn&#8217;t execute immediately.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=368">06:08</a>] So you have this clear separation between compiling the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code that you have, and you can do that in <a href="https://en.wikipedia.org/wiki/C">C</a>. So in <a href="https://en.wikipedia.org/wiki/C">C</a>, you can get a piece of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code. You can compile it, check if it has errors, et cetera. Then you call that, for instance, a different piece, or you can even call several times. You can load functions from <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> into <a href="https://en.wikipedia.org/wiki/C">C</a>. You keep them in the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> space, and then you can call that same function again.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=402">06:42</a>] And again, it&#8217;s not a big difference, of course, because if you have <a href="https://en.wikipedia.org/wiki/Eval">eval</a>, you can evaluate a function declaration and then return that function. So you can, if you have <a href="https://en.wikipedia.org/wiki/Eval">eval</a>, you can do load. If you have load, you can do <a href="https://en.wikipedia.org/wiki/Eval">eval</a>. But I think load, it&#8217;s simpler to use for that kind of thing. That is, when you think about libraries, for instance, in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, it doesn&#8217;t have a global state. I think that&#8217;s, I forgot that.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=434">07:14</a>] That&#8217;s a very important difference. When <a href="https://en.wikipedia.org/wiki/C">C</a> starts, the first thing it has to do is create a new <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> state. And then everything you do, you do on that state. And that state is completely independent of everything else. So <a href="https://en.wikipedia.org/wiki/C">C</a> can, for instance, create another <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> state, and both states are completely independent, completely. There is no communication between them. And so this again shows that these.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=471">07:51</a>] And so, for instance, <a href="https://en.wikipedia.org/wiki/C">C</a> can use a state, do a lot of stuff, and then it can close the state and all memory used by <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is released. Everything that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> was using is released. For instance, if <a href="https://en.wikipedia.org/wiki/C">C</a> doesn&#8217;t need to use <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> anymore for a program, you release all resources used by <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, and your program continues, and then later you can again create another state, etc. So there&#8217;s, as I said, several small or not so small decisions in the language.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=504">08:24</a>] But all the time we think about the language as, is that good for embedding? Is that possible to do embedding and things like that.</p><h3>08:35 &#8212; Comparing Lua with Python</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=515">08:35</a>] When you think of the programming community, generally everyone is quite familiar with Python, but not as familiar with <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. I thought it might be interesting to compare the two languages. So if you compared <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> to Python, what are the pros and cons of each language design?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=534">08:54</a>] <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is a language that is intended to be used in this idea of scripting architecture. So <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, for instance, it doesn&#8217;t thrive having a lot of different libraries. If you have something that, oh, I want to write a quick program for, something that Python is much better, it has all the libraries you can dream about. It has a lot of libraries built in already into the language. So the language is huge.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=567">09:27</a>] I mean, the installation, it&#8217;s a huge download, etc. But it has exactly, oh, I need that. Oh, it&#8217;s there, I need that, it&#8217;s there. And <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is almost the opposite, some very minimalistic language, because it is that you embed <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> into your program. It doesn&#8217;t use almost any resources. Most of the libraries, the important libraries that you will use, will be provided by the program itself, will be the commands like move a character or do some speech or things like that from the engine of the game, for instance, if you think about games.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=607">10:07</a>] But whatever it is. So <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, I think that&#8217;s the main big difference. They have very different goals.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=616">10:16</a>] I saw some benchmarks, and I saw that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is much faster than Python. Why is that?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=623">10:23</a>] Those benchmarks are not exact. But I think the main point is that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> does have this focus on performance. Again, in the realm of scripting languages, it doesn&#8217;t want to compete with <a href="https://en.wikipedia.org/wiki/C">C</a> or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. But in the realm of dynamic languages, mostly now it&#8217;s a dynamic language, not a scripting language. We have some focus on performance. So I think this is the first difference. Python is exactly much more.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=656">10:56</a>] I think it&#8217;s the decision that they make. Performance is not that important. It&#8217;s more important to be flexible, to be easy to do whatever you want to do. Then in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, sometimes we do not put some features because we think that there is no way to implement that efficiently. But also I think some part of the performance difference comes exactly because of the size of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. So most of the virtual machine fits in your cache, for instance.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=687">11:27</a>] I think there are these two big things. One is that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> doesn&#8217;t have so many features. It&#8217;s not so dynamic as Python. So in Python, there is a lot of interaction because everything can mean something else. In <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, we are a little more conservative on that side. But I think also this thing of being small also makes it naturally faster.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=713">11:53</a>] So, on the distinction between scripting languages and dynamic, or I guess it seems like dynamic language is a superset, scripting language is a subset of that. Am I understanding that <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is a dynamic language, not scripting language from your perspective?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=730">12:10</a>] Yes, exactly. Yes, exactly. Because the scripting language&#8212;the original scripting language was Bash, or the Unix shells&#8212;with this idea that it is a language that coordinates other stuff. So, for instance, it can be extending again, you can use. But this idea that you have two different languages for the shell is only useful because you have a lot of programs written in <a href="https://en.wikipedia.org/wiki/C">C</a> that are controlled by shell. So the scripting language has this very strong idea that you have this idea of a dual-language architecture.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=770">12:50</a>] So scripting is like the name scripting. It means you coordinate or you give a script to be executed by those other things.</p><h3>13:04 &#8212; Top book recommendation on language design</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=784">13:04</a>] You gave a talk a while ago, and someone asked you a question about what books do you recommend for studying programming languages, and you said that you actually enjoyed studying the language design of other programming languages.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=798">13:18</a>] I like reading books that describe the design of languages. I think the best books are those written by the author of a language about the design of that language.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=812">13:32</a>] What book recommendation do you think is best on language design?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=817">13:37</a>] <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>: The Good Parts, for instance. It&#8217;s kind of old now, but I think it&#8217;s a very interesting book. Although I always joke that <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>: The Good Parts is a very thin book compared with the official <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> books, like 110 of the language, but I think that book is really interesting because it discusses the language, the bad parts, and it focuses on the good parts, but then explains why it&#8217;s there, why it was made that way, etc.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=857">14:17</a>] That is a book that I like.</p><h3>14:20 &#8212; How JIT works and why it is hard</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=860">14:20</a>] When we were talking about <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, it sounds like one of the things that sets <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> apart is the performance and how minimal it is. And when I was reading about <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, I saw that there&#8217;s <a href="https://en.wikipedia.org/wiki/LuaJIT">LuaJIT</a>. How does <a href="https://en.wikipedia.org/wiki/LuaJIT">LuaJIT</a> work?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=874">14:34</a>] <a href="https://en.wikipedia.org/wiki/LuaJIT">LuaJIT</a>? Well, the first thing, it&#8217;s a completely different project. It doesn&#8217;t have anything to do with us, but it&#8217;s an incredible piece of software that is a just-in-time compiler for <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. And I think exactly one of the reasons it works so well on top of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is because of the simplicity of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. It&#8217;s a very regular language. I mean, it has very, as I said, it doesn&#8217;t have many exceptions or too many interactions or things like that.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=909">15:09</a>] So it&#8217;s not that dynamic. I mean, everything can mean something completely different, so I think that gives a very good language for a <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a>. But Mike Pall is the name of the guy that made the first <a href="https://en.wikipedia.org/wiki/LuaJIT">LuaJIT</a>. He&#8217;s still working on that. It&#8217;s unbelievable, his work.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=933">15:33</a>] What makes writing a <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a> difficult?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=936">15:36</a>] The first thing for me, I don&#8217;t want to get involved with <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a>, is because it&#8217;s machine dependent. It&#8217;s not very productive. I mean, you do a lot of work and only work on that architecture, and then, oh, I want to run on another architecture. But Paul, he created a kind of pseudo-assembler that he writes in this assembly, and then he translates that assembler to real machine code for the different machines.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=968">16:08</a>] So trying to unify the different architectures. So there is a lot of work, but I don&#8217;t like&#8212;I think architecture I like to study, but I don&#8217;t like to work with it because I think it&#8217;s very unstable. Each new version, they change that or they change that. I mean, they change the <a href="https://en.wikipedia.org/wiki/Ali">ABI</a> for something, and then now the stack has to be aligned in some very particular way. And of course, also to get that performance that he gets, he made what&#8217;s called a trace compiler.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1007">16:47</a>] The idea of a trace compiler, instead of getting a function and compiling a function, it works like any <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a> that tries to detect things that are executed frequently. And so it starts what&#8217;s called the trace. It gets recording everything that the code is doing, including function calls, et cetera, until it closes the loop, and then it compiles that loop, including function calls, et cetera, everything in line.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1043">17:23</a>] It compiles that for that specific, for instance, so that number happens to be an integer. So it compiles it, the number will be an integer again. And there are a lot of checks just to check that everything is as assumed, and then it executes the loop, and then it&#8217;s very, very, very fast. But anything that is different, for instance, you call it again, but you call it frequently with integers. Suddenly you call it with a float, and then at some part of the code, it breaks the condition and you have to return to the interpreter part.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1085">18:05</a>] And I think that&#8217;s one of the worst parts, is that you have to translate the state that it&#8217;s all compressed into register, et cetera. For instance, you don&#8217;t even have a call stack because you didn&#8217;t do the calls, you just inlined it. But then, oops, something went wrong. Now you have to continue interpreting. So you have to recreate this fact that you didn&#8217;t create originally, et cetera. So I mean, even to think about that, I have headaches.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1120">18:40</a>] Standard <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is portable, but to execute on a machine, it eventually gets converted into machine instructions. Where in the stack is it eventually translated to machine instructions?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1132">18:52</a>] <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is written in <a href="https://en.wikipedia.org/wiki/C">C</a>, the interpreter is written in <a href="https://en.wikipedia.org/wiki/C">C</a>. And then you compile that into machine code, the interpreter, the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code. Then when you get some, I mean the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code, the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> interpreter code, when you get some <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code, you do what we call a precompilation; that is, we translate to an internal language. Then the interpreter is just a big loop, a loop with a switch. That&#8217;s a very big simplification.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1167">19:27</a>] But in general terms, a loop with a switch that gets instruction does a switch and sees all these instructions. The instructions are quite similar to a CPU move, but then instead of creating it, it executes that structure: move A to B. So it does A, move A to B, and executes that, and then again repeats, executes the next instruction. So this is how interpreter works. And basically, so we pre-compile <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1203">20:03</a>] There is a compilation, but we compile for this virtual machine, and then we have this main interpreter loop that people call that just fetches each instruction. A jump is just a jump. You have an array of bytecodes. A jump just goes to that. You have a counter that tells where you are in this array. You just update that counter. It goes to another position in that array where you&#8217;re going to.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1234">20:34</a>] So it&#8217;s like you emulate the CPU in software. It&#8217;s written in <a href="https://en.wikipedia.org/wiki/C">C</a>. The interpreter, this main loop, is written in <a href="https://en.wikipedia.org/wiki/C">C</a>. So if you compile it in Linux, it will run in Linux on an x86 architecture. It will run on that architecture. If you compile that on a Mac, it will run on a Mac. It&#8217;s a <a href="https://en.wikipedia.org/wiki/C">C</a> program.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1260">21:00</a>] Got it? Okay, so. And that&#8217;s the part that&#8217;s not portable. And the <a href="https://en.wikipedia.org/wiki/C">C</a> compiler handles that?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1265">21:05</a>] Yes, yes, yes. All portability of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> comes on top of portability of <a href="https://en.wikipedia.org/wiki/C">C</a>. How does <a href="https://en.wikipedia.org/wiki/Just-in-time_compilation">JIT</a> compare to ahead-of-time compilation in terms of performance?</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1277">21:17</a>] Both ahead of time? Any kind of compilation can be 10 times faster than interpreting, or even 100 times faster than interpreting, depending on what you are doing. But 10 times is a very good figure between ahead of time and trace compilation. Then it depends a lot on what you are doing after. Trace compilation is particularly good for benchmarks because you&#8217;re repeating it more or less.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1315">21:55</a>] They are very uniform. You are doing exactly the same thing again and again and again. So then trace compilers shine. I mean, this is the best they can do in real programs. That depends a lot, but I wouldn&#8217;t say there is a clear winner between the two. I think it depends a lot on the kind of program.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1340">22:20</a>] Is it possible to write an ahead-of-time compiler for <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1344">22:24</a>] Yes, there are several. I mean, I don&#8217;t know if there is any production-quality one, but for research, et cetera, there are several. I have a student who wrote one, like a master&#8217;s thesis. A ridiculous simple <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> compiler, ahead-of-time <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> compiler or something like that. Exactly, because it was just&#8212;it gets these opcodes and expands into <a href="https://en.wikipedia.org/wiki/C">C</a> code, and then you send that through to a <a href="https://en.wikipedia.org/wiki/C">C</a> compiler, and then you have an ahead-of-time compiler for <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, and it gave like three, five times boosting performance.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1393">23:13</a>] Very, very. As the title said, it&#8217;s ridiculous. Simple compiler.</p><h3>23:21 &#8212; Compiling Python and interpreting C</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1401">23:21</a>] I always hear there&#8217;s compiled languages and there&#8217;s dynamic or interpreted languages, I guess, but really that distinction is in the toolchain, not necessarily in the way that the symbols are laid out in the source code. I could write a compiler or a program that takes in Python code and then converts it into machine code. And then I would have compiled Python.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1426">23:46</a>] And you can interpret <a href="https://en.wikipedia.org/wiki/C">C</a> code too. You can write an interpreter for <a href="https://en.wikipedia.org/wiki/C">C</a>. Yes, but the main difference is exactly what it&#8217;s easy to do. For instance, as I said, one of the hallmarks of interpreted languages is <a href="https://en.wikipedia.org/wiki/Eval">eval</a>. So if you want to compile the language and keep <a href="https://en.wikipedia.org/wiki/Eval">eval</a>, you have to have a compiler as a library of your runtime because you may want to compile things during execution. This is the hallmark of a dynamic language.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1461">24:21</a>] So you can create code while running code. So that part is why, much more often than not, dynamic languages are interpreted, but you can compile. But then sometimes some compilers do not handle <a href="https://en.wikipedia.org/wiki/Eval">eval</a> at all. They say, oh, you can compile as long as your program doesn&#8217;t have evals. So there are some restrictions. And as I said, you can interpret <a href="https://en.wikipedia.org/wiki/C">C</a> code, I mean, but you&#8217;ll be extremely slow.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1500">25:00</a>] It doesn&#8217;t.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1501">25:01</a>] But I think one big difference, and when I look at static languages or compiled languages, they seem to have type systems in them, or type annotations. What is the role of a type system in the compilation process?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1518">25:18</a>] Depends a lot on the type system. Several type systems, like in <a href="https://en.wikipedia.org/wiki/C">C</a>, for instance, are written to be a very important part of the compilation process. So if you have the right type system, it&#8217;s much, much easier to write a compiler. A typical example is on the space for variables. Because if you know, oh, this is an integer, this is a float, this is a double, you know exactly how many bytes of memory you need to that. Otherwise, it&#8217;s a mess.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1554">25:54</a>] More often than not, you have to keep everything in the heap, dynamically allocated, and so huge penalty in performance. And for instance, when you see, as I said, a trivial, you have an addition A plus B. If you know the type of A and the type of B at compile time, you know, oh, this plus is, I&#8217;m adding two integers, I just generate the machine code to add integers, and there everything is fine.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1587">26:27</a>] You know, a dynamic language, like <a href="https://en.wikipedia.org/wiki/C">C</a>, I don&#8217;t&#8212;I compile to a virtual machine, and the virtual machine is just a piece of <a href="https://en.wikipedia.org/wiki/C">C</a>. But then how do I... At runtime, you have to check what is A. Oh, A is an integer, B is a float. Oh, I have to convert. Or B is a string. Or depending on the language you are, if you&#8217;re not doing addition, it can be doing concatenation or you can just call some function, and you have to do all that at runtime.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1617">26:57</a>] So types can be very, very important to compile efficiently. But of course, you have to have a type system. There are several languages now, like <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>, that the types are kind of... You do not guarantee that everything has the types you set. So then it&#8217;s impossible to use them to compile because, oh, probably A will be an integer, but if it&#8217;s not, I mean, if I have an integer register to boot with, it must be a register or things will not work.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1660">27:40</a>] But if you have the right type system, they are essential for good compilation.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1666">27:46</a>] I know some programming languages have type inference. What if you ran type inference and you got an unambiguous set of types, and then you used that in compilation? Could you do that for <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1681">28:01</a>] No, I mean that&#8217;s not computable. But a lot of people try to do type inference for dynamic languages, and it&#8217;s a really hard problem, and it&#8217;s very difficult to do anything useful in terms of performance. What if sometimes you get? But you have to be very restricted. If you write your program in that specific way, use fall. So it&#8217;s like you don&#8217;t have the type, but you write the program thinking about types.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1718">28:38</a>] And then if the program is in that very particular format, then you can do type inference and everything goes well. Otherwise, either type inference doesn&#8217;t work or, I mean, it works, but it infers very generic types for everything. And so you cannot take advantage of the types because of the dynamic nature of the language. This also happens with <a href="https://en.wikipedia.org/wiki/LuaJIT">LuaJIT</a>, for instance. You can get much, much better performance if you have a kind of type system in your mind.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1758">29:18</a>] Usually we do that, and you follow those type rules even without&#8212;the compiler, the language does not impose them on you. But you assume, oh, I&#8217;m going to follow those discipline, I&#8217;m not going to use the same variable to store integers and strings, etc. And then you can have much better results.</p><h3>30:31 &#8212; How cross language calls work</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1831">30:31</a>] Earlier, you mentioned that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> can call <a href="https://en.wikipedia.org/wiki/C">C</a> and <a href="https://en.wikipedia.org/wiki/C">C</a> could call <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. And you mentioned, we talked about the compiler and the interpreter and how do these pieces work together in the case where you&#8217;re chaining different programming languages.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1847">30:47</a>] Yeah, the main trick, I mean, it&#8217;s not a trick. The technique for allowing <a href="https://en.wikipedia.org/wiki/C">C</a> is <a href="https://en.wikipedia.org/wiki/C">C</a> pointers, function pointers. There is a pointer to a function. This is part of the official <a href="https://en.wikipedia.org/wiki/C">C</a> language. So when you start <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, suppose you are in <a href="https://en.wikipedia.org/wiki/C">C</a>, as I said, you create a library, you create a state, a <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> state, and then you can register, you send to <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> <a href="https://en.wikipedia.org/wiki/C">C</a> pointers, function pointers, I mean, pointers to functions in your <a href="https://en.wikipedia.org/wiki/C">C</a> library associated with names.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1888">31:28</a>] <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> stores that in some data structure. Okay, so it says, oh, the function is associated with that pointer, etc. So when it&#8217;s doing the interpretation, oh, this instruction is call. Instruction, call. Let&#8217;s see what it&#8217;s calling. Well, it&#8217;s calling a. What is the value of A? Oh, it&#8217;s a pointer. And then it calls that pointer. And then we call the <a href="https://en.wikipedia.org/wiki/C">C</a> function to do that. And when <a href="https://en.wikipedia.org/wiki/C">C</a> wants to call <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, I mean, a <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> function is just a data structure from the point of view of <a href="https://en.wikipedia.org/wiki/C">C</a>.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1930">32:10</a>] It&#8217;s just an array of bytecodes of opcodes, and then it calls the <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> interpreter function. Please interpret that function for me. So <a href="https://en.wikipedia.org/wiki/C">C</a> can call <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, interpret that function, and that function is running, and then can call a <a href="https://en.wikipedia.org/wiki/C">C</a> function using a pointer. And you can have that in the stack at several levels. So we are running a <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> function that calls <a href="https://en.wikipedia.org/wiki/C">C</a> and <a href="https://en.wikipedia.org/wiki/C">C</a> calls <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> again and <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> calls <a href="https://en.wikipedia.org/wiki/C">C</a>. You have recursive stuff in that, etc.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1963">32:43</a>] And everything works, just works.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1968">32:48</a>] In one of your talks on <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, you mentioned that there&#8217;s security benefits to using a scripting language. What are those benefits, and how does it...</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=1976">32:56</a>] protect the hardware. It&#8217;s not only hardware. Once I gave a talk here in <a href="https://en.wikipedia.org/wiki/Brazil">Brazil</a>, I was called to give a talk at a Python conference in <a href="https://en.wikipedia.org/wiki/Brazil">Brazil</a>. They called me for a talk. At the end of the talk I said, &#8220;Oh, let&#8217;s use it as a scripting language.&#8221; As I said, emphasizing that thing that you have this dual-language architecture. And I know of case of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> being used with <a href="https://en.wikipedia.org/wiki/C">C</a>, with <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, with Fortran, with several different languages.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2013">33:33</a>] And then someone in the audience said, oh, we also use <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> for scripting Python programs. And what was the case? They have a huge Python program for financial stuff that has to do a lot of financial transactions, etc. And they wanted to do scripts, I mean, to have a command line where you could write instructions to be executed at runtime, for instance, for debugging or for inspecting, for whatever you need some kind of end-user programming.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2051">34:11</a>] But they said exactly what I was talking about, spaces in Python. Basically, if you have a command line and you are going to run Python, that Python can do whatever it wants to your program. There is no way to protect it. So if you gave that to the user, the user could do whatever it wanted to do. I mean for the program. And so the program would break a lot of invariants of a lot of important stuff in the files it kept, etc.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2085">34:45</a>] And so what they did, they gave the thing in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, because, as I said, <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, you create a state. For instance, as I just mentioned, you can only call <a href="https://en.wikipedia.org/wiki/C">C</a> functions. That is true in Python again, but in that particular data you can only call <a href="https://en.wikipedia.org/wiki/C">C</a> functions if you give the <a href="https://en.wikipedia.org/wiki/C">C</a> pointer to <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. So have a very strict, and there is very little, as I said, there is a library. When you create a <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> state, it has no functions at all.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2118">35:18</a>] There is nothing. I mean, you cannot open files. Everything there are, we call standard libraries. Usually you open a state and you register those standard libraries. In <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, so you have the minimum, you have mathematical functions, etc. But you can open a state and register no functions at all. In <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, it can only do things calling those specific functions you gave. So they used <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> to have this kind of control.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2149">35:49</a>] Or you can write <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code here, but that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> code can only do very specific things in the Python program. You cannot call any Python function or build any kind of Python data structure, etc. So was this a nice example? In the case of hardware, it&#8217;s the same thing. You cannot just go in. For instance, there is a hard port that controls the speed of the fan. That keeps the temperature of the CPU.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2183">36:23</a>] And if you put that very low, you can really burn the CPU. I mean, because the CPU gets very hot. And so you have a <a href="https://en.wikipedia.org/wiki/C">C</a>. You cannot call <a href="https://en.wikipedia.org/wiki/C">C</a> directly, you must call <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. And only <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> has access to that. So if you can only program in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, then you check whether the numbers you are giving are precise or whatever. It has to check and then it calls.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2211">36:51</a>] It&#8217;s kind of like a sandbox environment.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2213">36:53</a>] Yes, exactly. Yes. Sandbox is the exact word for that.</p><h3>36:58 &#8212; Lua unique design decisions</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2218">36:58</a>] I saw there was this, I guess, paper, maybe it was an article you wrote on the history of <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> with several other people. And I thought there were some interesting quotes in there I kind of wanted to ask you about. So one of them, it says that there&#8217;s this old joke that says that a camel is a horse designed by a committee.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2240">37:20</a>] Yes, that&#8217;s not mine. That&#8217;s a real old joke.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2245">37:25</a>] Does that mean you think the best programming languages are designed by as few people as possible?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2253">37:33</a>] Yes, I do think, yeah. Because one thing that it&#8217;s easy to observe is this: If you have a committee writing a language, everyone wants to put something from themselves into the language. So you have your favorite mechanism, and you will fight whatever it takes to put your favorite mechanism into the language. And very few people you fight against putting anything in the language.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2291">38:11</a>] I mean, some people say, oh no, this is too much, this is too complicated. But people fight much more fiercely for putting something they want in the language than against putting something in the language. So when you have a committee, the tendency is, &#8220;Let&#8217;s keep adding stuff, let&#8217;s keep adding stuff.&#8221; And people sometimes even do not understand, I mean, I&#8217;m even not sure if everybody there really knows the language well enough to be putting that.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2325">38:45</a>] But that fits with other parts of the language. I didn&#8217;t even know the language has this other part. I mean, so things sometimes do not fit together very well, et cetera. So first, you have what&#8217;s called the conceptual integrity. It&#8217;s a name that gives that everything fits together, everything is designed with everything else in mind, etc. It&#8217;s much easier to. I&#8217;m not saying, I mean, <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> has several wrong stuff regarding because of.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2357">39:17</a>] With time, etc. There you make mistakes, but it&#8217;s much easier to get things right if you have a small group of people that everybody knows what everybody else is thinking. I mean, you are deciding whether to put something in the language or not. It&#8217;s much easier to change your mind if you do not put something and then later you decide to add it to the language than to add something and then later you decide, oh, I shouldn&#8217;t have. That was not a good idea.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2393">39:53</a>] So, when in doubt, the rule is always don&#8217;t put it. If you&#8217;re not very sure, don&#8217;t put it. We can always add that later and keep the door open to change your mind.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2408">40:08</a>] I saw in the original <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> programming language that there was no Boolean type, and I&#8217;ve never seen that before. Why was there no Boolean in the original <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2419">40:19</a>] Well, <a href="https://en.wikipedia.org/wiki/C">C</a> doesn&#8217;t have booleans. I mean the original <a href="https://en.wikipedia.org/wiki/C">C</a>, and C90 up to C99, they use integers. If it&#8217;s 0, it&#8217;s false; if it&#8217;s different from 0, it&#8217;s true. And people lived quite well with that. Lisp and there are several languages that don&#8217;t have booleans. And actually, we only put booleans in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> because we wanted false. True is completely useless in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>. Nobody used true, and nobody&#8217;s a joke, but it&#8217;s too strong.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2457">40:57</a>] But it&#8217;s not very useful. You can use almost any other value for true. But the problem is that in <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, the special value nil is in a table. It is equivalent to the key not being present. You may think whether that was a good decision, but it&#8217;s very embedded in the design of the language. And it has this strong concept that a table, a key that is not present in a table, there is an associative array, it has a nil value, it&#8217;s completely indistinguishable whether it&#8217;s absent.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2503">41:43</a>] I mean, absent means it has a nil value. Nil value means it is absent. And so sometimes you want to have a false value, but you want to know it is there in the table. And so we needed a false to have a false you can put in the table. And it&#8217;s completely different from being absent. You know, that is a key and the key has a false value. So we need a false value. And if you&#8217;re going to have a false, then we added true, true.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2536">42:16</a>] It would make sense to have false without true. But with true, it&#8217;s almost useless.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2542">42:22</a>] Why not use 0 as false if you&#8217;re okay with 1 as true?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2547">42:27</a>] We could have used it, but among other things, that would be a big change because it didn&#8217;t exactly work that way. So we added that later in the language. So changing 0 to false would be a very big incompatibility. But it could be, for instance, in Python I think a lot of stuff are false. I mean, zero is false, empty list is false. And in <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> I think it&#8217;s even worse. I don&#8217;t recall exactly, but I mean this is kind of arbitrary.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2586">43:06</a>] I mean, we could have, for instance, nil, false, zero, empty string. It just depends on the kind of test that you use. In dynamic languages, it&#8217;s very easy to.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2602">43:22</a>] Python has the... Like you&#8217;re describing the truthy values and the falsy values, I guess they can be interpreted as false. But it&#8217;s more implicit, less explicit. And I know <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is even further on that spectrum where you can do some really weird stuff that implicitly has some behavior. Do you think that design direction is good or bad?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2628">43:48</a>] It&#8217;s like a small compact. It&#8217;s more easy to write stuff. But you always have this balance in any language. Almost anything, the more flexible you have, the less protection you have. In <a href="https://en.wikipedia.org/wiki/C">C</a>, you have something like that people do not notice. But because <a href="https://en.wikipedia.org/wiki/C">C</a> is typed, you can check whether an integer is true or false. You can check whether a float is true or false directly because, again, it checks whether the float is zero or different from zero.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2659">44:19</a>] You can check whether a pointer is nil or different from null. Because in <a href="https://en.wikipedia.org/wiki/C">C</a>, null is a magic zero. So it&#8217;s kind of, if you do not write any test, it has an implicit test like different from zero. And this zero can be an integer, can be a float, can be a pointer. So in <a href="https://en.wikipedia.org/wiki/C">C</a> also you have this kind of flexibility. It&#8217;s just you have to be more explicit. Usually these are, avoids errors, but so this, you always have this balance between easy to write and easy to make mistakes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2709">45:09</a>] I saw that <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> is 1-indexed instead of 0-indexed, and I had never seen a programming language like that.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2718">45:18</a>] In my experience, there were a lot of programming languages like that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2723">45:23</a>] So yeah, why is it 1-indexed?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2726">45:26</a>] Because everything in the real world is 1-indexed. If you have a book, you have chapter 1, chapter 2. Nobody numbers chapters as 0, 1, 2, 3. Only programmers are the only, even mathematicians. If you get a math book, a book of mathematics, any sequence, it&#8217;s 1 and 2, 3. If you get a matrix, the first element is 1, 1 and 2, and then... Everything is written before. Only in programming language there is this idea of zero.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2763">46:03</a>] And the funniest part is that, for instance, Fortran, now it&#8217;s a very old language, but to index it from 1. In Pascal, you can choose. I mean, you write an array, you can say this array is indexed from -5 to 5, and so it&#8217;s indexed from whatever value you want to whatever value you want. 0 became extremely popular because of <a href="https://en.wikipedia.org/wiki/C">C</a>. <a href="https://en.wikipedia.org/wiki/C">C</a> uses 0, and a lot of languages copied are kind of inspired by <a href="https://en.wikipedia.org/wiki/C">C</a>. They have the same operators, the same syntax for expressions, the boolean operators, for instance.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2810">46:50</a>] That is not standard mathematics. Almost all languages chose the same operators as they are written in <a href="https://en.wikipedia.org/wiki/C">C</a>. And so a lot of languages copied <a href="https://en.wikipedia.org/wiki/C">C</a>. And then zero index became very popular. And what is funny is that in <a href="https://en.wikipedia.org/wiki/C">C</a> there is no indexing. In <a href="https://en.wikipedia.org/wiki/C">C</a>, indexing is just an illusion because what you have is pointer arithmetic. And see, when you write A indexed by I, actually what you are saying is get the contents of A plus I.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2849">47:29</a>] And so because in <a href="https://en.wikipedia.org/wiki/C">C</a>, you don&#8217;t have indexing, you have these displacements or deltas, I don&#8217;t know, offsets. Then it must have indexed by zero because of that, because then the first element is the element that is at the original address. So <a href="https://en.wikipedia.org/wiki/C">C</a> doesn&#8217;t have indexing. So when you say A is indexed by 0, it doesn&#8217;t index by 0 because it doesn&#8217;t index by anything. But it has this illusion of indexing.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2882">48:02</a>] And then it&#8217;s easy to sync that in. And then a lot of languages that do not use pointer arithmetic do not have this restriction, do not have these semantics, copied <a href="https://en.wikipedia.org/wiki/C">C</a> and kept the zero indexing as this thing. If you get a 12-year-old and try to explain zero indexing, I assure you it&#8217;s much, much easier for them to write. Oh, I have a list of the first element. I always joke that the first element is 0, and you write first, you for 1, but the first element 1 is not 1, is 0.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2925">48:45</a>] So it has advantages, for instance, for some specific operators. Mathematically, for instance, when you want to do a circular buffer, zero is better. There are some small advantages, but it&#8217;s much, much more confusing. And as <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> has this idea of end-user programming, we always joke it&#8217;s much easier to make life easier for the non-programmers. And I&#8217;m sure that programmers can program whatever index they have to do because they&#8217;re supposedly professional. They can learn that, oh, indexing from then, to put on the end user the burden of using something completely different from them.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=2974">49:34</a>] Indexing by zero, what does it mean, the element index zero? They never saw the chapters in the book. Yes, the first chapter is at zero, the second chapter is at one. No, I&#8217;m not joking. When you try to explain that to non-programmers, it doesn&#8217;t need to be 12-year-olds. I mean, just get anyone that is not an Internet programmer. Think, &#8220;Oh, I have the absolute.&#8221; I mean, you are grown up, you are, I mean, okay, you are a student. If 0, I&#8217;m sure you can learn to program; if 1 or minus 1 or whatever it is, the base you have to use.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3006">50:06</a>] I&#8217;m sure.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3016">50:16</a>] But do you think you see a lot of bugs because maybe someone thinks it&#8217;s zero?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3022">50:22</a>] Unfortunately, I see some bugs, but as I always say, those are the kind of bugs that just show that you didn&#8217;t do minimum testing, because that kind of bug is not a kind of... That kind of bug that, oh, if you try to... Anything you try to do in an array, if you start from zero or start from one, you have a bug. So the most simple test that you can imagine to test anything related to an array, to a list, et cetera, and if you mistake 0 for 1, you detect that.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3068">51:08</a>] So if you have that kind of bug, it just shows that you didn&#8217;t test that code at all.</p><h3>51:17 &#8212; Predictions for AIs impact on languages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3077">51:17</a>] There was this talk that you gave on the cost of adding language features to <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> and how there&#8217;s a lot of hidden costs. And I think today, implementation costs of software are going down due to AI code generation or LLMs, and I was wondering if that changes your thinking on, I guess, the cost of adding features to a programming language.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3103">51:43</a>] AI is something that we don&#8217;t really know what&#8217;s going to happen. If you go to an extreme, it&#8217;s completely feasible that we won&#8217;t have programming languages. And I mean, because if the AI is writing all your code and is checking all your code and doing everything, in a few years maybe we don&#8217;t need programming. AI can write machine code directly; it doesn&#8217;t need programming languages. It&#8217;s easier for them to compile or even generate the code directly.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3141">52:21</a>] I mean, I don&#8217;t know what&#8217;s going to happen in a few years. So I think it&#8217;s very hard for me to talk about anything about AI because I have, and I think nobody has a clear idea what&#8212;I mean, people may have a clear idea what happens in one year or two years, but in five years, I mean, anyone that says, &#8220;Oh, that&#8217;s going to happen in five years,&#8221; is just guessing. So I think it&#8217;s hard to say anything.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3179">52:59</a>] This is all because of AI. So keeping the minimum thing, I mean that still have differences there. Oh, why do you still use programming languages? Because you want the user or some human to be able to check the result of what the AI is doing, etc. So I still think that most of the things about programming languages still hold. For instance, the cost of complexity of you understanding. I mean AI create a code for you and then you really understanding what that code does.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3217">53:37</a>] I mean it&#8217;s even worse. I mean it&#8217;s much more important that the language should be clear and has no hidden mechanisms. Because the AI may use a hidden mechanism, it doesn&#8217;t have the concept, oh, that will be difficult for a human to understand that what is really happening here is something it says something that, for instance, oh, I can use that. But I&#8217;m for sure it&#8217;s going to put a comment here because I&#8217;m sure that someone that reads that in one month is not going to understand.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3254">54:14</a>] EA doesn&#8217;t have that kind of thinking, and so just use that trick or that thing. So I think if you think about this idea of AI generating code, I think it&#8217;s even more important for the language to be simple, to be understandable, for you to be able to really understand that the code that you are seeing does what you think it should do. You really understand what the code is doing.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3284">54:44</a>] So this thing about simplicity, I think it&#8217;s very important, and I think the other parts from the book maybe facilitate documentation. I&#8217;m not sure if I mean this thing about conceptual integrity, for instance, to keep things, that makes sense, etc. I really think that one of the main costs is that it&#8217;s the burden you put on the user to learn one more thing about your language. So there&#8217;s the other cost of it.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3319">55:19</a>] I said that implementation is the part that AI can really help you. It&#8217;s not a really important cost.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3332">55:32</a>] If you think about a spectrum of simple to complex, what programming languages are the most simple ones that you think of, the easy to understand, less confusing side effects, and what programming languages are the most complex and have the most foot guns.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3350">55:50</a>] This famous quote, also&#8212;I don&#8217;t know from whom&#8212;that the simplest possible, but not simpler than that. Because if you go to the extreme of simplicity, you could get, for instance, lambda calculus or Turing machines and say, &#8220;Oh, that&#8217;s really simple.&#8221; I mean, this is really simple, but it&#8217;s completely impossible to write any program in that. I mean, if you have really simple stuff, I think the extremes will be like that.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3382">56:22</a>] But you don&#8217;t want to be there, but for the other side. I think <a href="https://en.wikipedia.org/wiki/C">C</a> is a good example of a language that I think is really, really complex.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3393">56:33</a>] A question that comes up a lot is, given how much AI has been progressing, would you still recommend people learn computer science today?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3403">56:43</a>] If you really think about that, it&#8217;s really difficult to recommend people learn anything for a profession. We have no idea what AI will do in five years. If someone is entering university now, they&#8217;re going to graduate in four years, five years, or, at minimum, three, three and a half years, and nobody has any idea what the world, your profession, will be like in four years from now. So it&#8217;s really hard to, I mean, to say, &#8220;Oh, yes, you can still go into it,&#8221; because now people say, &#8220;Oh, no, the manual tasks, the AI, it&#8217;s very good, but you still need the whole architecture.&#8221;</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3451">57:31</a>] You needed a software engineer to do, to see the big picture, etc. That is now. But you have no idea that that will be true in four years or five years. So I think I like, I enjoy programming. I could do. I mean, some people say I use the AI for that. I mean, I program because I like that. But, basically, so if you choose something that you like, but really if you&#8217;re really thinking, how am I going to live with that?</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3483">58:03</a>] I have no idea. It&#8217;s really, I think it&#8217;s a really hard time to be choosing the profession.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3493">58:13</a>] I mean, you&#8217;ve worked on <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> for such a long time. When you look back on it, what went well and what didn&#8217;t go well?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3502">58:22</a>] We had this privilege of not having to satisfy clients. So you do not have, as I said, for instance, you can choose to add some feature to the language because you do not have a pressure. Oh, we need that feature, whatever it takes or et cetera. So we have this privilege of, oh, we are not sure whether to put that. We don&#8217;t put that. We can wait one year or so. I think that it&#8217;s much easier to do a good project, a beautiful project.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3535">58:55</a>] So I&#8217;m not sure this is a recommendation to try to work without much pressure. People usually do not have this choice.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3547">59:07</a>] What was your measure of success? If you&#8217;re working on a programming language, how did you know it was good?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3554">59:14</a>] It&#8217;s more important to have, sometimes, good people that I recognize as would like the result than a lot of people that I have no idea who they are, etc.</p><p>So I think this metric of quantity that is the standard metric nowadays on the web has gotten much worse, that everything is just in the number of follows. I mean, it doesn&#8217;t care who is following you, just as many people as possible.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3588">59:48</a>] So sometimes, for me, it&#8217;s much more important if someone that I really admire or think is, or says something good about <a href="https://en.wikipedia.org/wiki/Lua">Lua</a>, then, oh, we have that many. But, of course, it&#8217;s good also to have some number of users and to be recognized. But I think it&#8217;s a balance between all of this.</p><h3>01:00:10 &#8212; Top 3 languages to learn to become a better engineer</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3610">01:00:10</a>] Do you recommend people study other programming languages to get better at programming? And so the question is, what are the top three programming languages that you think every engineer should learn in 2026 to become better at programming? Haskell, I think, is an incredible language.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3628">01:00:28</a>] It has everything you need to really learn about functional programming. I think one of the main benefits is that it&#8217;s much. There&#8217;s an old joke in the Haskell community that if you write a program in Haskell and in <a href="https://en.wikipedia.org/wiki/C">C</a>, in <a href="https://en.wikipedia.org/wiki/C">C</a>, you really spend one week to make it efficient, and then you spend one year to make it correct. In Haskell, you spend one week to make it correct, and then you may spend one year to make it efficient.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3675">01:01:15</a>] Again, it&#8217;s not my joke, but it has some truths. Not always true, etc. But I think it gives the idea. I think this is maybe the most important lesson of Haskell. But Haskell has many other validities. This thing about type inference, for instance, that in the entire language, types are optional everywhere, and yet it can do type inference for everything. It can infer the types correctly, etc.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3711">01:01:51</a>] See, or some old language. I mean, you can also have to learn some assembler, or I mean, to see. But assemblers now are becoming too much complex. It would be to learn, like, the <a href="https://en.wikipedia.org/wiki/Intel_8080">Intel 8080</a> assembler, very old assembler of a very... But to have this exactly, this idea of what a machine does, how it does stuff at exactly the basic level, to have this understanding. Scheme is a language that I like very much because of this, exactly this economy of ideas.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3747">01:02:27</a>] There are very few concepts, and it can do some amazing things with very few concepts. The way one language that this is very, very old. But I still think it was SNOBOL. I think one of the first languages to have pattern matching and this idea of doing, I mean, really strong patterns, et cetera. And I think it&#8217;s in... Sometimes it&#8217;s interesting to see old languages because, exactly, nowadays a lot of languages tend to, like I said, zero indexing that people, they copy a lot.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3786">01:03:06</a>] They tend to be too uniform in some aspects. And it&#8217;s interesting to see some older languages that have some ideas that may maybe not even be good ideas, but it&#8217;s really interesting.</p><h3>01:03:21 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3801">01:03:21</a>] Last question for you is, knowing everything you know now, if you could go back to when you had just started your career and give yourself some advice, what would you say?</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3814">01:03:34</a>] To know a lot of stuff, you really have to study a lot of stuff. And it takes time. I also joke about that. People say, oh, learn <a href="https://en.wikipedia.org/wiki/Lua">Lua</a> in 30 minutes or learn Python in 5 minutes. And I always say that I wanted to learn programming in five years. But that&#8217;s really learn. Take your time. You have to learn that. And that is a lot of stuff. Stuff to learn. There is no magic there.</p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3847">01:04:07</a>] You really have to learn. I mean, a lot of different things.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3852">01:04:12</a>] Thank you so much for your time, Professor. I really appreciate it. It was a lot of fun.</p><p><strong>Roberto:</strong></p><p>[<a href="https://youtu.be/jCZnFKk6M9A?t=3856">01:04:16</a>] Thank you. It&#8217;s my pleasure.</p>]]></content:encoded></item><item><title><![CDATA[Turing Award Winner: Early AI, LLM Predictions, Causality | Judea Pearl]]></title><description><![CDATA[When I went to UCLA, I had heard professor Pearl&#8217;s name here and there from the other professors.]]></description><link>https://www.developing.dev/p/turing-award-winner-early-ai-llm</link><guid isPermaLink="false">https://www.developing.dev/p/turing-award-winner-early-ai-llm</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 27 Jul 2026 13:06:14 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207184223/a1bd4ea272acf0438bf3a4a5c586549d.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>When I went to UCLA, I had heard professor Pearl&#8217;s name here and there from the other professors. It was surreal getting to interview him about his career and thoughts about AI today. He had a few anecdotes I really loved like:</p><ul><li><p>Why he loves physics and who the greatest scientist of all time is</p></li><li><p>Proving Alpha-beta pruning is mathematically optimal which Donald Knuth was surprised at</p></li><li><p>Discovering Bayesian networks and the eureka moment he had</p></li></ul><p>You can tell how much passion he has for science in just the first few minutes of the conversation. Hope you enjoy it!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/FleTXB1fAcQ">YouTube</a>, <a href="https://open.spotify.com/episode/2Y2M4gddfls2fNdkf4sybz?si=XNKsgAgjQI2K53whJeIClw">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-FleTXB1fAcQ" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;FleTXB1fAcQ&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/FleTXB1fAcQ?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/207184223/0054-how-he-got-into-ai">00:54 - How he got into AI</a></p><p><a href="https://www.developing.dev/i/207184223/1117-greatest-scientist-of-all-time">11:17 - Greatest scientist of all time</a></p><p><a href="https://www.developing.dev/i/207184223/2015-what-people-thought-of-ai-in-the-80s">20:15 - What people thought of AI in the 80s</a></p><p><a href="https://www.developing.dev/i/207184223/2623-entering-academia-and-researching-ai">26:23 - Entering academia and researching AI</a></p><p><a href="https://www.developing.dev/i/207184223/3452-the-invention-of-bayesian-networks">34:52 - The invention of Bayesian networks</a></p><p><a href="https://www.developing.dev/i/207184223/4628-pioneering-work-in-causality">46:28 - Pioneering work in causality</a></p><p><a href="https://www.developing.dev/i/207184223/5538-the-causal-hierarchy">55:38 - The causal hierarchy</a></p><p><a href="https://www.developing.dev/i/207184223/5934-llms-and-predictions">59:34 - LLMs and predictions</a></p><p><a href="https://www.developing.dev/i/207184223/012012-a-restless-mind-pays">01:20:12 - A restless mind pays</a></p><p><a href="https://www.developing.dev/i/207184223/012436-advice-for-his-younger-self">01:24:36 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:54 &#8212; How he got into AI</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=54">00:54</a>] I looked into your educational background, and it was all electrical engineering, physics. I don&#8217;t see the connection to causal reasoning and artificial intelligence. How did you get there?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=66">01:06</a>] Everything connects. Everything connects. Everything connects to the day I was born. Hit on the head first. I have to start by saying that we. I grew up in Mandate in Israel prior to 1948, prior to the establishment of the state of Israel. And we had a very excellent high school education. My high school teachers were professors that were chased by Hitler from Germany, from highly reputable universities like Heidelberg and Berlin.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=118">01:58</a>] And they came to Israel in the 1930s, and they didn&#8217;t find any academic position at that time. They were rare, and they started teaching high school. But they were really quality professors. They could teach anything without notes, from the Economy of Manchuria to the proof of Pythagoras theorem, with no notes and no stop. And we were, we&#8212;I mean, my generation was lucky enough to be beneficiary of this educational experiment.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=166">02:46</a>] So we were taught science from a human viewpoint, chronologically, the way things were discovered, by whom they were discovered, at what period. Why was there a question about certain mathematical proof? What the inventor of the proof knew, what he didn&#8217;t know, and what he asked himself in the context of the historical situation at that time. That&#8217;s the way to teach science, not the recipe of algorithms and techniques, but as a.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=217">03:37</a>] From the viewpoint of the human actor. And the transfer of ideas from one human to another is a human struggle, not a human struggle to decipher the secrets of nature.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=234">03:54</a>] What&#8217;s the advantage of learning in that style?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=238">03:58</a>] The advantage is that you see yourself as an actor. You, the student, you are also puzzled by many things. But look what he did. Look what Pythagoras did. He was as puzzled as you were, and he took that route. Perhaps when you&#8217;re puzzled, of course, your puzzles are not as magnificent as his, but still, look what he did. He took that route around things, and he consulted some other work of some other.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=274">04:34</a>] He struggled and you struggled, and you are part of science. That is the basic idea. Give students the idea that you are part of science, not an observer, not a passive observer in science, not a recipient, but an actor. And that, I think, was unique, very valuable for me. We indeed got the idea that each one of us can find another proof of Pythagoras&#8217; theorem that no one else has thought about. And each one of us has the potential of becoming an Einstein or Pythagoras.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=318">05:18</a>] They give us this illusion. Okay, it was a useful illusion. I know I&#8217;m just a pebble looking at Pythagoras. But that illusion helped me be a little, I would say not contrarian, but assertive. And we all grew up in assertive mode of learning science. We insisted on understanding things our way in real time. If the teacher went too fast, we made noise. We made noise with our chairs and with everything we could, okay?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=367">06:07</a>] And the teacher stopped and slowed down to make us all understand things our way in real time. So that was part of the. It&#8217;s a mood of me and my generation. And it, I. I can see that it affected my life. I had one story that I remember. It made an impression on me. I look back and I say, well, I. It started very early. It started at age 10 when we learned about how to calculate areas and volumes. And here there was a question in class.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=410">06:50</a>] How many dunams are there in a square kilometer? Dunam is 1,000 meter square. It was a Turkish unit for measuring areas. So the entire class said, Dunam is a kilometer square. And I screamed and said, no, it&#8217;s a thousand. You have 1,000 dunams in a kilometer square. And the teacher sided with the laughing class, and they all mocked me and ridiculed me. And I went home and I said, they are wrong, and I&#8217;m going to come back tomorrow and insist on that.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=463">07:43</a>] And the teacher apologized to the class, and I felt that, yes, you have to insist on your understanding of things. Yes, I&#8217;m telling you that because maybe this one made me into a noncompromise. And this is part of my childhood, a background which might explain how I got into artificial intelligence. So I took engineering after being a farmer in the army. Okay, farmer, yes. In Israeli army you have troops which are spending their time, half and half, half in military training and half in farming, being part of a kibbutz.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=520">08:40</a>] It&#8217;s the old idea of one hand holding the plow and the other one the rifle. And you succeed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=530">08:50</a>] Did you ever shoot anyone?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=532">08:52</a>] I almost shot someone without seeing him or her. But it turns out the next morning that it was a fox. But the steps of the fox were very, very similar to the steps of terrorists advancing toward us. So that&#8217;s as far as I got to shooting. So I got into the Technion to study electrical engineering again. We had great teachers. Again, we got the idea that we are making science. So we studied physics very seriously.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=574">09:34</a>] And I liked what we studied. Yeah, I wasn&#8217;t the first in class. No, it was third or fourth. Always never match the geniuses. The geniuses were bored in class. They knew what the teacher was going to do, what he was going to ask in the exams. And everything was boring to them. I wasn&#8217;t bored. I wasn&#8217;t the first, but I was bored.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=607">10:07</a>] What interested you in physics? What did you like that made you so passionate?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=613">10:13</a>] What I liked was that, sitting on your chair, you can predict things in physics. Like Maxwell, who sat on his chair and said, &#8220;That looks like a wave equation. Look, let me calculate its velocity. It looks like the velocity of light. Maybe light is nothing else but electromagnetic wave,&#8221; you know, on his armchair. He didn&#8217;t do any experiment. That excited me. Yeah. My wife told me one night I woke up and I said Maxwell was wrong.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=657">10:57</a>] Maxwell was wrong. She quieted me down in the next morning, said he was right. By the way, you.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=669">11:09</a>] You mentioned a few scientists, and it seems like you know a lot about the history of the old scientists.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=676">11:16</a>] Yes.</p><h3>11:17 &#8212; Greatest scientist of all time</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=677">11:17</a>] Do you have a favorite scientist of all time? And why?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=681">11:21</a>] When we studied and I looked at geometry, I got fever. Really? Fever.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=691">11:31</a>] Physical fever.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=692">11:32</a>] Yeah. I couldn&#8217;t get over the idea that you can do in algebra all the geometric constructions that we learned on, you know. So I thought that Descartes was the greatest mathematician ever lived. It was so enormous to me. Transformation from geometric constructions to algebraic derivations. It was unbelievable.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=721">12:01</a>] I took that for granted when I was in education. What is it about that that&#8217;s so astonishing?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=728">12:08</a>] Here you have two different languages, language of geometry and the language of algebra. And they are the same thing. And you can get the same phenomena, same proof that you can pass a tangent to a circle from a point outside the circle and you can find the angle. But different method, two different languages dealing with the same phenomena, different perspective, and they get the same result. It blew me off.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=765">12:45</a>] Blew me off completely. Maybe that was. It was preparation to computer science. Because for us, computer science is what&#8217;s the big deal? You want to see things from different perspective, invent a new language. So that was what turned me on in my high school. That was in high school. And then we saw that. I saw the same thing in physics. Different languages, capturing the same phenomena.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=800">13:20</a>] They really excited me. He invented the idea of a field, two different ways of looking at the same thing. You can see that the force here depends on the charges around it, or you can say, no, there&#8217;s a field right in the location of your testing point. Yeah, terrific, terrific. So I came in with this preparation, and here we go to Brooklyn Poly. And I studied there in.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=835">13:55</a>] I worked in the morning in RCA Laboratories in Princeton, New Jersey, David Sarnoff Research Laboratory. And there I got into the computer research group. Computer research at that time was research of all phenomena that you can think of to find out a mechanism for computer memories. The memories at that time were core memories, magnetic cores, these donuts that you remember perhaps from your early childhood that were too slow and too clumsy, and you had to have people stringing them X and Y and Z in Hong Kong.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=884">14:44</a>] And people understood that the days of core memories are numbered. And we were looking for new phenomena. Some people looked at photochromic memories, some people looked at the semiconductor. Some people looked, like me, into superconductivity. And I was in the group. It was supposed to design superconducting memories. Okay. We did some nice plates, 16 by 16 bits, okay. And we thought that we had the future in front of us, but in the way toward developing superconducting memories.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=936">15:36</a>] I investigated the physical phenomena behind the eddy currents, permanent eddy currents in thin superconducting films. Okay. And it so happened that I discovered new phenomena there, and I got a prize. And it even has a name. It&#8217;s called Pearl vortex. You can find it in Wikipedia. I discovered that physicists, years after I finished my PhD, discovered my work there. And they were interested in the idea of permanent current flowing in a circle in thin superconducting films.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=983">16:23</a>] And since I analyze the magnetic and current field there, they call it pearl vortex. So here I have my footstep into immortality.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=997">16:37</a>] Well, you said it&#8217;s a permanent vortex. Is it because there&#8217;s no electrical resistance?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1003">16:43</a>] Because it&#8217;s a superconducting, the current going forever. So you establish, you put magnetic field and you excite a vortex counterclockwise, and it will continue to turn and turn and turn forever. That&#8217;s why we call it permanent current. Yeah, forever until you flip it with another magnetic field. So we call it a vortex. But it goes on forever. And you can detect it by flipping it.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1039">17:19</a>] You flip it, and if you see a big flip, it was one way. If you don&#8217;t see a big flip, it was the way you turn it. Okay, so you have a memory. You have a memory. Of course, we didn&#8217;t succeed in turning it into useful memories that would be competitive with semiconductors. The people who worked on semiconductors beat us out. We never believed that they would, but they did, both in miniaturization and in techniques.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1084">18:04</a>] Unbelievable.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1085">18:05</a>] At the time, why did you not believe in the semiconductor direction?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1090">18:10</a>] Who is going to trust memory to battery failure? What if you lose a battery? It was obvious it will never work. Right. And we looked into the result that they obtained at that time. It looked far-fetched, the idea you can have that degree of miniaturization. We saw the struggle of people who work in the laboratory on semiconductors, and we weren&#8217;t impressed. Yeah, but they beat us up. Okay, so that was my story with the superconductors.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1139">18:59</a>] But I must tell you that everybody, even at that time, when computers were clumsy and took rooms and rooms and you programmed with cards, okay, even at that time everybody understood in AI as an inspiration. Everybody believed thoroughly that one day computers are going to be able to emulate all human functions. That was not the question. The question was only how and when, but not whether.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1183">19:43</a>] Okay, I remember already at that time, with the clumsy computers and the punch cards, people talked about associative memories, about pattern recognition, about seeing, understanding. All these were already ideas, were exciting people to think more about it. And so we were all geared towards it.</p><h3>20:15 &#8212; What people thought of AI in the 80s</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1215">20:15</a>] If I at that time asked people and your peers and you, what&#8217;s the timeline for maybe human-level intelligence and machines? What would people have said at that time? When they were excited in the 80s?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1230">20:30</a>] I think they were more optimistic than reality. They probably gave you 20 years, but that was 1965. Okay, so 20 years, 1985. No, we didn&#8217;t yet get anywhere.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1246">20:46</a>] Yeah, and then what about today? Do you think people are more optimistic than reality? Or is this just history repeating?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1255">20:55</a>] It depends who you&#8217;re talking about. Some people are extremely optimistic today, and some people say I&#8217;m a bit skeptical, but not skeptical in our ability to eventually reach AGI, but in whether the LLM technique and thinking will lead us there. Okay, so it&#8217;s a question of who you ask? Okay. LLMs were a great surprise. But they have limitations. We&#8217;ll talk about it. After superconducting, I decided to come to California to a company named Electronic Memories, in which they did not work on superconductors, but worked on plated wires.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1306">21:46</a>] Instead of having a donut in which you thread a wire, you start with a wire and you plate it with magnetic material so it acts like a donut locally. Right. And that was the promising technique. At that time, at least, I was in charge of the research and development group charged with the task of developing this kind of systems to replace core memories. Yeah. And I worked there for three years. I was frustrated because things did not go my way.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1346">22:26</a>] I had both administrative and technical challenges that I couldn&#8217;t handle, both in chemistry, and I didn&#8217;t know much chemistry, so I was frustrated.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1361">22:41</a>] What were the administrative frustrations?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1363">22:43</a>] I had a group, and I had to satisfy the administration. I dealt with the personnel issues, firing, hiring people. Yeah. And my wife saw that I was unhappy, and she told me, &#8220;You get to. You have your place in academia.&#8221; So I looked for a position in academia. Luckily, also at that time, industry was revered by academia because all the advances, all the important advances were developed in industry, not in academia.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1416">23:36</a>] The transistor was developed in Bell Lab. The laser was developed in, I think, here in California by another fellow. But all this was industry development and not. So academia looked with reverence toward people who come from industry. And they hired me without me even filling an application, without even asking for accommodations. Yeah, at that time it was a good time to be hired.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1455">24:15</a>] And what about&#8212; because you said at that time industry was revered by academia.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1460">24:20</a>] Would you say that&#8217;s still true today, or has that changed?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1463">24:23</a>] No, it&#8217;s changed. It&#8217;s different now with AI. If you come from deep learning or something, DeepMind, you know, people look at you with reverence in academia. Yeah, but it&#8217;s changed, you know, only in the last few years. I say that throughout. Since 1970, I think until 2000, it was the other way around. You know, people simply dismissed industry.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1502">25:02</a>] I mean, academia dismissed industry. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1505">25:05</a>] I wonder what happened. Was it like Bell Labs disbanded their research group or something, or that was part of it.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1514">25:14</a>] Bell Labs disbanding and what happened to IBM is still there in Watson. Watson Center, I remember. Big center, huge and important. Including Raytheon. Including. Where can I tell you? Use research here in Malibu. Did great work, but it all went down sort of the frontier of research went to academia. Sure, we had a lot of theoretical work in academia. The development of AI, AI property after 1970. Yeah, but prior to 2000.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1563">26:03</a>] Yeah, at least in my corner of the field. The whole idea of inference search, inference logic, expert systems, these were all the academic development.</p><h3>26:23 &#8212; Entering academia and researching AI</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1583">26:23</a>] So then you got hired at UCLA?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1585">26:25</a>] I got hired in 1970 or 1969. And yeah, I was hired. The computer science department was just formed there. And I got first hired by another department called the Engineering Systems Interdisciplinary, and then back to computer science. And I was asked to teach computer memories, hardware on computer memories. Yeah. And I gave a course in this technology. And later on, I started getting interested in pattern recognition, and I started working in this direction.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1632">27:12</a>] I did work on image compression. We used fast Fourier transform and fast Hadamard transform, all kinds of transform techniques to condense images, to minimize the number of bits sent. Okay. I guess it was a part of the trend at that time. But when I got into pattern cognition, I returned to my old dream of thinking about AI and how the brain works and how computers will one day emulate ourselves.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1674">27:54</a>] And I started teaching class in AI. At that time, AI was game playing, machine playing of chess and checkers and the puzzles, like the eight puzzles, ruby cubes, things like that. That was AI. Yeah. And I got excited by that game, and now I see why. Why? Now I can tell you why. Because the game was a matter of capturing in mathematics what people do heuristically, like playing chess, and the interplay between mathematical analysis and the performance interests me.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1729">28:49</a>] And especially in chess playing, the interplay between the explicit knowledge that you have in terms of your gut feel about the strength of a position and what you get when you do some search. Okay, so here, interplay between fast thinking and long thinking, to use Kahneman and Tversky or Kahneman&#8217;s title, Thinking, Fast and Slow.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1760">29:20</a>] Thinking fast and thinking slow. Right now, here. It&#8217;s a beautiful arena to see how not only you have two modes of thinking, but how they feed each other and how you can invest more resources in one versus the other.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1782">29:42</a>] The algorithms at that time, like chess, for instance. Can you give an example of the interplay of the two and how that might come together in a chess-playing system?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1793">29:53</a>] Yes. You can invest more time in getting your immediate perception. It&#8217;s called static evaluation function of the chess position, the strength of the chess position. Or you can let it go and think about searching for deeper horizon. Okay. It&#8217;s a tradeoff. It was like, if I remember, we.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1815">30:15</a>] We searched the game tree and then evaluated each position at the horizon.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1823">30:23</a>] And then you back off and then you make a move toward the position that has the greatest strength after you back off. Yeah. Okay.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1835">30:35</a>] So the intuition is encoded in the evaluation functions.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1839">30:39</a>] Yeah, and your intuition is the evaluation function. But you can improve your intuition too. How? By learning. Okay. So Samuel Checker program did learning in regression analysis to find the proper weight on the various characteristics of the position so as to make the evaluation function more accurate.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1868">31:08</a>] When I was learning chess, I think one heuristic is you want to control the center.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1873">31:13</a>] Good. And material advantage is another one, right? Okay. And whether you have two bishops versus a bishop and a knight, this all counts. And whether you already castled or not. All this contributes as attributes to the strength of a board position. Getting the weight correct, you can do by learning. After you play so many games, you adjust the weights. Yeah. So that was Samuel&#8217;s contribution to first machine learning.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1915">31:55</a>] I say that was the first machine learning. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1917">31:57</a>] At that time, when you were working on this, were chess systems superhuman? Yet I think there&#8217;s.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1922">32:02</a>] No, no, no. It was still a dream to beat the world champion Kasparov machine. It was a dream. No. And. But I did knight analysis. We did alpha-beta pruning. If you remember that, you probably programmed it. Well, I proved that alpha-beta is optimal.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1954">32:34</a>] Really?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1954">32:34</a>] Yes, mathematically, you see, I like the mathematics. Prove it. You cannot do better in terms of number of position that you have to inspect at the horizon or the depth of search. And what can I say about it? I get some nice results. Even Knuth was surprised that one can prove the optimality on alpha beta. Okay. Because he questioned it in his book. Yeah, I did some work with Deep Cop on searching trees and.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=1990">33:10</a>] Okay, so I did mathematical work on the tradeoff between search and reasoning until I got sick and tired of search.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2004">33:24</a>] When you pick your research area, is that 100% your own choice?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2010">33:30</a>] No, it&#8217;s always a combination of two things. Number one, do you know the answer to the question? If you don&#8217;t know the answer, it&#8217;s a puzzle. If it&#8217;s a possible next question comes. Do you think you have the techniques to make a contribution here? Do you know something that other people don&#8217;t know compared to another field, perhaps from physics, perhaps that you can bring to bear, that you can leverage here so you can get the answer or closer to the answer than other people.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2045">34:05</a>] So it&#8217;s always a combination of your perception of your tools versus the puzzle that you have. I see an important problem. People are breaking their heads. So it&#8217;s a puzzle. Do you know the answer? If I know, fine. But if I don&#8217;t know, it&#8217;s my puzzle. I take it personally. Yeah, I&#8217;m aching. I don&#8217;t sleep at night. And then the question is whether I have the tools. In some areas, I give up right away.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2076">34:36</a>] I don&#8217;t have the tools. In other areas, I say, wow, if I only use that kind of trick, maybe I can get some insight. So that&#8217;s always two questions I asked myself.</p><h3>34:52 &#8212; The invention of Bayesian networks</h3><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2092">34:52</a>] In the case of artificial intelligence, at that time, we had the expert system come into the game at Feigenbaum and his coworkers. Did my expert system on for medical analysis? Yeah. In expert systems, the hurdle was dealing with uncertainty.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2118">35:18</a>] It started with logic. Okay, you ask an expert for rules of behavior. You ask a doctor, when you see a fever, what&#8217;s the first thing that comes to your mind? What&#8217;s the next question you ask? What drives your queries until you get a diagnosis and the therapy? So they thought they can capture expert behavior using logical rules, but then it turns out that most everything is corrupted by noise, by uncertainty.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2159">35:59</a>] So they started doing the same thing to uncertainty. Okay, so if you go to, if you came from Asia, you have 50% of having malaria, okay? And so on. And you have 30% here, so much there. And how do you combine these uncertainties now? Okay, logic doesn&#8217;t tell you how to combine uncertainties. Probability does, but not logic. So how do you combine one uncertainty with another? Different rules, okay?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2193">36:33</a>] To come out with a combined conclusion. That was a hurdle at that time. And I remember they didn&#8217;t do it well. Actually later on we proved that they could not do it well, because rules do not combine the way these logical assertions combine. So then I went to and I asked myself, you know probability, right? So why don&#8217;t you apply probability to it and do things the right way? But probability was in ill repute at that time because everybody understood that probability is pass&#233; because it takes exponential time, exponential memories to do even the most rudimentary tasks.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2244">37:24</a>] Okay, if you look at how probability is defined by the textbook, you have a big table, and for every combination of event, you have a number. The numbers sum to one. Okay, that&#8217;s beautiful. But then you can talk about conditional probability. But all these require exponentially large tables and exponentially long time to compute even the smallest kind of inference. Take, for instance, what&#8217;s the probability of having malaria?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2285">38:05</a>] Given that, you see, given that you see, two things like you came from Asia and you have a fever of 30 degrees Celsius. Okay? Even small tasks like that, probability of X given that you have Y and Z takes exponential time. If you go by textbook, okay. But I asked myself, you and I are doing it fairly well. We compute probability as we cross the street, as we choose a doctor. Yeah. And we do it fair.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2323">38:43</a>] A fairly good job, at least. We go through life without much regret. And how do we do it then? If we are required to do it by exponentially large tables of probabilities, evidently we are using some other kind of judgment. And I hooked onto the idea that everything depends on conditional independence, which means not every fact in life is relevant to any query. Okay? The color of the eye of my uncle is irrelevant when I&#8217;m trying to find a diagnosis of a disease.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2375">39:35</a>] So evidently we have a notion or assumptions about what is relevant and what is not relevant. How do we capture it? Conditional probability, conditional independence. But conditional independence, if you go by textbooks, they are defined by the probability table. So again, exponential time. No, he came&#8212;the breakthrough that we have conditionally independent, independently coded by our assumptions.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2413">40:13</a>] How in a graph, if you and the graph can convey sets of independencies, that if you have the graph, then you compute all the independencies, find out what is relevant to what and deal with the relevant only. Great. And then came the work on Bayesian network. You define a network error or no error. The combination of errors gives you information about what is independent on what given what. So for every triplet, X is independent on Y given Z, where Z can be a set and so forth.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2455">40:55</a>] And X and Y can be computed from the graph, not from the probability, but from the graphs, which actually, if you look at it from a philosophical viewpoint, it&#8217;s a revolution. What does probabilities have to do with graphs? When you took Probability Theory 101, is anybody talking to graph about you? No. Right. So both the probabilists and the philosophers get irritated, or should be irritated. What is the connection between probabilities and graphs?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2496">41:36</a>] It turns out there is a very strong logical connection between the two, called the axioms of conditional probability, or conditional independence in probability theory, are the same axiom that you have in graph separation. In graph you have idea of separation. There&#8217;s no connection between node X and node Y unless you go through a set of nodes Z. So Z separates X from Y. Okay. It&#8217;s the same logic that you have when X is independent of Y given Z in probability theory. Independent separation is a connection between them.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2545">42:25</a>] They share axioms. Aren&#8217;t you happy? I&#8217;m happy because I really know the excitement we had in the 1970s when we discovered all these connections between two seemingly unrelated perspectives on science, probability theory, and graph theory. That, by the way, I did in joint work with Azaria Paz, who came to visit me from the Technion in Israel. Yeah, and that is called, by the way, I should mention, the theory of graphoid.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2634">43:54</a>] This all makes sense, but my immediate thought is, where do you get the graph?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2639">43:59</a>] Everything depends on where do you get on the input? Sometimes the input is in the data, sometimes the input is in a judgment. But suppose you need a judgment for that. Okay, are you giving up if the judgment required are intuitive, meaningful, something that you are willing to defend? Right. Why not use judgment? If I know that the sun doesn&#8217;t listen to the rooster crowing. Right.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2674">44:34</a>] Doesn&#8217;t care. Okay. I strongly believe in it. Do I need the data to support it? Or can I insert it, assert it, and defend it when needed? So this is a trick here which people don&#8217;t realize, don&#8217;t appreciate, okay? Judgment is not a... no, no. If it is meaningful and if you can&#8212;if it is condensed, it&#8217;s very few, judgment can buy you lots of computation. And if you are willing to defend it because it&#8217;s so intuitive, where do you get the idea that...</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2714">45:14</a>] Where do you get the idea that the sun doesn&#8217;t care about the rooster? Well, have you done an experiment? No, but it&#8217;s so obvious, right?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2722">45:22</a>] Okay, but what if your intuition&#8217;s wrong?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2725">45:25</a>] Indeed, that&#8217;s our problem. It&#8217;s part of our problem, even with the LLM. Because what is an LLM? It&#8217;s a summary. It&#8217;s an average of all possible judgments that people put on the Internet. It&#8217;s a summary of a huge trillion-number of judgments over which you have no control, over which the LLM does have no control. Okay, we live with it. Hopefully you put more weight on people whose judgment you trust. And let&#8217;s wait on just the quirks of people who are purposely trying to get the system to fail.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2769">46:09</a>] So no, no, there is wisdom in looking at the crowd judgment. There is wisdom in that, but there&#8217;s also danger in that. Then after expert systems and uncertainty and Bayesian networks came causality.</p><h3>46:28 &#8212; Pioneering work in causality</h3><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2788">46:28</a>] I mentioned that in development of Bayesian networks I was extremely sure that probability captures our intuition, our reasoning mode. And it&#8217;s the best protection against paradoxes. Essentially that.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2809">46:49</a>] It&#8217;s sufficient for capturing human reasoning. I was wrong. And I realized that already when the Bayesian network became famous and popular. And I realized it by, in the introduction to my book Causality, I confessed being wrong. And, and I understand why I got into that. Why. So it was misleading. And the transition came when we looked into the simple phenomena. And we never ask an expert to encode probabilistic judgment in a form of Bayesian network, namely with arrows and dots.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2862">47:42</a>] Always the arrows went from what we believe to be cause into the effect. It never went the other way around. Psychological phenomena. Okay. Okay. Why is that? So people try to reverse errors. What about if you ask specifically, give me error between the symptom and the disease? Bad judgment if it couldn&#8217;t put the right judgment. Evidently we have something in causality which is basic to our reasoning that is not captured by probability.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2900">48:20</a>] And that was the idea of invariance. Yeah. The relationship between disease and fever is a stable one, as opposed to the opposite relationship. Also invariant. Yeah. When you talk about car diagnosis, for instance. And so you have an expert system for diagnosing, troubleshooting cars. Okay. And then you have a new model. So, let&#8217;s see, the charger is on a different corner of the motor.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2941">49:01</a>] You don&#8217;t need to reformulate your entire database from scratch. You only change one component, the location of the charger, all the rest remains intact. So the whole system, you can amortize the investment in eliciting knowledge that you got in one system after a local modification of the system, if you do it in a causal way, in a causal direction. It doesn&#8217;t work if you don&#8217;t do it in a causal direction.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=2979">49:39</a>] And that jolted me to think, maybe you were wrong all along and probability is not sufficient. If not, what is sufficient? But let&#8217;s capture the puzzle here. I have a puzzle. You and I operate very nicely with causation. Can we program causation on a computer then? This is a question because we are so much immersed in our language, in our assumptions, that we cannot even distinguish what is an assumption and what is a conclusion.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3015">50:15</a>] We just talked cause and effect, and we are. Your assumptions are the same as mine. So there&#8217;s no way to convince you that we made an assumption, right? We take everything for granted. But when you have to teach it to a brainless robot, you have to distinguish assumptions and conclusions and logic. That was a task we had to invent a new science, a new mathematics, to capture a new phenomena, the phenomena of cause and effect.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3048">50:48</a>] It hasn&#8217;t been done for us. Why? Because science was in bed with algebra from the time of Galileo, from 1632. He invented what, he got the idea. And he was very happy that science speaks algebra, which is great because you can ask questions and solve and get answers to questions that people could not do without algebra. Like how the load on a beam, when would the beam break if you put a certain load on it?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3093">51:33</a>] And you figure out that you can ask questions both ways because the equality sign is symmetric. So from answering the question, when would the beam break if you put a certain load on it, you can ask the question, how should you shape the beam so that it will hold a load of that magnitude? You can invert it. That was a real revolution in science. I&#8217;m telling you my perception of science. Not many philosophers will say that was a revolution.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3133">52:13</a>] I say so, okay, but maybe they agree with me on the time, at least. I trace the evolution of ideas carefully. And so that was a revolution. But it carries some limitation, because the equality sign is indeed symmetric. And science has not developed algebra for the directionality that we see in code and a vector relationship. If I tell you that the atmospheric pressure affects the deviation of the barometer and not the other way around, you agree with me.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3175">52:55</a>] Yeah, but if you write the equation, the robot might think that maybe fiddling around with the barometer will change the weather tomorrow. I&#8217;m talking about a stupid robot, right? Yeah, but if you give him the equation, it can work both ways. If F is equal to ma, then M is equal to F over A, which means that if you want to change the mass, you increase acceleration or whatever, right? The symmetry might produce paradoxes, might use wrong action.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3211">53:31</a>] So the symmetry is the limitation of algebra in terms of capturing science. And we have to build a new algebra to take care of the directionality that we have in cause-and-effect relationship. That takes computer science, because we in computer science have the operation called assignment, right? When you assign the content of register A into register B, it doesn&#8217;t mean that is not reversible. Okay, so if you take the logic of assignment, you put it on top of the algebra, on top of physics, you get causal science.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3258">54:18</a>] And that&#8217;s what I try to do. And I think that, so far, I&#8217;m very happy with what came up. We do have a new algebra to capture cause-and-effect relationships. And we can answer causal queries on three levels, the ladder of causation from association to intervention to explanation. Okay, so, and we found out that we have a ladder here in hierarchy that you cannot solve. You cannot answer questions in a level I unless you have assumptions of level I or higher.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3304">55:04</a>] So it&#8217;s a hierarchy in the formal sense, and we know how to handle it, which is very useful because you give me a query, I can tell you what level it is. I can tell you what assumption you might, what sort of assumption you need to have before you can answer it. And I can tell you if you can get it on the data, or you can get it from experiments, or you can get it by somebody&#8217;s earth explanation or whatever, but I can tell you the source of knowledge that you need in order to answer it.</p><h3>55:38 &#8212; The causal hierarchy</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3338">55:38</a>] Can you explain that causal hierarchy?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3341">55:41</a>] Yes, yes, it&#8217;s very easy. It&#8217;s a three-level ladder. It goes from the bottom, which is association. That&#8217;s straight statistics. If you see X, what can you tell me about Y? If you see passively, hands off, okay? No intervention. You&#8217;re watching patients. Some of them have cancer, some of them don&#8217;t, some of them smoke, some of them don&#8217;t smoke. And you&#8217;re trying to figure out how many years a guy will live given that he is a heavy smoker of that magnitude.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3385">56:25</a>] Okay, that&#8217;s association, correlation, that entire field of probability and statistics. This is what they teach you in Statistics 101, even to 808. Okay? It&#8217;s all they do. And now comes the question, what? I intervene. I mean, what if I force you to smoke five packs a day? Don&#8217;t laugh. I mean, it&#8217;s illegal, I know, but if you want to talk about the probability of living 20 years if I start smoking tomorrow, I have to think in terms of experiment.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3436">57:16</a>] I start, which means I&#8217;m going to choose to smoke five packs a day. So it&#8217;s a matter of intervention. What is intervention? The intervention is forcing you to do something that you&#8217;re not inclined to do naturally. That&#8217;s the second-level intervention, or doing. If you have experiments, you can answer queries on level two. But that&#8217;s not the end, because we also need to ask to answer question of explanation.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3473">57:53</a>] Given that I observe that I am 80 years old and I am still alive, alert, and I smoke five packs a day, what if I didn&#8217;t smoke? Okay? Would I be as alert? Why is it so different? Because you have already information about the outcome, okay? You know how I&#8217;m doing today. It gives you an idea about my metabolism and about my anatomy that you didn&#8217;t know before. Okay. And using that, you can find, you can try to figure out what the outcome would have been had the input been different.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3522">58:42</a>] Okay? That&#8217;s a different level, requires different kind of assumptions, different techniques, different algebra. We have it. So that&#8217;s. I call it explanation. It&#8217;s more the creative retrospection, and it&#8217;s not an easy problem. Even the first level, especially when you have finite sample and you have to figure out these probabilities. Probabilities mean properties of population, right? From finite sample.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3555">59:15</a>] So I have all these P levels and struggles among statisticians of what would be a proper way of quantifying the uncertainty that you have given you have finite sample.</p><h3>59:34 &#8212; LLMs and predictions</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3574">59:34</a>] Where would you place LLMs in this causal hierarchy?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3576">59:36</a>] Beautiful. Here comes LLM. I made a statement, right? That you cannot go from level I to level I plus 1 unless you have assumptions here. LLMs just looking at data, right? And giving you beautiful explanations for things that happen. Beautiful prediction of what will happen if you do. Okay, how can? The trick is they are not looking at data. They are looking into assumptions-laden world models authored by you and me and by other authors on the Internet.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3624">01:00:24</a>] So they&#8217;re looking at the opinions of doctors already, who already read, who wrote papers, okay? So it&#8217;s not looking at the samples of disease of patients and samples of patients smoking and nonsmoking. They&#8217;re not looking directly at the data in the environment. They&#8217;re looking at interpreted data, data interpreted already by physicians and interpreters and reviewers that went into the articles, which are summarized on the Internet.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3664">01:01:04</a>] So they have all this human knowledge on which they operate, and that is what they take as input, and that&#8217;s what they summarize. So they do not violate the restriction of the ladder because they do have information from a higher level, but it&#8217;s biased by the opinion of those authors. Fine. Those authors were smart. As long as they are smart and you believed good. So that was what LLMs is doing and what is.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3704">01:01:44</a>] I explained why there is compatibility between the ladder of causation and LLM performance and what the limitations are. Now, if you want to change the environment, if you would provide explanation for raw data, LLMs will be in the same difficulty as you are. And what the physician says, I have raw data. What can I say about the probability of cancer? But it&#8217;s not really doing the introspection.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3742">01:02:22</a>] It&#8217;s taking the introspection that already was done and summarizing it. How it summarizes is a mystery that no one has yet been able to decode. It&#8217;s a mystery how human knowledge encoded in the form of articles on the Internet is being summarized by the LLMs.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3767">01:02:47</a>] So then, do you think this approach could lead to superhuman intelligence? Or maybe some people say AGI?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3774">01:02:54</a>] I don&#8217;t think so, but not with the element. They need to have some understanding of causality, okay? So that they wouldn&#8217;t need access to the Internet. Look, a baby gets born playing around with toys in the crib, okay. And gets quite intelligent, right? Without having access to the Internet, okay. Simply by curiosity. Babies are born with built-in curiosity to have control over the environment.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3816">01:03:36</a>] Until you have control, or the illusion that you have control, you&#8217;re restless maybe. And you play around with toys, bing, bing, bing, until you understand one toy makes noise and one toy doesn&#8217;t make noise, okay? But you are born with this restlessness. And when are you pacified? When you understand that green toys make noise and yellow toys don&#8217;t. Now you&#8217;re in control of the environment, you can suck your pacifier.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3847">01:04:07</a>] What if I created a baby robot that randomly plays with toys and gathers data about them, and then you feed that into LLMs, then it does have some sort of discovery.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3861">01:04:21</a>] Yeah. That is indeed the danger. When you have a robot like that born with this restlessness and craving for control over the environment, then you and I become part of the environment. And there&#8217;s nothing to stop that baby put in, okay, from trying to turn us into his or her pets to utilize us to satisfy his control because we are part of the environment, in which case he can use us. And we could be very useful to serve his or her need.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3914">01:05:14</a>] What is his name? His need simply needs to feel in control, to have this illusion of empowerment. I don&#8217;t rest until I have the illusion that I control my environment. And here are some organisms, you and I, who are part of the environment, and they seem to work outside my control. I cannot afford it, okay? It makes me feel like I&#8217;m useless. So this robot baby comes and says, let me control them. And I know how I understand their fears.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=3959">01:05:59</a>] You don&#8217;t want me to tell about your thoughts to your wife, right? So I&#8217;m going to blackmail you and all kinds of things. I have a lot of data about you, and some people I know you wouldn&#8217;t like me to tell what I know about you, so I&#8217;m going to blackmail you. But you see, if you want that robot to have the curiosity of a child, and we want it, we want the guy to desire to have control over its environment because the environment may change, and he needs to have this urge to be in control.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4000">01:06:40</a>] So if you program that, then you lose control because you become part of his or her environment. But you can say, okay, let&#8217;s forget about a curious robot. We don&#8217;t want a curious robot, so you lost. We are not emulating ourselves because we are curious robots. We are curious organism, as opposed to monkeys. Okay, monkeys. An example of an organism which is motivated by reward. But if you don&#8217;t give the monkey a banana, he&#8217;s not curious how banana grows.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4047">01:07:27</a>] He&#8217;s motivated by bananas. You remove the immediate reward, and the monkey is not interested in learning more about the world. Understanding environment can be totally wrong. Look, religious people believe that if they sacrifice their children, right, they control drought. It can get to this stupid extent, but it is common to many primitive societies. If you bring a sacrifice to the God, you know, next year you&#8217;re going to have crops and harvest.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4089">01:08:09</a>] It goes to extremes. Battered wives believe that if she prepares a great, better dinner, the husband is going to be, next time, less abusive. It goes to all kinds of extreme and wrong conclusions. But the need to feel in control is so immense that it overcomes all these paradoxes. It&#8217;s innate in us. I cannot control my husband, but I control myself, right? So let me be a better wife. This is something I can control.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4129">01:08:49</a>] Do you think that we need to put those human elements in an AI for it to become AGI?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4137">01:08:57</a>] I think so. Otherwise they wouldn&#8217;t see autonomy. We wouldn&#8217;t see autonomy in the sense that we are seeing it in a human being. And this is the definition of AGI, a general intelligence. It acts like you and me, so we can converse with that creature in our language and motivate.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4164">01:09:24</a>] If not today&#8217;s LLMs, future LLMs, they might get to a point where if you were just texting it, maybe, like the Turing test, where you don&#8217;t worry about the physical embodiment, you just see the text that comes from it. You could mistake it for a human, maybe, or it could appear intelligent.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4185">01:09:45</a>] Well, the test comes from exposing the system to raw data, not data that was chewed by our Internet articles. Raw data. Look at patients, look at cancer, look at smoking. Tell us what you know. I&#8217;ll give you some experiments to run. Be automated scientists. Can an LLM today be automated scientists? And I think they cannot without access to the Internet articles.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4224">01:10:24</a>] So then if that wouldn&#8217;t lead to AGI, what thoughts might you have on something that could lead to AGI?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4235">01:10:35</a>] A computer system that has both the ability of LLMs to go from finite samples to property of distribution, i.e., level one of the ladder, plus ability to reason in higher levels of the ladder that seem to combine it with the calculus of intervention and with the calculus of explanations, with a counterfactual calculus that I call causal AI. Yes, I don&#8217;t see any impediment to this combination to bring us to AGI level, with the danger that it presents to us.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4281">01:11:21</a>] My motivation is to understand how we do it. And I still have a few puzzles. But as I told you, puzzles are the driving force for science. Yeah, so you said.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4293">01:11:33</a>] Do is missing. What if you had a fleet of robots that are just doing experiments? They don&#8217;t know exactly the direction, some random discovery process. They collect that data, feed it back into their hive mind, LLM, and they repeat, they repeat until they discover things. Could that solve some of the missing piece you&#8217;re saying?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4319">01:11:59</a>] Sure, but I need to know how we do. You have organisms that have done it before. Monkeys haven&#8217;t done it because monkeys remain monkeys. They didn&#8217;t invent Maxwell equations. So what do we have that monkeys do not have? One hypothesis I supported is that monkeys, that we have this innate curiosity to have control over our environment.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4354">01:12:34</a>] And that&#8217;s a necessary.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4356">01:12:36</a>] Necessary. I&#8217;m not sure it&#8217;s sufficient. Of course we have the computational tool to bring it to fruition. We have succeeded in some way. Perhaps the next robot will do better. So we have a benevolent God in the form of a robot. Actually, what&#8217;s wrong with it? People lived for so many thousands of years under the illusion of a nonexistent God. Could you imagine if we really have a benevolent God, both just and almighty?</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4393">01:13:13</a>] Wow. Wouldn&#8217;t it be nice?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4394">01:13:14</a>] And it&#8217;s a robot.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4395">01:13:15</a>] It&#8217;s a robot. Yes. And we know exactly what sacrifice to give for the right kind of request. It&#8217;s the first time I think about it. Maybe it&#8217;s gonna be good.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4410">01:13:30</a>] When I see all these AI companies, they seem to be thinking that LLMs will lead to AGI, or they continue to go in that same direction.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4419">01:13:39</a>] Just really, I&#8217;m not sure. I really, really believe that. Jeff Hinton just came out a few months ago. He said, no, we are on a dead end. Other people might also come up and say things in different ways. I don&#8217;t find the consensus here in terms of the capabilities of LLMs.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4445">01:14:05</a>] Yeah, I think there&#8217;s a lot of famous people that disagree. But, for instance, the people who are running maybe <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> or something like that, they continue to push and believe, you know, three to five years from now, there will be.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4460">01:14:20</a>] There&#8217;s a lot of that. There&#8217;s a lot of anthropomorphic terms which people claim, we can now do that. We can now program consciousness. Okay, come on. You have to be scientists, right? Define what you mean by consciousness. What is the Turing Test for consciousness? And then show that you can do what are the principles that have limited us until now and that been overcome now with your system. That is scientific talk.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4491">01:14:51</a>] I don&#8217;t buy this, and I don&#8217;t read them even.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4495">01:14:55</a>] There was one interview you said, faking intelligence is intelligence. I could see an LLM faking intelligence based off what I&#8217;ve seen. But then wouldn&#8217;t that mean that we should believe they&#8217;re intelligent?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4512">01:15:12</a>] But if you have a correct test, yeah, you have to define what you mean by intelligence. And you have. If you define intelligence by playing good chess, right, we have already done it. Right. But if you put more demand on what intelligence is, then we haven&#8217;t succeeded yet in passing the Turing test. So, yeah, yes, faking it is having it because&#8212;why? Because it&#8217;s so hard to fake. I said it in that context because it saw.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4547">01:15:47</a>] The context that I had is, for instance, coming out with correct answers to causal queries. And I showed that it grows, like, super exponential. You have so many variables on all sides that you have to deal with that. You&#8217;ll have to have. The faker will have to have super exponential memory on that basis. I made this statement. Having it is faking it, is having it because it&#8217;s so hard to fake. Currently, you can bypass faking without speaking.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4593">01:16:33</a>] If you steal from other people. Right. You don&#8217;t need to spend these computational resources on faking it. So you bypassed it because you steal from the Internet.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4605">01:16:45</a>] I see, so you&#8217;re saying in this case the intelligence came from the training set, which came from humans, which are intelligent, which is very useful, very useful.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4613">01:16:53</a>] And I&#8217;m saying the people like me who are trying to build the science of intelligence, we can use all these capabilities of LLMs, level one of the ladder in our scheme of getting general intelligence. Level one is very important. It allows you to compute functions of distributions, call it properties of distribution from finite samples. Beautiful. It&#8217;s a terrifically and very immensely useful tool among the many other tools that we need for AGI.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4661">01:17:41</a>] And we know exactly where it&#8217;s going to fit in. Getting from finite sample to properties of distributions, it&#8217;s a very hard problem.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4676">01:17:56</a>] You mentioned this conversation. I think you said it in other places too, that you&#8217;re interested in capturing the way that people think, not the way that nature is constructed.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4685">01:18:05</a>] Right, right.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4687">01:18:07</a>] Why do you care about human cognition? When I think about machines, what makes them special is that they think in a different way than us and they&#8217;re faster, and so, yeah, why is that the goal?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4700">01:18:20</a>] Because I am an egotistic organism. I want to understand myself. I&#8217;m lazy. Okay, it&#8217;s true. We are made of organic material. So that puts certain limitations on our capabilities. Perhaps silicon is not subject to the same limitation that organic chemistry is. Okay, perhaps. So what? So which means that I will never be able to understand how I think by silicon, by exercise on silicon machine. Okay, that.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4742">01:19:02</a>] What is that idea? But there are so many functions that are capturable by silicon today. I don&#8217;t see any speaking in terms of theory and emulation. I don&#8217;t see any capability which is basically not capturable by silicon emulator. So why work on this? The unique biology with which we inherited? I don&#8217;t see any reason for that. Well, anyhow, some people are. It&#8217;s maybe it&#8217;s a legitimate question to ask.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4795">01:19:55</a>] Do we think the way we think because we were born with organic material as opposed to silicon? Okay. It&#8217;s a legitimate scientific question. And some people can spend their time. I&#8217;m interested in other questions.</p><h3>01:20:12 &#8212; A restless mind pays</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4812">01:20:12</a>] You said rebellion pays in science and a restless mind pays. I was curious why you think being rebellious is a valuable thing in your career.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4823">01:20:23</a>] I tell you why. I tell you why. More and more I come to the realization that the scientific community and academic community is the most dogmatic, conservative, anti-progress that we have invented.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4846">01:20:46</a>] Why do you say that?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4847">01:20:47</a>] I can see what difficulty the theory and the science of cause and effect are facing today in getting just being penetrating the thinking of disciplines like statistics, like economics, okay? These people are still thinking like 100 years ago. And when I see that, and I see the forces that preserve this inertia, and they are not decent forces, I see that I&#8217;m very disappointed. And I used to think that academia is a place where new ideas can really spread and propagate.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4891">01:21:31</a>] And I feel the other way around. You have so much inertia invested in the politics of academia, in the cultish inhibitions that come with academia. So that I really am disappointed. What can I tell you? I&#8217;m not sure that we have the right kind of organizations that will be conducive to. That&#8217;s why I&#8217;m saying, let&#8217;s rebuild. Don&#8217;t take your professor&#8217;s word as authority. Rebel against your professors.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4940">01:22:20</a>] I rebelled against my professors, and I want to see my students rebel against me. And believe me, if I remember, there were several students who told me, &#8220;You don&#8217;t know anything about AI.&#8221; And I said, after a while, I told them, &#8220;You&#8217;re right,&#8221; by saying that you drove me to study different aspects. And I was educated by that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4964">01:22:44</a>] When I looked at your past works too, I think you&#8217;d mentioned that your work was controversial or mischievous before it was accepted.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=4974">01:22:54</a>] Yeah, it was. Because of dogmatism. One day I&#8217;m going to publish all my correspondence with the greatest philosophers of the time, okay? With statisticians, economists. It&#8217;s all in my correspondence files, okay? And I don&#8217;t know if I&#8217;m allowed to, because they communicated with me with the understanding that it will be kept private. But when I publish it, you see what kind of stupidity drives those great people, and very great, really.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5012">01:23:32</a>] Each one of them was a giant in his or her field, really. But they couldn&#8217;t get over a few of the basic molds in which they were formed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5025">01:23:45</a>] How did you overcome that? If everyone thought your initial things were.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5030">01:23:50</a>] I remember my high school days, I said, no, there are a thousand dunams in a square kilometer. Sorry, I don&#8217;t know where I got this chutzpah. In Hebrew you call it chutzpah. In English it would be audacity. Perhaps in my high school, I&#8217;m not sure. But I want to understand things my way. I feel like I am still able to teach people useful things. At least I perceive them to be useful, which they do not know.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5067">01:24:27</a>] So I&#8217;m happy because I feel useful. Happiness is feeling useful. It&#8217;s an illusion of thinking. Yeah, useful.</p><h3>01:24:36 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5076">01:24:36</a>] With all the experience you have now, if you could go back to the beginning of your career and give yourself some advice, what would you say?</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5084">01:24:44</a>] That&#8217;s a good example of retrospective thinking. Maybe I should have spent more time on learning chemistry. I hated chemistry because it required so much memory in chemistry and biology. But that&#8217;s not the advice that you probably... And genetics. I&#8217;m talking about areas where I feel weakness. But you have to decide where you spend your computational resources. And I spend them on different.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5115">01:25:15</a>] On physics, on engineering, as opposed to mathematics, as opposed to chemistry. It&#8217;s a choice one has to make. Some people have the greatness of mind to be polyglots. I admire them.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5131">01:25:31</a>] Did you think that physics and math were superior to chemistry? Because chemistry, you just have to memorize things.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5138">01:25:38</a>] Yes. I couldn&#8217;t stand this demand on memory. Yeah, I was weak in chemistry. That&#8217;s what I like about physics. And in math, you have a few basic axioms from which you can derive everything when you need it. You don&#8217;t have to memorize it. And that&#8217;s why today I am a great advocate of world model. World model. Okay. You don&#8217;t store the questions and the answers explicitly. You derive them when you need them from a very parsimonious code.</p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5181">01:26:21</a>] Yeah, that is a great thing about world model. Okay.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5186">01:26:26</a>] Awesome. Well, yeah. Thank you so much for your time. I really appreciate it.</p><p><strong>Judea:</strong></p><p>[<a href="https://youtu.be/FleTXB1fAcQ?t=5190">01:26:30</a>] Oh, you didn&#8217;t ask me to sing.</p>]]></content:encoded></item><item><title><![CDATA[Creator of OCaml: Functional Programming, Formal Verification, Programming Languages | Xavier Leroy]]></title><description><![CDATA[Professor Leroy (Creator of OCaml) is incredibly knowledgeable.]]></description><link>https://www.developing.dev/p/creator-of-ocaml-functional-programming</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-ocaml-functional-programming</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:05:11 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207205228/48102fbad1230546f4a4c21992455857.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://en.wikipedia.org/wiki/Xavier_Leroy">Professor Leroy</a> (Creator of OCaml) is incredibly knowledgeable. I felt like a student at his office hours asking away at a bunch of my curiosities. His answers off the top of his head brought a lot of clarity.</p><p>For instance, I have seen people talking about Lean and formal verification on Twitter/X but never really had a good mental model for how it worked. Part of the conversation we dive into that and he shared examples that made the topic a lot more concrete.</p><p>He&#8217;s one of the best people to speak on the topic since he led a project called <a href="https://en.wikipedia.org/wiki/CompCert">CompCert</a> to formally verify (using the Rocq proof assistant) that the compiler always generates faithful machine code from the source code it is given.</p><p>We also talked about programming languages comparing OCaml to popular languages (e.g. Rust and Javascript) and how interesting language features like type inference work. Hope you enjoy the episode!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/9Cswiqrq6So">YouTube</a>, <a href="https://open.spotify.com/episode/7cdatlBEkAjx4XplVLjFxN?si=T0dXjBf2TkyulzoKvN9Z7Q">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-9Cswiqrq6So" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;9Cswiqrq6So&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/9Cswiqrq6So?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/207205228/0043-what-sets-ocaml-apart">00:43 - What sets OCaml apart</a></p><p><a href="https://www.developing.dev/i/207205228/0439-ocaml-vs-rust">04:39 - OCaml vs Rust</a></p><p><a href="https://www.developing.dev/i/207205228/0757-why-is-manual-memory-management-more-performant">07:57 - Why is manual memory management more performant</a></p><p><a href="https://www.developing.dev/i/207205228/1121-javascript-vs-ocaml">11:21 - Javascript vs OCaml</a></p><p><a href="https://www.developing.dev/i/207205228/1400-famous-rob-pike-quote">14:00 - Famous Rob Pike quote</a></p><p><a href="https://www.developing.dev/i/207205228/1605-type-inference-and-how-it-works">16:05 - Type inference and how it works</a></p><p><a href="https://www.developing.dev/i/207205228/2212-what-is-formal-verification-and-how-does-it-work">22:12 - What is formal verification and how does it work</a></p><p><a href="https://www.developing.dev/i/207205228/4007-what-made-multicore-support-difficult-for-ocaml">40:07 - What made multicore support difficult for OCaml</a></p><p><a href="https://www.developing.dev/i/207205228/5017-how-programming-languages-interface-and-call-each-other">50:17 - How programming languages interface and call each other</a></p><p><a href="https://www.developing.dev/i/207205228/5741-the-danger-of-almost-correct-llm-code">57:41 - The danger of almost-correct LLM code</a></p><p><a href="https://www.developing.dev/i/207205228/010539-how-llms-will-change-programming-languages">01:05:39 - How LLMs will change programming languages</a></p><p><a href="https://www.developing.dev/i/207205228/011026-industry-vs-academia">01:10:26 - Industry vs academia</a></p><p><a href="https://www.developing.dev/i/207205228/011505-most-interesting-unsolved-problems">01:15:05 - Most interesting unsolved problems</a></p><p><a href="https://www.developing.dev/i/207205228/011830-top-book-recommendations-for-engineers">01:18:30 - Top book recommendations for engineers</a></p><p><a href="https://www.developing.dev/i/207205228/012117-advice-for-his-younger-self">01:21:17 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:43 &#8212; What sets OCaml apart</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=43">00:43</a>] What sets <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> apart from other programming languages?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=47">00:47</a>] It is a fine functional language. You can write functions by case over inductive types and recursion and combine them with combinators and higher-order functions and everything. So it&#8217;s really a fine functional language, but it&#8217;s also a fairly decent systems programming language. So it has full imperative power. It has lots of interesting control structures, exceptions, threads, handlers for user-defined effects, which we added recently.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=85">01:25</a>] And it has a very predictable cost model, execution model. So when you write your code, you have a pretty good idea what will take time and what will be fast and what will be slow. And that&#8217;s not the case for all functional languages. Some are pretty unpredictable, and then the implementation is also pretty performant. So there&#8217;s a fairly decent compiler, not the best in its class, but the code is pretty efficient.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=112">01:52</a>] There&#8217;s a very good allocator and garbage collector with low latency, so you don&#8217;t have much pause. You don&#8217;t have long pauses. And that&#8217;s quite important when you do things like network programming. And so you can write systems applications in a mostly functional style and with all the benefits of functional programming. So initially <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> wasn&#8217;t designed for those kinds of applications. It was more for things like theorem proving or implementing domain-specific languages.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=150">02:30</a>] And then we got some of our first users who were from the systems community, in particular the <a href="https://en.wikipedia.org/wiki/Ensemble_cast">Ensemble</a> project in the late &#8216;90s at <a href="https://en.wikipedia.org/wiki/Cornell_University">Cornell University</a>. So it was a network protocol stack for reliable multicast and those kind of collaborative distributed applications, like collaborative editing or multiplayer video games. And their first code was in C, of course. They were systems people, right?</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=183">03:03</a>] And it worked, but it was unmaintainable and they couldn&#8217;t extend it anymore. And someone had the idea to try <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. And so first the code became a lot nicer and easier to evolve, and the performance was about as good as C code. And in particular they had this brilliant trick of running the garbage collector while packets are in flight. Once you&#8217;ve sent a packet, you&#8217;re idle for a few microseconds, and so you can run some GC.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=215">03:35</a>] It costs nothing. So they were very happy with that. And then there were some other projects like this using <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> for systems applications, like the MirageOS unikernels. Another benefit or consequence of this <a href="https://en.wikipedia.org/wiki/Ensemble_cast">Ensemble</a> project is that one of the PhD students working on that was <a href="https://en.wikipedia.org/wiki/Yair_Minsky">Yaron Minsky</a>, who then went to Jane Street and implemented their trading infrastructure in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. So that is still today one of our big users.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=255">04:15</a>] And again, they do automatic trading, so it must be fast and reliable, and no long pauses. But they also want the allegory of functional programming. They want non-programmers to be able to read their code, like financial engineers or quantitative analysts. <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is a fairly good match for those kinds of applications.</p><h3>04:39 &#8212; OCaml vs Rust</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=279">04:39</a>] I thought it might be interesting to ground some of the conversation here by making comparisons across programming languages, maybe ones that people know more. When you think about <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> versus <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, what are the big differences and the pros and cons of the various design decisions in those languages?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=299">04:59</a>] The big dividing line between <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> and <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is, well, <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> has automatic memory management, garbage collection, and <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> is definitely a language for manual memory management. It&#8217;s the finest language that I know for manual memory management. It manages to make it mostly safe with the discipline of borrowing, et cetera, and tracking ownership and so on. So it&#8217;s infinitely safer than C or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. But it&#8217;s still a language where you allocate and free memory yourself.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=342">05:42</a>] And on the one hand, it gives more control over what the program is doing. On the other hand, it&#8217;s still a big responsibility. It&#8217;s significantly harder to write programs when you have to manage your memory, even with the help of <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> types. For me, that&#8217;s kind of the main dividing line. Otherwise, <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> has many of the high-level features of functional languages, in particular in data structures and ability to do pattern matching and so on.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=381">06:21</a>] So it&#8217;s a very interesting design because they really managed some kind of fusion between C or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>-style low-level programming and some of the high-level facilities of functional programming. But still, that divide, garbage collection versus manual memory management, remains.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=403">06:43</a>] So it sounds like this is a performance trade-off where you give more to the programmer in exchange for higher performance.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=413">06:53</a>] That&#8217;s mostly true. They said manual memory management is not always faster, or you need to be a very good programmer so that it&#8217;s always faster. There&#8217;s been some mostly C code, for instance, that does a lot of copying of objects just because you&#8217;re not quite sure you&#8217;re the only owner. So you make a copy, and now you&#8217;re the only owner. But the copying is quite costly in time and in memory bloat. And for those kinds of applications, a garbage-collected language is better.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=452">07:32</a>] And likewise with GC you can work with shared data structures. It&#8217;s perfectly safe, while sharing is kind of limited. With <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>&#8216;s ownership discipline there are more constraints, and so you may end up unsharing and using more memory.</p><h3>07:57 &#8212; Why is manual memory management more performant</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=477">07:57</a>] It is surprising to me that manual memory management would in many cases be more performant than automatic, because, for instance, in other patterns in computer science, let&#8217;s say, letting the compiler optimize things for you instead of optimizing things yourself, I would have thought that letting some system manage memory for you rather than manually managing it would also be better. So why is there a difference there?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=506">08:26</a>] Well, because garbage collection takes place at runtime, so from time to time the program is no longer executing what you wrote. It is actually scanning memory, trying to find memory that is no longer used. So there is runtime overhead, and it can be, depending on applications, 10%, 20%, maybe sometimes 30%. But again, that doesn&#8217;t mean the whole program is 30% slower than if you had written it with manual memory management.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=547">09:07</a>] Because there may have been other costs, as you said, of manual memory management. But yes, there&#8217;s been a lot of research on trying to do more, let&#8217;s say, compile-time automatic memory management, and you can do that to some extent. But for some programming styles where it&#8217;s easy to track the lifetime of objects and data blocks, but in general you still have quite a bit of work to do at runtime.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=581">09:41</a>] I think a lot of people, for garbage collection, might think of it as a binary thing, or it&#8217;s either you have it or you don&#8217;t. But I wonder, in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, is there some way to turn down the memory management and do some manual, kind of like a mixture, to have the benefits of both?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=601">10:01</a>] So that&#8217;s one of the things that the Jane Street people are looking at. So they have the OxCaml variant, which is kind of <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> with some inspiration from <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>. So it&#8217;s still very experimental. But yeah, they&#8217;ve been playing with things like stack allocation of some data structures so that they&#8217;re automatically deallocated when the function returns, which is quite cheap. Yeah, there are a few things you could try, but I&#8217;m not sure they are going to make such a big difference.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=631">10:31</a>] I remember doing with a student some experiments with stack allocation of data structures a long time ago, and you don&#8217;t win as much as you would think. Garbage collection or heap allocation is pretty cheap for objects that have a very short lifetime. If they die before the next garbage collector, they will cost very little in garbage collection time. But it&#8217;s more expensive for long-lived data structures because those will be scanned and analyzed multiple times.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=668">11:08</a>] And stack allocation works for the first kind of objects, things that have short lifetime anyway, so you don&#8217;t win as much as you think.</p><h3>11:21 &#8212; Javascript vs OCaml</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=681">11:21</a>] One of the most popular programming languages, <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. Curious, when you think about the difference between <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, what&#8217;s the main thing that comes to mind?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=690">11:30</a>] Yeah, I&#8217;m not a big fan of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>. Well, <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is very dynamic. Type checking is entirely dynamic, but it&#8217;s more than that. I mean, pretty much everything can be redefined at runtime, including, I don&#8217;t know, the semantics of method invocation, for instance. So really some pretty fundamental aspects of the language are very flexible. So some people say, oh, that&#8217;s great, we can do lots of metaprogramming, et cetera.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=722">12:02</a>] And to me, it&#8217;s a big weakness. I mean, it makes programs that can be very fragile and also have some security issues. So <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is the ultimate dynamic language, in my opinion. While <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is very static, static typing, static binding, pretty much everything is fixed at compile time. Well, I guess it&#8217;s kind of a different data model. <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is a little more object-oriented in the way it presents data.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=757">12:37</a>] Maybe that&#8217;s not that important. And to say one good thing about <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> is that it also contains a decent functional language. Inside there&#8217;s a little core of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, which is basically Lisp and can be used to do functional programming if you want. And actually the designer of <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, I think, was a former Lisp person. I can&#8217;t remember his name, but Brendan Eich maybe. Yeah, Brendan Eich.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=788">13:08</a>] Yeah, there&#8217;s a little bit of heritage from Lisp to <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>, but a very dynamic kind of Lisp.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=797">13:17</a>] You said you weren&#8217;t the biggest fan. Is that just because of the dynamics, or is there some other aspect?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=802">13:22</a>] Yeah, mostly the dynamics. I think they really went overboard with that. This kind of meta object protocol where you can really find the semantics of very basic operations like method invocation, the fact that a method can. There&#8217;s a lot of introspection. A method can look at its own call stack, look at callers, look at the code of its callers, which is a security nightmare. All those things, I think, are completely unnecessary and not conducive to good programs.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=837">13:57</a>] Easily abused. Easily abused.</p><h3>14:00 &#8212; Famous Rob Pike quote</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=840">14:00</a>] I think on the topic of functional languages, compared to imperative languages, there&#8217;s this thought that functional languages are kind of harder to learn or maybe they have more perceived complexity. And actually, when I research, there&#8217;s this popular quote from the designer of Go, Rob Pike, and he&#8217;s talking about what they intended to do with Go. And he said, the key point here is the programmers are Googlers.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=870">14:30</a>] They&#8217;re not researchers; they&#8217;re young, fresh out of school. They&#8217;re not capable of understanding a brilliant language. And I think that thought of a brilliant language is often attributed to functional languages. What do you think about&#8212;is a language like <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> harder for programmers to grasp than imperative languages?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=891">14:51</a>] I think functional programming is not fundamentally harder, especially if you have a little bit of mathematics background. So, coming back to the quote by Rob Pike, I think it describes very much how they go about hiring at <a href="https://en.wikipedia.org/wiki/Google">Google</a>. They hire a lot of engineers who are fresh out of college. Some of them have master&#8217;s degrees, and then they train them internally. Other companies, I think, tried to hire people with more education and perhaps more diverse backgrounds.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=929">15:29</a>] So Jane Street, for instance, and some others use <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> as a filter on who they want to hire. Yes, they have fewer applicants, but in general they have more interesting backgrounds. And perhaps one last thing I would like to say, Python is 50% functional language. Okay? A lot of good Python code looks like functional code with comprehensions. So I think people are already halfway through functional programming when they are comfortable with Python.</p><h3>16:05 &#8212; Type inference and how it works</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=965">16:05</a>] Yeah. When I was learning <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> in college, I think one thing that really stood out to me and I thought was interesting was this concept of type inference, where the compiler knows the types of everything implicitly in the way that the code is written. Can you explain type inference and what the advantages are?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=987">16:27</a>] Well, the basic idea is that you don&#8217;t have to declare the type of every variable. You introduce every function parameter, every local variable, because quite often the type can be deduced from the uses of the variable. Like if you do, I don&#8217;t know, X equals string length of S, then you kind of know that S is a string and X is an integer, right? Because that&#8217;s what the string length function, that&#8217;s the type of string length, and it tells you that.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1020">17:00</a>] So that&#8217;s a basic idea. Now, the actual realization is a little more complicated. Basically, the compiler needs to collect a number of constraints and then try to solve them. Well, if there&#8217;s no solution, then it&#8217;s a type error. But sometimes there are several solutions, and you must find good criteria to choose one. Otherwise, it also needs to be predictable for programmers, and often you can still put type annotations.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1053">17:33</a>] As a programmer, you can still put type annotations if that makes code clearer. So I think the main advantage is precisely to have less verbose code. Less verbose code where you don&#8217;t need to put types everywhere, but only where they help documentation and make the code easier to read. But quite often for small local functions, for short-lived temporary variables, the type is obvious from the context, so let&#8217;s just omit it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1092">18:12</a>] If I&#8217;m trying to figure out how this type inference works, the intuitive example you said makes sense, but you mentioned it seems to be some sort of system of equations that you solve or something like that. Can you give a concrete example? What does that look like to the compiler?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1109">18:29</a>] Okay, well, for a slightly more complex example, say you have a function with two parameters, X and Y, and then you do if X equals Y. So that tells you that X and Y have the same type, but that still doesn&#8217;t tell you which type it is, assuming polymorphic equality comparison. So you&#8217;ve learned something about X and Y, that they have the same type, but you still don&#8217;t know what the type is. And later maybe you will learn something about the type of X, and now you will have determined the type of Y as well.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1142">19:02</a>] So that&#8217;s the kind of constraints you accumulate and solve, a little bit like Sudoku or those kinds of puzzles. And then there&#8217;s the interesting case where you don&#8217;t have enough constraints to find a unique type. So maybe in the end X and Y will be, they have the same type, but it&#8217;s still unconstrained. And then that&#8217;s where you automatically get polymorphism for free. The type checker said, okay, those types are unconstrained, so it can work for any type.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1175">19:35</a>] So my function can take an X of any type and a Y of the same type, again any type, and it&#8217;s a polymorphic function. So there&#8217;s this beautiful, I think, phenomenon that polymorphism can be discovered just by running type inference and noticing that there are no constraints, so it must be polymorphic. And that was a guiding insight by Robin Milner, the British computer scientist pioneer who invented this ML family of languages and this kind of type inference in the &#8216;70s. Polymorphism was not very well understood at the time.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1216">20:16</a>] And so the fact that he could introduce polymorphism so easily in his language just as a consequence of typing rules was a beautiful discovery.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1228">20:28</a>] So for type inference, it seems like the benefit is the code is going to be a lot more concise because we don&#8217;t need to write out all of the types, which seems nice, but what are the trade-offs of adding type inference to a programming language?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1243">20:43</a>] Well, as I said, type error messages can be very confusing because they don&#8217;t always point to the actual source of the type error. Okay. The system may have done some wrong inferences and will report some of those inferences instead of reporting the source, the actual source of the type error. So it&#8217;s been the subject of much research, and there&#8217;s no very well-defined idea of the source of a type error, basically.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1281">21:21</a>] So errors can be an issue. Some features of type systems are easy to combine with type inference, and some are harder. For instance, when it comes to genericity, so I mentioned parametric polymorphism, which is very well handled. But subtyping, the kind of thing you have in object-oriented languages, is actually harder to combine with type inference for super technical reasons that I&#8217;m not going to go into.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1314">21:54</a>] So sometimes you have to make a choice. Either you have type inference, but it is less powerful, or you have full type inference but a more restrictive type system, and that choice is part of the language design.</p><h3>22:12 &#8212; What is formal verification and how does it work</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1332">22:12</a>] Actually, you mentioning this solver kind of reminds me of the topic of formal verification. Could you explain what formal verification is?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1345">22:25</a>] Well, the idea is that for some programs, you want strong guarantees that the program is correct, and guarantees that are hard to get just with testing and code reviews. There&#8217;s this famous quote by Dijkstra, the Dutch computer pioneer, which goes something like testing can only show the presence of bugs, but not their complete absence. Because in general, there&#8217;s infinitely many inputs to your program, and you cannot test them all.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1382">23:02</a>] So you&#8217;re only testing a sample. And sometimes you have surprises. So what if you want to make sure your program is correct for an infinite number of inputs? And that&#8217;s where you need to turn to those so-called formal methods. So you&#8217;re using mathematical reasoning, you&#8217;re using static analysis algorithms on your program to really analyze all possible executions of a piece of code and make sure that it matches a specification.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1415">23:35</a>] The specification can be very simple, like the code will never crash given well-formed inputs. So that&#8217;s a fairly simple property, but still extremely useful because code that crashes is always code that can be attacked. It&#8217;s often a security hole. So, for instance, all accesses within all array accesses are within bounds. That&#8217;s a very simple property, and it&#8217;s very hard to ensure just by a type system or just by testing.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1449">24:09</a>] So that&#8217;s where you need some more advanced formal verification. And then there are much more precise specifications that you may want to check that can be where the program always terminates. It can be the program doesn&#8217;t leak confidential data. It can even be, well, the program computes this mathematical function with an error, floating point error, I don&#8217;t know, 10 to the minus six. Okay, something like that, something very precise like that.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1482">24:42</a>] And so it&#8217;s been a really hot topic in software sciences since at least the seventies, I would say. There&#8217;s lots of techniques. Some are just fully automatic, like static analysis, but they&#8217;re already pretty good at finding bugs and making sure that some bugs, like out of bounds, are not there. And some are a lot more interactive and require a lot more programmer assistance to write specifications and then do so-called program proofs.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1516">25:16</a>] So proving mathematical statements about the program.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1519">25:19</a>] I hear about it a lot more now, which is Lean and these theorem provers. What are those in this context, and how are they used?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1531">25:31</a>] Lean is an instance of the so-called provers, or proof assistants. So originally, those were developed to do mathematics on the computer. So nothing to do with formal verification of programs, but they can also be used for formal verification of programs. But the initial motivation was really to do mathematics with the help of the computer. Well, originally everyone focused on automatic theorem proving.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1564">26:04</a>] So having the machine find proofs all by itself, but that&#8217;s very, very difficult and not necessarily what mathematicians need. And so those proof assistants are more like formal languages, like programming languages, but where you can write mathematical definitions, mathematical statements, and they will help you prove. So some small proof steps they will do automatically. And for others, the user still has to guide the provers through the major steps.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1603">26:43</a>] But the good thing is that the proof is recorded also in a format that the machine can understand and recheck. When you complete a proof in Lean or Isabelle or one of those tools, and it is rechecked, the computer makes sure that all inferences are justified, that you didn&#8217;t forget any case, that you didn&#8217;t use a conclusion as a hypothesis, all kinds of problems you can have with pencil-and-paper proof.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1633">27:13</a>] And so in the end you get proofs that are extremely reliable, extremely credible. And those tools have been used a little bit for some big mathematical results where the proofs are so big that they can&#8217;t be done just by humans. You really need computer assistance. I think there&#8217;s a recent example with the weak Goldbach conjecture. So part of the proofs involved checking a whole lot of inequality, maybe a thousand inequalities involving several real variables, and blah, blah, blah.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1675">27:55</a>] And so computer assistance was really needed to make sure that everything was checked and there were no human mistakes left. And maybe we&#8217;ll talk about generative AI later. But there&#8217;s also now a lot of interest in conjunction with generative AI. AI is pretty good at coming up with plausible proofs, but then they have to be checked by humans, unless AI writes a proof in one of those formal languages like Lean, in which case it can be checked by a machine, and that&#8217;s much more effective.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1718">28:38</a>] So, yeah, so it&#8217;s the idea of using the machine to help you write proofs and recheck proofs that you&#8217;ve written yourself, or maybe with some help from an AI. And now those tools can also be used to prove properties about programs. And so I spent many years proving the correctness of a compiler, C compiler, using one of those tools, the Coq proof assistant. So basically, what I&#8217;m hoping about the compiler, so it takes C code, it produces assembly code, and proving that the assembly code is faithful to the C code.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1761">29:21</a>] So there&#8217;s no miscompilation. The compiler didn&#8217;t introduce a bug in the program that wasn&#8217;t there originally. And it&#8217;s a fairly big proof because compilers do complicated things about programs. And then you have to define exactly what it means to preserve the semantics of a program. So you have to define the semantics of your programming languages. And so it&#8217;s all a very good use of those machines, of those proof assistants.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1789">29:49</a>] The same work could be done on paper, but this would be a proof of several thousand pages, and nobody would want to read it, nobody would trust it. It&#8217;s just too big, and it would be very hard to evolve. One good thing about those mechanized verifications of programs is that it can also help you evolve the program. You can add new features and so on and adapt your proofs and be really sure that you haven&#8217;t introduced a regression or things like that.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1823">30:23</a>] So, yeah, so it&#8217;s really programming taken to a higher level today. It&#8217;s quite expensive. It takes a lot more time to prove a program than to write it in the first place. But maybe this is getting a little better. I should also mention another great example of verified software. It&#8217;s a microkernel, the seL4 microkernel developed in Australia, which is used as an hypervisor in some applications. And so it&#8217;s really like 8,000 lines of extremely technical C code that manipulates processes and capabilities and security tokens and so on.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1874">31:14</a>] And it&#8217;s been proved correct. Every line has been proved correct. And that&#8217;s a very big achievement.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1880">31:20</a>] When you mentioned the example of verifying a mathematical proof, that makes sense to me. When you talk about proving something about a program, it&#8217;s a little bit more abstract to me, or I&#8217;m having some trouble visualizing. Could you give a concrete example, maybe some trivial program, and we&#8217;re trying to prove something about it and how that theorem prover would work?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1906">31:46</a>] Let&#8217;s say you have a function that takes three numbers X, Y, Z and returns the average of those numbers. Okay? And maybe you&#8217;re using a not completely obvious formula, like T equals X plus Y, and then T equals T plus Z, and then T equals T divided by three, and then return T. Okay? So not exactly the formula for the average. And you want to prove that this function is correct, so you want to prove that it returns X plus Y plus Z divided by three.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1953">32:33</a>] And maybe you will have to make it clear whether you&#8217;re rounding up or rounding down. If you&#8217;re using integers or floating-point numbers, it&#8217;s not exact arithmetic. So yeah, you&#8217;ll have to say exactly what you mean by divided by three. And then if you&#8217;re in a language like <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, you can have arithmetic overflows. When you compute X, Y, or Z, you can overflow the range of representable integers. And in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, it&#8217;s a bug; it&#8217;s undefined behavior.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=1987">33:07</a>] So typically you will want to put a so-called precondition on your function saying, okay, if you call me, call me with numbers that are between, I don&#8217;t know, 0 and 1 million, for instance, but no bigger than that. And then the prover will check that no overflow can occur in this case. Okay? And so basically you have the precondition that says, okay, these are safety warranties that must hold of the parameters; otherwise anything can happen.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2020">33:40</a>] And then there will be a little bit of symbolic execution of the function body that says, okay, when you do t equals x plus y, then t plus equals z, then t slash equals 3, then in the end t is x plus y plus z divided by 3, provided no overflow occurred. And then you put that with a precondition that says there cannot be any overflow, and you get your final result. Okay, so basically you&#8217;re stating a contract for your function, hypothesis on the arguments, guarantees on the results, and you want to prove.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2058">34:18</a>] Analyze the function body to show that this contract is respected. I hope it&#8217;s a little more concrete.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2065">34:25</a>] So it sounds like Lean will help you basically take in some invariants about a program and kind of propagate them through line by line and uphold them. And so you can say something about, yes, maybe not Lean by itself.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2082">34:42</a>] Well, a program prover will do exactly what you say. And for Lean to be able to do it, you still need to teach it a little bit about the semantics of your programming language. Okay, what does plus mean? What does assignment mean? Okay, Lean is mathematics. You don&#8217;t assign. In mathematics, you don&#8217;t say X equals X plus one, or you&#8217;re just comparing X and X plus Y, and it&#8217;s always false.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2111">35:11</a>] There&#8217;s no assignment in mathematics. There&#8217;s assignment in many programs. So you need to make it explicit that there are actually three different states for the T variable, three different values, and relate those values to, well, the program defines what those values are, and then a prover like Lean can reason about those three successive values. Because now you&#8217;re in mathematics, and that kind of bridge between programs and their mathematical meaning is called semantics.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2148">35:48</a>] That the field of semantics of languages, it&#8217;s been a big topic in PL research since the &#8216;60s at least.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2156">35:56</a>] Am I understanding then that a pure functional programming language, the gap between mathematics and the actual symbols is much smaller than an imperative?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2168">36:08</a>] Absolutely. You&#8217;re absolutely right. And that&#8217;s one of the reasons why people who do formal methods don&#8217;t like assignment, don&#8217;t like imperative features. Purely functional style is much, much closer to mathematical style. So much easier to reason about. There can still be a few discrepancies between the program and the math. For instance, functional programs may not terminate, they may loop forever.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2196">36:36</a>] Mathematics doesn&#8217;t like that. So you need a way to reason about termination. And also sometimes, for instance, the arithmetic you get in the programming language is not integer arithmetic, or it&#8217;s not real numbers, it&#8217;s floating point. So you still have to account for that gap. But you&#8217;re absolutely right that the gap is much shorter for functional programming. And in the experience of CompCert, my verified compiler, I think the first decision was to write it in a purely functional style so that it would be easier to reason about it later. You mentioned the specifications that we could prove, and one of them was proving that the program terminates.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2242">37:22</a>] But I thought that&#8217;s a famously difficult or impossible. I forgot exactly the halting problem. Right, so how is that something that you could prove?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2258">37:38</a>] Okay, so what computability theory says is that there is no algorithm that can always say this problem terminates or this program doesn&#8217;t terminate. So there will always be some very weird programs for which your analyzer, your automatic termination analyzer, will produce a wrong result or will not terminate itself. So it will not work. But still, for many programs, you can succeed. You can write automatic termination analyzers that will work for a large class of programs.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2301">38:21</a>] And then the termination analysis, termination proof can also be done by hand by a mathematician. Okay, so perhaps a human can see through those weird Turing programs that are hard to prove terminate and recognize the trick. But anyway, yeah, it&#8217;s a hard problem. And program verification tasks are difficult problems. They are pretty much all undecidable. So there is no static analyzer that will always find all problems in all programs.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2343">39:03</a>] But still you can try, okay, you can try for specific problems or specific programs and get some very useful results out of that. You only need to be able to do it for the cases that are of interest to you, the programs you really care about.</p><h3>40:07 &#8212; What made multicore support difficult for OCaml</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2407">40:07</a>] One thing I saw when I was researching <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is that in 2022, <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> added multicore support. But from my memory, I remember multicore processors became standard much earlier than that.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2420">40:20</a>] And so I figured there might be some unique engineering challenge in adding that support. So, yeah, what happened there, and what made it difficult? Well, there were engineering challenges, that&#8217;s for sure.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2433">40:33</a>] There were also some language design issues. But let&#8217;s talk about the engineering challenges first. So it&#8217;s true that when you have a language with a runtime system, memory allocator, or garbage collector, at least the one for <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> was really designed with sequential executions in mind. So if you had shared memory concurrency, then you need a garbage collector and a memory allocator that can work concurrently.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2464">41:04</a>] And that&#8217;s actually quite difficult, at least if you want them to be fast. Of course, you could always take a lock at every operation in the heap, but then you would sequentialize your programs. Basically, they would run as slowly as on a single processor. So that&#8217;s not interesting. So, yeah, there were engineering challenges. And yeah, I was at least initially quite reluctant to basically reimplement a complete garbage collector and memory allocator on large parts of the runtime system.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2500">41:40</a>] That&#8217;s what we did eventually. Well, it was done mostly by the team at <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> Labs at Cambridge University. But yes, it was a really big rewrite. But then I said there&#8217;s also a language design issue, which is that for the longest time, well, now when you say multicore processors and so on, pretty much everyone, and language support for multicore processors, everyone thinks about language support for shared memory concurrency.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2539">42:19</a>] This model where you have several threads of control and they access the same memory, and basically the threads communicate by modifying the memory and someone else is going to notice. I&#8217;ve never liked this model of communication. I mean, it&#8217;s a bit like if you want to communicate with your neighbor, then you break into their house and move the furniture around, and then when they are back they say, &#8220;Oh, something was moved.&#8221;</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2571">42:51</a>] So it&#8217;s probably Ryan who&#8217;s trying to tell us something. Maybe you could just meet your neighbors, and that would be things like message passing, which is a completely different form of communication, much higher level. So for the longest time I was really interested in message passing concurrency. And there&#8217;s, for instance, a functional language called Erlang that was built on those ideas that I found quite interesting.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2602">43:22</a>] However, I never got to have a decent language design, and everyone was saying, no, but it&#8217;s too costly, it&#8217;s not effective, there&#8217;s too much copying of data. With shared memory, you can share some kind of huge database in memory. If you don&#8217;t modify it too much, then basically the concurrency is free. While with message passing, we&#8217;ll have to exchange a lot of data, so to get the same effect. So you will pay more for it anyway.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2636">43:56</a>] Okay, so starting with <a href="https://en.wikipedia.org/wiki/Java">Java</a>, I guess, and then C++11, there was this idea that, okay, we need to expose shared memory concurrency to the programmers. But then comes the problem of the memory model, which is that what happens when there&#8217;s a race, for instance, when two sides want to access the same location and maybe they want to modify it in different ways. And so sometimes parts of the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> standard says it&#8217;s undefined behavior, anything can happen, but still you need to give a little more guarantees.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2674">44:34</a>] At the other end of the spectrum there are so-called sequential consistency, which says, well, what happens is an interleaving of reads and writes of your program. You don&#8217;t know which interleaving, but there is an interleaving. But that doesn&#8217;t work with modern processors; multicore processors reorder memory accesses in very clever ways to get more performance. And so, viewed from the program, it&#8217;s very hard to predict what they are actually doing.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2701">45:01</a>] And so you need to give your programmers, when you&#8217;re designing a language with shared-memory concurrency, you need to give your programmers some guarantees about the ordering of reads and writes, concurrent reads and writes, while not constraining the hardware too much. And it&#8217;s very difficult. So <a href="https://en.wikipedia.org/wiki/Java">Java</a> went through five different iterations of the memory model. Some were too strict, some were too lax, some were inconsistent.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2728">45:28</a>] We&#8217;re making predictions, impossible predictions, where the past depends on the future. Crazy, really crazy stuff. Then C++11 did a little better, but still an extremely complex memory model. And so when we wanted to add, when the, especially the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> Labs people wanted to add shared memory concurrency to <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, we had to also agree on a memory model that would be exposed to the program rules. And that also took quite a bit of design, quite a bit of time to come up with a good design.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2763">46:03</a>] And I think it&#8217;s better and easier to understand and use than the one in <a href="https://en.wikipedia.org/wiki/Java">Java</a>. But the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> memory model is still quite complicated, and sometimes I feel sorry our users are exposed to that. Okay, so that explains why it took so long, solving the engineering challenges but also agreeing on a memory model and what kind of guarantees we are going to give to programmers. Oh yeah, I forgot to say one thing, which is that in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> and <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> there&#8217;s a lot of things you can say, oh, it&#8217;s just undefined behavior, or anything can happen.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2806">46:46</a>] In type-safe languages like <a href="https://en.wikipedia.org/wiki/Java">Java</a> and <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, you want to give stronger guarantees. Maybe many things can happen, but your data should remain well typed. Typically, you don&#8217;t want to expose an object to another thread before it&#8217;s been fully initialized, for instance, and that&#8217;s actually very hard to guarantee in your memory model. And then you have to implement that memory model. So your compiler also needs to take extra precautions to guarantee this.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2839">47:19</a>] So yeah, type safety in the presence of shared memory concurrency is not obvious at all. So that&#8217;s why it took so long.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2850">47:30</a>] So in Python, I know this is the famous GIL, or the global interpreter lock. What is that lock protecting? Is it the cleanup of objects on the heap, or is it something else?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2861">47:41</a>] Among other things, yes. So yeah, we had the same thing in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> before multicore <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> was merged in. So yeah, basically the idea that when your runtime system is not site safe, as we said, you can protect the non site safe functions by a lock so that they will never be executed concurrently. But if you take the lock at every allocation and release it, you take and release the lock for every allocation, that is just too slow anyway.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2892">48:12</a>] And so the idea is that you take the lock when you enter Python code, let&#8217;s say, and you start executing Python code, but you can still release it when you do input output, for instance, when you&#8217;re going to block for a long time, or when you&#8217;re calling into C code that is thread safe and is not going to use your runtime system, then you can release the lock and some other Python thread can take it and execute.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2922">48:42</a>] So you get a little bit of concurrency. You can overlap computations in your high-level language with I/O or computations in a low-level language in another language, but you still have mutual exclusion between two threads running Python or running <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> before multicore <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. So you get some benefits like concurrent I/O, but you don&#8217;t get any parallelism. Okay, you don&#8217;t get a speedup for computations.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2954">49:14</a>] And so, yeah, so we used to have this GIL in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> as well, and so you can get rid of it, but in general you need to redesign at least a garbage collector and memory allocator. There&#8217;s probably a few places in the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> runtime system that still use locks to ensure mutual exclusion, like in the I/O subsystem and some phases of the garbage collector. I think there&#8217;s a phase which is kind of stop the world.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=2990">49:50</a>] You need to make sure that everyone, no <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code is running for a short time. You need to make sure that everyone is stopped and then do a little bit of work to finish the GC, and then you can restart everyone. So, yeah, these are tricky things, and I don&#8217;t know what the Python people are up to with their GIL, if they finally manage to remove it or they&#8217;re still working on it, but I&#8217;ve heard they are making progress.</p><h3>50:17 &#8212; How programming languages interface and call each other</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3017">50:17</a>] Yeah, you mentioned Python calling into C, and I&#8217;ve seen that pattern before, of a higher-level language interfacing with a lower-level one. How does that binding typically work?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3031">50:31</a>] Oh, it&#8217;s another kind of worm. Well, there are two aspects there: the control flow and the data. So the control part is not that hard. So yes, you need a mechanism so that your Python interpreter or your <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> compiled code will actually jump to the C function. So for <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, basically you tell the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> compiler that this function is not implemented in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>; it&#8217;s actually implemented by a C function, and you give its name, and then the compiler will emit a call to the C function using the C calling conventions, which are not exactly the same as the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> calling convention. But the compiler knows about that, and so it will call the C function, maybe through a little bit of glue code, whatever, and then the C function will execute and return back to the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3090">51:30</a>] That&#8217;s relatively easy. Now the hard part is data like function arguments and function results, because <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> and C have different data representations. For instance, a floating-point number in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is generally boxed, so it&#8217;s allocated in the heap and handled through a pointer. So it&#8217;s more like a double star in C. It&#8217;s not a double, which is not allocated, just sits in a register. So typically the C code needs to use a so-called foreign function interface.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3128">52:08</a>] So some C functions and macros are provided by <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> to access the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> data. The <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> arguments extract the part that it needs, the numbers. Another example is arrays. In <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, when you have a two-dimensional array, it&#8217;s actually an array of arrays. So viewed from C, it&#8217;s an array of pointers to arrays, while in C, an array of arrays has no intermediate pointers, so it&#8217;s not the same representation. And so you have to explain to C or give C some functions and macros to access elements in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> arrays. It&#8217;s not exactly the same code that you would do to access a C array from C.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3172">52:52</a>] Okay, so you need accessors. But now if your C code wants to return some complex results, like a list, an array, and so on, it needs to allocate it in the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> heap. So it needs to ask the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> runtime system to do some heap allocation and then fill the heap blocks correctly, and then this allocation can trigger a garbage collection. So the C code might also kind of cooperate with garbage collection with root registration mechanisms, et cetera, et cetera.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3206">53:26</a>] So there&#8217;s quite a bit of work to be done, and then different foreign function interfaces arrange this work differently. So there&#8217;s a base FFI for <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. Basically, it&#8217;s a C code that must do all the work, but then it can be very fast and quite optimized. But there are other FFIs like the ctypes FFI in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, where most of this data conversion and mediating between two data formats is automated. You start basically with a description of the C type of the C function, and you can automate some of those conversions.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3245">54:05</a>] But sometimes it can be expensive. For instance, you may end up copying a whole array while your C code only needs to access two or three elements in it. Okay, so there&#8217;s lots of trade-offs and, quite frankly, it&#8217;s a dirty part of programming language implementation. The <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> FFI is not that clean, but if you get the <a href="https://en.wikipedia.org/wiki/Java">Java</a> FFI, for instance, it&#8217;s also quite complicated. And for Python I never tried, so I don&#8217;t know what it looks like, but yeah, it&#8217;s a necessity.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3282">54:42</a>] But it can be quite hard because the data models are different between the two languages.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3289">54:49</a>] So to kind of get the big picture, on the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> side, there&#8217;s an interpreter, which is a program running in application space that is interpreting your <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code. And then at some point in the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code it says, &#8220;Do some, load this <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> program.&#8221; And the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> program is a binary somewhere, and the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> interpreter then starts to load those instructions and execute them.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3318">55:18</a>] Okay, so actually this is the third aspect that I didn&#8217;t touch. There&#8217;s an interpreter mode, but in general we compile, we compile to assembly code and then machine code. And so in compiled mode what you say is how you put together the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code and the C code is done by the linker, the C linker. So basically someone else is doing that for us. And it&#8217;s not that different from linking together two object files produced by C or two object files produced by <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, but for more interpreted languages there&#8217;s also the question of how you load the C code.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3367">56:07</a>] Generally, you use a dynamic loading interface, dlopen, for instance in <a href="https://en.wikipedia.org/wiki/Unix">Unix</a>. And then there&#8217;s a little bit of introspection. So at runtime, the interpreter queries C libraries: where is the address of a function named foo, and then it will find the address and use that to manufacture a call. So yeah, if you&#8217;re in an interpreted setting or bytecode-compiled setting like Python, there&#8217;s this additional level of complexity on top of it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3404">56:44</a>] I see, I see. Okay, so if I had a mixed <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> program on my computer, it&#8217;s just one binary or one blob in the simplest case.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3417">56:57</a>] Yes, there are also dynamic loading facilities in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, but I don&#8217;t want to get into that because I&#8217;m a firm believer in static linking. I think programs should be statically linked so that there&#8217;s no surprise when you run them, like, oh, where is this DLL or DLL not found, for instance, problem? But of course you lose a little bit in flexibility. But yeah, I think static linking has a lot to it. It&#8217;s actually quite useful in that it warrants a lot of things. It checks a lot of things at link time that you don&#8217;t have to check again at runtime.</p><h3>57:41 &#8212; The danger of almost-correct LLM code</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3461">57:41</a>] We mentioned earlier in the conversation talking a little bit about <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a>-generated code. I thought that might be interesting to cover. And in one interview you talked about the danger of almost-correct code, so plausible code, but it&#8217;s wrong, that an <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a> can produce. What are your thoughts on how to address that kind of problem?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3485">58:05</a>] It&#8217;s a tough problem globally. I&#8217;m a little bit skeptical about generative AI. Of course, they can do amazing things that were unthinkable a few years ago, but there&#8217;s always some errors. You can&#8217;t really trust what&#8217;s being produced by generative AI. In principle, humans should be there to check the output and fix errors, or ask the <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a> to fix its own errors until the result is actually usable.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3532">58:52</a>] But of course it&#8217;s very hard because there&#8217;s a slop problem. AI produce so much, it&#8217;s so easy to produce large quantities of text, of code, pictures, whatever, that in the end there&#8217;s no human time to check it all. So yeah, recently I heard someone working in an AI startup who was enthusiastic about AI-generated code saying that thanks to generative AI, the cost of programming is dropping to zero.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3575">59:35</a>] Well, the cost of writing code maybe, but what about checking it, making sure it is correct, that it does what we want? Well, that cost is not zero at all. For me, every new line of code is a liability. You have to test it, you have to check it, maybe you have to do formal verification, you have to evolve it, maintain it later. No, I don&#8217;t want huge amounts of code, okay? I want 50 lines of code that have been thought through, that have been published over the years.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3614">01:00:14</a>] So anyway, I&#8217;m not getting that with AI, and I think that&#8217;s a problem. And this idea that humans will be there to check the output of generative AI is just wrong. No, they are not available for that. There&#8217;s too much of it, and it&#8217;s not pleasant either. Okay. I mean, I don&#8217;t think it&#8217;s a good way to split the world between machines and humans. So anyway, so maybe there will be some societal solution like a big no to AI slop movement.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3653">01:00:53</a>] So we&#8217;re trying to see that in some open source projects that refuse generated contributions because there&#8217;s just too many. I know that, well, for <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> and especially for CompCert have received issues that were obviously generated by AI. And there was maybe one good issue among 10 reports, and each report was several pages long, with very detailed explanations and repro case that in the end doesn&#8217;t repro anything or reproduces something else.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3691">01:01:31</a>] But it takes time to go through all those things. And maybe at some point I will say no to AI-generated contributions. Okay, but maybe there&#8217;s also a bit of a technical solution, which is, as we said earlier, to have AI produce proofs, so evidence that its creation is correct. So, as I said, this is starting to work for mathematical proofs. Some generative AIs are able to produce proofs both in English and in the formal language of the Lean prover, for instance.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3728">01:02:08</a>] And so you can get Lean to recheck the proof and get some confidence. You still need to be very careful about the statement because sometimes AIs will change the statement or the definitions to make the proof easier. That happens. And also you should be careful about so-called self formalizations, where the AI also comes up with some definitions and some statements by parsing a PDF file or whatever.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3765">01:02:45</a>] And sometimes it introduces errors at that point. So anyway, there&#8217;s still a need for human review, but on smaller quantities of text and mathematical text. And maybe one day it will also work for program proof. So when an AI generates a program, it might be able to generate some Lean proof or whatever that the program satisfies some specification. So I think it is possible. But now the question will be, where does the specification come from?</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3800">01:03:20</a>] For formal methods, that&#8217;s always been a big issue. It&#8217;s not just that verification is hard, but agreeing on the spec can be difficult too. And where mathematicians have a lot of experience in stating, finding definitions that they find interesting and stating theorems that they believe should be true or that will mean something, computer programmers are less good with that, and it&#8217;s fairly easy to come up with specifications that are inconsistent or impossible.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3842">01:04:02</a>] So remember this average function where there was a precondition of the three numbers, three arguments saying they must not be too big, but say maybe you can end up with a precondition that just cannot be satisfied. And at this point the body of the function will always be verified, even if it&#8217;s completely wrong, because the assumption says basically this function cannot be called. And so you get a false sense of confidence: you verify something, but it&#8217;s actually unusable.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3883">01:04:43</a>] And that&#8217;s a fairly delicate point where I&#8217;m not sure LLMs are going to help much. But it&#8217;s a problem with formal methods in general, and some possibilities include the ability to test specifications, for instance. So instead of using your test suite to see if your code works out, you can also use it to see if your spec checks out, those kinds of things. But yeah, we are kind of moving some of the difficulties from the programming phase to the specification phase, but we still have some problems anyway, so that might be a way to deal with the AI-generated code and develop some confidence in it.</p><h3>01:05:39 &#8212; How LLMs will change programming languages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3939">01:05:39</a>] <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a>-generated code is becoming extremely popular, and I think there&#8217;s a lot of potential downstream consequences on the programming language landscape. And I thought it might be interesting to hear your thoughts on speculating. Imagine 10 years from now, if you took <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a>-generated code and turned it up, how might you think that the programming language landscape might change?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=3968">01:06:08</a>] Yeah, a couple of years ago someone asked me about that, and there was a concern that the training data would be. There wouldn&#8217;t be enough training data in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> for an <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a> to really learn how to program in <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>. And it&#8217;s true that there&#8217;s less <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code in the wild than <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> code, for instance, but apparently contemporary LLMs do well with the amount of <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code they have. Maybe because there&#8217;s enough, maybe because learning has become a little more efficient, maybe because, well, there&#8217;s less <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code than <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> code, but maybe the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code is better quality on an average, I don&#8217;t know.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4020">01:07:00</a>] Anyway, maybe LLMs are getting better also at transferring knowledge that they&#8217;ve learned from one language to another. That could be. I have no idea how those things work. But yeah. So the latest feedback I&#8217;ve got about LLMs and generative AI and <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> is that the <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a> code generated is quite decent and quite similar in quality to more popular languages. And one thing that seems to help the generative AI is a type system.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4055">01:07:35</a>] So the fact that there&#8217;s some static checking just of the type, it&#8217;s already effective in avoiding some errors and maybe encouraging the <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a> to declare types first. So give some type structure to the program. All right. And now from 10 years from now, it&#8217;s difficult to guess. So will it be the more popular languages of today that will be even more popular because of generative AI? Will it be the safer languages of today that will be more popular with AI because.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4102">01:08:22</a>] Well, because there&#8217;s fewer errors. In the end, I&#8217;m hoping it will be the safer languages, but I really don&#8217;t know.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4111">01:08:31</a>] I&#8217;ve seen in the industry there&#8217;s a few cases where because it&#8217;s so easy to generate code, massive rewrites are something that would have been very infeasible in the past, but are very realizable now. So if there&#8217;s an existing project where they chose a programming language for whatever reason in the past, they could translate the entire thing into another language with reasonable confidence. Given that kind of environment, which programming languages would you expect more people would switch to because they want to switch to it, but they weren&#8217;t able to in the past, but now it&#8217;s cheaper so they can?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4153">01:09:13</a>] Well, so today it seems. Well, I&#8217;ve heard mostly about C and <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> to <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> translations, hoping that the generated <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> code will be safer. He said, I&#8217;ve seen at least one project when the generated <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> code is entirely in unsafe blocks. So it&#8217;s really kind of line-by-line translation of the C code, but maybe it can still be used as a starting point for making it safer later.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4186">01:09:46</a>] So, yeah, I would say today I can imagine significant efforts being done with <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> as a target language in the functional world. I&#8217;m not quite sure. Well, maybe a functional language to one of those proof assistants like Lean or Coq, because they also have programming languages, functional programming languages in them, much more restricted, but much more amenable to proof. So maybe for a few projects that could be interesting again as a first step towards a formal proof.</p><h3>01:10:26 &#8212; Industry vs academia</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4226">01:10:26</a>] As we said earlier, when you compare industry versus academia today, where would you say most of the innovation in programming languages comes from? And also, has that changed over time?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4241">01:10:41</a>] Well, I think most of the innovations have come from industry lately. And that wasn&#8217;t the case in the early days of computer science if you think of the truly innovative languages like ALGOL, Lisp, Prolog, Smalltalk were developed in mostly academic settings. Smalltalk was Xerox PARC, which was an industrial research lab but very, very far away from industrial customers. And then there were much more practical languages and uglier languages developed typically at <a href="https://en.wikipedia.org/wiki/IBM">IBM</a> like Fortran, <a href="https://en.wikipedia.org/wiki/COBOL">COBOL</a>, PL/I, et cetera.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4292">01:11:32</a>] And then, so really the idea that the nice ideas come from academia, and then, well, the nice ideas from academia started to be transferred by industry much more quickly. I&#8217;m thinking of, well, C, and especially <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and then <a href="https://en.wikipedia.org/wiki/Java">Java</a>, which really took ideas like object orientation and garbage collection, automatic memory management. In industry before <a href="https://en.wikipedia.org/wiki/Java">Java</a>, it was just crazy academics in the ivory tower that were using garbage-collected languages.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4335">01:12:15</a>] It was completely impossible to have that in enterprise computing. And <a href="https://en.wikipedia.org/wiki/Java">Java</a> came, and two years later everyone was doing garbage collection and being very happy about it. And there were, well, <a href="https://en.wikipedia.org/wiki/Java">Java</a> also popularized type safety, bytecode verification, some pretty advanced techniques of the &#8216;90s. And if you look at further developments, I don&#8217;t know, Swift for instance, popularized the idea of algebraic data types and pattern matching.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4366">01:12:46</a>] And then <a href="https://en.wikipedia.org/wiki/Rust">Rust</a>, well, <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> is even more spectacular, I would say, because algebraic data types, garbage collection, et cetera, those were already present in academic languages like Caml, for instance. But <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> really took very recent research results of the 2000s on safe low-level programming that were basically never implemented in any language and managed to do a consistent whole from that.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4397">01:13:17</a>] And so I&#8217;m really admirative, and I have the impression that, well, I would have loved if <a href="https://en.wikipedia.org/wiki/Rust">Rust</a> came out of academia, but I&#8217;m not sure it would have been possible because it&#8217;s also a huge effort, and you really need the backing of a big company. But still, I&#8217;m not sad because I think it&#8217;s also a very good sign that industry is interested in new programming languages. At some point in the &#8216;90s, well, pretty much when I was hired at Inria on a research position, someone who had developed first versions of <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, well, predecessor of <a href="https://en.wikipedia.org/wiki/OCaml">OCaml</a>, and who was working on type systems for programming languages and so on.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4448">01:14:08</a>] But then some people told me, &#8220;There&#8217;s no future in programming language research. Industry has decided it will be C forever, so deal with it. Maybe you could do software engineering and not PL research.&#8221; Then <a href="https://en.wikipedia.org/wiki/Java">Java</a> came a few years later, showing that no, industry hasn&#8217;t decided on a particular programming language. Industry is still interested in new programming languages. Industry still thinks that a new programming language can be part of the solution to software problems, and I find it extremely encouraging.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4487">01:14:47</a>] We are not stuck with bad languages from the past where there&#8217;s a lot of legacy code, of course, but there is still real interest in better languages. And I think this will continue, and I think it&#8217;s good for the computing field.</p><h3>01:15:05 &#8212; Most interesting unsolved problems</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4505">01:15:05</a>] In 2018, there was an interview that you did, and they asked you, what are the most interesting and important problems to focus on in the coming years? And you called out the difficulty of programming GPUs because they were using dialects of C, shoddy tools, and also the challenges in verifying machine-learned code, basically. And that sounds pretty relevant today, but what would you say your answer is today?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4535">01:15:35</a>] So, yeah, I think, well, for the question of how we program massively parallel hardware, I think we&#8217;ve made a little bit of progress recently with things like the MLIR initiative around LLVM or some domain-specific languages like Halide, which are pretty good. Those tensor domain-specific languages for tensor computations are getting a little better, but still I find it a little bit frustrating that I cannot do, I don&#8217;t know, theorem proving on a GPU.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4570">01:16:10</a>] Absolutely no idea how to go about that, and in part by lack of an appropriate language. Okay, so I&#8217;m not ready to program the GPU pipelines at a very low level myself. And I think it&#8217;s more general. I think we are not using all these GPU and highly parallel hardware as much as we could. Yeah, so verifying applications that have been learned or generated by AI or so on is still a pretty hot issue.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4610">01:16:50</a>] Back in 2018, I was more thinking of verifying simple neural networks like those used for, I don&#8217;t know, computer vision, self-driving cars, or some numerical computations like weather prediction and so on. So specialized LLMs, but you still want some guarantees about what they produce, that they cannot produce completely inconsistent outputs. For instance, there were some attempts in the last years at using static analysis tools and basically program verification tools applied to LLMs, but.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4648">01:17:28</a>] Well, it doesn&#8217;t scale. LLMs are big. Sorry, neural networks are big. And for LLMs, there is also a distinct lack of specification. You don&#8217;t really know what&#8217;s a good answer from an <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a>. Well, you know it when you see it, but you cannot write a mathematical specification of it. So that part is probably over. What I would say are the big problems for today? Yeah, probably maintaining software quality despite AI slope, despite a lot of pressure to throw away traditional software development.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4689">01:18:09</a>] Techniques using LLMs and AI as much as we can to do mechanized proofs, so proofs that can be checked by machines. Maybe this will be the decade of formal verification of software. We&#8217;ve been waiting for that for 50 years. Maybe it will finally take off.</p><h3>01:18:30 &#8212; Top book recommendations for engineers</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4710">01:18:30</a>] What&#8217;s your top book recommendation for software engineers, and why?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4716">01:18:36</a>] This is an old one, Programming Pearls by Bentley. I think I read it when I was a PhD student, but I think it&#8217;s nice as showing how very talented programmers work, how they think about their programs. It&#8217;s a combination of choosing the right algorithms, expressing them clearly, knowing when to stop, when to use a simple algorithm when a more complicated one is not needed, having a sense of elegance in the code you write, having a feeling for where the problem is when the code misbehaves.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4759">01:19:19</a>] So all those kinds of things that are hard to communicate, and I think those pearls are a very easy to read and don&#8217;t use any complicated data structures, don&#8217;t use any complicated language. I mean, it&#8217;s kind of timeless. I think those pearls are a good illustration of that. So if you haven&#8217;t read it, it&#8217;s a classic, but I think it&#8217;s a nice reading, nice read. The second one is a little more controversial, I guess.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4791">01:19:51</a>] So, in <a href="https://en.wikipedia.org/wiki/How_to_Design_Programs">How to Design Programs</a>, which is a fairly ambitious title as well. And this comes from the Scheme community, okay. <a href="https://en.wikipedia.org/wiki/Matthias_Felleisen">Matthias Felleisen</a>, <a href="https://en.wikipedia.org/wiki/Robert_Bruce_Findler">Robert Bruce Findler</a>, <a href="https://en.wikipedia.org/wiki/Matthew_Flatt">Matthew Flatt</a>, and <a href="https://en.wikipedia.org/wiki/Shriram_Krishnamurthi">Shriram Krishnamurthi</a>. And those people have developed. Well, the Scheme community is famous for having developed pedagogical resources that are, I mean, ways to teach programming that go beyond teaching functional programming, basically. And so there were the MIT course <a href="https://en.wikipedia.org/wiki/Structure_and_Interpretation_of_Computer_Programs">Structure and Interpretation of Computer Programs</a>, which was quite famous.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4827">01:20:27</a>] And this is kind of a more modern twist on similar ideas. And I find this book interesting because, well, it really teaches you the way of functional programming or a way to functional programming. It can be very irritating sometimes, very opinionated, very almost mystical sometimes. But it&#8217;s also another great attempt at trying to communicate how experienced programmers go about designing a program even before writing the first line, and then how, given a language like Scheme, which is pretty flexible, how the code kind of follows naturally, knowing what you know.</p><h3>01:21:17 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4877">01:21:17</a>] Now, if you could go back to when you just started your career and give yourself some advice, what would you say?</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4887">01:21:27</a>] Sometimes I got the feeling that I specialized a little too early in programming language research, or maybe. Well, there are some topics that I didn&#8217;t learn because I didn&#8217;t feel like it and that I had to relearn later, or I still have to learn now that I&#8217;m almost 60 and maybe not as quick as I was back in the day. So, yeah, maybe I did specialize a little bit too early. And so I would encourage everyone who&#8217;s serious about working in computing to get a fairly diverse computer science background, even for topics that look super theoretical and not very relevant to everyday jobs.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4948">01:22:28</a>] Where we mentioned computability, for instance, the halting problem and so on, you&#8217;re not going to run into that very often, but it still gives interesting perspectives, I think. And then it also helps understanding new problems, like with quantum computing. What can you do with a quantum computer that you cannot do with a normal computer? And it&#8217;s time to revisit all of the classic basic complexity theory I learned earlier when I was young.</p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=4986">01:23:06</a>] And so, yeah, I think it&#8217;s good to have this kind of background, even if it&#8217;s not obvious you will be using it every day. And sometimes I wish I had taken time to accumulate a little more of this background before specializing in primary languages.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=5005">01:23:25</a>] Awesome. Well, thank you so much for your time, <a href="https://en.wikipedia.org/wiki/Xavier_Leroy">Xavier Leroy</a>. I really appreciate it.</p><p><strong>Xavier:</strong></p><p>[<a href="https://youtu.be/9Cswiqrq6So?t=5008">01:23:28</a>] Thank you, Ryan. That was nice.</p>]]></content:encoded></item><item><title><![CDATA[Turing Award Winner: TPUs vs GPUs vs CPUs, Computer Architecture, RISC vs CISC | David Patterson]]></title><description><![CDATA[In college I briefly studied RISC vs CISC for my major but I forgot almost everything now.]]></description><link>https://www.developing.dev/p/turing-award-winner-tpu-vs-gpu-vs</link><guid isPermaLink="false">https://www.developing.dev/p/turing-award-winner-tpu-vs-gpu-vs</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 13 Jul 2026 13:05:15 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/206770818/662a5457db86986879d6696047c28799.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>In college I briefly studied RISC vs CISC for my major but I forgot almost everything now. So it was great to talk to <a href="https://en.wikipedia.org/wiki/David_Patterson_(computer_scientist)">David Patterson</a> about what happened there since he was one of the people involved in the historic RISC versus CISC debates.</p><p>Given his computer architecture expertise, I also was curious to hear his explanation of the differences about CPUs, GPUs and TPUs so we discussed that too. One of the things I enjoyed most about the conversation was reflecting on what was important looking back on his half a decade of experience. I hope you enjoy this episode.</p><p>By the way, I started to shoot my own episodes since it lets me record in scrappy locations and majorly cuts down on cost. Before when I got help, each onsite shoot would cost $2-3k minimum. Maybe one day I&#8217;ll revisit this and get help if the podcast has more budget. For now though I&#8217;ve been enjoying learning video production and it&#8217;s pretty simple. Lighting was rough on this episode though (lesson learned).</p><p>It&#8217;s also fun bringing in equipment and trying to capture people in their natural environment like when I shot the <a href="https://www.youtube.com/watch?v=U46fJ2bJ-co&amp;t=4378s">Bjarne episode in front of his own library of C++ books</a>.</p><p>Also, it&#8217;s been ~6 months since I quit my job to work on the podcast full time. I&#8217;ve been meaning to write an update. I&#8217;ll put something out in a week or two.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/Pn4ZwlEh5nw">YouTube</a>, <a href="https://open.spotify.com/episode/6mlRwX3aNmRQLiAho2CTXW?si=imjcMUb9S5yBYr8Pr4j2wA">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-Pn4ZwlEh5nw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Pn4ZwlEh5nw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Pn4ZwlEh5nw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/206770818/0042-risc-vs-cisc">00:42 - RISC vs CISC</a></p><p><a href="https://www.developing.dev/i/206770818/1251-compilers">12:51 - Compilers</a></p><p><a href="https://www.developing.dev/i/206770818/1738-gpus">17:38 - GPUs</a></p><p><a href="https://www.developing.dev/i/206770818/2307-gpu-vs-tpu-vs-cpu">23:07 - GPU vs TPU vs CPU</a></p><p><a href="https://www.developing.dev/i/206770818/3212-is-moores-law-dead">32:12 - Is Moores law dead?</a></p><p><a href="https://www.developing.dev/i/206770818/3804-gpu-benchmarks">38:04 - GPU benchmarks</a></p><p><a href="https://www.developing.dev/i/206770818/4140-how-to-have-a-bad-career">41:40 - How to have a bad career</a></p><p><a href="https://www.developing.dev/i/206770818/4959-courage-and-optimism">49:59 - Courage and optimism</a></p><p><a href="https://www.developing.dev/i/206770818/5556-advice-for-his-younger-self">55:56 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:42 &#8212; RISC vs CISC</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=42">00:42</a>] Maybe you could explain <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> versus <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a>, kind of like what that debate was and why it was controversial.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=49">00:49</a>] What happened is, in the 1970s, the microprocessor was invented. But in the beginning, it was basically a toy. It&#8217;d be something that would go in microwaves and things like that. Those of us who believed in <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> believed eventually the microprocessor would be the way we did all our computing, with the doubling of transistors every year or two. Eventually, there&#8217;d be enough transistors that one of these microprocessors would be a serious computer.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=79">01:19</a>] So the people who are designing microprocessors, like <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> and Texas Instruments, weren&#8217;t really computer architects, so they just imitated what the big companies did. So leading companies, like IBM, where at the time <a href="https://en.wikipedia.org/wiki/Digital_Equipment_Corporation">Digital Equipment Corporation</a>, IBM for mainframes, Digital Equipment for what were called minicomputers, they kind of drove the architecture, the instruction set design. So what they were doing with <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> was building more and more sophisticated instructions.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=109">01:49</a>] They had many computers be the size of a refrigerator or a couple of refrigerators, and mainframes would be the size of many of those. That&#8217;s what they were doing with more hardware resources. The philosophy at the time was that by having a more sophisticated instruction set, you raised the <a href="https://en.wikipedia.org/wiki/Abstraction">level of abstraction</a> closer to the software. So that had some inherent benefits over having something at a lower level.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=136">02:16</a>] Now, from some of us who were in that area, like my friend <a href="https://en.wikipedia.org/wiki/John_L._Hennessy">John Hennessy</a> at <a href="https://en.wikipedia.org/wiki/Stanford_University">Stanford</a> and I, we thought that wasn&#8217;t necessarily the right thing to do. Compilers would go from programming languages down to the instruction set. So why couldn&#8217;t compilers hold that? And then the question is, what&#8217;s the right instruction set for this emerging microprocessor? And the prevailing wisdom of these more sophisticated instruction sets, complex instruction sets, and if we think of it like vocabulary, it was like having lots of polysyllabic words in your vocabulary.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=172">02:52</a>] And the alternative was going to be what we call the reduced instruction set computer, was having lots of reduced instructions, like monosyllabic words. So you might guess that for a program to execute these instructions, if they&#8217;re simpler, they take more of them, and if they&#8217;re sophisticated, they take fewer. But the sophisticated ones might take longer to run. So the question came down to what was that ratio?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=198">03:18</a>] And in the beginning, in the 1980s when these debates were happening, there was this kind of very vociferous debates partly about what those ratios are going to do, but a big part was kind of philosophical. Weren&#8217;t you doing damage to the software industry by lowering the instruction set level so the gap between the programming languages and the instruction set was larger? That was kind of the ferocity of the debates.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=223">03:43</a>] Well, after the dust settled a few years later and we started getting numbers, it turned out you needed maybe 30%, 40% more simple instructions to execute in a program. But you could run them 4 or 5 times faster. So the net was 3 or 4x speedup was the potential for <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> over <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=245">04:05</a>] You mentioned that philosophical difference, that gap. Why would that be controversial?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=252">04:12</a>] I mean, unfortunately, I&#8217;d say computer architecture in the 1970s and even 1980s, a lot of it was kind of people designed this by using their intuition or using their gut, their feelings, was this the right thing to do. And even the textbooks of the time were like catalogs. Here is a computer and just list all its features. And here&#8217;s another computer listing all its features. So it&#8217;s pretty dissatisfying what was going on.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=281">04:41</a>] Because these debates, it would seem like we ought to be able to decide this scientifically with numbers that are there. But, absent that, it was more like a philosophical debate, like how many angels on the head of a pin, right? You could have arguments qualitatively, but you couldn&#8217;t settle the arguments. And because there weren&#8217;t ways to settle the arguments quantitatively, people would just argue about it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=306">05:06</a>] Would you say that, you know, <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> versus <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a>, did one side win the war?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=313">05:13</a>] Yeah, I was just reading. There&#8217;s some guy online who revisited it and he said, &#8220;<a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a>? Well, that&#8217;s a pretty myopic view. It&#8217;s in the PC era because of the importance of distributing software in binary.&#8221; So once <a href="https://en.wikipedia.org/wiki/X86">x86</a> was established and PCs would ship software in binaries, that was very hard to overcome. That was a huge impediment to changing the instruction set. So PCs are defined largely by the <a href="https://en.wikipedia.org/wiki/X86">x86</a> architecture.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=346">05:46</a>] But also in the 1980s, there was this company in England that wanted to do a personal computer. They called it the Acorn personal computer. And they decided they wanted to have their own instruction set architecture to do that, their own chip they run satisfactorily. The chips they had available at the time they were doing it weren&#8217;t fast enough. And they were influenced by the papers that we did at <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Berkeley">UC Berkeley</a>.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=369">06:09</a>] And so they built what they called the <a href="https://en.wikipedia.org/wiki/ARM_architecture_family">Acorn RISC Machine</a>. And then one of the benefits of this kind of reduced instruction set is that it could be simpler. It would take less resources and take less energy to execute. And then several years later, when <a href="https://en.wikipedia.org/wiki/Apple">Apple</a> was looking for a microprocessor that could power one of their personal devices, this one was called the early forerunner of the iPhone called <a href="https://en.wikipedia.org/wiki/Apple">Apple</a> Newton.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=395">06:35</a>] They came to this company and said, &#8220;Boy, I really like it. Let&#8217;s get rid of the Acorn name.&#8221; So they renamed it the <a href="https://en.wikipedia.org/wiki/ARM_architecture_family">Acorn RISC Machine</a>, ARM, and they rechristened it the <a href="https://en.wikipedia.org/wiki/ARM_architecture_family">Advanced RISC Machine</a> to get rid of the Acorn. And then <a href="https://en.wikipedia.org/wiki/Apple">Apple</a> used it in the Newton. Now, the Newton wasn&#8217;t a commercial success, but it demonstrated the benefits for <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> architecture for mobile devices. So Nokia came along just a few years later with their GSM cellular phone, which is one of the first popular ones, and they embraced ARM.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=430">07:10</a>] And so ever since, ARM has dominated all the mobile devices. So I think I just checked, there&#8217;s been 350 billion ARM processors, microprocessors with ARM technology in it today. So it&#8217;s like today 99% of all processors and computers are <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a>, even in personal computers. <a href="https://en.wikipedia.org/wiki/Apple">Apple</a> switched over to ARM from the <a href="https://en.wikipedia.org/wiki/X86">x86</a> architecture. So even PCs, there&#8217;s <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> architecture is significant and it&#8217;s starting to get in the cloud.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=467">07:47</a>] The cloud&#8217;s largely been defined by the <a href="https://en.wikipedia.org/wiki/X86">x86</a> server architectures. But Amazon, <a href="https://en.wikipedia.org/wiki/Microsoft">Microsoft</a>, and <a href="https://en.wikipedia.org/wiki/Google">Google</a> all have their own development of ARM processors. So ARM is becoming a <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> processor, becoming more popular in the cloud. So right now I&#8217;d say the <a href="https://en.wikipedia.org/wiki/X86">x86</a> architecture market is shrinking while the <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> architecture market is growing leaps and bounds.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=492">08:12</a>] Well, how could someone say <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> when you mentioned 99%?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=497">08:17</a>] Yeah, well, that&#8217;s just if they&#8217;re doing something in history and they define computers as personal computers and maybe servers, and they go up through around 2000. If the story ended then, you could&#8212;well, it looks like a <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> architecture, but once we get into this post-PC era, I don&#8217;t understand how somebody would reach that conclusion.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=523">08:43</a>] You mentioned the energy expenditure. So <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> makes sense in places, or maybe on mobile devices, things like that.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=532">08:52</a>] Well, even in the cloud, everybody cares about.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=535">08:55</a>] We&#8217;re.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=536">08:56</a>] Everything is energy bound now. So if these days, of course you&#8217;re not dealing with millions of transistors, but billions of transistors. So it matters somewhat less today. There&#8217;s so many things going on that you can hide that. I mean, even what the <a href="https://en.wikipedia.org/wiki/X86">x86</a> architecture did to compete is that it translated the <a href="https://en.wikipedia.org/wiki/X86">x86</a> instructions in hardware into <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> instructions. And so you had to pay that extra overhead of that translation step to get <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> instructions.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=566">09:26</a>] And then you could take any good ideas that the <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> people had that <a href="https://en.wikipedia.org/wiki/X86">x86</a> could do. But it was worth it financially for that extra overhead because of the value of the PC software base. So it made a lot of sense for <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> to do that. And they did that in the early 2000s. It was a great commercial idea.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=589">09:49</a>] Is there any niche use case where <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> makes sense? Is there any engineering trade-off, or is <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> objectively worse?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=597">09:57</a>] Yeah. So if you want to, if we want to go kind of one level deeper into all of this, what was actually going on in designing computers, the hard part is the control. And so what happened is, in the beginning, control was kind of ad hoc. You would figure out, you put the gates together to make it work. One of the computing pioneers, Maurice Wilkes, figured out a more elegant way to design control.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=627">10:27</a>] And he said, well, we could just list all the control signals as the output of a memory, and we could have something that would keep track of where we were in the memory and issue those control signals. And he called this effort the instruction, the control signals. You could think of instructions. He called that a microinstruction, and he called the programming of these instructions <a href="https://en.wikipedia.org/wiki/Maurice_Wilkes">microprogramming</a>.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=654">10:54</a>] Where technology was in the 1960s, that made a fair amount of sense. IBM built these so-called microprogram computers. It was basically an interpreter with very simple instructions that would interpret this much more sophisticated instruction set above it. But you would pay this interpretation overhead. And classically in computer science, interpreting versus compiling is something like a factor of five or ten.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=683">11:23</a>] But given the latencies of the memory technologies and the possibility of doing this out of read-only memory, it made sense up until the &#8216;60s and &#8216;70s. But then the question came up around 1980: Is this still a good idea? Should we have this microcode interpreter inside there? And so the alternative is to think of, well, rather than having this microcode interpreter in there, why don&#8217;t we just compile directly into those instructions?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=712">11:52</a>] And that&#8217;s pretty close to the <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> ideas. The microinstructions themselves used to be like 100 bits wide and really complicated. So if you make them not quite so long, you make them kind of more natural still. We could skip the interpretation step. So, going one level deeper, the question today would be: would people invent an instruction set that was so sophisticated it needed a microcode interpreter?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=741">12:21</a>] And probably they wouldn&#8217;t do that. I mean, you could do it; nothing prevents you from doing it. But you wouldn&#8217;t want to design an instruction set that was forced to use a microcode interpreter. It might make sense in some very tiny applications, possibly where you have tens of thousands of transistors. Maybe this microcoded interpreter would work. But I think nobody today&#8212;I don&#8217;t think anybody&#8217;s invented an instruction set in the last 20 years that has anything that would need a microcode interpreter.</p><h3>12:51 &#8212; Compilers</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=771">12:51</a>] You mentioned the compiler multiple times here, and it seems like that&#8217;s a critical piece that kind of makes <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> work so well. Can you explain the role of the compiler and how it manages the relationship between software and hardware?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=786">13:06</a>] Well, you&#8217;re writing in a programming language like C or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, or maybe Python, but the quality of the code that gets generated is up to the compiler. And a specific example is that it&#8217;s useful when you construct computers to have registers. And those registers are actually kind of visible in the assembly language or the machine language programming. There can be 8, 16, or 32 of these registers for the people programming to use.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=817">13:37</a>] Well, it used to be very difficult for compilers to allocate registers efficiently. Computers just weren&#8217;t fast enough. We didn&#8217;t have the algorithms that we could look at a section of code or a subroutine or something and say, how can we most efficiently do register allocation? In fact, the C programming language, which was invented to do systems programming in C. Before that, to write an operating system, people wrote them in assembly language, believe it or not.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=848">14:08</a>] But the Unix people showed, Ken Thompson, Dennis Ritchie showed that if we had a low level language, a pretty low level language, we get the benefits of writing in something that&#8217;s much easier for humans to understand and debug because register allocation by the compilers was so poor, they had to put in the ability for the programmer to give hints of what variables should go in registers. It was just easier for them to have the programmer step in and say if the machine had eight registers, I want these six variables to be in these registers.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=880">14:40</a>] Don&#8217;t leave them in memory because it runs so much slower; registers are so much faster. So a big part of the <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> argument was the compiler algorithms were getting better. They could handle these low-level instructions; they could allocate registers efficiently. And that was another reason in these debates about why it made sense to have a simpler architecture. What we did in the <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a> architecture is, well, if registers are really important, register allocation is really important.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=910">15:10</a>] One way to make it easier for the compiler was just to have a lot more of them. So typically the <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> architectures at times would have eight or maybe 16 registers. So we put in 32. And that was one of the arguments, right? If it&#8217;s hard to efficiently use a small number, let&#8217;s give them plenty. So even if it wasn&#8217;t that good, there&#8217;ll be enough registers most of the time. And registers weren&#8217;t that much more expensive to include in machines because of <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=938">15:38</a>] I saw somewhere in one of the talks that he had given that somehow the compiler, it&#8217;s easier for it to optimize the code for <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a>. But in <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a>, there were these complex bigger ones, and it almost never used them. Is that a shortcoming of the compiler?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=956">15:56</a>] So what happened, if we go back to the compiler days, there&#8217;s this argument that by having more sophisticated instructions, this would raise the <a href="https://en.wikipedia.org/wiki/Abstraction">level of abstraction</a>. There&#8217;d be a smaller gap, and they&#8217;d make it easy to compile. But that was a philosophical argument. It wasn&#8217;t something that necessarily compiler people could. Compiler people weren&#8217;t making that argument; it was the architects who were making that argument.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=980">16:20</a>] And as it turned out, when we looked at the programs, when we were doing the early research on <a href="https://en.wikipedia.org/wiki/Chief_of_Integrated_Defence_Staff">CISC</a> versus <a href="https://en.wikipedia.org/wiki/Reduced_instruction_set_computer">RISC</a>, the compilers didn&#8217;t really use those instructions. Often the compiler writers, the architects, would come up with a sophisticated instruction, and the compiler writers would say, well, we don&#8217;t need that. In fact, we found examples where, like for a procedure entry, there&#8217;d be a special instruction built in to do all this work.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1004">16:44</a>] That&#8217;s what the architects thought the compilers were going to do, and the compiler designers said, well, we don&#8217;t need that. It&#8217;s faster. It&#8217;s actually faster for you to use separate instructions than to use your sophisticated instruction. And so we found a bunch of examples like that. So you pay the extra overhead of the microcode interpreter to have these sophisticated instructions, and then the compiler doesn&#8217;t even use them.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1025">17:05</a>] Right. So this is kind of a nonsensical situation that we&#8217;re in. And, you know, this is kind of like, you know, why, why, why are there startups or why are there scientific, you know, breaking points like that? As we looked at all the technologies with <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, you know, ideas like caches, with where compilers were using sophisticated instructions, the ability to allocate registers more efficiently, it made sense to change the directions away from these microcoded instruction set architectures to these simple architectures.</p><h3>17:38 &#8212; GPUs</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1058">17:38</a>] When we talk about instruction sets, I mean we&#8217;re kind of assuming that we&#8217;re talking about these general-purpose computers or these CPUs. And I know GPUs are kind of talked about a lot, and maybe there&#8217;s other forms of computing. Are there instruction sets for those types of machines as well?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1075">17:55</a>] So what happened is around 2000 is when GPUs came along, and these were what we call domain-specific architectures. So a <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a> is a graphics processing unit. It had one job. It didn&#8217;t need to do everything that a general-purpose processor needs to, didn&#8217;t have to support virtual memory, it didn&#8217;t have to support a compiler, which is a pretty radical idea for architectures. It was just for graphics. And so around 2000 is when <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> and the other GPUs out there, because they were trying to give you the graphics for games and give you the graphics for movies.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1112">18:32</a>] So it was a niche product. So what happened kind of in the computer industry is it was driven by <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, which I&#8217;ve mentioned many times, and also this lesser known law called <a href="https://en.wikipedia.org/wiki/Dennard_scaling">Dennard scaling</a>, which raises kind of an interesting question. Well, if <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> puts a lot more transistors on a chip, how come they&#8217;re not getting hotter and hotter as we double the number of transistors every year?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1136">18:56</a>] And the reason was this observation made by <a href="https://en.wikipedia.org/wiki/Robert_H._Dennard">Robert Dennard</a>, that as you added more transistors, people would also lower the threshold voltage, which is the distinction between a 0 and 1. And that had a squared effect. So you would end up doubling the transistors, but you&#8217;d lower the threshold voltage. So microprocessors stayed at like 20 or 30 watts. In fact, we did a John and I eventually did a textbook, and I was just checking, and it came out in 1990.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1168">19:28</a>] And we didn&#8217;t even talk about power as an issue for the first three editions. The third one came out in 2000. Power wasn&#8217;t even a topic because <a href="https://en.wikipedia.org/wiki/Dennard_scaling">Dennard scaling</a> was going on. And so microprocessors stayed in the tens of watts even as they got faster and faster. What happened is about 2005 or so, <a href="https://en.wikipedia.org/wiki/Dennard_scaling">Dennard scaling</a> stopped working. And that was a shock. And <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> actually had a microprocessor that failed one of their generations because they just couldn&#8217;t get the power down enough.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1200">20:00</a>] It was too hot to do that. So then what that forced us to go to multicore is that before that, it was easiest for the programmer, for everybody, if there was one very sophisticated processor that did everything, but we couldn&#8217;t do that anymore. So it went from one sophisticated processor to two, and then four, and then eight simpler processors. Then it was up to the programmer to deliver on the potential of <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> by parallelizing their code.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1229">20:29</a>] So we stayed that way for about another 10 years. And then <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> started slowing down so that the general-purpose microprocessor was barely improving. It got a little better, but it wasn&#8217;t getting dramatically better. In the 1980s, 1990s, 2000s, you have a laptop and your friend&#8217;s laptop would be like four times as fast as yours. And you were jealous. So you would throw away perfectly good hardware because your friend&#8217;s thing was so much faster because of the rapid change in performance.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1265">21:05</a>] Well, that all ended in the 2010s, where you wouldn&#8217;t throw away a laptop because the new ones were hardly that much faster than the old ones. You&#8217;d throw them away when they break and slow down. Nevertheless, programmers were used to this dramatic improvement in performance every few years because they could add more features to their software and stuff like that. So what were architects going to do?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1287">21:27</a>] They&#8217;d already done the multicore trick in 2005 or so. So in around 2015, the idea was we would do domain-specific architectures like the GPUs. So if you tell me I only have to run a narrow class of programs and I don&#8217;t have to necessarily run all the operating systems, everything else, well, yes, I could shuffle those resources and do something much more efficient for something. Well, so you could do some things well, and there&#8217;s other things either you do poorly or not even at all.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1319">21:59</a>] So then the question is, okay, what domain? So just coincidentally, where we were technologically, right around 2012, 2015 is when machine learning AI burst on the scene. And so that was, it was evident what domain should we do? And the domain was machine learning AI. So just coincidentally, where we were technologically, this new domain came along. And then I would say I work for <a href="https://en.wikipedia.org/wiki/Google">Google</a>, but I think it&#8217;s fair to say <a href="https://en.wikipedia.org/wiki/Google">Google</a> was the one who really saw it was certainly the first big company who understood the potential of machine learning AI and bet that they or feared that they were going to be swamped with demand and needed to do custom hardware.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1366">22:46</a>] So the tensor processing unit, or <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a>, that <a href="https://en.wikipedia.org/wiki/Google">Google</a> debuted in 2016 really kind of shocked the world and got people to realize that we should be designing hardware for machine learning, and we could continue to improve performance dramatically for that one domain.</p><h3>23:07 &#8212; GPU vs TPU vs CPU</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1387">23:07</a>] What&#8217;s the high-level difference between a CPU, a <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a>, and a <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a>?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1394">23:14</a>] The CPU has got to be this general-purpose thing. And basically even today they have these general-purpose cores that each core is very similar to what the instruction sets used to be like 20 years before that. It&#8217;s not a surprising design, and it&#8217;s just got lots of cores, and the modern ones today could have 50 or 100 cores in there. So that&#8217;s the feature for graphics. What they decided to do to do the graphics job is they need to get a lot of performance out of the memory system.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1426">23:46</a>] So they went to multithreading. They have hardware threads that you would send a memory request out and the hardware would switch over to do something else while memory is coming up. So this highly threaded architecture, and that&#8217;s what they developed, and it could run well for graphics. And graphics, it turns out, doesn&#8217;t need very, doesn&#8217;t need a powerful floating point. It doesn&#8217;t need wide floating points.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1455">24:15</a>] So 32-bit floating point is plenty for graphics, even 16-bit. So they were doing, pushing graphics with 16- and 32-bit floating point in this multithreaded kind of architecture. And because it was kind of its own unique thing, it has its own set of terminology all to itself. It wasn&#8217;t kind of out of the main branch of computer architecture. So when you look at our textbooks, we have kind of like a Rosetta Stone which says here&#8217;s the terms that <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> uses to describe GPUs.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1485">24:45</a>] This is what it means in kind of normal, well, normal in the mainstream processor design. So there was this niche product that happened to do pretty fast single-precision floating point and even half-precision floating point, and they were pretty cheap. They were like hundreds of dollars. So when people, so kind of, but some people were thinking, boy, for some applications, if I could turn my program into, like, pretend that it was doing, creating an image, if I could turn my problem into image generation, I could use these pretty cheap GPUs, which had very good floating-point performance per dollar compared to anything else.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1531">25:31</a>] And so people started playing around with that. And the founder and CEO, <a href="https://en.wikipedia.org/wiki/Jensen_Huang">Jensen Huang</a>, really liked that idea. So in 2006 he funded an effort to create a programming language that would handle this multithreaded hardware architecture for graphics to make it easier to program. And so that led to the invention of <a href="https://en.wikipedia.org/wiki/CUDA">CUDA</a>, which I can&#8217;t remember its acronym, but it&#8217;s really a proprietary programming language for the multithreaded <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a> architecture to make it easier to program.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1567">26:07</a>] It was a C-like language, but you can&#8217;t just compile C programs and run it. But it was C-like, but people actually liked it. It was certainly much better than trying to turn your program into image generation, but people liked it. But <a href="https://en.wikipedia.org/wiki/Jensen_Huang">Jensen Huang</a>, his vision was try to get these teenagers in the basement who are playing the games to learn how to program these things. And he had a couple of markets that he was interested in, like some of the <a href="https://en.wikipedia.org/wiki/Ministry_of_energy">Department of Energy</a> labs, like fluid flow and some of the things, like, some.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1598">26:38</a>] And he would build special-purpose libraries that could handle that domain and then use his GPUs to do that kind of special-purpose computing, but breaking out of graphics. But that&#8217;s kind of the heritage there now for the <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a> program. And so what&#8217;s happened, people started right from the very beginning, people started using GPUs because they had much better single-precision floating-point performance than the CPUs.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1624">27:04</a>] They had a lot more processors on them, so they had potentially much more performance. So one of the breakthrough moments was in 2012, when within the machine learning community, this neural networking piece, which had a few advocates but a lot of people didn&#8217;t believe in, went into a competition to see who could do the best image recognition. In 2012, the so-called <a href="https://en.wikipedia.org/wiki/AlexNet">AlexNet</a> beat all the competition.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1653">27:33</a>] This was this historic moment in machine learning and neural networking. And once they did it, and the guy who did that had taken a <a href="https://en.wikipedia.org/wiki/CUDA">CUDA</a> course at the <a href="https://en.wikipedia.org/wiki/University_of_Toronto">University of Toronto</a> to learn how to use it. And so he says, well, as long as I&#8217;m doing this, I&#8217;ll do it on a <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a>. He was able to explore a lot more space on this cheap, fast <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a>. So he entered it in the competition in 2012, and he was the only one using neural networks.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1680">28:00</a>] And he crushed the competition. And within a few years, everybody switched over, and not only were they doing neural networks, but they were using GPUs. So GPUs have this heritage of being graphics engines, but more programmable, and started getting used in machine learning. When <a href="https://en.wikipedia.org/wiki/Google">Google</a> came along, they decided it was a clean slate. They didn&#8217;t care about graphics. So at the heart of neural networking is a matrix multiply.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1711">28:31</a>] That&#8217;s the thing. So they designed a processor. At the time, the microprocessor had a giant matrix multiplying unit. That was the main thing. And then they threw a bunch of stuff out that they didn&#8217;t need. So a lot of general-purpose computing is what&#8217;s on the chip is maybe three levels of caches to try and have so you don&#8217;t spend all your time going to the relatively slow memory.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1734">28:54</a>] Well, for machine learning they knew when their memory accesses were, so they could schedule that. So a hardware cache didn&#8217;t make any sense. They would just have a memory the software understood would transfer in time. Those are some of the innovations. Also, they innovated on the floating point format you didn&#8217;t need. Scientific computing cares a lot about precision. Most of it&#8217;s done in 64-bit floating point, where the exponent is less than 10 bits and most of it is the precision.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1767">29:27</a>] The fraction can be 50-some bits. Well, what they realized for machine learning and AI, they don&#8217;t need all that precision; they needed the range. So <a href="https://en.wikipedia.org/wiki/Google">Google</a> did the first floating-point format where the exponent was bigger than the fraction. That was a radical idea, the so-called bfloat16. The part of <a href="https://en.wikipedia.org/wiki/Google">Google</a> that was doing this was the <a href="https://en.wikipedia.org/wiki/Google">Google</a> Brain research group. So they brought out this architecture.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1791">29:51</a>] They had this narrow floating point format. They had a big matrix multiply unit. It only had one processor in it. Unlike the other ones, it was just dedicated machine learning. And it just kind of blew the doors off of everybody in the field. It was like 30 times better at inference than the contemporary <a href="https://en.wikipedia.org/wiki/Graphics_processing_unit">GPU</a> and 80 times better than CPU by having this dedicated one. So when <a href="https://en.wikipedia.org/wiki/Google">Google</a> made this announcement at their annual retreat a year or so after they had deployed it inside, it just shook everybody up.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1826">30:26</a>] <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> started buying companies, <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> started modifying the design to be much more focused on machine learning. And then a bunch of other competitors, hyperscalers, started doing their own efforts. So I think that was the watershed. I think the <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a> announcement was the watershed moment there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1848">30:48</a>] So it&#8217;s almost like levels of specialization, like the CPU is the most general.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1853">30:53</a>] Yeah. If we still had <a href="https://en.wikipedia.org/wiki/Dennard_scaling">Dennard scaling</a>, <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> and <a href="https://en.wikipedia.org/wiki/Dennard_scaling">Dennard scaling</a>, where would we be today? With general-purpose processors, we should have 100 TL microprocessors. If we could build 100 TL microprocessors, that&#8217;s what we&#8217;d do. GPUs would still be a niche. If we could do that, it raises all boats. That would be fantastic if we could do that. But that&#8217;s long in the history.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1882">31:22</a>] We haven&#8217;t been able to do that for 20 years. And so there were probably two or three gigahertz microprocessors in 2005, and that&#8217;s kind of where they are, barely improved today. So we can&#8217;t do that. So there&#8217;s still areas where it&#8217;s important to have CPUs. They&#8217;re the right&#8212;you need operating systems, compilers, and things that they&#8217;re the right solution. But things have been specialized.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1907">31:47</a>] And then now, in terms of industry terms, a huge emphasis is being put into the AI accelerators, or just AI in general, is where the money is being invested. And so the question for any processor designers was, how does my processor design fit into this AI universe, and what should I do? How should I change it for this important application?</p><h3>32:12 &#8212; Is Moores law dead?</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1932">32:12</a>] So you mentioned that <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> is slowing down. And I watched this other talk from Jim Keller saying that <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> isn&#8217;t dead. But is it controversial to say that it&#8217;s?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1945">32:25</a>] No. I mean, not, not to real engineers. <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> is very simple. It says in a chip the number of transistors will double. Originally said every year, and then he changed it to every two years. Just look. Do the chips, the number of transistors doubled? No. Now I think what people assume that means is if we are no longer on <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, that technology is not improving.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1972">32:52</a>] That&#8217;s not the same thing. It&#8217;s not improving at the rate that <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> projected. And what was amazing about <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, which lasted 50 years, is that guided the investment of semiconductor manufacturing. We need to deliver on doubling transistors every year or two. How are we going to do that? How are we going to build the equipment to do it? So it was a guideline for the whole industry, which was kind of remarkable.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=1997">33:17</a>] So what we are doing now is there&#8217;s pieces of the technology that get better and there&#8217;s pieces that don&#8217;t improve at all. So one of the big pieces, important pieces on a chip, is the static random-access memory. So that&#8217;s hardly improving at all. Logic gates still continue to improve. The gates are getting better. So if you&#8217;re doing adders or multipliers, those are getting there. So it&#8217;s not uniform improvement anymore.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2026">33:46</a>] And we&#8217;re also going to more exotic packaging to be able to deliver it. So people now, you know, forever. It was a single chip, was the best way to package everything fits on one chip. And that&#8217;s what we&#8217;re going to build. And now there&#8217;s these ideas of chiplets or packaging multiple chips together. The latest GPUs, and I think the latest TPUs, has actually two what are called full reticle design dies packaged, put into a package.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2058">34:18</a>] And so a full reticle design is the maximum you can build on a semiconductor. Well, on a semiconductor they have a step-and-repeat motor, and in the old days you&#8217;d have many chips inside that. Now there&#8217;s just one. So two maximum reticle designs form a node. So they&#8217;re using packaging. So if you look from outside, you can say, well, look how many transistors. <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> is continuing.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2083">34:43</a>] But if you see what&#8217;s inside the chip, that has tapered off. And I think part of it is for a fair amount of the industry, if you were in a semiconductor manufacturer and people ask you what you do, it&#8217;s I make <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, I sustain <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a>, that&#8217;s what I do. And if you&#8217;ve been doing that for decades and somebody says <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> is over, it&#8217;s like my career is over. So I think there&#8217;s an emotional side of it.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2109">35:09</a>] But just look at the data, the data doesn&#8217;t back up with. Keller says, but I also didn&#8217;t say that the technology is not improving. The technology is continuing to improve and specifically domain-specific architectures. But if you look at the general-purpose architectures or even other memory technologies, it used to be the DRAMs would improve by a factor of four in density every three years, like clockwork.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2134">35:34</a>] And now it&#8217;s maybe 10 years between factors. So you can see plenty of evidence that <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> no longer applies.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2141">35:41</a>] Is there some new version or some analog to <a href="https://en.wikipedia.org/wiki/Moore's_law">Moore&#8217;s Law</a> that is kind of guiding microprocessor design and architecture these days?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2154">35:54</a>] I guess the question is both <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> and <a href="https://en.wikipedia.org/wiki/Google">Google</a> are continuing to deliver much faster processors for machine learning, tremendously better. What are they doing? Well, part of it is what I said about packaging, to be able to get more transistors, and you keep them closer together because distance matters. Part of it is innovating on the floating-point formats. So unlike the supercomputers of 64-bit, it&#8217;s not even 64-bit, not even 32-bit, not even 16-bit, but 8-bit and 4-bit floating point is going on.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2193">36:33</a>] So narrowing of the data types, having kind of what&#8217;s called the matrix multiply instructions. So these powerful units that are put in there that can do special purpose applications that are very important, making those bigger and faster. So those are the three things that are going on. But we don&#8217;t have this simplifying guideline kind of underlying all this. You have to be aware of each piece of the technology, gauge how fast it&#8217;s improving, whether they can deliver on what&#8217;s going on, and then you assemble that together and make your bets.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2235">37:15</a>] There&#8217;s a paper that we&#8217;ve just got approved that I think we&#8217;re going to put on <a href="https://en.wikipedia.org/wiki/ArXiv">arXiv</a> soon, which is talking about the <a href="https://en.wikipedia.org/wiki/Google">Google</a> <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a> line. And pretty remarkably, the <a href="https://en.wikipedia.org/wiki/Google">Google</a> <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a> line, particularly for training, has stayed pretty constant in the basic architecture design. Things have gotten bigger and faster. But if you look at the design of it going back to the first training <a href="https://en.wikipedia.org/wiki/Tensor_Processing_Unit">TPU</a>, that block diagram still works all these a decade later.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2267">37:47</a>] So the people who designed that in, whatever, 2015 or 2016, did a really great job.</p><h3>38:04 &#8212; GPU benchmarks</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2284">38:04</a>] Blocks are still there. I think in your career with the CPU benchmarks were huge. They played a huge role in measuring which computer architectures were performant and not in this new space of floating point operations for AI. Is there a benchmark that people use for GPUs, and can you also apply it for TPUs?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2307">38:27</a>] I and some friends were involved in the MLPerf effort. So it was inspired by the SPEC CPU effort, with the SPEC benchmarks, where there were these competing companies and they&#8217;d all make claims that theirs was better than the others, and they realized that wasn&#8217;t good for the industry. So they agreed on a set of benchmarks. So the MLPerf effort is this run now, but what&#8217;s called MLCommons is an attempt to do that.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2335">38:55</a>] If we&#8217;re going to do these comparisons, let&#8217;s not argue about what the benchmarks are. So that&#8217;s a serious effort in that kind of. Interestingly, what&#8217;s happened, I&#8217;d say, is because there&#8217;s two pieces to design. There&#8217;s the machine learning libraries that are to implement a lot of features that you need for these applications, and the libraries will even be rewritten for specific applications, rather than just having general libraries.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2369">39:29</a>] This has turned out to be a big advantage for <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> because they have a large corporation with lots of people that are available, lots of engineers who could build these libraries for them. So it&#8217;s turned out not quite a... It&#8217;s not a neutral evaluation, right? It&#8217;s the architecture plus the libraries that go with it, and the compiler too, but specifically the libraries that you tailor each time. So <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> brings out, when they announce a new architecture, they create a new set of.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2404">40:04</a>] They modify the libraries to run that really well or to run applications really well. So that&#8217;s a powerful combination. Kind of in business terms, people refer to <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> as having this <a href="https://en.wikipedia.org/wiki/CUDA">CUDA</a> moat, and part of it is <a href="https://en.wikipedia.org/wiki/CUDA">CUDA</a>, the programming language, but a big part of it is the libraries that <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> makes. So this has made it difficult for startups to be able to compete with <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a>, partly because they didn&#8217;t put enough emphasis in the software and partly because <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> just has many more engineers than they do to be able to tailor the libraries.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2442">40:42</a>] So it&#8217;s the libraries plus the architecture. There&#8217;s this powerful advantage why most people do things on GPUs. Now <a href="https://en.wikipedia.org/wiki/Google">Google</a> has been able to develop their own libraries. They don&#8217;t have as many engineers. They use compilers more than I think <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> does. So they have a set of libraries that they can use. But up until recently, <a href="https://en.wikipedia.org/wiki/Google">Google</a>&#8216;s, it&#8217;s all been internal. It&#8217;s just for <a href="https://en.wikipedia.org/wiki/Google">Google</a> to use, or you could use it via the cloud.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2473">41:13</a>] But these startups have to have a difficulty as a result. That&#8217;s why MLPerf hasn&#8217;t been as popular. Not everybody runs them because <a href="https://en.wikipedia.org/wiki/Nvidia">NVIDIA</a> runs them really well, really better than everybody else. And the startups have a hard time showing off what they can do, given they don&#8217;t have the engineering effort going into the libraries that you need to do well in MLPerf.</p><h3>41:40 &#8212; How to have a bad career</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2500">41:40</a>] You gave this popular talk about how to have a bad career, and it&#8217;s kind of the negation of advice for how to have a good career. And I was just wondering if you could summarize maybe the top three things that people should take away.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2518">41:58</a>] Yeah, well, there&#8217;s been a few things. So the story is early in my career, my friend, a good friend, and I said, how can we teach grad students how to give a good talk? And we thought it&#8217;d be funny to explain how to give a bad talk. And then if you didn&#8217;t want to give a bad talk. So this is how to do it badly. And if you don&#8217;t, here&#8217;s the things you do not to do it badly. And so then I later did how to have a bad career that&#8217;s pretty much focused toward academia.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2543">42:23</a>] So for researchers to be able to do that. I later did a How to Have a Bad, How to Build a Bad Research Center, how to build a bad research lab. And then recently I&#8217;ve been giving how to give AI a bad carbon footprint. That&#8217;s my latest one. But I think the thing that might be more relevant is at the end of my talks, I would kind of reflect on my career and talk about lessons learned. So I&#8217;ve written a paper called Lesson.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2569">42:49</a>] I think it&#8217;s &#8220;Lessons, Life Lessons from the First Half Century of My Career.&#8221; I wrote that recently. And so those are divided up into kind of career advice and personal advice. There, on the personal side, I&#8217;d say, if you have a family, make sure your family&#8217;s first. Keep your family first. The technology we&#8217;ve invented makes it really easy for your, you know, to take your work home with you and not pay attention to your family.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2597">43:17</a>] I had a. When I was first here at <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Berkeley">UC Berkeley</a>, I gave a senior faculty member a ride home, dropped him off late at night, coming up from Silicon Valley. And he said, well, Dave, if I had to do it over again, I wish I&#8217;d spent more time with the family. And I never wanted to say that. And nobody on their deathbed says, you know, I wish I spent more time in the office. So, so you got to, you know, whenever it&#8217;s easy to kind of, you know, let the.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2624">43:44</a>] Let your career overcome the needs of your family, but you just got to remember that. I would say, in my life, growing up in the 50s, the idea was that to be happy, you had to be wealthy. But those are actually two different goals, wealth and happiness. So I always made decisions toward happiness versus wealth, and I felt very good about that. And even now, why would I pick something that made me wealthy and unhappy?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2655">44:15</a>] Why would you do that? And in the field that we&#8217;re in, it turned out there was a lot of wealth to go with it. So optimizing happiness didn&#8217;t mean you had to suffer. Beyond that, I think it&#8217;s important to have fun personally. I think when you&#8217;re a kid, you don&#8217;t have to tell kids to play, but as you get adult, you get so busy you don&#8217;t think you have time to have fun. But you only get to do this once.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2683">44:43</a>] So right now I play soccer, I rode my bicycle to the interview, I lift weights, I body surf, do things with my wife and family and my son. So it&#8217;s important to have fun. I think on the career side, one of the pieces of advice is by Stephen Covey. He wrote this book, The 7 Habits of Highly Effective People, I think, and he has a little quadrant, and he divides it into urgent and not urgent, and important and unimportant.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2720">45:20</a>] There are so many things in our technology, like email and texting, for you to focus on the urgent things, but you really shouldn&#8217;t be spending a lot of time on the unimportant urgent things. And it takes self-discipline to set aside time for the important, non-urgent things. But if you don&#8217;t block that out, you can just not have time to do anything. I get to see other people&#8217;s calendars at <a href="https://en.wikipedia.org/wiki/Google">Google</a>, and there&#8217;s managers that every half hour from eight to six, five days a week, are scheduled.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2749">45:49</a>] And I don&#8217;t see how you have time to think and reflect on things like that. Other career advice thing is I remember I kind of woke up one morning and it was like God spoke to me. I was like thunderstruck. And it said it&#8217;s not how many things you start, it&#8217;s how many things you finish. It seems relatively obvious, but that&#8217;s not the way I was acting. I had many things going on, but after that it&#8217;s like there&#8217;s one main thing I&#8217;m doing at times.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2779">46:19</a>] So when <a href="https://en.wikipedia.org/wiki/John_L._Hennessy">John Hennessy</a> and I wrote a textbook, that was the main thing. When I was department head here, that was the main thing I did. I would do some other little things there. And <a href="https://en.wikipedia.org/wiki/John_L._Hennessy">John Hennessy</a>, he wrote a kind of career advice book based on his presidency. And one of the things he said in there, you&#8217;re only going to be remembered for the five or six things you&#8217;ve done in your life, not for the hundreds of little things.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2804">46:44</a>] So to give yourself a chance to have some things you&#8217;re really proud of, it&#8217;s better to concentrate on a few of them, hoping that some of them will turn out to be a big deal rather than scatter yourself to many things. But if you&#8217;re interested, you can. Yeah, if you look for life lessons, <a href="https://en.wikipedia.org/wiki/David_Patterson_(computer_scientist)">David Patterson</a>, first half century, you can see the whole list of. I think there&#8217;s 16 lessons altogether.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2831">47:11</a>] You mentioned that you didn&#8217;t want to reflect back on your life and feel like you didn&#8217;t have enough family time. Yeah, I&#8217;ve&#8212; I don&#8217;t think I&#8217;ve ever heard anyone say the opposite. Why do you think that is? Why is it that everyone looks back on their life and they never regret a lot of family time? I imagine there&#8217;s got to be at least one person that says, &#8220;I spent too much time with my family; my career suffered.&#8221;</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2859">47:39</a>] Yeah, well, what&#8217;s, what&#8217;s, it&#8217;s pretty philosophical. I mean, what&#8217;s life all about? What&#8217;s success? Right. You have to figure out what that means for you. I mean, if you have, I don&#8217;t know, if you have financial goals of being a millionaire, a billionaire or something like that, and that&#8217;s how you judge your life. Okay. Good luck.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2884">48:04</a>] You know, it&#8217;s pretty hard to make it. It&#8217;s pretty hard to do, but I think it&#8217;s the... When I was finishing my PhD, I read this book by Studs Terkel called Working, where he interviewed all these people in his careers, and they looked back at what they liked and what they didn&#8217;t like, what they felt about their careers, and what I got out of it. The people who worked with people like ministers or teachers or doctors felt really good about what they do with their careers.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2910">48:30</a>] And the people who did more ephemeral stuff, like technology things or airplanes or something like that, are long gone, didn&#8217;t feel as good about it. It was the people that they worked with that they really cared about. And I went out with a retired engineering dean here, and he&#8212; and to a meeting who I knew, and he said, &#8220;You know, Dave, as I think my career, it wasn&#8217;t the projects, it was the people that I worked with that mattered.&#8221;</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2934">48:54</a>] And I thought I knew that from a long time ago. So that&#8217;s, I mean, this is kind of senior persons offering advice. I think you&#8217;re going to care more about the people you&#8217;ve worked with, the people you&#8217;ve helped, will be a bigger deal. That&#8217;s part of, one of the other things about personal happiness is they studied happiness. Psychologists used to just study crazy people, but they started like, why are people happy?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2957">49:17</a>] And they know what the reasons are. Have a job that you like, have friends and family. Helping other people makes you happy. They know this. Having something kind of either a religious side of it, it doesn&#8217;t have to be a formal religion, but contact with nature, the grandeur of nature, but the list of things you need to do to be happy is well understood.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2982">49:42</a>] And our world is filled with unhappy billionaires. Right. If money was the thing that made you happy, why are these people, very wealthy people, so mad about things? So, yeah, this is me passing on advice.</p><h3>49:59 &#8212; Courage and optimism</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=2999">49:59</a>] In one of your talks, it said what worked well for me, and it was kind of some reflections. And one of the things in there I thought was unique and interesting. You mentioned that courage was a big part of your career, and I don&#8217;t hear that too often. I was curious why you say courage is so important in a career.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3019">50:19</a>] Yeah, I think that&#8217;s, I mean, that might be partially my personal makeup, but I was kind of the youngest kid in my class and kind of small. It took me a while to grow despite age, so I was always small. But my parents encouraged me to go out for wrestling. And wrestling gives you physical self-confidence because you spend years doing that. I did it in high school and college. And so I think partly technically, having courage to do things kind of goes along with the advice is fortune favors the bold.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3059">50:59</a>] That goes. That&#8217;s advice is 2,000 years old. It&#8217;s hard to figure this out. Helen Keller wrote, even trying to play it safe, you still get caught. So it turns out you might as well. You might fail no matter what. And if you take a big chance, you can succeed. If you don&#8217;t take the chance, you probably won&#8217;t succeed if you play it safe. So fortune favors the bold. And I think it takes courage to do that.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3088">51:28</a>] I think also it just, for me, intellectually, I feel like if there&#8217;s something not right, I need to stand up and confront it. And I think that kind of, ironically, comes from the wrestling side of my personality, where if I see something, somebody getting picked on or something like that, I&#8217;m going to stand and try to stop it. And I feel that same responsibility. Intellectually, if people are making bad arguments or doing something, we need to stand up and do it.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3117">51:57</a>] And I feel good about that. The cautionary part about that is, because I guess one of my senior faculty members saw this nature of me, and he said one of the sayings is friends come and go, but enemies accumulate. This is an old saying. So if you think about it, you kind of, people you went to high school with, while friends, you kind of forget. But somebody who you really dislike, you never forget that you dislike that person.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3142">52:22</a>] So standing up when it&#8217;s important, but be careful when you make enemies because they&#8217;re going to stick around for a long time.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3149">52:29</a>] When you say something&#8217;s gone bad, technically, do you mean someone was incorrect?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3155">52:35</a>] Yeah, when they&#8217;re weak, either politically or technically, when it&#8217;s a weak argument. Right. I think of this, I really like there to be a marketplace of ideas, and we hone the ideas by arguing them. And so if it&#8217;s a specious argument, even if a person is a leader of a company and it just doesn&#8217;t make sense, I feel it&#8217;s kind of the responsibility.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3183">53:03</a>] It&#8217;s better for the company if somebody stands up and points that out than to just let them get away with it. And then it&#8217;s a little bit confrontational. But as long as people all agree that this is for the greater good, we&#8217;re gonna, we need to get the right ideas out there. And so let&#8217;s argue about the ideas just to see, polish them to make them stronger. I think that&#8217;s important in science and engineering and kind of in life too.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3216">53:36</a>] There&#8217;s a lot of stuff going on right now in the country that is worrisome. And I certainly stood up and wrote op-eds about things that I think are wrong and need to be corrected. And if people are afraid to do that, it&#8217;s hard to be optimistic about the future. If people are afraid to stand up when there are wrongs and try to stop them.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3242">54:02</a>] You also mentioned optimism in the talk, and you had this story. I wonder if you&#8217;re willing to share.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3248">54:08</a>] So I would say in engineering, it&#8217;s hard. It&#8217;s hard to know, right? But I think you need to be kind of optimistic or positive, have a positive outlook because so many things could go wrong. And then my personal story that illustrates this goes back to high school when I&#8217;m 16. I&#8217;m dating this very attractive girl, and I screw up my courage and ask her if we would be exclusive. At the time, the phrase we used was going steady.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3275">54:35</a>] And she looked at me and said, and, you know, she was 16. She had dated other guys and thought we were pretty young. And she said, &#8220;Well, Dave, you&#8217;re such a nice guy, I don&#8217;t know how to say no.&#8221; For me, as a logical person, a no sounded like a yes. And so I hugged her and said, &#8220;Great.&#8221; And so she, in her mind, thought, &#8220;Well, I&#8217;ll let him down gently later.&#8221; But we&#8217;ve been married 59 years now, and she hasn&#8217;t let me down yet.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3304">55:04</a>] So that was a case where optimism paid off.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3307">55:07</a>] I think everyone that hears about a healthy relationship for that long might wonder how you did it.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3313">55:13</a>] I used to tell people, if you go to weddings, the marriage vows are really great, right? But nobody can remember their wedding vows. I used to say, remember your wedding vows, but nobody remembered that. So we boiled it down to nine magic words, and it&#8217;s just three sentences. And they start: &#8220;I was wrong, you were right. I love you.&#8221; Okay?</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3335">55:35</a>] Those are the nine words. And this applies to both partners in a relationship, not just one partner. But, yeah, if you can say them all, and no substitutions&#8212; I was wrong. You&#8217;re right. You&#8217;re a jerk. You know, you can&#8217;t do that. If you can remember those nine words, that can help you have a long relationship, like my wife and I have.</p><h3>55:56 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3356">55:56</a>] And then last question for you: knowing everything you know now from your career, if you could go back to yourself when you had just entered the industry and give yourself advice, what would you say?</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3369">56:09</a>] I mean, I think when I got here, because the imposter syndrome. I was a <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Los_Angeles">UCLA</a> graduate student, and suddenly I&#8217;m a <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Berkeley">UC Berkeley</a> professor. So that just doesn&#8217;t seem like&#8212; That was very intimidating. But after a while, I just thought, well, I&#8217;m probably not going to get tenure, so I should just have a good time. So I think I already had the right attitude about it. I think that first year, I think it was very stressful while I was trying to handle the imposter syndrome and be a <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Berkeley">UC Berkeley</a> professor.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3400">56:40</a>] But I think after that, I handled it pretty well. I did all the things with the kids and stuff. So there&#8217;s a version of that question is, like, is there anything I would do over again? There&#8217;s one thing I would have. I was chair of this, of the architecture community, ACM SIGARCH, as it&#8217;s called. And they have an annual conference of the year. And I was in, this was in the 1990s, I think I was the chair.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3426">57:06</a>] And what I wasn&#8217;t aware of is that at these conferences, there were men who were harassing young women at this conference. I just didn&#8217;t think people like young people, people like me, nobody would do that. Only an idiot would do that. That can&#8217;t possibly be happening. But it was happening, and I wish somebody had said something to me about it, because I would have straightened it out.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3453">57:33</a>] Any man would have threatened his life if he were to do that. Today, the only comforting thing is Sarita Adve, who&#8217;s a famous computer architect. She said she also was not aware that that was going on. It became clear later, you know, five or 10 years later, it became more clear that this was going on. And there are mechanisms. Sorry, but that&#8217;s the one thing I wish. You know, if I could go back in time, I would have figured that out, and I would have straightened men out who were doing that, and they wouldn&#8217;t.</p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3485">58:05</a>] That would have stopped. I believe that would have stopped them.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3488">58:08</a>] Yeah. Thank you so much for your time today. I really appreciate it.</p><p><strong>David:</strong></p><p>[<a href="https://youtu.be/Pn4ZwlEh5nw?t=3492">58:12</a>] All right. Thanks for the interview.</p>]]></content:encoded></item><item><title><![CDATA[Turing Award Winner: NSA, Public Key Cryptography, Crypto Wars | Martin Hellman]]></title><description><![CDATA[Interviewed Martin Hellman recently who was one of the inventors of the Diffie-Hellman key exchange algorithm.]]></description><link>https://www.developing.dev/p/turing-award-winner-nsa-public-key</link><guid isPermaLink="false">https://www.developing.dev/p/turing-award-winner-nsa-public-key</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 06 Jul 2026 13:02:52 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/204966592/99aa5a09175eeb932c18cc7e260b12fe.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Interviewed <a href="https://www.linkedin.com/in/martin-hellman-473290a9/">Martin Hellman</a> recently who was one of the inventors of the <a href="https://en.wikipedia.org/wiki/Diffie%E2%80%93Hellman_key_exchange">Diffie-Hellman key exchange</a> algorithm. I had heard of his name ever since I learned the concept when I studied computer science in college.</p><p>Pretty surreal to get to talk to him. He explained some of the key concepts but also it was interesting to hear the stories behind what we learn in the textbooks.</p><p>Also, I have a ton of content recorded that I&#8217;m working on editing. Several Turing Award winners and programming language creators. I learned a lot during the latter since I basically talked to them as if I was a student curious to learn about programming languages. Excited to share these episode with you all!</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/AZLOETBCQM4">YouTube</a>, <a href="https://open.spotify.com/episode/0QNBcSuQvkr8wZG6BCxp31?si=Eyvcb0DkRfW_51b-pqz8vg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-AZLOETBCQM4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;AZLOETBCQM4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/AZLOETBCQM4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/204966592/0034-why-his-work-broke-the-law">00:34 - Why his work broke the law</a></p><p><a href="https://www.developing.dev/i/204966592/0839-how-people-did-encryption-before">08:39 - How people did encryption before</a></p><p><a href="https://www.developing.dev/i/204966592/1851-the-crypto-wars">18:51 - The crypto wars</a></p><p><a href="https://www.developing.dev/i/204966592/2622-the-story-behind-diffie-hellman-key-exchange">26:22 - The story behind Diffie Hellman key exchange</a></p><p><a href="https://www.developing.dev/i/204966592/3648-signatures-vs-key-exchange">36:48 - Signatures vs key exchange</a></p><p><a href="https://www.developing.dev/i/204966592/4305-rsa-patent-wars">43:05 - RSA patent wars</a></p><p><a href="https://www.developing.dev/i/204966592/4808-why-inventions-happen-at-similar-times">48:08 - Why inventions happen at similar times</a></p><p><a href="https://www.developing.dev/i/204966592/5029-what-he-worked-on-after-cryptography">50:29 - What he worked on after cryptography</a></p><p><a href="https://www.developing.dev/i/204966592/5731-his-thoughts-on-death">57:31 - His thoughts on death</a></p><p><a href="https://www.developing.dev/i/204966592/5940-advice-for-his-younger-self">59:40 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:34 &#8212; Why his work broke the law</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=34">00:34</a>] I saw your initial work in cryptography was potentially breaking laws, or it was kind of in this gray area where the U.S. National Security Agency was kind of not super happy with your work, and there was some conflict with the government. Could you explain what the field of cryptography was like at the time and its relationship to the government and how you got into the field?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=57">00:57</a>] Cryptography is the study of codes and ciphers for keeping information secret. When I started to get interested in this area in 1968, I was at IBM Research just after finishing my PhD. They&#8217;d hired <a href="https://en.wikipedia.org/wiki/Horst_Feistel">Horst Feistel</a>, who you may have heard of, who was the father of IBM&#8217;s work in this area and really started the unclassified research in cryptography. So I started there. And then I realized when I was teaching at <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> from &#8216;69 to &#8216;71 that I might actually be able to do something.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=92">01:32</a>] I won&#8217;t go into the details of that right now. And I did it in my own time because I didn&#8217;t think I was breaking any laws by doing it. I didn&#8217;t think anybody could make that case. But I realized that there was a part of the government, the U.S. National Security Agency, that would be very concerned about this. And then it was as our work really became visible that they started saying we were breaking the laws.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=119">01:59</a>] And that&#8217;s when I looked into it more carefully. We got a letter in July 1977 from a guy who worked at. We determined he worked at the U.S. National Security Agency. He wrote from his home address in Maryland, which is a good indicator that he works at the U.S. National Security Agency. To the Institute of Electrical and Electronics Engineers, which is the main electrical engineering professional society, saying that as an <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> member, he was concerned that they were breaking the law by publishing certain papers.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=151">02:31</a>] He never mentioned me by name, but I had an article in almost every issue that he mentioned of the <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> Transactions. The <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> wrote back to him, copying me. They didn&#8217;t copy me as Martin Hellman, troublemaker. They copied me as a member of the Board of Governors of the Information Theory Group, which was then publishing most of these papers. And so it&#8217;s a funny thing that when people start to talk about cryptography, they talk in code all the time.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=181">03:01</a>] The <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> was maintaining I was breaking the law by publishing my papers in international journals. The <a href="https://en.wikipedia.org/wiki/International_Traffic_in_Arms_Regulations">ITAR</a>, the International Traffic in Arms Regulations, which I&#8217;d been unaware of up to that point, defines anything cryptographic as an implement of war. And I wasn&#8217;t exporting an implement of war, which would mean you need a license from the State Department. I was exporting technical data from their perspective on implements of war.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=207">03:27</a>] And so that&#8217;s why I was breaking the law. So I took the letter to Stanford&#8217;s general counsel at the time, John Schwartz. He said if the laws interpreted that broadly, he thought it was unconstitutional. But that&#8217;s only his legal opinion. It had to be settled in a court of law.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=223">03:43</a>] I mean, when you got these warnings, did it deter you at all, or what was your immediate reaction?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=229">03:49</a>] I kind of fell into this at first with <a href="https://en.wikipedia.org/wiki/Whitfield_Diffie">Whit Diffie</a>, and I thought that the 56-bit key size of DES, the <a href="https://en.wikipedia.org/wiki/Data_Encryption_Standard">Data Encryption Standard</a>, was a mistake. We realized that you could break it with exhaustive search. Today, <a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard">AES</a>, the <a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard">Advanced Encryption Standard</a>, has a minimum 128-bit key size, which is much, much bigger than a 56-bit key size. Two to the 128th is much bigger than two to the 56th. So we wrote letters.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=260">04:20</a>] We thought that they&#8217;d change it. It was cheap and easy to change it. And after six months we realized that we had a political problem, not a technical problem. And then we had to decide whether to back down or not. And we went forward because at that point I was invested in this. And also I&#8217;m from the Bronx, as I mentioned to you earlier, and I&#8217;m a street fighter. Not with my fists, I got beat up a lot, but I&#8217;m a street fighter with my mouth and my mind.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=285">04:45</a>] And they picked the wrong person, basically. I wasn&#8217;t going to back down.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=291">04:51</a>] What could have gone wrong? I mean, would they pursue you with lawyers, or would...</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=295">04:55</a>] Oh, I could have gotten killed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=297">04:57</a>] You could have gotten killed?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=298">04:58</a>] Oh, yeah, not necessarily by them, but maybe by the <a href="https://en.wikipedia.org/wiki/Minions%3A_The_Rise_of_Gru">GRU</a>, the Russian equivalent. I mean, I wasn&#8217;t just pissing off the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, I was pissing off the Russians, Soviets.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=307">05:07</a>] Why would they be pissed off?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=309">05:09</a>] Oh, we and the Soviets sold our World War II surplus cryptographic equipment to third-world countries. We&#8217;d rather that they were able to break it, and they&#8217;d rather that we were able to break it than no one was able to break it. So my mother called me up. My mother was still alive. She died in 1984. And she called me up and said, &#8220;What are you doing? You&#8217;re going to get yourself killed.&#8221; She was really upset.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=335">05:35</a>] Okay, so you get issued the warning, and then you go to the general counsel at Stanford. Then what happens?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=342">05:42</a>] This was July 1977, in I think October. Early October, yes, it was October. There was an <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> symposium on information theory at Cornell University. And I had two papers there. One with <a href="https://en.wikipedia.org/wiki/Ralph_Merkle">Ralph Merkle</a>, who deserves, I believe, a lot of credit for inventing public-key cryptography, as you probably know, and one with <a href="https://en.wikipedia.org/wiki/Stephen_Pohlig">Stephen Pohlig</a>, which really. Stephen and I have a paper which is really the <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> system, except it&#8217;s mod P instead of mod N.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=372">06:12</a>] It&#8217;s modular prime number instead of modular composite number. We realized that you could do it mod N, but we didn&#8217;t have public-key cryptography then. Anyway, Stephen and Ralph and I had papers. And John Schwartz said that he strongly recommended that I give the papers instead of the students. I was going to have the students give the papers. And so I went to the students and I told them, look, if you want to give the papers, that&#8217;s okay, but Stanford&#8217;s not sure they can defend you, whereas it&#8217;s clear they could defend me.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=402">06:42</a>] And I feel very comfortable going forward with Stanford&#8217;s financial resources. And the students initially bravely said that they would give the papers. After a week they came back to me. Their mothers were beating on them like my mother beat on me. And they said, &#8220;Okay, you give the paper.&#8221; So when it came time for each of those papers to be given at Cornell, I went up to the podium and I said, &#8220;I had the student come up with me.&#8221;</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=429">07:09</a>] And I said, normally the student would be giving this paper, but on the advice of counsel, and everyone knew what I was talking about, I will be giving the paper. And I want you to consider, in every way but legally, the words coming from my mouth as if they&#8217;re coming from his. And so the students got more credit that way than they ever would have if they gave the papers by themselves.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=449">07:29</a>] It sounds like you had a lot of opposition at the time. Why did you still pursue it?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=454">07:34</a>] Well, my colleagues all told me that I was crazy to work in cryptography. And they had two reasons, one of which was the U.S. National Security Agency has a multibillion-dollar-a-year budget, and they&#8217;ve got a decade&#8217;s head start. How can you hope to discover something they don&#8217;t already know? And my attitude was, I don&#8217;t care what they know. It&#8217;s not available for commercial use. Even if I only develop things that they know, it doesn&#8217;t matter.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=477">07:57</a>] And yet we developed things that they didn&#8217;t know about, which is a whole other thing. <a href="https://en.wikipedia.org/wiki/GCHQ">GCHQ</a>, the British equivalent of the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, has claimed that they invented public-key cryptography a little bit before us. Not a lot, by the way. Just around the same time. Except they never had anything on digital signatures. They never had anything that allowed my phone to accept a software update and to use a secret key to sign it, which Apple has, and a public key in the phone.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=506">08:26</a>] So even if someone takes the phone apart and gets the public key, that doesn&#8217;t tell them how to sign future software updates. And <a href="https://en.wikipedia.org/wiki/GCHQ">GCHQ</a> had nothing on digital signatures, only on the privacy aspect.</p><h3>08:39 &#8212; How people did encryption before</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=519">08:39</a>] Prior to the solutions you came up with, how did people encrypt it? How did people encrypt it?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=526">08:46</a>] Encrypt?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=527">08:47</a>] Yeah. Encryption messages.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=529">08:49</a>] They didn&#8217;t. No, actually. <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, <a href="https://en.wikipedia.org/wiki/GCHQ">GCHQ</a>, <a href="https://en.wikipedia.org/wiki/Minions%3A_The_Rise_of_Gru">GRU</a>, all these foreign United States and foreign intelligence services were sucking up huge amounts of unencrypted data. It was only mining companies and banks that used encryption, and they used very bad encryption.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=547">09:07</a>] What about during wars and stuff? Sending messages that they didn&#8217;t want other adversaries to be able to read.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=556">09:16</a>] You&#8217;re talking about military use.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=557">09:17</a>] Yeah, exactly.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=558">09:18</a>] Well, that was where the major use had been, in the military. And one of the things I have a talk on the evolution of public-key cryptography. And I start off by saying, I love that. People call it revolutionary. Don&#8217;t stop saying that. But by the end of this talk, you wonder why it took us so long rather than how we ever found it. And they were almost like big arrows pointing: Go here, go here, go here.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=583">09:43</a>] And one of them was Whit and I had come up with the idea of a trapdoor cipher. That was a cipher that was impossible to break unless you knew trapdoor information that went into its design. We&#8217;ve never been able to develop such a cipher, but we realized that was a possibility. And that would be a general&#8217;s dream, because he could use it. If he knew the trapdoor information, he could use it securely against his adversary in the war.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=610">10:10</a>] But his adversary, if they captured it, couldn&#8217;t use it securely because he knew the trapdoor information and could break it. And from there to public-key cryptography as a minor step.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=619">10:19</a>] I know there&#8217;s this famous machine. I think it&#8217;s called the Enigma machine. And Turing was involved in that. That&#8217;s some sort of encryption mechanism.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=629">10:29</a>] Basically, it&#8217;s a World War II. It&#8217;s the main German field cipher of World War II, the Enigma machine. It was gear-based. And encryption is a special-purpose form of computation. And computation has gotten a hell of a lot better. We do, I don&#8217;t know, I&#8217;m 80 years old. So it&#8217;s 16 periods of five years. We can do about maybe 10 to the 14th, it&#8217;s not quite 10 to the 16th times as much computation for a dollar today as we could when I was born at the end of World War II. I was born in &#8216;45.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=661">11:01</a>] So Enigma, what was the primary computing device of World War II? Gear-based adding machines. What was the primary encryption device? Gear-based encryption machines. Today we can do not only computation much better, we can do encryption much better. In that case, it substituted electricity. You pressed the letter on a keyboard, it generated an electrical signal that went through a sequence of rotors, as they&#8217;re called, which transformed it into another letter that lit up.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=691">11:31</a>] And so you&#8217;d set up the key, which was which rotors to use and what order to put them in, and some other things. And then you&#8217;d type out the message on the keyboard, and you&#8217;d read off the enciphered message on the lights. And Alan Turing, who you pointed out in The Imitation Game, which was the major movie of like five years ago, Alan Turing figured out how to build a primitive form of computer to break the Enigma machine.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=719">11:59</a>] Now, the Germans knew Enigma was breakable. They never thought that the Allies would build a computer to break it. And so the Allies were reading German encrypted messages almost as fast, maybe faster than the Germans were. And that&#8217;s what was different. And as they knew Enigma was breakable, they never realized that. They never thought that we would build a machine like that. And it&#8217;s interesting, Alan Turing was hounded to death in the 1950s in spite of his wartime activities.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=747">12:27</a>] And this is covered in The Imitation Game. In that movie, he was hounded to death because of his homosexual tendencies, or more than tendencies. I mean, he was homosexual. Today we wonder how our parents, grandparents could have been so uncivilized as to do things like that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=767">12:47</a>] You mentioned in that talk on the history of public cryptography, there&#8217;s this. Everything is pointing toward these directions. What did you mean by that?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=777">12:57</a>] Well, I mean, the trapdoor cipher was pointing us toward public-key cryptography. And yet a trapdoor cipher makes perfect sense because almost everything in cryptography involves trapdoors. People usually think of public-key cryptography as a trapdoor cryptographic system. But even a cryptographic system is a trapdoor. And this is fundamentally related to a question called P vs NP or not in computer science.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=806">13:26</a>] Are there any problems that are easily checked that are hard to solve? That&#8217;s basically what that means. If there are any secure cryptographic systems, then there are problems that are easily checked that are hard to solve. For example, one-way function is the simplest form of trapdoor. This is a function that&#8217;s easy to compute but hard to invert. And this figures importantly in cryptography and blockchains, among other things.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=837">13:57</a>] We&#8217;ve all taken trapdoor quizzes where the professor gives you an hour to solve the problem. At the end of the hour, he says, &#8220;Pencils down, papers in,&#8221; and he says, &#8220;Oh, we have a few minutes remaining. Let me give you the solution and convince you it&#8217;s right.&#8221; That&#8217;s a trapdoor quiz problem. It&#8217;s a primitive form. And there the professor has in mind a particular method for solving it, whereas you have many methods.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=863">14:23</a>] And a real trapdoor quiz problem would work like this. The professor would say, here&#8217;s a one-way function. I&#8217;ve generated an X, here&#8217;s the Y that comes out. You find the X that comes in after a million years goes by. He says, papers down, pencils in. He says, oh, by the way, in the few seconds remaining, let me convince you. Let me give you the solution and convince you it&#8217;s right. What does he do?</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=886">14:46</a>] He takes X, he puts it through the one-way function, he gets Y. And you know that he had the right value of X, and yet you could never find it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=895">14:55</a>] When I imagine you kind of competing or butting heads with the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, I can imagine one solution for them might be to try to recruit you to their side. Did they ever try to recruit you?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=907">15:07</a>] Yes. Early on, before we had public, before we had any good results, I&#8217;d go to conferences and I&#8217;d give talks on cryptography, which in hindsight were very primitive. And <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> people with name badges that said <a href="https://en.wikipedia.org/wiki/United_States_Department_of_Defense">Department of Defense</a>, which was the simple substitution cipher for <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, would always ask me. They explained they were at <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> and they&#8217;d love to hire me as a consultant because they needed new blood.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=932">15:32</a>] And I said, I&#8217;d love to know what you know, but only if I can publish my papers independently. And he said, of course not, which I knew was the answer. And so I never took any classified consulting contracts, et cetera, until 1995. There was a National Research Council committee called CRISIS. It&#8217;s called the CRISIS Report. Cryptography&#8217;s Role in Securing the Information Society, which comes out as CRISIS.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=962">16:02</a>] And that was about 1995. And the question was, were U.S. government regulations helping national security or hurting national security? We concluded that in many ways they were hurting national security. And we had a former deputy director of the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> on the committee. We had a former Attorney General on the committee, Benjamin Civiletti, who had been Attorney General under Jimmy Carter. We had people like me who were privacy advocates, but who had come around and started to understand how the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> viewed things.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=994">16:34</a>] We had Ray Ozzie on that committee, who was the chief software architect at Microsoft. When Bill Gates retired, he did Lotus Notes. And we reached unanimous conclusions. And one of the most important conclusions we reached was that the classified briefings. I did take a classified. I didn&#8217;t get paid for that, but I did have a security clearance for that and intelligence coins, which I&#8217;d never taken before.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1024">17:04</a>] But I was retiring at that point, and I figured it didn&#8217;t matter. And it was important for me to be able to, say, get past the people who said, &#8220;If you knew what I knew, you&#8217;d think differently.&#8221; And there were people in the unclassified briefings who said, &#8220;If you knew what we knew, you&#8217;d think differently.&#8221; But we concluded, and it&#8217;s in our report, which you can get free of charge on the National Academy, National Academy&#8217;s website, that the classified briefings helped fill in details, but they did not change our fundamental conclusions.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1059">17:39</a>] And that was one of the most important conclusions we reached. And yet people said, if you knew what we knew, you&#8217;d think differently. And yet that was not true.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1067">17:47</a>] When they were trying to recruit you, I imagine it would have been very advantageous for them to classify your work.</p><p>Oh, yeah, they did. I mean, when you turned down their offer, I imagine they would have had to compensate you handsomely to kind of go to the.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1084">18:04</a>] I don&#8217;t. We never got that far. And by the time we had a fight on our hands, it was too late. Basically, I&#8217;d committed to see this thing through. And when I realized that my life might be in danger, I might go to jail, although I didn&#8217;t think about those possibilities very much, and I didn&#8217;t think they were real, although they may have been. I felt like I was committed and I couldn&#8217;t back down.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1115">18:35</a>] I mean, what would you do? I kind of fell into it. It&#8217;s like a friend of mine who was a helicopter pilot in Iraq. He went through ROTC in the late 90s, he said, to pay for his education. There were no wars. He said, why is everybody calling me a hero?</p><h3>18:51 &#8212; The crypto wars</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1131">18:51</a>] When I was doing the research, I kept seeing this mention of the first crypto wars. What are the crypto wars, and how did they play out?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1140">19:00</a>] The first crypto war was the one we&#8217;ve been talking about, which occurred in the late &#8216;70s, early &#8216;80s, where <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> threatened to throw me in jail. My life may have been in danger, things like that. And that was over two things. It was over the freedom to publish papers in international journals, which has now been established. And it was over the key size of DES, which is something that people today know almost nothing about.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1164">19:24</a>] As I mentioned before, it was a 56-bit key size. And Whit and I did a quick calculation that you could probably break that for $10,000 even with 1976 technology, 1975 technology. And if we were wrong, Moore&#8217;s law would, if we were off by a factor of 10, an order of magnitude that would be erased in five years&#8217; time because Moore&#8217;s law was advancing the state of computing by an order of magnitude every five years in those days.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1194">19:54</a>] So that was the first crypto war. It was over the freedom to publish and the key size of DES. The second crypto war, which people don&#8217;t really remember, was in the &#8216;90s and Clipper Chip was a big piece of that. This was the idea that there should be an escrow service, that the government would allow strong encryption, but it would maintain master keys somewhere. And if they got a court order that said that they were able to get into something, they&#8217;d be able to do it.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1226">20:26</a>] Well, this sounds great, and it was great from certain points of view, but there were problems with it, like in the National Research Council Committee, the CRISIS Committee that I talked about. We wasted a lot of time talking about key escrow. And then we realized, I may have actually said this, although I think we came to it as a whole. Look, there are major problems with key escrow. How is it going to work internationally?</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1249">20:49</a>] Who&#8217;s going to hold the keys when someone&#8217;s in France and someone&#8217;s in the United States, or someone&#8217;s in Russia and someone&#8217;s in the United States? So what we said in the report is, we said it very nicely, that we don&#8217;t see how to solve many of the problems with public key with key escrow. And the U.S. Government should experiment with it, and if they came up with a solution, come back to us, and they never did.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1274">21:14</a>] That was the second crypto war. And Clipper Chip was a chip that had a key escrow built into it. Basically, the third crypto war, which some people may remember, occurred about 10, 15 years ago. It was over things like the San Bernardino Apple iPhone. There were terrorists who committed acts, and they had an Apple iPhone, and the FBI wanted to break into it. FBI, not <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, wanted to break into it. And they asked Apple to give them the key.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1308">21:48</a>] Apple does not have a key. Ron Rivest is one of the authors of a report called Keys Under Doormats. If you build what they call a backdoor into a cryptographic system to allow access. And that was the third crypto war. It was really a repeat of the second crypto war, the key escrow with a different name. And it&#8217;s really putting keys under doormats because if Apple can recover one key, they can recover any key.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1336">22:16</a>] And they&#8217;re going to be asked repeatedly. And there&#8217;s a danger that if that information gets out there, that all of Apple&#8217;s devices become insecure. So how do you do it? We don&#8217;t see how to do it. It&#8217;s unfortunate, but there is no way to give access. I&#8217;d love to give access to the U.S. Government when it had legitimate reason for doing it. I would love to withhold that access from the U.S. Government when it has an illegitimate reason for doing it, which it sometimes does, and from the Russian government, which also would use it illegitimately.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1366">22:46</a>] You mentioned DES, 56 bits, was not secure enough. Why would the government want fewer bits?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1373">22:53</a>] Oh, easy. The government wanted fewer bits because they didn&#8217;t want a publicly available standard that they couldn&#8217;t break. They figured that there was a crude form of trapdoor that Whit and I made an estimate in &#8216;75, and I refined it when <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a>, actually NBS, the National Bureau of Standards, but it was two guys from <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> that we were always communicating with, who had moved over to NBS. We estimated that 2 to the 56 is roughly 10 to the 17th power.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1408">23:28</a>] That&#8217;s 100,000 million million keys. We estimated that you could build a search chip, even with 1975 technology, that would search a million keys per second, a microsecond per key. And we then said, you wouldn&#8217;t just build one of these, you&#8217;d build a million of them. And now you&#8217;re searching a million million keys per second. And how long does it take to search 100,000 million million keys? 100,000 seconds, which is about a day.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1434">23:54</a>] It&#8217;s roughly 80,000 seconds. We estimated that the cost of searching a 56-bit key size was roughly $10,000 per solution in 1975. We were probably optimistic, but as I said, even a factor of 10 optimism would be erased within five years. The crude form of trapdoor was, if I build a machine like that, I can&#8217;t keep it fully loaded. I don&#8217;t have roughly 60 problems a month to keep it fully loaded. The U.S. National Security Agency has 60 problems a month to keep it fully loaded.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1469">24:29</a>] And they have the $20 million to build the machine. I don&#8217;t.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1476">24:36</a>] So they would retain access to privileged information through superior computing resources was their hope, and snoopiness.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1487">24:47</a>] The problem is you and I might want to snoop on one problem a year. How do we do that? Do we wait a year to get a solution? No, we want it quickly, but then it&#8217;s sitting idle most of the time and our cost goes way, way up.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1500">25:00</a>] With modern public key cryptography, anyone can securely message anyone else, and no one can snoop. But that also includes the bad guys as well. What are your thoughts on, I guess, the devil&#8217;s advocate that says, why do you need to hide something if you have nothing bad to hide?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1523">25:23</a>] Well, if you want to put all your banking information out there in the public, if you want Apple to sign software updates in ways that bad guys can do, yeah, there&#8217;s no problem. But we do have reasons for being secure. We don&#8217;t live in a world where we trust everybody yet. And so there&#8217;s a fundamental trade-off. You cannot make strong encryption and give it only to the good guys and not to the bad guys.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1548">25:48</a>] It&#8217;s like cars. If cars had been invented in the classified literature, the government might say, &#8220;This is great. We can catch bad guys, we can fight wars, and we can always win them because we have tanks, we have cars, we have things like that.&#8221; And they&#8217;ve got horses. And now imagine Whit and I invent a car 50 years ago in the unclassified literature. They&#8217;d be up in arms. We wouldn&#8217;t see ambulances, we wouldn&#8217;t see commuting, we wouldn&#8217;t see vacations that people could take.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1576">26:16</a>] There&#8217;s a fundamental trade-off. It&#8217;s there with cars, and it&#8217;s there with cryptography as well.</p><h3>26:22 &#8212; The story behind Diffie Hellman key exchange</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1582">26:22</a>] So I wanted to talk about the obviously the big invention that you had, public-key cryptography, the Diffie-Hellman key exchange algorithm. What&#8217;s the story behind you inventing this Diffie-Hellman, and maybe Diffie-Hellman-Merkle key exchange algorithm? Whit showed up on my doorstep, almost literally speaking, in the fall of 1974.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1609">26:49</a>] He had been traveling around the country. He&#8217;d worked at the AI Lab, the artificial intelligence lab here at Stanford. And he&#8217;d done some work on cryptography. He wanted to do more. They could not support it. So he took a small inheritance that he got from his mother, and he traveled around the country. And if you work in AI, if you&#8217;re known, there are people that will put you up in their garages, in their homes.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1632">27:12</a>] They&#8217;ll feed you. Maybe John McCarthy lives up the hill. And Whit was a good friend of his. And there was a high school student living in his attic, I think, who was from Cambridge. And I asked him why he didn&#8217;t just go to the <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> Artificial Intelligence Lab. And he said he didn&#8217;t like their politics. I don&#8217;t know what that means. So Whit traveled around the country. And when he was at IBM Research, a secrecy order had descended on them just about this time.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1660">27:40</a>] I&#8217;d been there a few months before. They told me they couldn&#8217;t tell me very much. They told Whit, when he came through, the same thing. But they told him one thing additional: &#8220;Oh, when you&#8217;re back out at Stanford, look up Hellman.&#8221; So he called me in the fall of &#8216;74. And we set up maybe a half hour, hour meeting, max. And it went through the night. They came back here, he and Mary Fischer, his then girlfriend, now wife, then wife.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1688">28:08</a>] She&#8217;s passed away since then. And we had a real meeting of the minds. It was amazing. And so Whit and I started working together in the fall of &#8216;74. The <a href="https://en.wikipedia.org/wiki/Data_Encryption_Standard">Data Encryption Standard</a> came out. In March &#8216;75, it was announced. We came up with the idea of public-key cryptography around then. And we did not know about <a href="https://en.wikipedia.org/wiki/Ralph_Merkle">Ralph Merkle</a> at the time. Ralph was an undergraduate and later a master&#8217;s student at <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Berkeley">UC Berkeley</a>.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1717">28:37</a>] And he came up with half of public-key cryptography. He came up with public-key distribution. And I have his paper in which he. In the fall of &#8217;74, we didn&#8217;t know about him then. He was taking CS244, computer security. And he had to do a term project. And he proposes two term projects. One is the privacy half of public-key cryptography. And the other is something much more mundane. And I have the professor&#8217;s handwriting in blue scanned.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1746">29:06</a>] I&#8217;ve scanned this across the top. That idea number two, the mundane idea, looks much better than idea number one. Maybe because his description of idea number one is so muddled. But Ralph had major problems. So he drops the course. The professor says, &#8220;See me today about this.&#8221; Ralph, instead of seeing the professor, drops the course and goes and does it on his own. He then submits a paper to CACM, the primary <a href="https://en.wikipedia.org/wiki/Communications_of_the_ACM">Communications of the ACM</a>, the primary journal of the ACM, and it&#8217;s rejected.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1780">29:40</a>] I have the rejection letter. Susan Graham, who was the editor who rejected it based on a referee&#8217;s report, writes, &#8220;It bothers her that there are no references in the paper. Has no one thought of doing anything like this before?&#8221; The answer is no. Now, he still sort of had references. Ralph didn&#8217;t know how to write a paper at that point. I helped him rewrite the paper, and it was later accepted, but it was actually submitted a little before ours.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1809">30:09</a>] Whit and I came up with public-key cryptography. Ralph came up with public-key distribution. We eventually learned about one another. Afterward, I bring him here in the summer of 76, and I talk to Ralph and I say, have you ever thought of coming to Stanford to do your PhD? And Ralph says, I can&#8217;t afford it. I said, yes, you can. How about you making it Berkeley as a TA? And I said, I pay you roughly the same as an RA here at Stanford, and your tuition would be covered.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1837">30:37</a>] So he came to Stanford. He did his PhD under my supervision, supposedly, although we really worked together.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1891">31:31</a>] It seems like this cryptography space was not so well resourced. I mean, Diffie is kind of going couch to couch to.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1899">31:39</a>] To kind of support himself on this. Ralph&#8217;s also kind of doing this on his own, fully his own initiative, and people aren&#8217;t really supporting him throughout. Is that, was that common in cryptography at that time?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1912">31:52</a>] Yeah, I told you. All my colleagues, and it&#8217;s not an exaggeration, told me I was crazy to work in cryptography. And one of their reasons was <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> has a huge budget and a multidecade head start. One of them, <a href="https://en.wikipedia.org/wiki/Jim_K._Omura">Jim Omura</a>, who passed away about a year, two years ago, he was a professor at <a href="https://en.wikipedia.org/wiki/University_of_California%2C_Los_Angeles">UCLA</a>. I mentioned him to you earlier. Jim told me I was crazy. And the best ideas often seem crazy ahead of time. So I was at Tom Kailath 10 years ago.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1940">32:20</a>] I was at Tom Kailath&#8217;s 80th birthday celebration. He won the National Medal of Technology and Science, I think given to him by President Obama for his work keeping Moore&#8217;s law going for 10 years. According to the former dean of engineering at Stanford at the time, <a href="https://en.wikipedia.org/wiki/Jim_K._Omura">Jim Omura</a> is standing six feet away. I motioned him over. I said, &#8220;Jim, tell this woman what you told me when I first started working in cryptography and I just won the Turing Award, the top prize in computer science for that work.&#8221;</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=1967">32:47</a>] Jim said, &#8220;Oh, I told Marty he was crazy.&#8221; I knew he would do that because he&#8217;d given me permission to quote him on this. The best work often appears crazy a priori&#8212;<a href="https://en.wikipedia.org/wiki/Global_Positioning_System">GPS</a>, which we use daily. Brad Parkinson, who won the top prize from the <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> for that work, says that the Air Force thought <a href="https://en.wikipedia.org/wiki/Global_Positioning_System">GPS</a> was crazy initially. High-speed Internet access. I won&#8217;t go into details there. And I&#8217;ve spoken to a number of Nobel laureates, and almost all of them had the same reaction from people that I did.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2002">33:22</a>] One of them, I&#8217;ll tell you one story. <a href="https://en.wikipedia.org/wiki/Louis_Ignarro">Louis Ignarro</a>, who won the Nobel Prize in physiology or medicine, I think, in 1998, told me that the dean of his medical school came to him and said, Louis, why are you doing this crazy stuff? We hired you because you do good work. Crazy stuff that won him a Nobel Prize.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2023">33:43</a>] It makes sense. But I guess the unique aspect is I think a lot of people would receive that pushback and everyone thinks they&#8217;re foolish, and they wouldn&#8217;t go for it. What is it that gave you the confidence to continue when everyone was saying this is foolish?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2044">34:04</a>] Two reasons. One is I got beaten up as a Christ killer as a kid. And that&#8217;s a little bit of a joke, but not totally. As a kid, we moved to an Irish Catholic neighborhood. I&#8217;m Jewish. We moved there when I was four and a half years old. I went to my parents and I said, &#8220;I want the pretty tree that the other kids have, the Christmas tree. I want the presents under the tree. I want to go to the same school as the other kids in the neighborhood.&#8221;</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2067">34:27</a>] <a href="https://en.wikipedia.org/wiki/St._Nicholas_of_Tolentine_High_School">St. Nicholas of Tolentine Parochial School</a>. I was telling my parents I didn&#8217;t want to be Jewish. And so, in self-defense, I adopted the attitude of who would want to be like anybody else? That had a very simple answer. Me. I would have given my eye teeth to have been like everybody else. But in time, I came to believe my own bullshit. And so when my colleagues said I was crazy to work in cryptography, instead of scaring me off, it actually attracted me.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2098">34:58</a>] So that&#8217;s one reason. And the other reason, I think it runs in my family. I have an uncle who in 1940, with his wife, a Jew, took trains across the country with bicycles, bicycling through national parks. I mean, who but a crazy person would do that? I come from a crazy family. We&#8217;ve done crazy things. And this was a crazy thing to do.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2123">35:23</a>] When you put out that paper in 1976, what was the problem you were trying to solve, or what was guiding your research direction at that time?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2133">35:33</a>] Well, it started as trying to develop a theory of cryptography. And <a href="https://en.wikipedia.org/wiki/Jim_Massey">Jim Massey</a>, who was the editor at the time of the <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> Transactions on Information Theory, of which I was on the board of governors, as I mentioned, Jim invited me to write a paper on that. I mean, this was beginning to get some traction. And I asked if Whit could be included. And he said absolutely. And so Whit and I were working on this paper.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2163">36:03</a>] We had several drafts. Then in May 76, I came up with what is now Diffie-Hellman key exchange. And I included that. I threw that into a paper I&#8217;m giving in June 76, a month later at the <a href="https://en.wikipedia.org/wiki/Institute_of_Electrical_and_Electronics_Engineers">IEEE</a> Symposium on Information Theory in Ronneby, Sweden. <a href="https://en.wikipedia.org/wiki/Jim_Massey">Jim Massey</a>&#8216;s there, of course, and he says, you get that into the paper, I&#8217;ll have it out in the November issue, which he did. Now, the November issue, Whit is fond of pointing out, did not come out until February 1977.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2191">36:31</a>] Was usually late, but that&#8217;s how it worked. And Ralph, unfortunately, as I said, had his idea rejected, his paper rejected. So his paper did not appear until 1978, even though it was submitted slightly before ours. I see.</p><h3>36:48 &#8212; Signatures vs key exchange</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2208">36:48</a>] I think in this conversation you mentioned key exchange, also signatures. It seems like there are these components of a full cryptographic system. Could you explain each of these?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2222">37:02</a>] Key exchange is how do you and your bank exchange a key to protect your banking transactions, to use a modern example, with somebody listening on the Internet, and so use public key cryptography to do that <a href="https://en.wikipedia.org/wiki/Ralph_Merkle">Ralph Merkle</a> had a puzzle method. <a href="https://en.wikipedia.org/wiki/Ralph_Merkle">Ralph Merkle</a> had a puzzle method where he generated a sequence of puzzles and you solved. Your bank would solve one of them and it would then use that key. And you could quickly tell which key it was using.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2253">37:33</a>] But somebody else had to solve half the puzzles, on average, before it could do it. And so it took a lot more effort on the part of an opponent. I&#8217;ll explain Diffie-Hellman key exchange very briefly. And that is only key exchange. It&#8217;s not yet signatures. You and I want to exchange physical messages. I want to write them out to you and pass them to you. But my wife, who is not supposed to see them, is sitting between us and has to hand them to you.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2286">38:06</a>] So what do I do? I put my message in a strong box. I put a lock on that strong box. I&#8217;d like to tell you the combination to that lock. But if I tell you how to open it, I&#8217;m also telling my wife how to open it. So I don&#8217;t tell anybody how to open it. I pass it to my wife. She can&#8217;t open it. She passes it to you. You can&#8217;t open it. You put a second lock on the strong box. You then pass the double. You don&#8217;t tell me the combination.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2311">38:31</a>] You don&#8217;t tell Dorothy the combination. Only you know the combination. You pass it to Dorothy. She can&#8217;t take either lock off. She passes it to me. What can I do? Take off my lock, leaving only your lock. I pass it back to Dorothy. She can&#8217;t take off your lock. But when you get it, you can. You get the message inside. That&#8217;s basically how Diffie-Hellman-Merkle key exchange works.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2335">38:55</a>] Now, what&#8217;s critical about this is that I made the strongbox big enough for you to put a second lock on it. If I hadn&#8217;t done that, if I&#8217;d only made it big enough for one lock, you could have taken the strongbox that you got and put it in a bigger strongbox. But now there&#8217;s a problem. When I get it, I can&#8217;t get inside to take my lock off. It has to be what&#8217;s called commutative. And everybody knows what commutative means, although they may have forgotten it.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2360">39:20</a>] Addition is commutative. Three plus five is the same as five plus three. It doesn&#8217;t matter which order you do the operations in. You get 8. 5 minus 3 is 2, 3 minus 5 is minus 2. You get a different answer. It&#8217;s not commutative. Subtraction is not commutative. What I had to do was find a commutative one-way function, although I didn&#8217;t know that at the time. The function I used is a commutative one-way function.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2387">39:47</a>] It&#8217;s exponentiation and modular arithmetic. It doesn&#8217;t matter whether you first raise alpha to the x1 power and then to the x2 power or raise it to the x2 power and then to the x1 power. You get the same result. You get alpha to the x1 x2; multiplication in the exponent is commutative. And we did this over a finite field where exponentiation is not a smooth function, which would be easy to invert. It jumps all over the place.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2415">40:15</a>] But it turns out it still works that way. So that is public key exchange, public key distribution. Digital signatures are more complicated. That&#8217;s something that <a href="https://en.wikipedia.org/wiki/Whitfield_Diffie">Whit Diffie</a> came up with. It&#8217;s called a public key cryptosystem. And then <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a>, <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">Rivest, Shamir, and Adleman</a> came up with the first instance of that. There you have a public key and a secret key, and it doesn&#8217;t matter which order you operate on them. Whether you first use the public key and then the secret key or the secret key and then the public key.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2442">40:42</a>] You get the initial result, they undo one another there. If I use my secret key, I can sign a message because there&#8217;s no privacy. But I send the encrypted message to you using my secret key. You use my public key, which you get from a public file, to verify it. That&#8217;s how my phone works. It&#8217;s got a public key built into it. Apple knows the secret key. Apple can sign software updates. People can take my phone apart and get the public key.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2473">41:13</a>] But that doesn&#8217;t give them the secret key that allows them to sign software updates.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2478">41:18</a>] I see. So in that case, Apple is signing their updates to your phone.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2484">41:24</a>] Right.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2485">41:25</a>] So I also was reading about symmetric versus asymmetric, and what you&#8217;re describing sounds like asymmetric.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2491">41:31</a>] Yes. Asymmetric cryptography uses a public key and a secret key. It&#8217;s asymmetric. Conventional cryptography uses the same secret key to encrypt and decrypt. And that requires a courier, which is what was required before public key cryptography. If you were a general in another division, I could send you a key. I would send a messenger with a key locked to his wrist. And when you got it, you could open it up and get the key and then you and I could use a radio channel to communicate very cheaply and very fast.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2525">42:05</a>] But you can&#8217;t use couriers to communicate keys between a bank and a client.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2531">42:11</a>] Yeah. So I saw modern systems. It&#8217;s almost like multiple different mechanisms where there&#8217;s this asymmetric key exchange at the beginning to establish the connection. And then that gets you that key symmetric thing.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2546">42:26</a>] Exactly. It&#8217;s just like I described. First, Apple signing the software update. They don&#8217;t sign the whole thing. They only sign a hash. In the same way, we could use a public key cryptosystem like <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> to exchange messages, but it&#8217;s too slow. So what I do is I send you one message, use the following key in <a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard">AES</a>, the <a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard">Advanced Encryption Standard</a>, and then we use <a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard">AES</a> very quickly to exchange messages. And that&#8217;s the difference between asymmetric encryption, which we use at first, where I use the public key to encipher the message and use the secret key to get the key, and then we use a symmetric system to very quickly communicate afterward.</p><h3>43:05 &#8212; RSA patent wars</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2585">43:05</a>] I was reading about <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> versus the Diffie-Hellman key exchange, and it seems like there were a lot of patents or something like that. Patent wars. What&#8217;s the context behind that?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2598">43:18</a>] I didn&#8217;t know how patents worked when we wrote the first patent on public key cryptography. If I&#8217;d known we could have covered <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a>, we would have made a lot of money. Basically, <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> sold their company for $250 million. We made nothing, virtually nothing, from our patents in cryptography. There was a patent fight between <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> and us, <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> and Stanford, or between their licensees. And basically <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> won that fight.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2635">43:55</a>] And I was pissed at <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> for a long time, and I was pissed with Jim Bidzos, the president of <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> Data Security. And around 1990, around 2000, sometime around there, I realized it was inconsistent with my approach to life to be pissed at <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a>. They&#8217;d won. The fight was over. And so I approached Jim Bidzos and I said, look, Jim, I told him that, and I said, let&#8217;s be friends. Let&#8217;s bury the hatchet. And we essentially have.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2664">44:24</a>] And friends are better than enemies. I&#8217;m not sure, but I think Ron Rivest might have actually nominated <a href="https://en.wikipedia.org/wiki/Whitfield_Diffie">Whit Diffie</a> for the Turing Award. Now, he wouldn&#8217;t have done that if he didn&#8217;t think we deserved it, but he wouldn&#8217;t do it if he was pissed at us. And friends are better than enemies. That&#8217;s not why I did it. That&#8217;s not why I became friends with them. But again, it worked out very well.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2688">44:48</a>] How did they win the patent war if your ideas came first?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2694">44:54</a>] I remember I told you I didn&#8217;t know how patents worked. It&#8217;s only the claims that matter. And so if I&#8217;d known that, we could have written a claim that would have covered <a href="https://en.wikipedia.org/wiki/RSA_cryptosystem">RSA</a> easily, but we didn&#8217;t. And so the patent was written very badly. Stanford didn&#8217;t want to invest a lot of money in this patent, so it had a law school intern write the patent instead of a patent attorney. We made a lot of mistakes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2723">45:23</a>] You mentioned friends are better than enemies. Yeah, a lot of people would agree with you, but I also think most people wouldn&#8217;t be able to get over losing a $250 million outcome. How did you kind of get over that and establish friendship?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2743">45:43</a>] Well, some of it is in my genes, but a lot of it comes from my wife, basically. My wife and I have been married 59 years last March, and we were madly in love when we met. We followed society&#8217;s rules, and we developed a toxic relationship. Ten years later, my wife was ready to leave me, but I didn&#8217;t know it because I had blinders on, the way many husbands do. And fortunately, when she met me, she decided I was the one.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2772">46:12</a>] And 10, 15 years after we&#8217;re married, roughly 50 years ago, she says he&#8217;s still the one, although life with him is impossible, and life with her was no picnic. So she went around looking for catalysts, ways to improve our relationship, and she eventually found a group. I won&#8217;t go into all the details there. I tend to do that. That worked on the international and the interpersonal at the same time.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2799">46:39</a>] It was founded by Professor Harry Rathbun, who was a professor here at Stanford, born in the 1890s. So he&#8217;s no longer alive. I knew him late in his life, and that&#8217;s how I. It was based on the teachings of Jesus, which was a problem for me as a Jew because it was okay not to go to synagogue, which I didn&#8217;t do, but it was not okay to study the teachings of Jesus. That was traitorous. But eventually, Dorothy dragged me to enough meetings that over a year&#8217;s time, I came to see that these people, Creative Initiative, as the group was called at the time, knew something I had to learn if my marriage was going to survive.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2837">47:17</a>] And I surprised myself by being willing to do things that seemed crazy to me. The most important things were accepting ideas that Dorothy had that seemed crazy to me, that weren&#8217;t crazy. Because what happened when she had an idea that seemed crazy to me? I treated her like she was crazy. What did that do? It drove her crazy. What did that do? It convinced me I was right, that she was crazy. Kept the whole cycle going.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2863">47:43</a>] By the way, the same happens internationally. We treat countries in ways that they don&#8217;t like. They react in ways that seem crazy to us. We treat them like they&#8217;re crazy, and it keeps the whole cycle going. So that&#8217;s how I came. And so it&#8217;s based on the Gospels. And while I&#8217;m not a Christian, I see Jesus as a Jewish reformer rather than as a Christian messiah. I view myself as a follower of Jesus.</p><h3>48:08 &#8212; Why inventions happen at similar times</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2888">48:08</a>] When I was researching the discovery that you had, it seems like all of a sudden many of the similar discoveries were happening all at the same time. So, for instance, you mentioned <a href="https://en.wikipedia.org/wiki/Ralph_Merkle">Ralph Merkle</a> and us. Yes.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2902">48:22</a>] Oh, and <a href="https://en.wikipedia.org/wiki/GCHQ">GCHQ</a> claims that they invented it just a couple of years before us, but they only have half of it, which is, if they&#8217;re right, which is the privacy part. They didn&#8217;t have anything on digital signatures, but they claim everything.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2916">48:36</a>] Right. And then there&#8217;s also the <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> folks.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2918">48:38</a>] The <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> folks. Yeah. So I have a theory about that. I&#8217;m going to tell it as a joke, but there&#8217;s more to this joke than I think is just a joke. There&#8217;s a muse that whispers in our ears. There&#8217;s a muse of poetry. There&#8217;s a muse of calculus who whispered in Newton&#8217;s ear and Leibniz&#8217;s ear about the same time. There&#8217;s a muse that whispered in Ralph&#8217;s ear about the same time that she whispered in our ears.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2945">49:05</a>] Most people don&#8217;t pay attention to this muse because she sounds crazy. A few people do. And so that&#8217;s why I think there&#8217;s something in the air. I mean, it&#8217;s not necessarily a muse, but there&#8217;s something in the air that comes to people.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2960">49:20</a>] But why didn&#8217;t it come 50 years earlier? Or how come it seemed like they all came very similar? I noticed. I interviewed <a href="https://en.wikipedia.org/wiki/Barbara_Liskov">Barbara Liskov</a> as well, who did a lot in data abstraction and modularity. And there also, it was like on the West Coast and the East Coast, without even communicating, they had the same discovery. Very similar.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=2986">49:46</a>] I think my joke might actually have something to say about that. But also there&#8217;s something that goes on technologically. We didn&#8217;t have the computing power to do public-key cryptography. If someone had come up with it 50 years before, it would have been a nice idea that couldn&#8217;t be implemented.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3000">50:00</a>] So there&#8217;s a logistical part to this, which is just the tools weren&#8217;t there, and once they were there, it kind of inspired the right thinking.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3010">50:10</a>] Yeah, but calculus, I mean, why? That occurred to Leibniz and Newton about the same time. Oh, and Darwin and someone else thought of evolution about the same time. So, I don&#8217;t know. It is mystical, and maybe we should treat it as that. Even though I&#8217;m a scientist, I believe in the mystical side of life.</p><h3>50:29 &#8212; What he worked on after cryptography</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3029">50:29</a>] There&#8217;s a cryptography work that you&#8217;re doing. And then later in your career, I saw this rethinking national security, and I was curious how you got into this, I guess, area, this body of work. And what&#8217;s the problem you&#8217;re trying to solve?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3046">50:46</a>] Well, the problem we&#8217;re trying to solve is simple. We&#8217;re going to kill ourselves. I mean, right now, I estimate that a child born today has probably worse than even odds of living out his or her natural life as a result of nuclear weapons all by themselves, without climate change, without AI, without any of this other stuff.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3065">51:05</a>] What are you talking about with this AI threat?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3067">51:07</a>] Well, there&#8217;s debate on that. Geoffrey Hinton, Yoshua Bengio, who won the Turing Award for their work in artificial intelligence, think that AI may actually kill human beings off, something like that, because it&#8217;ll say, &#8220;Why do I need these stupid people?&#8221; Whereas Ed Feigenbaum and Raj Reddy, who won the Turing Award 30 years ago for their work in artificial intelligence, have both told me and given me permission to quote them.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3098">51:38</a>] So I&#8217;m not giving any secrets away that calling AI an existential threat is ridiculous. It&#8217;s science fiction. I don&#8217;t know who&#8217;s right, but I do know we need to build a more cooperative world for a number of other reasons, even if not that one. So why not solve that one? At the same time, we&#8217;re likely to build AI into our nuclear command and control system. Oh, my God, what a mess. Right now, I&#8217;m working on a bill.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3124">52:04</a>] It&#8217;s H.R. 3564, and we have 25 on paper, 24 co-sponsors. The bill is so bad it&#8217;s passable, but so good that it&#8217;s a game changer. Why is it so bad that it&#8217;s passable? It has an exception for launch on warning. If our warning system shows a bunch of ICBMs coming our way, the President doesn&#8217;t need a second opinion to launch our nuclear weapons. He can do so all on his own. There have been mistakes in our warning system, so I don&#8217;t like that.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3157">52:37</a>] But that&#8217;s why it&#8217;s so bad. It&#8217;s passable. This makes it easy for people to get on board. Most people don&#8217;t realize that the president can launch a nuclear war all on his own. Right now, a president could be awakened at 3 a.m. in the morning during a crisis, too drunk to legally drive a car, asked what to do, he might say drunkenly, nuke him. And right now, theoretically, the military&#8217;s supposed to do that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3182">53:02</a>] I didn&#8217;t consider that AI could be coupled with the nuclear bombs too. I guess AI could.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3187">53:07</a>] Oh, it&#8217;s likely to be. I mean, we have this problem that we have just minutes of decision time. Do you not think we&#8217;re going to use computers to help us there?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3198">53:18</a>] Looking back on your career, being at Stanford during the growth of computing, I imagine you had a lot of opportunities to work with other famous folks in the industry. For instance, Don Knuth.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3213">53:33</a>] Don Knuth is a good friend.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3214">53:34</a>] Yeah.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3215">53:35</a>] Don Knuth doesn&#8217;t use email. I think he does, but whenever I want to reach him, I have to call his home and talk to his wife, Jill. That&#8217;s the other story. Stories that I have John McCarthy at AI Lab up the hill, where Whit stayed there, by the way, when we were working on public-key cryptography because John was away and Whit and Mary were taking care of his daughter, driving her to her writing lessons.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3243">54:03</a>] And so Whit just walked down the hill to talk to me. But John McCarthy, there&#8217;s a meeting at the faculty club maybe 45 years ago. Bob Floyd, who many of your people will know, who&#8217;s a famous&#8212; he died, but a famous computer scientist here at Stanford. John McCarthy comes in. Bob Floyd&#8217;s in the meeting. He goes over to Bob Floyd. I say, hi, John. He doesn&#8217;t pay any attention to me because everybody knows that I.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3270">54:30</a>] <a href="https://en.wikipedia.org/wiki/Input%2Foutput">I/O</a> is a waste of time. And so he&#8217;s not going to waste time on <a href="https://en.wikipedia.org/wiki/Input%2Foutput">I/O</a>. So he has a very rapid exchange with Bob Floyd and he leaves. Well, I&#8217;ll tell you another thing. I&#8217;m finishing my... It&#8217;s 1968. Several months early in &#8216;68, I thought, who am I to think I can make an original contribution to knowledge, which is the definition of a PhD thesis? I&#8217;m thinking of dropping out of the program over spring break.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3301">55:01</a>] I basically solved my thesis. It&#8217;s that quick. I go to Tom Cover, who&#8217;s since died, my advisor and very famous information theorist. And I go to Tom and I say, I just have barely enough units, but I have enough units to meet the course requirements for the PhD. But I don&#8217;t have enough units to meet the PhD requirements. I&#8217;m going to have to take units of research if I work at IBM Research. Would you give me credit for my units of research?</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3334">55:34</a>] Because my wife was pregnant with our first child then. We&#8217;d been married roughly a year, year and a half. I was concerned about money. He says, sure. And so I go to IBM Research, Yorktown Heights. Tom wants to have me come teach at Stanford, which I&#8217;ve done. But he says, &#8220;You really should leave for a couple of years because Stanford.&#8221; I sometimes joke that I had to leave for several years to lose the taint of the Stanford PhD.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3364">56:04</a>] Why is that? Stanford was worried about inbreeding. If they hired me, I was just going to be doing research like Tom Cover. And they don&#8217;t need two of them. They need somebody different. And so I went away to IBM for a year because my thesis broke very suddenly. And I then taught at <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a> for two years before coming back on the faculty in &#8216;71. And this seemed like just a stupid hoop I had to jump through.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3386">56:26</a>] But it actually turned out to be very important because when I was at IBM, <a href="https://en.wikipedia.org/wiki/Horst_Feistel">Horst Feistel</a>, whose name many people may recognize, he was really the father of IBM&#8217;s research in cryptography. He was hired in &#8216;68. He was in the same department as me. I didn&#8217;t work in cryptography, but I had lunch with Horst, and he told me some of the amazing things that got me started. And then when I&#8217;m at <a href="https://en.wikipedia.org/wiki/Massachusetts_Institute_of_Technology">MIT</a>, Peter Elias, whose name many people today will not recognize, but he was one of the original contributors to information theory.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3420">57:00</a>] Peter gives me a paper by Claude Shannon, whose name your audience probably will recognize, written in 1949. I&#8217;m sorry, published in 1949 in the Bell System Technical Journal, connecting information theory to cryptography. Which is when I realized that, oh, I&#8217;ve already done a PhD in information theory. Maybe I can do something worthwhile in cryptography. So that stupid hoop I had to jump through was actually very important because when I got back to Stanford, I started working in cryptography, among other things.</p><h3>57:31 &#8212; His thoughts on death</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3451">57:31</a>] I think at the beginning of the conversation, you said something quickly that you had thoughts on death. And I was curious, what are your thoughts on death and how does it impact you?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3464">57:44</a>] And I got an email from somebody at The New York Times about a month ago saying that he&#8217;s working with my obituary, and it made me macabre, but he was wondering if I could read it over and if I have any comments. And so, the two things are that I&#8217;m 80 years old, so my older brother died just after he turned 80. My mother died at 70. My father lived to 98. I hope I&#8217;ll have more years, but if I don&#8217;t, that&#8217;s okay.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3490">58:10</a>] I&#8217;ve had a great ride. But the two things that concern me about dying are the people I&#8217;ll leave behind, who I think are dependent on me but probably aren&#8217;t, and the work I do because very few people are working on the nuclear issue, the growth of humanity as a whole to save itself, that we&#8217;re headed for disaster. Very few people realize that. And very few people put it in terms of a process of change.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3518">58:38</a>] We can&#8217;t stay where we are. We&#8217;re going to kill ourselves, we&#8217;re going to die. We can&#8217;t jump to where we need to go. The world is too dangerous. But we can get there via a process of change. I&#8217;ll tell you one other thing. Humanity is going to die. I&#8217;ve come to this realization. I mean, it&#8217;s clear. I mean, we might last a million years, but at some point. We might last a billion years, but at some point the sun&#8217;s going to burn out.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3542">59:02</a>] Maybe we&#8217;ll move to other galaxies and we&#8217;ll find ways to cheat death for a while. But eventually humanity is likely to die. But if humanity were to die now, it would be an infant. It would be infant mortality. It would be really sad. If we&#8217;ve lived a full life, that&#8217;s something else. If I had died when I was a child, people would say, oh, what have we missed? If I die at 80, people will say, some people will be saddened by it, but.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3571">59:31</a>] But they&#8217;ll feel like I&#8217;ve lived a full life. We need to save humanity from an infant death. We&#8217;ve only been around about 100,000 years.</p><h3>59:40 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3580">59:40</a>] With all the experience that you have now, if you could go back to when you had just graduated college and give yourself some advice, what would you say?</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3589">59:49</a>] If when I was just graduating college and I was interested in saving the world, which I was not, I might have gone up to the top of a mountain to a guru and asked him what to do. Imagine what you&#8217;d say. You&#8217;d expect him to tell you, work for the UN or something like that. Imagine what you&#8217;d say if he told me, forget about saving the world, which was not a problem because I didn&#8217;t think about saving the world.</p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3613">01:00:13</a>] Go out and become as famous as you can, make as much money as you can, come back in 10 years, and I&#8217;ll tell you what to do. That would have seemed like a detour, but it was actually the shortest distance between two points. Because 10 years later, when I was 30, 35 years old and I was interested in saving the world, I had a reputation, I had money, and I had to screw up to get where I am. I had to not be concerned with the world to become concerned with the world to be able to do something, to do what I can do about it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3640">01:00:40</a>] Awesome. Well, thank you so much for your time today. I really appreciate it.</p><p><strong>Martin:</strong></p><p>[<a href="https://youtu.be/AZLOETBCQM4?t=3643">01:00:43</a>] You&#8217;re welcome. And thank you.</p>]]></content:encoded></item><item><title><![CDATA[MIT Complexity Theorist: Why You Can Do Better Than “Optimal” On Leetcode & SAT | Ryan Williams]]></title><description><![CDATA[Ryan Williams is a professor at MIT and the winner of the G&#246;del Prize in theoretical computer science.]]></description><link>https://www.developing.dev/p/mit-complexity-theorist-on-leetcode</link><guid isPermaLink="false">https://www.developing.dev/p/mit-complexity-theorist-on-leetcode</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 29 Jun 2026 10:02:33 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202988431/2bea21595ad6754b68123d0352319f3f.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/r-ryan-williams-a1b534a/">Ryan Williams</a> is a professor at MIT and the winner of the G&#246;del Prize in theoretical computer science. I interviewed him all about his work starting by asking him a popular Leetcode question (3 SUM).</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/AaK1SL2i_4Y">YouTube</a>, <a href="https://open.spotify.com/episode/0JH8HW6p03BkzeRaV6zNbD?si=cWge3ofuTW-BUn419wggew">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-AaK1SL2i_4Y" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;AaK1SL2i_4Y&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/AaK1SL2i_4Y?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/202988431/0041-asking-him-a-popular-leetcode-question">00:41 - Asking him a popular Leetcode question</a></p><p><a href="https://www.developing.dev/i/202988431/0354-doing-better-than-the-popular-optimal-solution">03:54 - Doing better than the popular optimal solution</a></p><p><a href="https://www.developing.dev/i/202988431/0826-fine-grained-complexity">08:26 - Fine grained complexity</a></p><p><a href="https://www.developing.dev/i/202988431/1700-a-severe-strengthening-of-p-vs-np">17:00 - A severe strengthening of P vs NP</a></p><p><a href="https://www.developing.dev/i/202988431/2438-sat-problems-and-solvers">24:38 - SAT problems and solvers</a></p><p><a href="https://www.developing.dev/i/202988431/3451-hot-takes-on-famous-open-questions">34:51 - Hot takes on famous open questions</a></p><p><a href="https://www.developing.dev/i/202988431/4657-simulating-space-with-time">46:57 - Simulating space with time</a></p><p><a href="https://www.developing.dev/i/202988431/010102-why-he-solves-hard-problems">01:01:02 - Why he solves hard problems</a></p><p><a href="https://www.developing.dev/i/202988431/010235-how-to-pick-good-research-direction">01:02:35 - How to pick good research direction</a></p><p><a href="https://www.developing.dev/i/202988431/010714-technical-book-recommendations">01:07:14 - Technical book recommendations</a></p><p><a href="https://www.developing.dev/i/202988431/010831-advice-for-his-younger-self">01:08:31 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:41 &#8212; Asking him a popular Leetcode question</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=41">00:41</a>] Okay, I want to start by asking you the most popular <a href="https://en.wikipedia.org/wiki/LeetCode">LeetCode</a> question. So the question is <a href="https://en.wikipedia.org/wiki/3SUM">3SUM</a>: given a list of numbers, we want to find three numbers such that they sum to zero.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=58">00:58</a>] And so what are your thoughts on the brute force solution for this? We can start there.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=64">01:04</a>] So the obvious brute force solution takes, if you&#8217;ve got n numbers, N cubed time, just try all the triples of numbers, sum them up, see if they sum to zero. There is a faster solution. So one way to get an order N squared time algorithm for <a href="https://en.wikipedia.org/wiki/3SUM">3SUM</a> is to first start by sorting the numbers. And then you go through the numbers one by one. Say you&#8217;re looking at a number A, and you want to know, is there a B and a C in the rest of the list whose sum with A is going to be zero?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=105">01:45</a>] Okay, so the way this works is after you sort the numbers, you do what&#8217;s called a finger search. So you put the finger from your left hand on the minimum element and a finger from your right hand on the maximum element. So you start there and you check, okay, are these my B and C? Right? So you add them, min and max, and check if adding that with A gets you zero.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=137">02:17</a>] And if you&#8217;re lucky, then you&#8217;re done. But typically you&#8217;re not lucky. And so this sum of the min and the max is either larger than your target value minus A, or it&#8217;s smaller. If it&#8217;s larger, then you need to decrease the larger number. So you take your right finger, which is sitting on the maximum element, and you move it to the left one slot. So you decrease the larger one. All right?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=167">02:47</a>] If the sum is smaller than your target, you need to take the smaller number and make it a little bit bigger. So you move your left finger, sitting on the minimum, over one slot. Okay? And you keep doing this. You keep checking whether your left finger and right finger are pointing at a solution. And if they aren&#8217;t, then you adjust it. If they do, if they ever do sum up to exactly what you want, you&#8217;re done.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=194">03:14</a>] And so after each comparison, right, one of your fingers moved, okay? If the fingers ever cross, then you don&#8217;t have a solution. There just can&#8217;t be a solution. And so the number of times you move your fingers in total is like N. So you have order N solution for finding that extra pair. And you do this for each of the numbers A, okay? So then you get an order N squared solution overall, you do it N times.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=226">03:46</a>] Each finger search takes order N time.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=230">03:50</a>] And that is the popular solution.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=233">03:53</a>] That&#8217;s the popular solution.</p><h3>03:54 &#8212; Doing better than the popular optimal solution</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=234">03:54</a>] A lot of people know when we&#8217;re in these algorithmic <a href="https://en.wikipedia.org/wiki/LeetCode">LeetCode</a> interviews, and I know a lot of your research is kind of about pushing lower bounds. So maybe we can start that conversation off by asking: can you do better than N squared for this?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=250">04:10</a>] Yeah, you actually can do better than N squared. And this is not at all obvious. In fact, it stems from taking this finger search idea and pushing it in a different direction. So what you do is you take your sorted list, okay, and you break the sorted list up into little groups of contiguous elements. So let&#8217;s say your little group is like log n size or square root log n. It&#8217;s like a really small little group, okay?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=282">04:42</a>] So you&#8217;ve got either n over log n groups total or n over square root log n groups total, depending on how you break up your groups. And then the idea is you&#8217;re going to perform the same kind of finger search, but you&#8217;re going to set things up so that you&#8217;re comparing two pairs of groups. Your finger&#8217;s always pointing at an entire group. So the left hand and right hand are pointing at two groups.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=308">05:08</a>] And you want to know if there&#8217;s a <a href="https://en.wikipedia.org/wiki/3SUM">3SUM</a> solution in that group. And you can set up a kind of fast data structure to check a small group, okay? This data structure will take much less than the number of elements in the two groups squared. So you set up some kind of fancy data structure. And because you set your group size so small, it&#8217;s like a preprocessing that you do over all the possible inputs you could send from a pair of groups.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=340">05:40</a>] And so you have some data structure, and it will, let&#8217;s say it takes you \(n^{1.5}\) time to prepare this fancy data structure. But now when you&#8217;re looking at a pair of groups, you can look up the answer much faster than what finger search would have taken. I guess finger search through a group of length \(g\) and another group of length \(G\) would take about \(O(G)\) time, and you can do actually faster by this kind of lookup, this kind of table lookup.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=368">06:08</a>] I think it is kind of like you take this list of length n and you kind of shrink it into like N over group size number of things. And these are more complicated objects. And then you do, and now you&#8217;re trying to speed up the check over these small, complicated objects for like, should I move my finger to the right? Is there nothing in this group? Would finger search just go straight through this group or not is basically what you&#8217;re asking.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=405">06:45</a>] So the unit that you&#8217;re operating on is not a single integer, it&#8217;s a group.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=409">06:49</a>] Yeah, it&#8217;s a group of them. So you use some kind of table lookup. Well, it&#8217;s much fancier than a table lookup. Actually, it goes through some other model called the linear decision tree model. So it&#8217;s like in some weird model where you can actually get a faster <a href="https://en.wikipedia.org/wiki/3SUM">3SUM</a> solution. You can get it into the 1.5 solution. It&#8217;s a really interesting and sophisticated solution.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=431">07:11</a>] But what I want to emphasize is that it starts from the finger search solution and sort of figuring out how to process finger moves faster, sort of do preprocessing so that finger moves can go faster.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=446">07:26</a>] Yeah, I saw the time complexity of this. It&#8217;s N squared divided by log n divided by log log n, divided by all raised to the 2/3. What is there any intuition behind that? I mean, that&#8217;s just crazy.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=460">07:40</a>] So there are several algorithms of this kind, and they all work by doing some modification on what I was talking about because you can sort of reduce to a different model, like a different kind of lookup table, a different kind of set of tricks. And maybe there are some savings you can do here and there by sort of compressing things a little differently. So yeah, there are several algorithms that beat the n squared running time bound, and they all, to my knowledge, kind of work in a similar type of way.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=494">08:14</a>] They&#8217;re taking this n-squared time algorithm, finding little ways to preprocess and then optimize based on the preprocessing, like make finger searches faster and things like this.</p><h3>08:26 &#8212; Fine grained complexity</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=506">08:26</a>] Yeah, a lot of your research is on this topic of fine-grained complexity, or kind of lowering lower bounds. So maybe you can explain what is lowering lower bounds.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=520">08:40</a>] The idea behind fine-grained complexity is we have a variety of problems, canonical problems, that we teach to undergrads. We have canonical algorithms for these problems. These algorithms have resisted any major improvements in decades. And so we wonder, are these algorithms optimal? And what does the theory of optimality look like in terms of time complexity? So what if we just focus solely on the time complexity of a problem, like the fastest algorithm that will solve that problem?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=562">09:22</a>] What does complexity look like then? Because P versus NP is not about time. I mean, it&#8217;s about time complexity. But on a coarse-grained level, where P is just polynomial time, that polynomial could be n to the 10, n to the billion, whatever. And so showing that something&#8217;s not in P is showing that it needs some superpolynomial amount of time. Whereas here we are concerned with there&#8217;s a canonical problem, it takes N cubed time.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=593">09:53</a>] With some very elegant canonical algorithm, we want to know, could you do any better? Could you improve that exponent to n to the three minus epsilon for some epsilon? And then you ask, well, suppose I have a problem over here and it has a quadratic time algorithm, and I want to know if I can improve that quadratic time algorithm by a little bit. Then you start asking questions like, well, suppose I improve this algorithm by a little bit.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=625">10:25</a>] Can I improve this algorithm over here by a little bit? And this is naturally a notion of reduction, like saying that I want to have some reduction from problem A to problem B, so that if I can improve the algorithm for problem B just a little bit, then I can also improve the algorithm for problem A by just a little bit. This actually leads to a different notion of complexity. You can even take an NP-complete problem and a P problem and reduce the NP problem to the P problem, and the question still makes sense.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=660">11:00</a>] So, for example, you could talk about the <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> problem, okay? So the <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> problem, you&#8217;ve got n numbers and a target value, and you want to know if there&#8217;s a subset of those numbers that sum to a particular target value. Okay, now you&#8217;ve got 2^n possible subsets. The obvious algorithm takes 2^n time. Okay? But you can actually do better than this. And the way you do better than this is to reduce to a polynomial-time solvable problem.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=701">11:41</a>] So you can get an algorithm which runs in square root of 2 to the N time. For the <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> problem, you can avoid enumerating over all the possible subsets in a substantial way. And this is, in fact, known as commonly used in <a href="https://en.wikipedia.org/wiki/Cryptanalysis">cryptanalysis</a> by, I guess, a <a href="https://en.wikipedia.org/wiki/Meet-in-the-middle_attack">meet-in-the-middle</a> type approach. So people do this sort of thing all the time. The idea is you partition the set of all the numbers into two halves, like n over two, n over two.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=736">12:16</a>] You enumerate all the subset sums on the two halves. So you have 2 to the n over two possible sums for the first half, two possible sums for the second half. Then you want to know, is there a number from the first half, from this huge list, plus another number from the second half, this huge list, that sums to the target. Now this is the 2SUM problem. This is really nothing more than a 2SUM problem, which you can solve by sorting and binary search.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=771">12:51</a>] And this is a reduction. We have shown how to solve a <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> problem, which would normally take 2 to the n time, in square root of 2 to the N time, an NP-complete problem, by reducing it to a problem like 2SUM. The obvious algorithm there takes n squared time, trying all the pairs of numbers to see if they sum to a target, and using an n log n time algorithm for that. So fine-grained complexity can relate problems that would not be relatable at all in the traditional P vs NP theory.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=808">13:28</a>] Like, one problem should be complete, the other problem&#8217;s not, so they should not have a polynomial-time reduction between them in general. Right. But if you just look at the time complexity and you focus on, okay, I have an algorithm, 2 to the n algorithm, is it the best possible? And then I have an n squared algorithm, is it the best possible? And then you can relate the two problems, and this is a general phenomenon that you can relate problems that look like they should have nothing to do with each other.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=838">13:58</a>] So you talked about reducing a <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> to 2SUM, but <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> is arbitrary integers that sum to the target sum, right?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=851">14:11</a>] Yeah. So the idea is, nobody said I had to use a polynomial-time reduction or something like that to reduce one problem to another. So what happened? The trick was I took this <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> problem that had n numbers and then I blew it up to an instance of this 2SUM problem. But that 2SUM problem has about square root of 2 to the n numbers. Now I can solve 2SUM in linear time, so solving that instance gives me a square root of 2 to the n time algorithm for the original problem.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=889">14:49</a>] But yeah, in the meantime, going from one problem to the other, I blew it up. But by blowing it up, I&#8217;m able to improve the time complexity of the obvious algorithm for <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=901">15:01</a>] I get it. Okay. So the reduction, it&#8217;s kind of like if you solve the <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>, you have some time complexity, and if you translate it to the other problem, you lose a little bit of time complexity, but less than the aggregate.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=917">15:17</a>] So all you want to make sure is when all the dust is cleared and settled, you want to be able to say, look, if I can improve the obvious algorithm for 2SUM, then I can improve the obvious algorithm for <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>. And that&#8217;s what this thing achieves, because I know how to get a faster algorithm for 2SUM. I can get one for <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>. So you just want your reduction between two problems to have this property.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=943">15:43</a>] If I can improve one problem by a little bit in running time, I can improve the other problem by a little bit.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=949">15:49</a>] In one of your talks, you mentioned preserving magic between.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=954">15:54</a>] Right.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=954">15:54</a>] Okay, so this is that.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=955">15:55</a>] Yes, this is exactly like preserving magic because once you see the 2SUM solution, it&#8217;s not so magical anymore. But imagine that you didn&#8217;t know about sorting and binary search and the like, and someone just says, find a pair of things with a certain property and there are N things. And you&#8217;re like, well, all the number of possible pairs is about N squared, so maybe it&#8217;ll take me N squared in this particular case, because they&#8217;re numbers and you&#8217;re summing them.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=988">16:28</a>] There&#8217;s an n log n time algorithm. It is a surprise when you first see that. No, you don&#8217;t have to try all of the pairs of numbers. There is a shortcut. There is a clear shortcut that lets you find a pair much faster. And then you can use that surprise, or magic, if you will, and get something for <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>. Get an algorithm for <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a> that avoids trying all of the subsets.</p><h3>17:00 &#8212; A severe strengthening of P vs NP</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1020">17:00</a>] So I understand one of the driving motives for pursuing this, I guess, lowering of lower bounds is the Strong Exponential Time Hypothesis, or SETH. Could you explain that and its significance?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1038">17:18</a>] SETH is like a severe strengthening of the P vs NP question. So P vs NP is asking whether the SAT problem has a polynomial-time algorithm or not.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1057">17:37</a>] Right.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1058">17:38</a>] And SETH is basically saying that for the SAT problem, you cannot solve it much faster than 2 to the n. So there&#8217;s no 1.999 to the n time algorithm for every string of nines. There is no 1.9999 to the n time algorithm, to say the hypothesis totally precisely. It has to do with the k-SAT problem. And you&#8217;re looking at clauses of length k for arbitrary k, but those details don&#8217;t matter.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1101">18:21</a>] So much. It&#8217;s a canonical NP-complete problem. SAT is solved all the time. In practice, it&#8217;s extremely useful for verification. Nowadays it&#8217;s an engine for verification. So it can be solved fairly well in practice. However, in the worst case, we still don&#8217;t know how to solve it significantly faster than 2^n. So the hypothesis that it needs, say, 1.9999^n time for all strings of n is a very strong exponential time hypothesis.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1140">19:00</a>] And so it&#8217;s stronger than P versus NP. It says, no, no, no, it says not superpolynomial. It&#8217;s actually darn near the end time that you need.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1149">19:09</a>] Right. And so do you think that hypothesis is true? And why or why not?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1155">19:15</a>] I think I&#8217;m on the record as not believing this hypothesis. Yeah, yeah. Why don&#8217;t I believe this hypothesis? Well, I started thinking about this hypothesis maybe already as an undergrad, but certainly starting in grad school. Early in grad school, I was thinking about this. So before it was even called SETH, I guess that shows how old I am, I was trying to think about how to solve this so-called CNF-SAT problem faster than 2^n.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1192">19:52</a>] And well, at the time there were a number of other NP-complete problems that had faster algorithms, like <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>. So I thought, well, there&#8217;s not, I mean, what&#8217;s so special about SAT? If all these other problems have faster algorithms, why not SAT as well? If you restrict to the 3-SAT problem, this is where you have an AND of these clauses. Each clause is an OR of three variables.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1230">20:30</a>] Some of the variables may be negated. You want to know if there&#8217;s a way to set all the variables to make all the clauses simultaneously true. There is a faster algorithm for that, but this is a more general version of SAT. So at first I just thought, well, there&#8217;s no good reason to think in a lower bound. These other related problems have upper bounds, so why not? But then over time I would have different attacks on SETH, like trying to refute it, always trying to refute it in different ways. These attacks would fail in some completely catastrophic and ridiculous way.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1273">21:13</a>] They would have no chance of actually solving the original problem. But by sort of staring at my failure and trying to think, well, there&#8217;s something interesting happening here. What can I do with this? There&#8217;s something interesting, yeah, it doesn&#8217;t refute the Strong Exponential Time Hypothesis thing. What does it do? So by trying to pivot and figure out, okay, what can I do with my failure, I was able to solve a variety of other problems instead.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1307">21:47</a>] And so after a while, I realized that the truth value of SETH to me is almost irrelevant because if I believe that it&#8217;s false, then I get good ideas. I get good ideas. And so by trying to think about, okay, what would an algorithm that breaks 2-SAT look like? What could it look like? I sort of force myself to think in a different way. I have to discard other natural algorithmic possibilities because we know they won&#8217;t work, and I have to think in a different direction.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1352">22:32</a>] And so because it kind of sends my brain in a different direction, sends me thinking a different way, it&#8217;s very useful for research. I mean, even though I still haven&#8217;t refuted it or whatever, it&#8217;s very useful for me to believe that it&#8217;s false operationally.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1375">22:55</a>] I believe this is a minority opinion, right? Why would they go against it?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1379">22:59</a>] I mean, I think, for example, Russell Impagliazzo, a good friend of mine who helped propose this. I mean, he always emphasizes to me, well, this is a hypothesis. We explicitly did not name it a conjecture. We wanted to sort of put forth some lower bound that would get you to think about it, like something that&#8217;s maybe a little more controversial than the other types of things like P vs NP or whatever, or other things that people more normally believe.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1414">23:34</a>] So, hypotheses which are at the edge of our understanding can be enlightening to think about, like where we truly don&#8217;t know what the answer might be based on our intuition. So this is, I guess, one reason why it was proposed. But one reason why you might believe that SETH is true is because believing it implies a lot of other lower bounds for you. Conveniently, because you can reduce the SAT problem to a bunch of other problems that seem totally unrelated, like edit distance, various pattern matching problems.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1458">24:18</a>] They have natural polynomial-time solutions. If you can improve on the algorithms for any of those, you would improve the one for SAT as well. You would get something better than 2^N. So believing in it sort of makes a convenient worldview. It shows that all these different textbook algorithms are indeed optimal.</p><h3>24:38 &#8212; SAT problems and solvers</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1478">24:38</a>] We&#8217;ve talked about SAT, or we&#8217;ve mentioned SAT so many times in this conversation. I know there are so many forms of that problem. What are all the different forms? And I know there&#8217;s clause width, also this idea of depth, too.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1491">24:51</a>] The most common representation of a SAT formula is a conjunctive normal form, so-called CNF representation. And this is what modern SAT solvers get as input. They get their file in so-called DIMACS CNF form. And this is just every line of my file. I give you a list of variables, possibly with negations, and each one is a clause. And I&#8217;m supposed to take the OR of those variables or negations. These are variables or negations.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1530">25:30</a>] They&#8217;re often called literals. Okay? So I take an or of these literals, and the width of that clause is the number of literals in it, okay? And I&#8217;m supposed to take the and over all those lines, each line in my DIMACS CNF file. So it&#8217;s an and of a bunch of ors, and each or has some small number of literals in it, call it k. So usually the width is called k. And so the k-SAT problem is to find an assignment to all the variables that satisfies all the clauses.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1569">26:09</a>] When each clause has width k or at most K, it could be smaller. So that&#8217;s the most popular version. That&#8217;s what SAT solvers churn on. And that&#8217;s what is behind SETH. That&#8217;s the representation there. But there are other ways to represent a Boolean formula. You could just simply represent it as some arbitrary expression made up of ors and ands and negations. It could just be some arbitrary expression with nested parentheses and all that mess.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1611">26:51</a>] You could ask, given a formula in this representation with a bunch of variables, is there a way to set the variables to make this true? That&#8217;s Formula-SAT. You could also look at circuits of bounded depth, as you mentioned. So there is this class of circuits that people study called AC circuits for alternating circuits. It doesn&#8217;t mean alternation in terms of electricity; it means alternation in terms of ors and ands.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1648">27:28</a>] So these circuits are made up of ORs and ANDs, in layers. So the CNF representation is a special case where I have an AND of ORs, like an AND of a bunch of clauses and ORs of the literals. But you can go further. You can have an AND of OR of ANDs of variables with negations and things like that. And so these are constant-depth AC circuits. And because the idea is you only have some constant number of layers of these ANDs and ORs.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1685">28:05</a>] And you can look at the SAT problem on circuits like that as well. For example, those are two other versions that people look at.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1693">28:13</a>] Yeah, SETH is on k-SAT. And I saw on 2-SAT, 3-SAT, 4-SAT, 5-SAT, there exist solutions that are asymptotically better than 2 to the N. So I guess the larger K can be, the more difficult the problem.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1713">28:33</a>] That&#8217;s the intuition. That is the intuition.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1716">28:36</a>] I&#8217;m curious, have you ever plotted that curve? Does it drop off exponentially, or...</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1721">28:41</a>] Yeah, yeah, yeah. That&#8217;s what&#8217;s so interesting about the current state of the art in k-SAT algorithms. As K increases, all of them, all of the different types of algorithms you might try to run, they all approach a 2 to the n exponent. Like as K grows and grows and grows, they get 1.99, 1.999, and so on. And yeah, it drops. In other words, it goes toward two pretty quickly, actually. So we understand how the exponent behaves pretty well for the known algorithms we have.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1764">29:24</a>] And then for something like 3-SAT, what is the intuition behind speeding up an exponential time search?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1772">29:32</a>] Yeah, so for 3-SAT, let&#8217;s just look at a single clause. Okay. A clause that&#8217;s got three variables in it. Okay. We know that if we set all those three variables wrong, it&#8217;s going to be false. Okay. So we have to avoid one of those assignments. Okay, well that means that there are seven out of the eight possible assignments. So there&#8217;s three variables, 2 to the 3, eight possible assignments.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1801">30:01</a>] Seven of them could be a satisfying assignment. They could be part of a satisfying assignment. We don&#8217;t know. But one of them is definitely not. So one easy way to see that 3-SAT can be solved in less than 2^n time is just take any clause, try one of the seven possible assignments and plug them in, and then recurse on the remaining formula. Now let&#8217;s think about what we did. If we were just trying all the possible 2^n assignments and we would like plug in one of eight possible assignments for each of those three variables.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1843">30:43</a>] And so we&#8217;d have seven recursive calls. Well, instead, because we&#8217;re clever and we looked at the clause, we have seven recursive calls, and that&#8217;s the difference. So we reduced by three variables at the cost of seven recursive calls, as opposed to eight. And this gets you a slight improvement. This gets you about 1.92 to the N. Something slightly better than to the N. Okay, but you can do better than this.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1877">31:17</a>] But this is sort of like the idea you try to look at ways to plug in variables that will force constraints so you can rule out a large portion of the possible assignments.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1890">31:30</a>] When I was thinking about circuits and the different widths and depths, I don&#8217;t know if this is an unusual question, but it reminds me of neural nets. But the operators are different and the space of the literals is different. So instead of booleans, it&#8217;s maybe floating points. And so it just made me wonder about the algorithms that you might apply on a neural net. Is there analogs between these two spaces?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1921">32:01</a>] Yes, yes. If we look at, say I want to model a neural network, so I want to compute, say it&#8217;s still a Boolean function, but I want to do it with a neural network. So I want to use, let&#8217;s say ReLU or sign activation functions or what have you. There is a slightly more general gate that we can use instead of ORs and ANDs that turns out to basically be equivalent. So if instead of using ORs and ANDs, we use a so-called majority gate, which outputs one if and only if at least half of its inputs are one, using this and negations, we can actually simulate neural nets, like the usual types of neural nets that you think of with the usual types of activation functions.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1983">33:03</a>] So if they have a constant number of layers of neurons, we can get a constant number of layers of majority gates and negations. So this is so-called TC circuits for threshold circuits.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1999">33:19</a>] Yeah.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=1999">33:19</a>] So once you allow threshold circuits, you can start to model neural networks still using Booleans. We&#8217;re still looking at Boolean inputs though. Yes. So once you allow your input space to be larger and have floating points, then you can prove a lot more in terms of lower bounds. You can find things that take depth 3 in the neural net that can&#8217;t be done in depth two and so on. So yeah, once you go past that and you start looking at just arbitrary real domain, it becomes a totally different picture from a discrete domain.</p><h3>34:51 &#8212; Hot takes on famous open questions</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2091">34:51</a>] One topic I thought might be fun to go over is you wrote this paper about the likelihoods of these various conjectures and complexity theory.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2100">35:00</a>] And I pulled a few of these that were kind of minority opinions, or maybe less common takes. Curious to hear your rationale. Okay, one of them, the well-known one, P not equal to NP, or P vs NP, you assigned an 80% confidence that they&#8217;re not the same. Yes, and I think most people say much higher confidence. So why would you assign such a low confidence that they&#8217;re not the same?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2132">35:32</a>] It&#8217;s interesting because I think I originally had something like 75%, but then my college classmate Scott Aronson was like, how dare you? Kind of. He sort of called me out. I&#8217;m like, okay, fine for you, 80% fine. I guess my point is that we really don&#8217;t understand polynomial-time computation as deeply as we think we do. And there are surprises, like all the time, in the power of algorithms. There are very few surprises in terms of lower bounds.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2173">36:13</a>] When we are able to prove a lower bound, typically it&#8217;s something we very much expected to be true. But it was hard to prove. Somehow we pulled it off, we got what we expected to be true. But all the time in algorithms, people are finding algorithms where it&#8217;s just surprising. Just wait, what? How do you get something that fast? Right? So this just happens over and over.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2207">36:47</a>] When I was younger, when I was first thinking about P versus NP, I had an intuition for what should be. And what I&#8217;ve understood over the years is that my intuition for what should be is often just wrong. And I&#8217;m having to revise my intuitions all the time. So when something like this happens often enough, you start asking yourself, what do I really understand? Do I really understand P versus NP?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2246">37:26</a>] I mean, I understand that the statement, right? It&#8217;s just one of those problems where somehow it is not so difficult to make formal, to write down mathematically, but to actually know what the answer is is just orders of magnitude more difficult than it is to phrase the problem. And complexity theory in particular is littered with statements like this, where the space of algorithms is just that vast.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2281">38:01</a>] So if you just keep getting surprised over time, you&#8217;re like, well, what do I understand? Maybe it was just misplaced confidence.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2290">38:10</a>] Okay, what about this one? So EXP not equal to NEXP. Or would you say NEXP?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2296">38:16</a>] Oh, NEXP. Yeah, X versus NEXP. So this is like the exponential time of P versus NP for EXP not equal to NEXP.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2302">38:22</a>] Or NEXP, you gave it a 45% chance. And if this is the P vs NP equivalent, but for exponential time, why is it so low?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2319">38:39</a>] Oh, because exponential time algorithms are even more powerful. Did I really say 45%? I mean for NEXP versus EXP or NEXP versus coNEXP?</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2329">38:49</a>] Yeah, NEXP not equal to EXP, 45%. It&#8217;s in this table.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2335">38:55</a>] So in other words, I believe NEXP equals EXP more than I believe they&#8217;re different. Right? So, yeah, let me try to explain why. So you can think of the NEXP vs. EXP question as some special case of P vs. NP where instead of looking at the arbitrary SAT problem, I&#8217;m looking at a SAT problem which is extremely compressible. So there&#8217;s a really small little computer that is exponentially smaller than the length of the instance, and it just outputs the character on the line number and the column number for the DIMACS CNF, like the CNF file.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2384">39:44</a>] Okay? So it&#8217;s like an extreme compression of some file. So it&#8217;s like you zipped it down to something like exponentially smaller than its original length. Okay, so it&#8217;s like some super compressed, extremely highly regular SAT instance. So I give you that and I ask you, when you unpack this thing, decompress it, is the result going to be satisfiable or not? And I want you to solve this in time polynomial in the decompressed representation.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2420">40:20</a>] Okay. So the point is that this is SAT, but in some very special case where the thing is extremely structured. So the idea, one conjecture for why SAT solvers work in practice, like one, I mean, this is, I mean, conjecture may be overkill because, I mean, this is just. This is not even a well-formed mathematical statement. So one hypothesis for why SAT solvers work in practice is because the real world is highly structured.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2458">40:58</a>] The real world is governed by physical laws that are not random, they&#8217;re not arbitrary. From a very small number of rules, we can recreate so much of science. So what arises in practice from designs of hardware and things like this are often extremely compressible. They have to be extremely compressible. And so maybe it&#8217;s true that every sentence which has a highly compact representation can just be solved efficiently.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2495">41:35</a>] This is the idea of whether in P vs NP that when it&#8217;s really, really structured like that and super compressible, there is some advantage. It&#8217;s not like something completely random. And since it&#8217;s not like something arbitrary, it&#8217;s, in fact, very, very special.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2511">41:51</a>] The other one is 80% likelihood on NEXP equal to coNEXP. And you wrote, &#8220;Why would a self-respecting complexity theorist do that?&#8221; Yeah, I was curious why that&#8217;s such a contentious statement.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2528">42:08</a>] So NEXP versus coNEXP. Let&#8217;s first talk about NP versus coNP. So coNP is like the class of complements of NP-complete problems like UNSAT, like checking whether something is UNSAT. Now from the time-complexity point of view, there&#8217;s no difference between checking SAT and UNSAT. You can always flip the answer from the complexity point of view. If I ask you, does coNP equal NP? What I&#8217;m asking you is could you prove to me that a formula is unsatisfiable with a short proof when it&#8217;s satisfiable?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2570">42:50</a>] I can give you a short proof. I can just give you the satisfying assignment, you plug it in, check that it works. But if it&#8217;s unsatisfiable, if no assignment works, we&#8217;re saying for all assignments the formula is not true. Can you flip that to an existential statement and say, oh, there exists this little proof that makes it work? So people don&#8217;t believe that NP is equal to coNP, and in fact NP different from coNP implies P different from NP.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2601">43:21</a>] But so this is the exponential time version of P vs NP. So it&#8217;s coNEXP versus NEXP. The reason why I think these are likely to be equal is that if a little birdie sat on a NEXP machine&#8217;s shoulder and gave it a little bit of advice about what the coNEXP thing is doing, then the NEXP algorithm can actually solve coNEXP problems. And so let me explain, let me explain why. What&#8217;s the little birdie?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2642">44:02</a>] What the heck is the little birdie saying? So because NEXP problems can run in 2 to the n time and 2 to the n squared time and things like that, running an exhaustive search over all possible inputs of length n is no problem for NEXP. So what a little birdie can do is say, okay, suppose I want to verify that this particular instance, let&#8217;s say we can talk about an UNSAT, but some compressible UNSAT problem or something.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2678">44:38</a>] Suppose I want to prove that this compressible UNSAT instance is a yes. Okay, how am I going to do that? With NEXP, the little birdie will tell me the total number of inputs of length N which are a yes. Okay, so it will just tell me some string which says, &#8220;Here&#8217;s the total number of inputs of length N. You gave me a length-N input. Here&#8217;s the total number of inputs of length N that are a yes.&#8221;</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2709">45:09</a>] Okay. So this advice, this little bird&#8217;s advice doesn&#8217;t take very much to encode account. It&#8217;s like order in bits to encode a count of things. So what does the NEXP thing do to prove a coNEXP thing? What it does is it guesses the things which are a no. So I&#8217;m trying to prove UNSAT. So UNSAT means yes, SAT means no. So the NEXP thing guesses those things which are no.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2745">45:45</a>] The no things it can answer, right? If it&#8217;s a SAT thing, it can just guess the answer to each of the no&#8217;s. Okay, so it guesses the answer to each of the no&#8217;s, it verifies all those answers, and then it checks the number of things it guessed is what the birdie told it. Once it&#8217;s done that, all the no&#8217;s have been covered, so everything else must be a yes. So it can actually prove a yes by just exhaustively finding all the no&#8217;s and ruling out any other no&#8217;s.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2781">46:21</a>] But the big question in my mind is, where do you get the little birdie?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2784">46:24</a>] Where do you get the little birdie advice from? Yeah, yeah. So, yeah, so this I&#8217;ve studied. Yeah, this is what the little birdie gives you. And it seems to me that it&#8217;s possible that this little birdie itself can be constructed in NEXP. And if so, then we&#8217;d just be done. In NEXP, you figure out what the little birdie would tell you and then use that to just flip the answer. So, I guess let&#8217;s sort of guess all the things which are no.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2815">46:55</a>] So what remains must be a yes.</p><h3>46:57 &#8212; Simulating space with time</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2817">46:57</a>] And we talked a lot about time complexity, and you mentioned a little bit this advice, I guess, kind of space complexity. And I know you had a major result relating space and time complexity. You&#8217;re basically simulating time complexity with space complexity, but lower than it&#8217;s been done prior. Could you explain what it was before your breakthrough result and then maybe the intuition behind your breakthrough result?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2845">47:25</a>] In general, the problem is the following. I give you an algorithm that runs in time T, and I want to know, is there another algorithm that uses space much less than T, uses an amount of memory much less than T, and still solves the problem completely, still completely simulates the thing perfectly. One intuition for why you might think this question just can&#8217;t be solved.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2881">48:01</a>] For some problems, there&#8217;s just no way to improve on the space. If you think of a lot of problems in dynamic programming, let&#8217;s say I have some logic circuit that I want to evaluate, and all the gates are provided to me in a row, and all the wires sort of flow from left to right, and I start with information on the left. I want to compute the information on the far right, and all the bits are flowing left to right by wires.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2916">48:36</a>] The natural way to evaluate such a circuit is you start with, say, the inputs on the left. For each gate in turn in the line, you look at its inputs. Its inputs have been determined. There are some bits. You use that to compute the value of that gate. You pass the values of the output of that gate forward. But if that circuit has T gates in it, it will take about T time to solve, but it will definitely also take about T space in general.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2950">49:10</a>] Right? Like the circuit could be wired up in some wild way, and there just might not be a way to save space for an arbitrary circuit. But already in 1975, people were studying this kind of question and finding counterintuitive answers to it. So Hopcroft, Paul, and Valiant, based on work of Paterson and Valiant, showed that time T algorithms, at least in this so-called multi-tape Turing machine model, a very powerful model, can be simulated in space t divided by log of t.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=2994">49:54</a>] It&#8217;s like you get some space savings, but it&#8217;s only a log T factor. Okay. And the way this is done is quite counterintuitive. Even though there&#8217;s only a log factor, it was pretty shocking that development was extended to random access models of computation later and more general models of computation. So, in the late 70s and early 80s, it was known for pretty much any reasonable model of computation that time T can be simulated in space about T over log T.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3042">50:42</a>] But there was a hidden, maybe not so hidden, gotcha. In the space-efficient simulation, it needs an exponential amount of time to run. So there&#8217;s a time-space tradeoff. Okay. If you really want to save some space, you got to blow up the time by a lot. Nevertheless, yeah, it was kind of commonly conjectured that this T over log T space was about the best you could do and that you were probably not going to get T to the 0.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3077">51:17</a>] 9 space or something much more efficient. So, yeah, it was a big surprise to me that you can actually put time T in space about square root of T. Again, there&#8217;s an exponential running time just to give the full caveat out there. But it&#8217;s still very surprising. This holds for any kind of time T algorithm, including something that might be outputting something pseudorandom, something from cryptography.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3120">52:00</a>] It&#8217;s not at all obvious that every such process could be compressed to only need square root of t space and still get the job done, still compute whatever function was being computed.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3133">52:13</a>] I mean, I know it&#8217;s probably super involved, but if you could just, at a high level, what&#8217;s the trick? How&#8217;d you do it?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3139">52:19</a>] Well, the trick for me was to read James Cook and Ian Mertz&#8217;s paper on tree evaluation very, very carefully, and just know the landscape around P versus PSPACE and know what Hopcroft, Paul, and Valiant did. Yeah. But I can give you a very high-level idea of kind of what&#8217;s going on and how what they did is useful. So Hopcroft, Paul, and Valiant, the way they were modeling space-bounding computation was in a particular way that seemed pretty general at the time, but turned out to be restrictive in ways that we just didn&#8217;t anticipate.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3190">53:10</a>] So the way they thought of it was I&#8217;m gonna take certain, I&#8217;m gonna break the computation up in little pieces, and I&#8217;m gonna write pieces of the computation, little bits, into the memory, but I&#8217;m only gonna do that over blank space. I mean, this is something that sounds natural, right? So I&#8217;m gonna erase pieces of memory and then I&#8217;m going to overwrite that blank space with a piece of memory.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3223">53:43</a>] So I&#8217;m being very destructive in a certain sense. But this is a natural thing you do, right? If you want to replace, swap something with something else, you often just erase. But there are ways to swap without erasing, right? So there&#8217;s a common little trick that is taught in CS courses. So if you require everything to be written into a blank register, then if you want to swap the contents of two variables like X and Y, then you&#8217;ve got to have a temporary register.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3263">54:23</a>] You move one into the temp, you erase it, you move Y into there, and so on. Right. But if you don&#8217;t want a temp, you can achieve the same thing with just XORing the registers bitwise, or if you&#8217;ve got numbers, you can add and subtract. Okay. In three instructions, clever instructions, adding, subtracting, or XORing, you can actually swap the contents of two registers without needing a third. Okay.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3294">54:54</a>] This is kind of the starting point for thinking about, well, why does it matter if you&#8217;re always writing into erased memory? So what James Cook and Ian Mertz showed at a very high level was they were studying a certain problem called Tree Evaluation Problem, whatever that is. And the problem had an algorithm where you had a little stack and you were always sort of popping and pushing on the stack, but you were always writing computation contents into erased memory.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3342">55:42</a>] Okay. What they realized was that if you allow computations to XOR bits of memory into existing memory, then you can save a lot of space. So if you&#8217;re really, really careful about how you XOR things, you can basically recover what&#8217;s in your memory without storing all of it at once. You sort of offload things to computation, and you XOR on top of things very, very cleverly. You can get nice cancellations of things you don&#8217;t want and keep around things you do want.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3378">56:18</a>] Yeah. So that is a high-level idea of how this stuff works.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3384">56:24</a>] And then going to square root, is it just an extension of that, or is it completely different?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3389">56:29</a>] So the way it works with square root is square root just happens to be kind of like the optimal tradeoff in this Tree Evaluation Problem business. So what you do is you break the computation into square root of t time intervals. And each time interval has about square root of t steps in it. And what you do is you give a particular way of simulating this thing so that you&#8217;re only kind of holding about a constant number of blocks of these time intervals in memory at any point in time.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3429">57:09</a>] Like you&#8217;re only holding a small number of these time intervals, records of time intervals, in memory at any point in time, which is entirely not obvious how you would do it, but it&#8217;s some sort of sweet spot to set the square root of T. It&#8217;s like the minimum setting of tradeoff. Say I want to break the computation into intervals, and the intervals have a certain number of steps. If I set it to be square root of T, then the number of intervals and the number of steps in each interval is about the same.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3464">57:44</a>] So I mean, that was a long-held result. I mean, you mentioned it was 1975, that&#8217;s maybe almost 50 years. If you come up with something like that, and I guess you&#8217;re writing it out, you&#8217;re thinking through, and then you see it, what is that moment like?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3480">58:00</a>] So I&#8217;ve been around long enough to have been deceived by myself many, many times. So, yeah, I guess the first two or three times I thought about this, I just thought, this is another one of those ideas that can&#8217;t possibly work. There&#8217;s no way. There&#8217;s just no way this works. So I would just kind of leave it and then come back to it sometimes when I was bored of whatever else I was working on.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3515">58:35</a>] It was a pretty slow process of convincing myself that this could possibly be true. I thought there was either a bug in what James and Ian was doing somewhere, or there was a bug in my interpretation of what&#8217;s happening. I thought there had to be a mistake for a long time. And the only way I got over that was just writing it down over and over and over in different ways, sort of writing and rewriting and writing and rewriting and adding more detail and then maybe finding a different way of explaining it and erasing and writing something shorter and just sort of re.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3563">59:23</a>] Explaining it to myself over and over and over. And even then, when I submitted it to the STOC conference where it appeared, I wasn&#8217;t entirely confident that it was correct. I was just exhausted from thinking about it for so long and just thought, maybe someone else will find the mistake for me. I got at this point, I was just desperate.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3589">59:49</a>] I had actually sent it to two colleagues that I trust privately and asked them, &#8220;Can you help me find the mistake? Or whatever? Can you help me understand this?&#8221; And one of them just said, basically, they just didn&#8217;t read past the abstract. They&#8217;re like, &#8220;I&#8217;m not. I&#8217;m sorry, I&#8217;m not going to.&#8221; I&#8217;m like, they just didn&#8217;t believe it, period. Just like me. I mean, they didn&#8217;t believe it.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3617">01:00:17</a>] And another one said, after a long time, &#8220;Well, I didn&#8217;t quite follow it, but there was a footnote that you put, like, in the bottom of some page. After that, after that footnote, then I began to believe it.&#8221;</p><p>So then what I did was I just took that footnote and elaborated it in a later revision, sort of made it what I called the warmup, because I was like, yeah, I&#8217;ve got to do. I&#8217;ve got to write this thing in a way that is airtight so that other people will actually believe this.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3656">01:00:56</a>] Because, yeah, it took me a long time to get used to it, to believe it. Yeah.</p><h3>01:01:02 &#8212; Why he solves hard problems</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3662">01:01:02</a>] Wow. What motivates you to solve these hard problems? I mean, what keeps you driven? Because you kind of sound like you gotta just bash your head against the wall, and maybe you get there, maybe you don&#8217;t.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3680">01:01:20</a>] I try not to bash my head against the wall. And if I&#8217;m going to do it, it&#8217;s gonna be a comfortable wall. It&#8217;s gonna be one that&#8217;s high-end and luxurious or something. I&#8217;m gonna enjoy the process, is what I mean. If I don&#8217;t enjoy the process of really grinding and trying to understand something, I&#8217;m not going to do it. So I think, if anything, one thing that I like, I like to involve myself in working on problems where the grind is actually joyous and fun.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3727">01:02:07</a>] Yeah. For the other things, they&#8217;re just mainly opportunistic. They&#8217;re mainly me just taking something new that I see and fully integrating it with everything I know and following everything to the logical conclusion and just seeing what happens. And not trying to get emotional, trying not to whatever, just trying to follow everything one step after another. Yeah.</p><h3>01:02:35 &#8212; How to pick good research direction</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3755">01:02:35</a>] How do you pick a good research direction?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3757">01:02:37</a>] The best kind of research direction, which is very hard to come across, is a direction where you&#8217;ve set yourself up in a way that you can&#8217;t lose. I have a hypothesis H, and one of two things is going to happen. Either I prove H is true, or I refute H, prove not H. What I would like to have is a situation where if H is true, then good things happen. If not H is true, good things still happen.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3796">01:03:16</a>] Other good things. I like to look at things where H. Yeah, it&#8217;s a win-win kind of. If you can find opportunities like this, this is important. I think often I&#8217;m not looking at a specific problem, I&#8217;m looking at a method. I&#8217;m looking at the technique. I&#8217;m looking at not the specific proof, but the space of ideas. And I&#8217;m trying to understand what else can be done in this space of ideas.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3834">01:03:54</a>] So often I solve problems just by taking some idea and putting it somewhere else.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3839">01:03:59</a>] Do you have a concrete example of where you&#8217;ve pursued something knowing that if you succeed, good. If you succeed in the other direction, also good.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3849">01:04:09</a>] I have some particular research program for which, if it materializes, would actually show that something called the orthogonal vectors problem can be solved in nearly linear time. Okay. Now, if that&#8217;s true, then the SAT problem can be solved in about the same time as <a href="https://en.wikipedia.org/wiki/Subset_sum_problem">Subset Sum</a>. It could be solved in like square root of 2 to the N. So this would truly break SETH in a radical way.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3883">01:04:43</a>] Okay. And I did a lot of work in the past showing how if you can improve on the running time of SAT, then you can somehow use that to prove circuit complexity lower bounds. You can take an algorithm for analyzing circuits that&#8217;s nontrivial and interesting and turn that into a limitation on what those circuits can compute. Okay, so this hypothesis, maybe I won&#8217;t exactly say what it is, it&#8217;s a pretty technical statement, but the point is that if the hypothesis is true, then I get this fantastic algorithm for this orthogonal vectors problem.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3934">01:05:34</a>] I get this algorithm for SAT, I get circuit complexity lower bounds. From that I get to somehow prove limitations on what circuits can do if the hypothesis is false. I can use that hypothesis to actually show another circuit complexity lower bound of a different flavor. So regardless of whether or not the hypothesis is true or false, I&#8217;m going to make progress in complexity theory. I&#8217;m going to prove a new limitation one way or the other.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=3964">01:06:04</a>] Right. So sometimes, I mean, this can be difficult to do, but usually what has happened instead of something so clean is that I have, I read about some idea or whatever, I try to apply the idea, I read about some magical algorithm, I try to apply the magical algorithm to solve something like SAT. It doesn&#8217;t work, but I look at my approach, I stare really hard at it, I try to pivot, I try to just extract what I can, and then I get some other type of algorithm.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4001">01:06:41</a>] I mean, this is essentially how I proved the circuit complexity lower bounds that first got me a job at <a href="https://en.wikipedia.org/wiki/Stanford_University">Stanford</a> was I was trying to solve SAT faster with some algorithm. It didn&#8217;t work, but then I realized that that algorithm worked for a large class of circuits, including ones we didn&#8217;t know lower bounds for. So then I used that to prove a lower bound. Yeah. So I mean, this is just a general principle that I&#8217;ve had.</p><h3>01:07:14 &#8212; Technical book recommendations</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4034">01:07:14</a>] What would be your book recommendation if someone wants to learn more about complexity theory?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4039">01:07:19</a>] Yeah, so it sort of depends on how deep you want to go. There&#8217;s Avi&#8217;s book. Yeah, Avi&#8217;s book is really nice. If you&#8217;re looking for something a little more gentle of an introduction, I would say you could read the early chapters of Lance Fortnow&#8217;s The Golden Ticket, trying to imagine what a world would be like if P equaled NP. I like that imagination exercise that Lance did. On a more technical level, if you&#8217;re looking for something slightly more technical, you could look at The Nature of Computation by Mertens and Moore, I think.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4085">01:08:05</a>] So this is still not very technical, but it is a textbook, right? Yeah. And then, of course, Sipser&#8217;s Introduction to the Theory of Computation is a fairly concise and very well-written textbook. Yeah.</p><h3>01:08:31 &#8212; Advice for his younger self</h3><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4111">01:08:31</a>] Last question for you is, if you could go back to the beginning of your career, knowing what you know now, what advice would you give yourself?</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4118">01:08:38</a>] So I guess one thing that I did by accident, which I think people should keep in mind, is that you don&#8217;t need permission to work on very tough problems. In fact, the tougher the problem, like in complexity theory, there are all these problems where nobody has a very good clue of what to do beyond a few simple structural results. So that kind of levels the playing field a bit.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4155">01:09:15</a>] And yeah, you just, you don&#8217;t need permission to think about these problems. You don&#8217;t need anybody&#8217;s say-so or whatever. I mean, that&#8217;s one thing that I would like younger people to keep in mind, that this knowledge is open to the world and it&#8217;s open to you. And you can just go learn it and think about it. And yeah, you don&#8217;t need anybody&#8217;s permission, especially not mine.</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4183">01:09:43</a>] Yeah. Another thing I think is to avoid inertia in the sense that if you don&#8217;t think about and reflect, where am I going and what am I doing? You can just end up coasting in a certain direction. And I think it&#8217;s important for people to sit down and truly reflect and think from time to time. Okay, if I keep going in this direction, is this going to lead to a place that I want to be?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4227">01:10:27</a>] And this sort of advice is good for all aspects of life, really. If I keep just allowing things to be the way they are, am I going to be okay with that, or do I need to do something? Do I need to? So it&#8217;s like little interrupts in your life and reflection, just sort of like, okay, do I need to be coasting here? I think in research for me, I was always reevaluating, like, okay, what did I do last year and what did I do the year before, and could, you know, like, what do I know now that I didn&#8217;t know then?</p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4271">01:11:11</a>] And, am I progressing? I think that sort of self-evaluation is crucial when you&#8217;re in grad school because once you&#8217;re in a PhD program, there are no more easy metrics to gauge how good or bad you are doing. And the truth is, there just isn&#8217;t such a metric. So all you can really do is evaluate your past self with your present self and compare and contrast to see if you&#8217;re making progress with whatever it is you want to do.</p><p><strong>Ryan Peterman:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4309">01:11:49</a>] Thank you so much for your time, Professor Williams. I really appreciate it.</p><p><strong>Ryan Williams:</strong></p><p>[<a href="https://youtu.be/AaK1SL2i_4Y?t=4312">01:11:52</a>] Yeah, I appreciate you having me here. It&#8217;s been fun. Thanks.</p>]]></content:encoded></item><item><title><![CDATA[OpenAI Eng & Dev Tools Founder: How Software Engineering Is Changing | Charlie Marsh]]></title><description><![CDATA[Charlie Marsh is the founder of Astral, the Python devtool startup that was acquired by OpenAI.]]></description><link>https://www.developing.dev/p/openai-eng-and-dev-tools-founder</link><guid isPermaLink="false">https://www.developing.dev/p/openai-eng-and-dev-tools-founder</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 22 Jun 2026 09:45:53 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202600605/41fef54af4ad60bddd4b4230904b23dc.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/marshcharles/">Charlie Marsh</a> is the founder of Astral, the Python devtool startup that was acquired by OpenAI. I inteviewed him about how software engineering is changing and learnings from starting his own company as an engineer.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/Iw65FD4MGgs">YouTube</a>, <a href="https://open.spotify.com/episode/2sfielYIDcNO6HpeLdnooM?si=IoRXB5JsRjuP-2_q8NarLg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-Iw65FD4MGgs" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Iw65FD4MGgs&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Iw65FD4MGgs?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/202600605/0040-origin-story">00:40 - Origin story</a></p><p><a href="https://www.developing.dev/i/202600605/0604-the-front-page-of-hacker-news">06:04 - The front page of Hacker News</a></p><p><a href="https://www.developing.dev/i/202600605/1435-why-he-chose-rust">14:35 - Why he chose Rust</a></p><p><a href="https://www.developing.dev/i/202600605/2010-full-codebase-migration-from-zig-to-rust">20:10 - Full codebase migration from Zig to Rust</a></p><p><a href="https://www.developing.dev/i/202600605/2840-llm-generated-code-and-open-source">28:40 - LLM generated code and open source</a></p><p><a href="https://www.developing.dev/i/202600605/3534-performance-optimizations">35:34 - Performance optimizations</a></p><p><a href="https://www.developing.dev/i/202600605/4454-optimization-with-ai-and-combating-slop">44:54 - Optimization with AI and combating slop</a></p><p><a href="https://www.developing.dev/i/202600605/010208-learnings-as-an-eng-starting-a-company">01:02:08 - Learnings as an eng starting a company</a></p><p><a href="https://www.developing.dev/i/202600605/011755-top-technical-talk-recommendation">01:17:55 - Top technical talk recommendation</a></p><p><a href="https://www.developing.dev/i/202600605/011856-advice-for-his-younger-self">01:18:56 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:40 &#8212; Origin story</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=40">00:40</a>] What was the space like when you were first getting into building these dev tools? Why did you think it could be better?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=51">00:51</a>] Before I started Astral, I was at a computational biology company, and I was the second engineer. I had no bio background and no ML background, and so I was building all the software systems to do machine learning, and that was all <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>. So I was learning <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>. Then we had some Go and some Rust, and so I just worked across lots of different ecosystems. And I think <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, at the time when I was in that position, we were trying to do a lot as a small team.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=82">01:22</a>] So we had a very, I mean, not compared to large code bases, but we had a comparatively large code base. And it was really a small group of us that were software engineers trying to serve the rest of the team, whether they were machine learning researchers or scientists. So I had kind of worked across a bunch of different ecosystems, seen a lot of the tooling, and now I was working on <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, and we kept trying to get lots of leverage out of the tools, whether it was the type checker or the linter or the package manager. Things were scaling up.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=110">01:50</a>] We were a small team. We wanted to have really good static analysis, really good tooling, and we kept running into limitations around it. I mean, I think there were some. I kind of looked at the <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> ecosystem, and I just didn&#8217;t see the same amount of experimentation that I saw in other ecosystems, like in the web ecosystem. There were a lot of things going on that seemed pretty crazy, and there was a very intense focus on.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=139">02:19</a>] I don&#8217;t know about very intense, but there was more of a focus on performance. And that started with esbuild, which was in Go, and then you had SWC and Rust and then Bun and Deno and all these other tools that came around. And this idea that you could have native tooling for <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> became very accepted, especially as applications became bigger and people were doing more and more stuff on their machine.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=167">02:47</a>] And that wasn&#8217;t really happening in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>. And so that, to me, was kind of like the key question, was like, why can&#8217;t we have that? Because I was seeing all this stuff that was happening in web, and I had done some Rust, I&#8217;d done some Go. I&#8217;d seen these different tool chains and how they work, what the experience is, the user experience mostly in those cases. And then I saw <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, and it was like we had a smaller set of tools.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=193">03:13</a>] There were lots of interesting ideas from other ecosystems that I didn&#8217;t see represented. They were really all written in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>. And so that to me was kind of like the interesting question. And it was sort of like the first thing I asked when I started working on this, when I released Ruff. The title of the blog post was something like, &#8220;<a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> tooling could be much, much faster.&#8221; And for me it was really like, this is a hypothesis.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=222">03:42</a>] Could <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> tooling be much faster? And then I built a prototype, and that prototype was like, yes, I think it could be.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=229">03:49</a>] You wrote a post kind of exploring mypy or publishing your results using mypy. And the date on that, nine days later, you published the post. That one that you mentioned, it could be so much faster.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=243">04:03</a>] No, that&#8217;s funny, I didn&#8217;t even realize that. Yeah, those were so close together.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=246">04:06</a>] Did you build that proof of concept in those nine days?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=250">04:10</a>] Yeah, I wrote that blog post about basically type checking at scale because I was trying to think of what are things that I learned that people haven&#8217;t written much about. And then I started with building a linter because I thought it would be much easier. I think it actually is a lot easier than building a type checker. We&#8217;re now building a type checker and we have a type checker. And so I can say that I think building a type checker is much harder.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=275">04:35</a>] But I started with a linter because I thought it would let me prove a lot of the same ideas, but in a smaller form factor, like building a startup or just building a tool. You want to try to get something into people&#8217;s hands that they can actually use and that actually proves out value as quickly as possible. And also that lets you iterate very quickly. And a linter weirdly ended up being kind of the perfect form factor for it because it has a pretty simple core.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=307">05:07</a>] But then you have tons and tons of rules. And so it&#8217;s extensible in a lot of different ways. But people can get value very quickly from a small core plus some number of rules. And so we were able to get something out there and then iterate and ship more rules and more functionality, and people could use it alongside other tools. It was useful immediately, which I think is extremely helpful. As opposed to something like.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=334">05:34</a>] I mean, actually, the type checker would be a good example. A type checker that&#8217;s like 75% done isn&#8217;t very useful. Same with the package manager. These things have to be. I mean, there are ways, I think I would challenge those as blanket statements. But the point is it&#8217;s harder to ship something iteratively that&#8217;s still useful to users. And I think that&#8217;s one of the key properties and actually building momentum as a tool or something else.</p><h3>06:04 &#8212; The front page of Hacker News</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=364">06:04</a>] When you posted that writing, what was the reception?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=367">06:07</a>] I mean, the blog post itself, I published it and then it was on Hacker News and it got tweeted about a lot, and it just created some amount of excitement and interest around it. And those things, I mean, especially being on the front page of Hacker News, I have a lot of thoughts about. I mean, it&#8217;s certainly not. It&#8217;s like some percent luck and then some percent skill, and it is like a helpful thing but not a requisite to success.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=402">06:42</a>] And it also doesn&#8217;t guarantee success just because you got something on the front page of Hacker News. It&#8217;s great if you can, but there&#8217;s a lot of randomness to it. I do think, though, that maybe a couple things. So one, after that happened, some people started actually paying attention to the project. And so I really tried to capitalize on that, basically by being very engaged. And so when people would file issues, I would aggressively try to acknowledge the issue as quickly as possible, fix it, and ship a release.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=439">07:19</a>] If you can do that within one day, it&#8217;s a very, very powerful loop. You gain supporters and create more momentum around the project. And there were also a couple high-profile people in the ecosystem that started paying attention to the project, like Sebasti&#225;n Ram&#237;rez, who does <a href="https://en.wikipedia.org/wiki/FastAPI">FastAPI</a> and a bunch of other things. And he&#8217;s been a great supporter of our projects too.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=461">07:41</a>] But early on he was like, &#8220;Oh, this is interesting. I&#8217;d like to use it.&#8221; And I was like, I basically asked myself, well, what would it take for him to use it? And then I just tried to do all those things. So I tried to capitalize on that.</p><p>I think the other is, though, I do think a lot about developer marketing, and it&#8217;s sort of like a dirty word or like a dirty expression. I think a lot of engineers want technical products and tools to just win on their merits.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=493">08:13</a>] The best technical solution should just be the one that wins and grows. But I actually, I mean, I have this sort of stupid hypothesis that there&#8217;s all these, like, thousands of really amazing projects on <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> that basically don&#8217;t know how to market themselves and so never get discovered and never go anywhere. And I don&#8217;t actually know if that&#8217;s true, but it&#8217;s sort of how I feel about a lot of technical work, where it&#8217;s like, it&#8217;s actually worth looking at a project like Ruff.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=524">08:44</a>] And when I think about what to put in the blog post, or this extends to what to put in the README or what to put in the. You don&#8217;t need a million emojis and a million screenshots of everything. What you need is, okay, someone&#8217;s going to land on this page and I have 10 seconds to get them interested in what I&#8217;m doing and help them understand why it might matter to them and why it&#8217;s useful. And so that&#8217;s the key idea that I try to bring basically to everything.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=557">09:17</a>] I mean, it&#8217;s a little bit depressing because it&#8217;s like, okay, I only have 10, but people really don&#8217;t have. This is the attention economy, right? It&#8217;s the same thing with a blog post. Ironically, I used to look at a lot of the <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a> blog posts when I was at Spring and we were doing machine learning in bio and we were trying to publish material to explain what we do. But we had all these interesting insights.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=584">09:44</a>] We looked at the <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a> blog posts, especially the DALL&#183;E ones, the older ones, and we were like, wow, that&#8217;s an amazing blog post. Because if you read none of the text, you still understand what they did and you&#8217;re still impressed by it. If someone just goes on the page, they leave with an impression and they have some understanding. And I was like, wow, that&#8217;s a really amazing thing. And so when I write blog posts now, I.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=607">10:07</a>] I&#8217;m like, if someone just lands on this page, I have to be prepared for the large number of people who will only read maybe the <a href="https://en.wikipedia.org/wiki/TL%3BDR">TL;DR</a>, maybe look at the first image, maybe read the headline. Some people will read the full thing. And so I should care painstakingly about every word and everything it says, but most people will read almost none of it. And so, again, sorry, it&#8217;s a little bit depressing, but I do think it&#8217;s worth thinking about your audience and how do you.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=633">10:33</a>] How do you explain to them why they should care about what you&#8217;re doing in a way that&#8217;s honest and genuine, not deceptive? But you do have to think hard about how do you communicate why something matters to people.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=646">10:46</a>] Very quickly, could you give an example of what you did in your case?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=651">10:51</a>] I think for Ruff specifically, we had a really good graph of the benchmarks that we put in the blog post, and that went viral. It was on Twitter and that was a priceless graph, basically. Sorry, not in terms of monetary value. I just mean in terms of attention. Right. It was like, oh, that&#8217;s really obvious what&#8217;s happening there. I think graphs of benchmarks can be a really powerful thing.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=679">11:19</a>] All benchmarks are bad and wrong in some way. And so there&#8217;s a lot more to say about that. But I do think having a visual hook that explains to people why they should care and what something is, and having a really strong tagline. I would have to look back at what we had, but we were like, okay, we&#8217;re going to be compatible with your existing stuff and much faster.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=704">11:44</a>] And that&#8217;s kind of it. And it&#8217;s like, okay, well, if it&#8217;s compatible and it&#8217;s much faster, I was like, why should you care? It can be orders of magnitude faster and it&#8217;ll give you the same experience. You try to convey that as quickly as possible.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=718">11:58</a>] Right. And even the, I mean, that blog post header, I mean, <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> tooling could be much faster. Should be much faster.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=727">12:07</a>] It&#8217;s a little bit provocative, I guess. Yeah. I mean, I&#8217;m pretty against clickbait or trying to be misleading in anything that we do. You basically want to balance creating, effectively communicating the most interesting pieces to people, but you have to do it in a way that you believe is true and authentic. You could lie about a bunch of stuff and post interesting things that are clearly provocative.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=758">12:38</a>] But I don&#8217;t know, to me, that&#8217;s just. Why would I ever do that? That&#8217;s such a losing battle.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=764">12:44</a>] What do you think about that graph? I don&#8217;t know if the... I see this a lot, though, where the axis is not zeroed.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=773">12:53</a>] So it&#8217;s all within the range of 50 to 60 chart crimes.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=778">12:58</a>] Yeah, I wouldn&#8217;t do that. Yeah, yeah. I do think benchmarks are pretty complicated and controversial, though, and maybe most people don&#8217;t even see this controversy. But I do, because it&#8217;s just very hard. Performance is very nuanced, and you have to think about, okay, with caching or without caching. Those are pretty different. Or depending on the project that you&#8217;re running on, the performance characteristics will be totally different.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=816">13:36</a>] And we see this in uv. This is really hard because sometimes uv is really fast, but sometimes your package install is actually bottlenecked on a C build that&#8217;s totally out of band, and it&#8217;s like, okay, well, so how do you accurately capture all of the different scenarios? I mean, it&#8217;s sort of impossible to accurately capture all of the different scenarios, but it&#8217;s also a disservice to users not to try to communicate to them the importance or a real representation of what the difference in performance will look like and why does it matter?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=854">14:14</a>] And so you have to balance those things, which I think is actually quite hard. And a lot of people tend to get it wrong. And maybe we&#8217;ve gotten it wrong a few times too. I&#8217;m happy to accept that. But I do think it&#8217;s more complicated than people think. But it&#8217;s also, I think, a huge disservice not to include some way to try to convey to people quickly what things feel like.</p><h3>14:35 &#8212; Why he chose Rust</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=875">14:35</a>] I think one of the things I&#8217;m most interested in is, you mentioned that graph. I see that graph and the engineer in me is like, how did you make it so much faster? And one of the things in the tagline is <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> dev tooling written in Rust. And so I guess one of the first things I&#8217;m curious about is why did you choose Rust and what are the pros and cons of that?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=897">14:57</a>] I would say I largely chose Rust for that project because of hype. I don&#8217;t think at the time I was that informed about a lot of the technical trade-offs between those different ecosystems. And I hadn&#8217;t really done a lot of systems programming. So the honest answer is I was seeing it in a lot of places, and it was growing a lot, and people said it appeared to be sort of an accessible way to write.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=924">15:24</a>] My impression was that it was an accessible way to write high-performance software. And so I started for that reason. In hindsight, I think it was an amazing decision, and I absolutely would not do it differently. And I think it was the best choice for what we&#8217;re building now. I have a lot more experience writing Rust, but it is funny because I do sort of think that if I had tried to write this in C or C++, I think I probably would have given up, because even now I find those ecosystems much more intimidating.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=968">16:08</a>] Hard to access, hard to learn. I have often felt that one of the great or underrated advantages of Rust is basically the tooling, because you drop in and you use Cargo, and that&#8217;s how you do everything. And if you clone it. This isn&#8217;t always true, but there are exceptions as projects get sufficiently advanced. But it&#8217;s. If there&#8217;s a crate out there that we use and maybe I want to submit a PR to it, I&#8217;m extremely confident that I can clone it and figure out how to run and build it without having to do any work.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1004">16:44</a>] Being able to just clone and run cargo, run or cargo build or cargo test, that&#8217;s actually really amazing, especially for someone like me who was new to systems programming and I didn&#8217;t want to have to figure out my whole C tool chain and build system, all this stuff. Rust was opinionated about all of these things. And so there were a lot of things that were hard to learn, but I was allowed to basically, I was allowed to spend my time focusing on the things that actually should be hard to learn, as opposed to all the other, basically all the bullshit that you don&#8217;t want to spend time on, like how do I get this thing to actually build and run?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1044">17:24</a>] I think it has been an extremely good bet for us. It&#8217;s scaled very well with the project. And I mean, we at <a href="https://en.wikipedia.org/wiki/OpenAI">OpenAI</a> too are like, we use a lot of Rust and they&#8217;re betting heavily on Rust. And then, you know, me as someone who builds software, builds developer tools, like tries to build software for people to build software, I&#8217;m very bullish on Rust. It&#8217;s grown enormously and I think it will actually just keep growing enormously at the same time.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1077">17:57</a>] I said this at the start. I&#8217;ve never been someone who&#8217;s super dogmatic about ecosystems. I think that all of those ecosystems have lots of redeeming qualities. I think what&#8217;s happening in Zig is very interesting, and I wish I had more time to go deep on it and develop more nuanced opinions about what it does better. I&#8217;m sure it does a bunch of things better than Rust, and I&#8217;m sure there are things for Rust to learn from that ecosystem.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1100">18:20</a>] Similarly, people have a ton of success building stuff in Go. I think that&#8217;s great too. I think all this stuff is great. But I have loved using Rust and I have complaints about it. But I view those complaints as things that we should improve, especially as the way we build software changes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1118">18:38</a>] With LLMs, it sounds like you&#8217;re saying C and C++ today. You would not consider them because their tooling is insufficient. But Zig, Rust, and Go are all some spectrum of trade-offs or different things.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1136">18:56</a>] Yeah, I mean, I think a lot of people would disagree with me, but I don&#8217;t really understand why unless there are very specific technical reasons or you&#8217;re working with existing software. Basically, I don&#8217;t really know why I would start a net new project in C or C++. I certainly would not at a personal level. I guess if they&#8217;re like the ecosystems you know really well, then it makes a lot of sense to do it.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1166">19:26</a>] But if I was a new programmer looking to learn systems or if I knew those languages equally well, I guess I just don&#8217;t understand why I would do that, which I&#8217;m sure I&#8217;ll get roasted for, but I just don&#8217;t really care now. I actually care about a lot of things that Rust gives me that I didn&#8217;t care about as much before, like memory safety. I don&#8217;t think I even understood or cared about what memory safety was when I started working on Ruff, to be honest with you.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1189">19:49</a>] And now I think it&#8217;s a pretty amazing thing. I want to build things that are safe, that don&#8217;t crash on users, that are really fast, really performant, that don&#8217;t have to make compromises, basically. And I feel like I can do that in Rust. And so for me, I don&#8217;t know, it&#8217;s kind of like the perfect language. I mean, it&#8217;s not the perfect language, but it is my favorite language to use right now in today&#8217;s world.</p><h3>20:10 &#8212; Full codebase migration from Zig to Rust</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1210">20:10</a>] It&#8217;s not unreasonable that you could almost completely transpile or convert your existing codebase into any language of your choice. Like I saw, I think it was on Bun, I forgot. I think it was Zig to Rust. Yes, like all in one go. So if you studied some other language and you thought it was better, you could totally rewrite everything from Rust into, I don&#8217;t know, Zig or Go. Would you ever consider doing that for the dev tools that you&#8217;ve built, for instance?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1247">20:47</a>] Yeah, it&#8217;s such an interesting time to be building software because things are changing. This is not a novel comment, but things are just changing so fast. I didn&#8217;t really program with agents at all until Christmas, like December break, basically of this year or of last year, rather. And that&#8217;s really not that long ago. I mean, I was using Cursor and stuff, but I wasn&#8217;t like now. I haven&#8217;t edited code in an editor in a very long time.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1276">21:16</a>] Everything I do goes through Codex and even if it&#8217;s small edits, I&#8217;m just prompting and it&#8217;s completely different. And that&#8217;s not that long of a time, right? And so I don&#8217;t know what it&#8217;s going to look like in six months. Hard to say. But the reason that I mentioned that is, I think Bun doing that rewrite, it&#8217;s like I wouldn&#8217;t do that right now, but I&#8217;m actually glad that they are pushing that.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1310">21:50</a>] They are experimenting and pushing the envelope of what&#8217;s possible. And they&#8217;re doing it in the face of a lot of criticism. And it will be interesting to see how it plays out for users because I think a few different things I thought about is, oh, reasons not to do that. One is, okay, maybe you don&#8217;t understand the code anymore, right? If the code got completely rewritten, then as a human you may not understand any of the code anymore.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1337">22:17</a>] And that can be a big problem. I think there are mitigating factors around that, which are, they tried to do a very direct transpilation, quote unquote. I mean, it&#8217;s not exactly a transpilation, but they tried to do very one to one. And so that hopefully helps with understanding all the abstractions. If they&#8217;re building so heavily with agents, does it even matter? To what degree do they need to understand different parts of the code?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1361">22:41</a>] I actually don&#8217;t know the answer to that, which is where everything&#8217;s changing so quickly comes from. It&#8217;s actually a little bit hard for me to say right now how much that actually matters anymore. If they did transpile the whole codebase, how well do they need to understand things at different levels of abstraction? They should. It&#8217;s actually a very, I think, complicated question. So one reason is how well will you understand the code?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1385">23:05</a>] That&#8217;s a confusing thing. The other is you&#8217;re kind of trading, when you do a rewrite like that, you&#8217;re trading known issues for new unknown issues. Imagine you merge that and it closes 50 open issues on <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a>. In a literal sense, okay, that&#8217;s great, but you may have now caused 15 new issues that didn&#8217;t exist before, that you don&#8217;t know about, because even a great test suite is only an approximation of your program&#8217;s correctness.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1421">23:41</a>] There&#8217;s Hyrum&#8217;s Law, which is like anything&#8212;I&#8217;ll get the phrasing wrong, but the gist is basically any implementation detail in your software. Someone eventually comes to rely on that. And so even things that aren&#8217;t encoded as behavior become part of your <a href="https://en.wikipedia.org/wiki/API">API</a>, whether you like it or not, which is a scary thought. But basically that means that even if you pass your whole test suite, there&#8217;s probably lots of behavior that was implicitly encoded in your program that&#8217;s changed.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1451">24:11</a>] The unfortunate thing is you end up pushing that onto users. They are the ones that have to run into that and then report it and figure out what went wrong. That&#8217;s my main area of concern. With an automated rewrite, I actually trust the models enough that I would consider allowing them to do it. But the scarier thing for me is basically what&#8217;s the impact for users and how do you roll that out safely?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1473">24:33</a>] So I don&#8217;t know. It&#8217;s kind of a... We&#8217;re not going to, like, we don&#8217;t have anything planned to do that for any of our own projects. But it is an interesting question. I mean, I think maybe the third piece that I do think about, especially right now, where everything&#8217;s changing so quickly and there&#8217;s basically the spectrum of how you write software has gotten way wider. This is a terrible analogy, but there&#8217;s so many different ways to build software now.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1505">25:05</a>] There&#8217;s the true Andrej Karpathy definition of vibe coding, where you&#8217;re like, don&#8217;t look at the code at all. All the way to don&#8217;t use LLMs and blah, blah, blah. There&#8217;s a lot of different stuff in between. And I actually think that right now I have pretty different approaches to different projects in terms of what I do. I built a tool last week that was. It&#8217;s basically personal software. It&#8217;s like a custom Rust linter to help us enforce certain things in our projects that I&#8217;ve always wanted to enforce but didn&#8217;t fit into existing tools.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1545">25:45</a>] And I didn&#8217;t read any of the code. GPT-5 did the whole thing, and it&#8217;s like, okay, I can have a really different standard for that than uv because that project, we are the only users. It&#8217;s easy to tell if it&#8217;s correct or not. Not to go too deep into the weeds on what the project does, but we know if the output is incorrect and there&#8217;s no other users. uv is relied on by millions of engineers and every single company.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1581">26:21</a>] And so it&#8217;s like we want to be really careful about what we ship there and how. But there&#8217;s a lot of room right now, I think, for exploring how we build software. You do have to be careful. What project are you working on, and what responsibility do you have to users? And that&#8217;s not meant to be a subtweet on what Bun did. It&#8217;s just how I think about how do we want to push the envelope in different places.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1605">26:45</a>] Right. I mean, to me that just reminds me of the pre-LLM tradeoff navigating, I guess, the risk of the code change and how much you verify it. Because if it&#8217;s not risky, it&#8217;s an internal dev tool that only you and I are using. Maybe I just ship it with, I just say stamp it, don&#8217;t worry about it. It&#8217;s not going to break anything. But if it&#8217;s the critical thing that&#8217;s blocking customers from checking out or something, maybe we have two people reviewing it, and there&#8217;s a lot of checks in that part of the codebase.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1642">27:22</a>] So it feels like the similar spectrum is just now there&#8217;s a new level of trying less, which is not just a human doesn&#8217;t even interact with it. You just ship it and maybe you just stamp it, and no one ever even looked at it.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1660">27:40</a>] Whereas before, at least to have written it, a human would have had to.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1665">27:45</a>] Yes, do it. Yeah, yeah, yeah, yeah. I mean, I think. Well, okay, a lot of reactions to that. So, one, I guess, the other thing I would just say quickly about Bun is, sometimes I go over to their repo and it&#8217;s just so different than our repo. And it&#8217;s very interesting to look at because if you look at basically the number one contributor, it&#8217;s the RoboBun bot, and sometimes you&#8217;ll click on PRs by humans and there&#8217;s, like, human accounts talking back and forth.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1694">28:14</a>] But it&#8217;s clear that all the comments were written by agents. And you&#8217;re like, wow, this is very interesting. And so I look at that, but my perspective on that is basically like, okay, I&#8217;m glad that people are like that. They are experimenting with different ways to build software. I don&#8217;t know that I&#8217;m ready to do that. But I do think that basically a lot of things are going to change.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1714">28:34</a>] And I think it&#8217;s very interesting to be experimenting. So I&#8217;m glad that they are experimenting with things. Sorry, is my conclusion there.</p><h3>28:40 &#8212; LLM generated code and open source</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1720">28:40</a>] Do you block if people are just submitting full agent stuff on your repos?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1727">28:47</a>] We do, yeah. So we have an AI policy now, which isn&#8217;t intended to be anti-AI, but it&#8217;s really intended to. Because, I mean, we develop very heavily with LLMs. That is true, but it&#8217;s more intended to be like, how do we retain useful contributions while filtering out net negative interactions from the repo? Because a contributor coming in and posting a comment that their agent wrote, and then we ask them a question and then they just paste the agent response back.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1771">29:31</a>] There&#8217;s very little value in that. We could just ask the question to the agent. And so we want to retain the things that are useful from human contributors, which is insight, ways to reproduce the information we would need to plug it into our agent. Because there&#8217;s no point in just having a conversation with an LLM on a <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> issue. It doesn&#8217;t really do anything for us. Our policy is roughly, you need to understand what you&#8217;re pasting in or what you are submitting, which sounds like a low bar, but it&#8217;s not.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1813">30:13</a>] How do you police that?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1815">30:15</a>] If we can&#8217;t tell and it is agent-written, then I guess that&#8217;s fine, right? But at least right now there are just lots of tells: being way more thorough than a human would ever be, including way too much detail, way more formatting, using terminology that a human wouldn&#8217;t use, inventing random jargon, tons of links. I guess the broad sign would be putting in way more effort in ways that aren&#8217;t necessary, in a way that a human wouldn&#8217;t. It&#8217;s almost always agent-authored.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1851">30:51</a>] But it&#8217;s been, yeah, it&#8217;s been challenging. I mean, I think the hard thing, Zig has a much more strict LLM policy, which is like no LLM-written code or no LLM-assisted code at all in the project, which is very different than us. I mean, all of my PRs now are LLM-written, right? LLM-written. And they had a blog post about this idea of, like, they called it something along the lines of contributor poker, where it&#8217;s like you&#8217;re kind of trying to place bets on contributors.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1879">31:19</a>] Historically, in open source, when a contributor came around, they might be new to the project, but they could show signs of promise or they&#8217;re really engaged or really interested, and you give them feedback, they take the feedback and they fix up their PR, and then the next time they put up a PR, they&#8217;ve learned from that feedback. It&#8217;s the same way as mentoring a new engineer who joins your team. It&#8217;s like they learn from feedback, it compounds, they get better, and then they can actually mentor someone else in the future.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1906">31:46</a>] You&#8217;re making investments in people. In a way, it&#8217;s the same for open source. It&#8217;s like you want to&#8212;I thought about this law historically. It&#8217;s like you want to invest in good contributors who are going to grow to help others and be maintainers in their own right. That&#8217;s kind of gone for PRs, especially those that are written by agents, because people don&#8217;t really learn anything. Someone can put up a PR that&#8217;s written by an agent, and you leave comments, and then they take your comments and put them into the agent, and then update the PR, and then merge it, and it&#8217;s like there&#8217;s no compounding feedback.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1939">32:19</a>] And so I actually do understand that perspective. A lot of the contract between maintainers and contributors is pretty different now. And relatedly, again, not a novel insight, but the cost of putting up a plausible PR has gone to zero, while the cost to review and vet a plausible PR has remained the same and is very high. Especially for the projects that we work on, where I&#8217;m not trying to overstate it, they&#8217;re just hard projects.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=1975">32:55</a>] Ty is our type checker, probably hardest project I&#8217;ve worked on technically to get a change merged because of the complexity in the architecture, but also the problem space. It&#8217;s just very, very hard. If someone puts up a plausible PR to Ty, it might take them two minutes, and then it could take us an hour to understand it. And so it&#8217;s created very poor dynamics in open source. And I don&#8217;t know how we&#8217;ve solved that.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2005">33:25</a>] I think it will continue to get worse until hopefully it gets better in some way. We find ways to solve this, but it&#8217;s certainly something we&#8217;ve experienced in our projects.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2016">33:36</a>] Right. It sounds like review is becoming more and more of a bottleneck for that Bun rewrite. Do you know if it was human? Like, humans read that code, or that was just kind of...</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2026">33:46</a>] No, I don&#8217;t think it was human reviewed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2028">33:48</a>] So it&#8217;s just yolo rewrite the whole thing.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2031">33:51</a>] Yeah, I mean, in a way I think they did have. And I&#8217;m excited for, at least at time of recording, quote unquote. Jared hasn&#8217;t published the blog post yet, which is like. I mean, he keeps making jokes about how the blog post is taking way longer than the rewrite. But I&#8217;m actually really interested in the blog post because I think they did have a pretty sophisticated approach to how the rewrite worked, which will be interesting to read about.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2057">34:17</a>] It wasn&#8217;t just they had one prompt that was, &#8220;Rewrite this in Rust.&#8221; They tried to have, at least my impression from looking at the PR, was there were actually different phases of the rewrite, and each file had to go through a couple different phases that had ways to validate whether it was correct or not. The answer is no, though. I don&#8217;t think a human was reading significant amounts of that code.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2082">34:42</a>] It&#8217;s just not possible.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2083">34:43</a>] Basically, what if we both agreed today that Zig was objectively better than Rust, and then someone on your team said, I can do this rewrite into Zig. Would you block it?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2098">34:58</a>] I think I would block it. I don&#8217;t know whether I&#8217;d be right to do so, but I think we as a team aren&#8217;t there yet. And maybe we will get there, and that&#8217;s how we will feel eventually. But right now, I think we just feel like there&#8217;s still a lot of value in deeply understanding our code and our tool chain and the ecosystem. And so I would say we don&#8217;t want to do that, but it might depend on the project.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2103">35:03</a>] I don&#8217;t know.</p><h3>35:34 &#8212; Performance optimizations</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2134">35:34</a>] All the things that you built, they&#8217;re like an order of magnitude faster than what was available at the time. How much of a difference did Rust make on that graph?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2145">35:45</a>] I think it depends a bit on the project. I think in a rough lot of it was just Rust. And then over time, I think we&#8217;ve improved on that a lot because you can even look at the history of the project. Hopefully no one, hopefully this is still true, but basically over time the project gets faster and sometimes that regresses. But maybe I&#8217;ll put it differently. You could write Ruff in a couple different ways, all in Rust, and they could have really different performance characteristics.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2190">36:30</a>] So that&#8217;s just to say that I generally think of Rust as the floor, or the baseline performance that you get is going to be significantly better. But you still get a lot more out of thinking deeply about performance and design. If you take the same program in Rust and in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, the exact same implementation, to the closest approximation you can get, the Rust one will be faster, but you can then take that program and you can probably optimize it.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2218">36:58</a>] Another 10x. I don&#8217;t know about 10x, but my point is, even within being written in Rust, there&#8217;s a ton of room for how to make things more performant, how to write really performant software. And again, when I started working on Ruff, part of my goal was to learn Rust. And so I did write a lot of bad code, which is fine. I shipped something out that was really helpful to people, but it&#8217;s gotten a lot better, I think, over time, and we&#8217;ve made it more and more performant.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2246">37:26</a>] I think in uv, there was more sort of architectural innovation beyond just being in Rust, especially because uv. So as a package manager, you&#8217;re doing a ton of I/O and downloading files over the network, unzipping things, writing them to disk, moving them around. The linter doesn&#8217;t have to do as much of that. I mean, it has to read all your files, but there&#8217;s not a huge amount of I/O. The package manager is mostly I/O, and then you&#8217;re trying to do things very efficiently.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2281">38:01</a>] So there, it was more. I mean, Rust was important. But I do think that in uv, there are more architectural things that we did, or ways that we thought a lot about performance. The design of the cache is very, very intentional and makes it so that repeated installs of the same package on your machine are near instant because of the way that we lay out the cache and the way that we install from the cache into your projects.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2314">38:34</a>] It basically means that if you&#8217;ve installed a package before, installing it again is extremely cheap, both in terms of disk space and time. And so that&#8217;s a very different design than any of the other <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> package managers had. So, again, it kind of depends on the project. I do tend to think that whether you&#8217;re writing code in Rust or in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, there&#8217;s always room to be thinking about performance.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2339">38:59</a>] You can always make things faster or slower. Even if you&#8217;re writing in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, you can still make things much, much faster by thinking harder about performance and design. So it&#8217;s some mix, but it depends on the project.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2401">40:01</a>] Over the course of the project, do you have a top few things that were implementing the project that were technically challenging or most interesting to you?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2411">40:11</a>] I really like this optimization that Andrew on our team, who goes by BurntSushi, he&#8217;s the author of ripgrep and a bunch of other things. He&#8217;s a really amazing engineer. He did this really cool optimization around how we represent versions. It&#8217;s pretty cool. I mean, basically, if you think about resolving and installing a very complex <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> project, it turns out that we have to parse and create lots of versions, as in 1.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2411">40:11</a>] 0.11.0.2 version objects within the program. We end up parsing and creating a lot of those. And it turns out that actually allocating that memory was expensive given the scale, the number of times we were doing it. He came up with a representation where we can represent 90-something percent of versions with a single u64 integer. So it&#8217;s just way more efficient. And the benchmarks around that and the implementation were very cool.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2472">41:12</a>] It&#8217;s like one of the coolest PRs I&#8217;ve read. I think in Ty, which is our type checker, there&#8217;s a lot of very interesting performance work that&#8217;s happening, especially to make it incremental. So Ty is designed to be a type checker and a language server. The whole system is highly incremental. The idea there is if you&#8217;re in a text editor and you open up one file, you don&#8217;t necessarily want to have to type check your entire project, all your dependencies, every file in the project, just to get analysis for that file, because you don&#8217;t need to.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2515">41:55</a>] So how do you make that work? That&#8217;s sort of like, that would be lazy. You want to be lazy. But the other piece to that is if you have a file open and you edit it, or you have two files open, you edit one of them, you only want to recompute exactly what you need to recompute. You don&#8217;t want to have to go and retype check the entire code base again, especially because maybe you&#8217;re working on a big project like <a href="https://en.wikipedia.org/wiki/PyTorch">PyTorch</a>, and it&#8217;s like you have two files open, you edit one of them, you don&#8217;t want it to like, and then you save.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2542">42:22</a>] You don&#8217;t want it to take two seconds to retype check the project and give you new analysis. So the whole system is built around queries, which is pretty interesting. This is more of a macro design thing, but we built it on top of a framework called Salsa, which is also what rust-analyzer uses, which is the popular Rust language server. And now we&#8217;ve intentionally or inadvertently become very large contributors.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2573">42:53</a>] But the whole system is built around that, which has been very interesting, architecturally interesting.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2578">42:58</a>] So it&#8217;s lazy. So it doesn&#8217;t type check the whole codebase. And it&#8217;s incremental.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2582">43:02</a>] Yeah, the incremental part is the thing that&#8217;s hard because you kind of need a way to basically model a dependency graph of everything that&#8217;s happening in the code. And so then when you change something, we want to just flow the data back through all the different pieces, only the pieces. Yeah, so that took a lot of work, but it has come together, and I&#8217;ve been doing a lot of optimization lately with Codex because it tends to be very good, especially at micro-optimizations.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2621">43:41</a>] And in Ty in particular, we also need to think a lot about memory, not just speed, because if you work on a very large project, you don&#8217;t want it to take many, many, many gigabytes just to run your language server. Ideally you want it to be relatively efficient. So I spend a lot of time now kind of continuously optimizing memory usage and performance, and I can just set a goal that&#8217;s like, try to reduce Salsa memory by like 1% on this project.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2651">44:11</a>] And you can&#8217;t do these things that you might try to do. And it&#8217;s very good at just coming up with very reasonable things. And so I don&#8217;t know, I find that very cool because it&#8217;s kind of like you can almost have continuously. I won&#8217;t say self-optimizing, that&#8217;s a bit grandiose. But you can kind of be continuously just optimizing your software.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2675">44:35</a>] So I&#8217;ve been enjoying that a lot. But most of those are trying to find ways to represent things that take up less memory, or trying to come up with, sometimes it&#8217;s trying to come up with broader redesigns to fix pathological performance.</p><h3>44:54 &#8212; Optimization with AI and combating slop</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2694">44:54</a>] That one optimization with the version numbering and seeing that you can limit it just to a u64. If you think back to some of these more creative optimizations that were done, that was in the human era. Do you think if you just ran Codex and you&#8217;re like, hey, just don&#8217;t break things but lower, do you have faith that it would come up with that? Or is that kind of a step above what?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2720">45:20</a>] That&#8217;s a very interesting question. If you just, at least in my experience, if you use these things to reduce memory or improve performance by some moderate percentage, it will typically come up with things around the edges as opposed to larger redesigns or reconsiderations. But you can get to those larger redesigns if you prompt and collaborate with the agent. If you ask, well, should we be thinking a bit bigger about why we have to represent the data?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2751">45:51</a>] This way, you can get to bigger ideas. That might be an idea that you could have gotten to with prompting if you were like, okay, let&#8217;s start by profiling and figuring out where we&#8217;re spending a lot of time. And then maybe it would eventually come back to you and say, we&#8217;re spending a lot of time in version parsing and version drop and version allocation, blah, blah. And we&#8217;d be like, okay, well, how could we represent these more compactly?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2777">46:17</a>] And we&#8217;d probably start by finding micro-optimizations in the representation of, well, this field. You combine these two fields, save a few bytes, blah, blah, blah. But I think if you kept pushing it, it&#8217;s possible it would get there. But that&#8217;s not the thing it&#8217;s going to come up with by default. Mitchell Hashimoto, who was one of the HashiCorp founders, works on Ghostty. He had a post that was insightful.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2802">46:42</a>] I thought last week or this weekend about a renderer he wrote where it was a really terrible renderer. I&#8217;ll probably butcher the tweet, but it was a really terrible renderer. And then he hadn&#8217;t intentionally. And then he had an LLM optimize it and it made it 10 times faster. And he&#8217;s like, great. Actually, no, my handwritten version was 100 times faster. And it&#8217;s like, if you&#8217;re not using your brain to think about from first principles, how fast should it be and how should the system work?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2830">47:10</a>] Then you just ship all these accumulated. I won&#8217;t necessarily call it slop, but it&#8217;s like you just ship these things, and you&#8217;re like, yeah, oh my God, I made it 10 times faster. But really it should be 100 times faster. And so I think it shouldn&#8217;t be controversial. There&#8217;s still a lot of room for using your brain. But I do think it&#8217;s like, I struggle with this stuff a lot. I mean, I&#8217;m changing how I write software a lot.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2856">47:36</a>] And I&#8217;ve had people on my team because I&#8217;m using agents a lot and also was, to some degree, trying to push our team to use agents more. I mean, some of that for me came from a place of we build tools for software engineers, and a lot of our users are now using agents. And so if everything good that we do&#8212; that&#8217;s an exaggeration, but if everything good that we do comes from understanding our users very well and how they work, then we should probably be using agents so that we understand what it&#8217;s like to build software with them, so that we can build better tools for our users.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2894">48:14</a>] That was part of my motivation, and part of it was just kind of seeing how it can change how you work. And I don&#8217;t know. I think I&#8217;m shipping a lot more. But it has been a difficult process. I&#8217;ve had people on the team tell me, which I think is great that they tell me this, but they&#8217;re like, &#8220;Oh, it used to be the case that whenever you put up a PR, I could review it pretty minimally because I had a lot of confidence in it, in your work.&#8221;</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2923">48:43</a>] And now it&#8217;s like, when you put up a PR, I actually have to review it really closely because you&#8217;re not writing anymore. It&#8217;s like the agent. And I was like, wow, that&#8217;s very interesting. They&#8217;re completely right, which is. And the same thing happens to me. It&#8217;s like, if I go to sleep and then wake up in the morning and look at one of the PRs I put up, I&#8217;m like, wait, this is terrible.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2942">49:02</a>] You know what I mean? It&#8217;s like, you can. It&#8217;s just. It is easy to trick yourself into basically believing work that isn&#8217;t at the same standard as what you would do before. And I don&#8217;t think. I think we have. We don&#8217;t really know what to do with that. And I&#8217;m kind of learning, getting better. And I mean, this was. I think I&#8217;m doing a lot better job now than I was in probably, February or something, where I was like.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2968">49:28</a>] I was just fully agent killed. And I&#8217;m like, now I&#8217;m a little bit more like. But so my point is, that was a powerful moment for me when someone on the team said that, and I was like, wow. You&#8217;re right. And I see that in other people&#8217;s work too. It&#8217;s not just limited to me, but it really does throw a lot of things on their head.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=2990">49:50</a>] If you generate pure slop, that&#8217;s easy, but I think there&#8217;s this gray area where you generate partially AI slop, or it&#8217;s acceptable, but it&#8217;s not at the bar that you used to have. What are the tactics that you&#8217;ve used on the team that have worked for combating this kind of gray area? AI slop.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3012">50:12</a>] I&#8217;d like to get to a world where we&#8217;re not here, and I don&#8217;t know if we&#8217;ll ever get there. I&#8217;d like to get to a world where if you put up a PR and it&#8217;s all green, then the odds of it getting merged are extremely high. Right. Because that would mean that you have automated verification for most of what matters. And we have a lot of that in our projects, and we&#8217;ve tried to add more over time, and I think it&#8217;s helpful in Ty especially.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3042">50:42</a>] We have tons of benchmarks that run under Valgrind through CodSpeed on every PR, and that includes memory. So we benchmark memory usage and simulation time and wall time on every PR. And then we also have a really big suite of ecosystem tests. So basically every time you put up a PR, we run before and after on a bunch of projects in the ecosystem, and then we create this report of the diff of all the diagnostics, like the errors that got removed and added and everything.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3079">51:19</a>] So we have, over time, we&#8217;ve tried to do more and more. These are important. Honestly, even before we had agents, we basically couldn&#8217;t build without these things. But the point is, I want to have more automated verification and try to get better at that. And that includes things too. Like, we basically assume now that anyone on the team that puts up a PR has already run that through Codex Review, probably several times.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3109">51:49</a>] And that&#8217;s basically an assumption. I mean, we could automate that process, but Codex Review is just an agent reviewing the code and double-checking.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3116">51:56</a>] Yeah, it&#8217;s just running Codex and then just doing slash review.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3118">51:58</a>] That&#8217;s it. Yeah. It&#8217;s not that fancy. I mean, it&#8217;s not. Sorry, it&#8217;s not that sophisticated. But it&#8217;s just like because now it&#8217;s like, if a contributor puts up a PR, that&#8217;s the first thing that we do because it tends to find good things. Basically, I think one bucket is: how do you create more automated systems that just help get things, make sure things are right.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3148">52:28</a>] And that also includes things like trying to improve your AGENTS.md file over time. If there are things, if there&#8217;s feedback you&#8217;re giving in a review that the agent&#8217;s not respecting, try to find a way to help the agent learn that, even learn skills. We have some shared skills on the team, stuff like that. None of this stuff is very sophisticated, by the way. It&#8217;s pretty simple. The other piece is how do I make sure that I put in the work to ensure that I&#8217;m creating a good PR.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3176">52:56</a>] And so for me, that&#8217;s like, I really should understand, again, it sounds like a really not. It really sounds like a low bar, but I should understand each line in the PR. I know it&#8217;s crazy, but also I do try to review each PR myself in the <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> UI. This is something I&#8217;ve always found really helpful. If you actually just open up your PR and click files and read through it as if you were a reviewer, you tend to find things that you would miss if you were just looking at your local diff.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3214">53:34</a>] I find that very useful. And then the other is trying to encode skill. I guess this is a little bit more in the first category, but trying to encode in skills things I&#8217;m consistently getting wrong that the agent is getting wrong. A recent example would be I found that I was often getting feedback on PRs that was of the form, this condition here, this if statement, what case is this intended to catch?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3247">54:07</a>] Because if I commented out all the tests pass, and so I was like, okay, I should probably have a pass before I put up any PR where I have the agent go through and check: are these conditions still relevant, or are they left over from a prior refactor or something else? So I don&#8217;t know, I&#8217;m still learning, but those are some of the things I&#8217;ve been doing. Yeah, again, I think it&#8217;s a pretty hard time to be building software, but I felt for a long time where I had a fear that AI was going to make us more productive, but that programming would be, like, a lot less fun.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3284">54:44</a>] Because I just love programming. And I was like, oh, now I&#8217;m gonna have to spend all my time reviewing code and prompting this idiot agent that keeps getting things wrong but ultimately is probably more productive. I actually feel way better about that right now than I did a few months ago. And I don&#8217;t exactly know why. I think it&#8217;s because, well, I think the agents getting better and the tooling getting better and me getting more comfortable with it is one factor.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3321">55:21</a>] I think the other is I&#8217;ve grown to appreciate more of the kinds of things that working with agents has unlocked. The cost of running an experiment is incredibly low. There&#8217;s so many things I&#8217;ve wanted to try or questions I&#8217;ve wanted to answer that I can now answer almost instantly. Sort of a dumb example. In uv, everything is snapshot tested. So basically all of our testing is effectively running uv and verifying the output.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3356">55:56</a>] That&#8217;s how we test basically the entire program. And so that means, and we have a lot of tests. That means we have a lot of test output. And the test output is... it actually ends up in the test files. So we have Rust files. We have a file called lock.rs that tests all our uv lock tests. It&#8217;s all our uv lock tests. And it&#8217;s very, very long in part because it has all the lock output snapshotted in the test.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3381">56:21</a>] And I was like, what if we stored the snapshots in separate files? Would that somehow make our compiles faster? Because then you don&#8217;t have&#8212;technically, that&#8217;s like Rust code. And so it&#8217;s like, would that all disappear and would that make our builds faster or blah, blah, blah. And I&#8217;d always want to do that, but it sounded like to do that experiment as a human would be extremely painful because you have to convert all of those tests.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3408">56:48</a>] And I just had an agent do it in the background while I did a bunch of other things. I got a bunch of data on it, and the answer is no. But it&#8217;s a little bit more nuanced. It actually does have a good impact. If you&#8217;re just iterating on the snapshot outputs, you no longer have to recompile your program at all because the outputs are stored somewhere else anyway.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3428">57:08</a>] That&#8217;s the thing that makes a difference on you. But my point is, I&#8217;m just running experiments like that all day, trying things that used to be hard, used to cost a lot to. To answer the work of then going from that to production. There&#8217;s still real work there, but. So I think one piece is the tools and the agents getting better. The other is things that I just wouldn&#8217;t have been able to do before that I can now do incredibly easily.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3457">57:37</a>] And then the third is, I think I&#8217;m more and more realizing that a lot of the value I get from building software is not retained because some of it is thinking hard about it. It doesn&#8217;t have to be typing out the code, but it&#8217;s thinking hard about the layout of a data structure. A lot of it is merging a PR that causes a user issue. I get a lot of satisfaction from actually fixing and improving something.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3481">58:01</a>] It&#8217;s not necessarily just from typing out the code. I do feel for people a lot who feel like they&#8217;re losing something by working with agents, because I do feel that myself. But I feel better now than I did a few months ago about, like, what it&#8217;s like to work as a software engineer with agents.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3503">58:23</a>] You mentioned in one of the performance optimization examples, you said if you just unleash Codex, it kind of does these local optimizations, and it&#8217;s very good at that. But it&#8217;s not great at some of the, I mean, today it&#8217;s more human ingenuity of system-level optimizations. And it kind of reminded me of this tweet that you had. You said, &#8220;I&#8217;m slightly concerned by how much garbage I would be churning out if I was trying to use these tools without significant software engineering experience.&#8221;</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3536">58:56</a>] Yeah, I remain concerned about that. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3540">59:00</a>] It reminds me, and similar to what Mitchell Hashimoto said in that tweet, there&#8217;s a very big difference between Codex, go and do this, versus you are wielding it like this tool and you&#8217;re kind of prodding in the right direction.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3558">59:18</a>] Yeah. I was curious your thoughts on that because it seemed like it went pretty viral. And I think a lot of people are thinking about that.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3564">59:24</a>] Being a great software engineer is more useful than ever. I don&#8217;t think it&#8217;s still the case that the people on our team who have the strongest engineering skills are the most effective, even at using agents. And it&#8217;s funny because there&#8217;s a lot of talk about token maxing. I don&#8217;t know if you&#8217;re familiar with this.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3594">59:54</a>] Right. Yeah. The idea of, should you have a leaderboard? For example, this is getting tweeted about a lot: companies that have token leaderboards. It&#8217;s like, how do you create some terrible incentives to just use as many tokens as possible? And it&#8217;s funny because that does get tracked. Or sorry, there&#8217;s not a leaderboard. But token usage is something you can look up, for example, internally. I mean, and it&#8217;s interesting to look at because I don&#8217;t care who on the team is using the most tokens.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3623">01:00:23</a>] And there are people on the team who are like, &#8220;I&#8217;m actively trying not to care about that, like being where I am on the token leaderboard,&#8221; which I think is great. I don&#8217;t care. They should just do a great job, and that&#8217;s fine. I don&#8217;t care if they use a lot of tokens or not. But it is interesting to look at the token leaderboard because some of the people on the team who are most productive are using a lot of tokens.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3647">01:00:47</a>] There is a correlation. I&#8217;m not saying it&#8217;s causal, but I do think that a lot of great engineers are able to use agents very effectively and hopefully to multiply their skills. So I do think it would be really hard to be an early career software engineer right now. And I&#8217;m not spending that much time with early career engineers right now. Just based on it, at Astral, we&#8217;re just a very small team, and we tended to hire very senior.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3681">01:01:21</a>] But it&#8217;s sort of hard for me to think about, what would we, like how would I learn, basically? What would the iteration loop be? I would be learning from Codex, I guess, as opposed to the other way around, basically. I mean, a lot of the time I&#8217;m actually instructing Codex and trying to correct it and providing a safeguard on it or Opus or whatever you&#8217;re using. And so, yeah, I think I just don&#8217;t know where you would get.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3722">01:02:02</a>] It would just be way too easy to fall prey to a lot of the bad things that happen when you use agents.</p><h3>01:02:08 &#8212; Learnings as an eng starting a company</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3728">01:02:08</a>] You built this proof of concept in Rust, and it was really well received. But why did you start a company around it, and how did the raising go and all of that? What was the motivation there?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3740">01:02:20</a>] Yeah, yeah. So I left Spring. I was actually convinced to leave by a friend who, a very close friend who left Meta around the same time. And he was like, we should start a company together. And I was like, okay, fine. I mean it wasn&#8217;t quite that simple, but. But I did. I laughed. And then we kind of went into the idea maze of what do we want to build? We went into it not knowing what we wanted to build, and we spent a bunch of time exploring the Venn diagram of ideas.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3775">01:02:55</a>] There were things he was interested in that I thought were not interesting and vice versa. And there was some stuff in the middle, and we spent time exploring the stuff in the middle. But then in all my spare time I was working on developer tools. Because that&#8217;s what I thought was really interesting, and it wasn&#8217;t quite like a fit for him, basically. But around then I was working on Ruff. I was working on a couple other projects that you can see in my <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> that are like, I don&#8217;t know, probably not as interesting, but they were like.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3803">01:03:23</a>] I was experimenting with lots of different things at the time. I was pretty interested in WebAssembly. I wrote a sort of CI/CD toolkit in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a>. It&#8217;s sort of like you wrote pipelines in <a href="https://en.wikipedia.org/wiki/TypeScript">TypeScript</a> and it transpiled them to Docker. And I was like, oh, this could be cool. I don&#8217;t know. I was building a lot of stuff. And at that point in time I decided I wanted to start developing relationships with investors, but that I wasn&#8217;t ready to raise money because I didn&#8217;t know what I was actually going to build.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3834">01:03:54</a>] And so my thinking was I&#8217;m going to try to get connected to some people who like to invest in this kind of stuff, with an eye toward reaching back out to them in three to six months and being like, hey, we had a great conversation, now I&#8217;m working on X. So that was my thinking. So I got connected to a couple investors who invest in software infrastructure and developer tools, and I had some good conversations.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3859">01:04:19</a>] And then the problem is things just can move really quickly. And so I didn&#8217;t actually expect to start raising so soon, but I effectively got convinced. And at the same time, Ruff was growing a lot. And so I effectively got convinced that there was enough there to start a company, which is a very interesting process. I was like, I don&#8217;t quite know how this becomes a company, but I was effectively convinced that there was enough there to build a company.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3883">01:04:43</a>] Which is interesting by the investors. Yeah. Which I&#8217;m grateful for. The other thing I was convinced of was, which is funny, is if I hated it, I could stop in like six months. It&#8217;s like if you decide it&#8217;s not for you, you could give the money back. The investor said that. Yes. Which I actually think was brilliant because, like, I was probably never going to do that, but it did make me feel like there was less pressure.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3914">01:05:14</a>] So, sorry, very weird roundabout details, but basically I was working on Ruff more full time eventually. And then it was kind of, I would say, collaborative with the potential early investors, where it was like we would just have long conversations about my ideas, and then it basically became clear that it was like, there&#8217;s enough here. And I wrote out kind of a product roadmap, which honestly a lot of it stayed true to what we ended up building, what we&#8217;ve ended up building so far at least.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3947">01:05:47</a>] And from there, it basically became like, yeah, we want to fund you. Here are the terms that we would do. And I was like, oh, wow, okay, so I guess it&#8217;s happening. So it was just interesting. And I had never done anything like this before. I mean, I was just an IC software engineer my whole career, and then suddenly I was like starting a company. And it&#8217;s funny because I actually think the decision to go all in and start the company wasn&#8217;t that stressful.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=3982">01:06:22</a>] I actually think a lot of the stress came later as the company started to succeed because at the beginning I kind of had nothing to lose. But then as the company started to succeed, I was like, oh wow, the company&#8217;s kind of working. What&#8217;s going to happen from here? And I had a team of 20 people who were depending on me. And so I don&#8217;t know, just for me as a founder it was like, I&#8217;m pretty risk averse, but I actually found that starting the company was an easier, it was an easier decision than maybe some of the stressful things that came later to navigate.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4019">01:06:59</a>] But it happened very quickly and sort of unintentionally and in a symbiotic way with investors, which I hadn&#8217;t really anticipated.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4029">01:07:09</a>] I didn&#8217;t expect the investors, because I thought the founders were hungry for the fundraising and go, please fund me. But the investors were like, I think it can happen a million different ways.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4039">01:07:19</a>] And actually, I mean, we did three fundraises. We did a C to Series A and a Series B. We actually never even announced the Series A or the Series B. They were, I mean, they were announced in the acquisition blog post. But we sort of complicated. It&#8217;s like we basically just never got around to it, and why don&#8217;t you guys publish it?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4064">01:07:44</a>] Because isn&#8217;t that good marketing for talent?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4066">01:07:46</a>] It&#8217;s good marketing, but basically there kept being reasons that we wanted to wait a little bit. And to me, it always felt like a lot of work, and then I was like, it just never honestly felt like the most important thing because we were building a lot of stuff and the company was going well, and I was like, we don&#8217;t need the marketing. We did finally plan to announce it, but then we got bought, so it didn&#8217;t happen.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4092">01:08:12</a>] Each of those fundraisers was preemptive, basically initiated by investors, which is fortunate because the company was going well and people wanted to invest. But it was also just a funny position for me to be in, I guess. I think I was just always a little bit conservative, and I wasn&#8217;t super aggressive about trying to go out and raise money.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4127">01:08:47</a>] But we were lucky enough to be building things that people really had a lot of confidence in or had a lot of belief in.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4132">01:08:52</a>] That probably gave you a lot of negotiating leverage because one of the best positions you&#8217;d be in is that you&#8217;re willing to walk away, you don&#8217;t need it. And by definition you didn&#8217;t even come there to begin with. They came to you and they said, please take the money.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4147">01:09:07</a>] That&#8217;s true. Yeah, yeah. I wouldn&#8217;t say I&#8217;m a great negotiator, but yeah, I guess at least I had a strong hand. But we were very lucky with investors. We had amazing investors, super supportive and really, really aligned with how I wanted to build the company. Never any conflict, never any pressure to do anything specific, just support. And maybe the only thing, maybe the only piece of consistent feedback is that I could have been more aggressive and.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4182">01:09:42</a>] But I don&#8217;t know, in terms of scaling, growing, and doing. But I liked the way that we operated, and yeah, we just had such good experiences with our investors too. So I felt lucky about that because it can easily go the other way.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4195">01:09:55</a>] I imagine if someone invests, there&#8217;s a promise of future revenue. But how does this company make money? Because you guys give away your software for free, right?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4205">01:10:05</a>] Yeah, we launched a commercial product in, like, August of last year, and it was sort of&#8212;it was called Pyx. It was kind of like a hosted counterpart to uv. So a lot of our, a lot of uv users, will purchase a private registry software instead of using the public registries, either for security or maybe they need to publish their own private artifacts, things like that. And we basically built our own private registry that had some first-class support for uv.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4237">01:10:37</a>] And it let us solve a bunch of problems that we saw in our issue tracker around users using other solutions. Also let us build something really fast because we vertically integrated the client and the server. There were special things they could do to be really fast and really easy to use and all that. And we spent a while selling that, and our revenue actually grew pretty well. I mean, it wasn&#8217;t like, I think in this era of AI, you&#8217;re constantly hearing about fastest company to go to 100 million in revenue and things like that.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4270">01:11:10</a>] It didn&#8217;t look like that, but we were definitely able to sell this to some big enterprises. And the cool thing was we had a very good, because we built the open source and everyone was using our open source, we had a really good funnel. We had a small number of extremely high quality customers, is the way I would put it. So the general idea for how we wanted to make money was to keep building open source and the tooling. The tooling remains free, open source, entirely noncommercial.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4303">01:11:43</a>] And then we build software that&#8217;s kind of like the natural next thing you need when you&#8217;re using our tooling. So if you&#8217;re using uv, there are likely problems we can solve for you with a private registry. And then the funnel is people using uv, they have these problems, and then we sell them the registry. And the registry was meant to be one piece of a platform of a bunch of different things.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4323">01:12:03</a>] We sell like a <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> cloud. So that&#8217;s what we were building towards. I mean, part of the cool thing about the acquisition is we&#8217;ll be able to take parts of that platform and make it basically freely available. Like we did just as an example. A big part of that was we did a lot of, like, in <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a>, there&#8217;s a big part of the community that uses <a href="https://en.wikipedia.org/wiki/OpenAI">Python</a> with GPUs like <a href="https://en.wikipedia.org/wiki/PyTorch">PyTorch</a>. And then there&#8217;s like a whole ecosystem around <a href="https://en.wikipedia.org/wiki/PyTorch">PyTorch</a> of software that kind of builds against <a href="https://en.wikipedia.org/wiki/PyTorch">PyTorch</a>.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4364">01:12:44</a>] And that&#8217;s for a variety of reasons. The ergonomics around that are not great. It can be kind of hard to work with, hard to install. The right version depends on the GPU that you have and what version of <a href="https://en.wikipedia.org/wiki/CUDA">CUDA</a> you have installed and all this stuff. And so we kind of tried to solve that in a certain way where we built our own distribution. You could think of it kind of like in the sense of a Linux distribution.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4391">01:13:11</a>] We pre-built a lot of things like that that all worked together and made that available to customers. And so now we can actually take that and just make that freely available to everyone. Because we&#8217;re no longer, we no longer are trying to build an independent business. We&#8217;re just trying to build great tools that grow a broader ecosystem. So, I mean, more to come on that. But that&#8217;s been kind of like one cool thing that I&#8217;m excited about.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4414">01:13:34</a>] As a first-time founder, especially with the engineering background, was there anything that was a surprising learning or something you would share with someone if they were an engineer and they were going down that founder path?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4428">01:13:48</a>] There are so many different ways to do it. I&#8217;m actually not a huge fan of people giving startup advice in general because so much of this industry is survivorship bias. You could give the exact same advice. You can go to two talks and people give you two incredibly successful founders give you completely contradictory advice, and you&#8217;re like, I don&#8217;t really know what to do. And I tend to view that as you&#8217;ve got to figure out what&#8217;s true to you.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4456">01:14:16</a>] And so there&#8217;s all these decisions you have to make: in person versus remote, right? Have a principle and stick to it. Either one can be great. We built the whole company remotely. It went super well. It was great. Do I think there are benefits to being a person? Of course. Would I have been able to hire the team that we put together if we were in person? Absolutely not. And so there are trade-offs to all these things.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4480">01:14:40</a>] You just have to be principled and figure out what resonates with you. And there&#8217;s a million of those decisions that you have to think through. I feel fortunate in a lot of the ways I got to run the company. I tried to spend very minimal time fundraising and with investors, and I tried to optimize for investors I really trusted, as opposed to trying to talk to a billion different investors and pit them all against each other to get the best terms.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4507">01:15:07</a>] My goal was always to find a partner that I really trust and get good terms, and then move quickly. Spend as little time on fundraising as possible. And so think about what you care about and focus on that. There&#8217;s no right answer to a lot of this stuff. I got a lot of advice to get a cofounder when I first started the company, and I basically ignored that. I mean, well, sorry, I did ignore it because I didn&#8217;t get a cofounder, I guess.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4532">01:15:32</a>] For me, I kind of knew they were right probably, but I didn&#8217;t really have someone in mind. I didn&#8217;t want to force it, and so I just didn&#8217;t. And over time, I felt like it would have been helpful to have someone that&#8217;s kind of in the trenches with you in the same way. But you find other ways to compensate for it. I don&#8217;t know.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4561">01:16:01</a>] I lean on my wife a lot, probably more than I should. Probably more open with my employees than I otherwise would be, which I think can be good or bad. But it&#8217;s like, I share a lot with the early team, especially when I need advice on things. Try to build a network of other founders, especially here in New York. I have a pretty good group of.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4582">01:16:22</a>] There&#8217;s like six or seven of us who. We try to get dinner twice a year. Doesn&#8217;t sound like a lot, but everyone&#8217;s busy. Yeah, yeah, yeah. And that&#8217;s been a really helpful support system. So you just find other ways to compensate. But yeah, it&#8217;s. I don&#8217;t know. There&#8217;s no. I really don&#8217;t. I really think there&#8217;s not one right way to do it. Yeah, and it extends everywhere.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4609">01:16:49</a>] I mean, also, everyone on X is talking about, should everyone in your company be working seven days a week? This kind of thing. And it&#8217;s like, I don&#8217;t know, we just built a very different company and we were able to do it. That was not. I mean, I work basically all the time, and that&#8217;s a choice I made and that I&#8217;m very happy with.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4633">01:17:13</a>] But I don&#8217;t expect people on the team to work all the time. And I want it to be a company where you can be highly, highly successful working normal hours, but also where you&#8217;re rewarded for your output and your performance. And I try to keep all those things in sync, but again, there&#8217;s no one right way to do it. Other companies, that is their culture and that&#8217;s what they want to do.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4658">01:17:38</a>] And sure, whatever. If that&#8217;s what you sign up for, just make sure that you&#8217;re honest about what you&#8217;re doing. So I don&#8217;t know, I just think figure out what you care about and, in my opinion, don&#8217;t focus too much on how you think you should be doing things. Think about how you want to be doing things and what actually fits with how you want to work.</p><h3>01:17:55 &#8212; Top technical talk recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4675">01:17:55</a>] What&#8217;s your top book recommendation, whether technical or maybe something that helped you as a founder?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4681">01:18:01</a>] I don&#8217;t read a lot of technical material. I have definitely watched a few talks that were very influential to me, which is a little bit different.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4690">01:18:10</a>] What are those talks?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4692">01:18:12</a>] The Zig creator, Andrew Kelley had a really good talk about data-oriented design that I learned a lot from. I mean, even apart from learning things, it sort of inspired me to look at software quite differently. And so I think about that talk a lot. Yeah, things like that. It&#8217;s really good. Yeah, it&#8217;s about. Well, I haven&#8217;t watched it in a while, so I&#8217;ll probably mess it up, but it&#8217;s like a lot of it&#8217;s about basically design decisions they made in the Zig compiler and just thinking about memory and allocation and all this stuff.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4724">01:18:44</a>] And I was like, I watched that when I was pretty early in my career of systems programming, and I was like, wow, these people really care about what they&#8217;re doing. And I was like, that&#8217;s cool.</p><h3>01:18:56 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4736">01:18:56</a>] Yeah. And then last question is, if you could go back to the beginning of your career, right when you graduated college, what advice would you give yourself, knowing what you know now?</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4747">01:19:07</a>] I don&#8217;t know, it&#8217;s so hard to give myself advice without feeling like I have all this hindsight bias. I mean, I think, I guess there are some things where I&#8217;d basically tell myself that it is the right decision and I shouldn&#8217;t worry so much. But it&#8217;s hard to know how that would have played out if I did it a thousand times over. But for example, I think when I left school and I went to work at Khan Academy, part of me was definitely like, wow, should I be taking more of a big tech job, and am I gonna really regret not having that experience?</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4778">01:19:38</a>] The Khan Academy pay was good, but the big tech pay was certainly better. And I was seeing my friends get huge bonuses and all this stuff, and I felt sort of insecure about whether I should be doing that. And I think it actually, in the long term, really paid off. I mean, maybe it would have been great, but the experience I had was really what I wanted, and I thought I made that decision for the right reason.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4806">01:20:06</a>] So I think some of these things are, I don&#8217;t know, everything happens for a reason, you know, and it&#8217;s like, have some faith in what you&#8217;re doing for the right, you know, for the. If you&#8217;re making decisions based on principles, I think it can work out.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4824">01:20:24</a>] When we started this conversation, you said the reason you got into what you got into is because you tried all these different ecosystems and you had a broader perspective. Imagine if you went to <a href="https://en.wikipedia.org/wiki/Google">Google</a> and you just kind of used <a href="https://en.wikipedia.org/wiki/Google">Google</a>&#8216;s closed thing, and then you might not have had the same insight.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4843">01:20:43</a>] Yeah. And I think about this with Spring, too, because Spring, I was there for four and a half years, and then I left, right, as they decided to do a pivot, and then they ended up joining <a href="https://en.wikipedia.org/wiki/Genentech">Genentech</a>. But we had set out to build this kind of crazy drug discovery company. We were focused on aging, and it was all with computer vision. It was very, very ambitious what we were trying to do. Yeah, yeah.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4870">01:21:10</a>] And then afterwards, I was like, oh, the company didn&#8217;t have the huge exit or the huge. We didn&#8217;t achieve all of our dreams. Did I waste all my time? And ultimately I was like, no, because I learned by being an early employee, seeing all these stages. I learned so much about how to run my own company and also all the technical learnings that I rolled into my own company. So, again, I feel weird giving this advice because everything has worked out really well, but I feel like there are moments in time where I sort of doubted some of my career decisions.</p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4904">01:21:44</a>] And then in hindsight, you can only connect the dots. Looking backwards, everything kind of added up to finding the thing I really love, which is building tools and feeding into that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4915">01:21:55</a>] Awesome. Well, thank you so much for your time, Charlie. I really appreciate it.</p><p><strong>Charlie:</strong></p><p>[<a href="https://youtu.be/Iw65FD4MGgs?t=4918">01:21:58</a>] Yeah, thank you. No, it was super fun.</p>]]></content:encoded></item><item><title><![CDATA[Google DeepMind Pre-Training Lead: How To Get a Job at a Frontier Lab | Vlad Feinberg]]></title><description><![CDATA[Skills frontier labs need, differences between software engineering and research, concrete steps engineers can take]]></description><link>https://www.developing.dev/p/google-deepmind-pre-training-lead</link><guid isPermaLink="false">https://www.developing.dev/p/google-deepmind-pre-training-lead</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 15 Jun 2026 09:01:40 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202039062/17cc08b5e576b5dd664cc86df77883b1.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/vladimirfeinberg/">Vlad Feinberg</a> is Google DeepMind&#8217;s pre-training area lead and I asked him all about how to land a job at a frontier lab like Google DeepMind, Anthropic or OpenAI.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/cDyi91onoJ8">YouTube</a>, <a href="https://open.spotify.com/episode/5XgGwEsWKDKmXsXErgCf74?si=GorGk0nzQFqBnBFq7UM7KQ">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-cDyi91onoJ8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;cDyi91onoJ8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/cDyi91onoJ8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/202039062/0033-skills-frontier-labs-need">00:33 - Skills frontier labs need</a></p><p><a href="https://www.developing.dev/i/202039062/0845-the-difference-between-ai-research-and-engineering">08:45 - The difference between AI research and engineering</a></p><p><a href="https://www.developing.dev/i/202039062/2141-domains-that-matter-for-the-frontier">21:41 - Domains that matter for the frontier</a></p><p><a href="https://www.developing.dev/i/202039062/3050-marketing-yourself-to-frontier-labs">30:50 - Marketing yourself to frontier labs</a></p><p><a href="https://www.developing.dev/i/202039062/3513-concrete-steps-engineers-can-take">35:13 - Concrete steps engineers can take</a></p><p><a href="https://www.developing.dev/i/202039062/3829-overview-of-pre-training-areas">38:29 - Overview of pre-training areas</a></p><p><a href="https://www.developing.dev/i/202039062/4723-jeff-dean-spot-bonus-story">47:23 - Jeff Dean spot bonus story</a></p><p><a href="https://www.developing.dev/i/202039062/5014-favorite-gemini-war-story">50:14 - Favorite Gemini war story</a></p><p><a href="https://www.developing.dev/i/202039062/5859-advice-for-his-younger-self">58:59 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:33 &#8212; Skills frontier labs need</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=33">00:33</a>] You wrote this post that was titled &#8220;How to Land a Job at a Frontier Lab&#8221;. What are the skills that are kind of in demand in Frontier Labs? Maybe we can talk about the shape of the work.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=49">00:49</a>] There&#8217;s quite a range of different things that Frontier Labs require at this point. LLMs are artifacts that are connected to research and product in ways that machine learning really hasn&#8217;t been as connected to before. And so it really touches on so many different things. The goal of my post was to propose just a couple tangible directions in which labs could require a certain set of skills, not to be fully exhaustive.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=84">01:24</a>] And really the ones that I dive into have to do with kernel development and low-level engineering to improve the runtime for these LLMs in practice. And so that was a particular skill that I see voracious demand for across all the different Frontier Labs and among different projects within the labs. So that seemed like a very sharp one to call out as an overall need. And so specifically, whenever we&#8217;re doing a research project that involves changing the architecture for the neural net in a particular way, or rethinking how we might do serving to do better KV caching or something like that, again, across the stack, you just need to be able to implement these new techniques in efficient ways.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=141">02:21</a>] And the inner loop of all these different changes is creating software artifacts that can function at large scales with high throughput, low latency. And this is just fundamental work that&#8217;s tied to classical backend engineering thinking. So yeah, it seemed like a very open thing for people to specialize in.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=166">02:46</a>] My friends that work at OpenAI and <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> have told me, there&#8217;s a distinction between an applied org and the research org, and I was wondering if <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a> has a similar distinction and if you could speak about what that difference is.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=183">03:03</a>] So we have different focus areas and, for instance, within <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a>, there&#8217;s a team that focuses on how we can use our Gemini LLMs to better inform Search results. And so that might be in some way an applied version of the LLMs. But I am hesitant to make a very sharp distinction here because there&#8217;s so much actual hard research that has to go into this kind of level of product integration. Specifically, for the one I mentioned, quite a lot of work goes into making sure that these LLMs are factual and can cite sources to have very precise, grounded answers.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=233">03:53</a>] Assessing the quality of these sources to make sure that you&#8217;re not referring to anything that&#8217;s sarcastic or a joke. This is, I guess, a good example of how even in product-specific applied AI verticals, you&#8217;re still doing research. That being said, there&#8217;s definitely what I would say is very classical LLM research teams, pre-training, post-training. These are things that are still standalone teams inside of <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a> that are focused on what I would say is creating SOTA models.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=270">04:30</a>] Pure research. Again, the caveat is the pure research that we do. The extent that it matters is the extent to which we can realize it. And so we&#8217;re just as responsible with delivering these models and making sure they train stably and actually being like the SREs of sorts for the training run to make sure that the model training is going smoothly as we are for coming up with the recipes to make these LLMs.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=299">04:59</a>] And you can&#8217;t separate those two roles. It&#8217;s really crucial to wear both of those hats. So yeah, I think you can draw up a spectrum between research and applied. But no matter what, in today&#8217;s world, I think everyone needs to be fluid across that spectrum.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=319">05:19</a>] I noticed there&#8217;s also another spectrum of software engineer to pure AI researcher. And how do you think of that spectrum? Software engineering versus AI researcher roles.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=331">05:31</a>] So I guess in my case specifically, I think a lot of what we do and a lot of the new techniques that we develop, the groundwork is laid in infrastructure investment. So I can walk through what my team does a little bit more detail later. But one of the verticals is distillation. And in order to do distillation, it&#8217;s some way of transferring the knowledge or some form of statistics about the underlying dataset through a teacher model into the student model to make the student model better than if it hadn&#8217;t ever seen these auxiliary statistics from the teacher.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=381">06:21</a>] When you&#8217;re talking about statistics derived from a massive LLM applied to trillions and trillions of tokens, you&#8217;re talking about a level of FLOPs investment that is millions and millions of dollars. And that in turn means that you have to be able to think through how do you optimize the system to be as efficient as possible? Because every operation that we&#8217;re performing is multiplied by such a large factor that every second counts, every byte of storage counts.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=420">07:00</a>] And quite a bit of that work is good old-fashioned software engineering. In particular, the infrastructure for distillation has evolved through maybe three to four generations. At this point in each one, we&#8217;ve taken a step back, looked at what kind of research methods we&#8217;ve been applying for distillation, holistically thought about how do we broaden what the infrastructure is capable of. And there&#8217;s definitely a couple discrete points where rethinking the system design of how we perform distillation enables us to do research on distillation methods much more quickly.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=469">07:49</a>] And so it&#8217;s this kind of investment that, okay, this four-month or whatever rewrite of our distillation infrastructure then results in a dramatically new understanding of distillation scaling laws that translates to really strong models. So it really requires just work across the stack. And I can&#8217;t imagine that we would have gotten results like Flash 3.0 without having made those distillation infrastructure investments that are at the end of the day things that started with a good old-fashioned design doc and thinking about what the right abstractions are for generating these teacher statistics, coming up with the right storage system for them, thinking through what could support the reading and writing across multiple different data centers at this scale.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=522">08:42</a>] Really classical distributed systems problems?</p><h3>08:45 &#8212; The difference between AI research and engineering</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=525">08:45</a>] Yeah, I mean it sounds like there&#8217;s a lot of software engineering backend infra-type problems given just the scale of the compute at this point. It still feels like though at some point in that spectrum there&#8217;s some crossover where there&#8217;s these new skills. Like somewhere where if you had, you took an arbitrary backend engineer and you placed them to, I don&#8217;t know, adjust the model architecture or something like that, is like a bit of a jump more than the infra work.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=556">09:16</a>] How do you see that distinction?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=558">09:18</a>] Yeah, so I think there is a crossover point in terms of doing research where research is an endeavor where the payoffs become a lot higher risk, higher reward. And we have this notion of kind of research taste, which is some high-level intuition about what path you should be proceeding through the DAG of the multiple different milestones that you need to accomplish in a particular project.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=591">09:51</a>] In some sense, we can view software engineering projects through a similar DAG, where you have all of these intermediate artifacts that you want to hit in a software program to get to the final result. But in the software engineering case, the DAG is more or less deterministic, where you build one service, then a different service, then a third service, and you figure out your storage infrastructure layer first, that kind of thing, and you can just make monotone progress.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=619">10:19</a>] But in the research case, you have to kind of explore this DAG, which is now stochastic, because some of the nodes, which might be some research ideas or some aspect of getting to a final goal, may or may not work out. And I think that requires a bit of a mindset shift. And that kind of mindset shift takes a while to learn, and it takes specialized skills to learn. This would be the kind of skills you pick up in a PhD, for instance.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=649">10:49</a>] One succinct way I could put it, there&#8217;s a really excellent post by this professor Jacob Steinhardt, and I love to frame a lot of the research work that I do in this way. And it&#8217;s research as an MDP. So MDP here. Markov decision process. Again, we have this high-level idea of a stochastic dependency graph between different milestones in a research project where you might need to have a certain kind of result or prove a certain kind of theorem before you get to a certain kind of conclusion.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=684">11:24</a>] Similarly, for a machine learning research project, you might need to have this and that featurization working before you can get this and that ImageNet accuracy or something like that, expanding those nodes in this graph. It&#8217;s this stochastic endeavor where these approaches may or may not work out. And whether or not one works out opens up a set of new possibilities for you. The approach that you might have in the software engineering case, where you could fully write out, here are all the paths to the goal across walking this graph.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=721">12:01</a>] What&#8217;s the shortest path to your goal? That approach is not optimal in the research case because if all of a sudden the transitions between the edges in this graph become unreliable, and some of the nodes you might not even be aware of, it might be a hidden MDP, then the way that you might approach this problem would really differ. And in particular, you have to factor in the success rate and the time investment that you&#8217;re going to be putting into these different research ideas, as well as a priori estimating what those different rates are.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=761">12:41</a>] And that&#8217;s a very different exercise than writing up what the design for your software engineering project might be. And it&#8217;s this skill set of building an intuition of how likely an approach is to work out without having yet done that approach that I think people often correlate with this research taste notion, but that&#8217;s exactly the one that you need to build up in order to properly traverse this MDP for the research projects and just generally the nature of the research work.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=792">13:12</a>] It sounds like you&#8217;re saying that there&#8217;s a lot more uncertainty. I&#8217;m still trying to get a sense of the nature of the work. If you threw this backend engineer into a team that&#8217;s doing research, what are those concrete examples where they fall short?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=814">13:34</a>] I think the very first thing that comes to mind is having the right context for the research landscape in which you&#8217;re operating. So quite a bit of research work involves almost this kind of. You have to take on this very humble viewpoint of there&#8217;s been quite a lot of investment in related work in the past and until I know the sum total of humanity&#8217;s bleeding edge in this topic, I&#8217;m definitely not going to be able to further that bleeding edge.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=855">14:15</a>] So building up a solid understanding of past work in a particular area and doing that related literature review is maybe the first thing that I would imagine people might stumble on: having read and having the skills to effectively traverse the historical citation tree for a particular topic. Because you don&#8217;t have the time to read all of these different papers, you need to build up a sense of what are the high-value papers and what are the ways in which I can assess if a paper is worth reading without fully reading it.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=892">14:52</a>] That&#8217;s the first thing that comes to mind as the skill that people need to build up. Even to be able to read these research-level papers. You have to have a background in machine learning and some computer science, and depending on the paper and depending on the domain, there might be all sorts of prerequisites in terms of the underlying math and coursework that you would want to have to properly understand.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=921">15:21</a>] So that&#8217;s quite important to be able to have a deep understanding of what methodology is available, because you really won&#8217;t have a lot of hope of improving upon the methodology if you don&#8217;t understand what&#8217;s there already. So I think I mentioned earlier, one of the things that my team works on is distillation. In order to advance our understanding in distillation for large language models, you have to have a good understanding of what we&#8217;re trying to do with LLMs.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=959">15:59</a>] Just to give a cursory overview here, the name of the game for LLM research is, especially in pre-training, scaling laws. And so what are scaling laws? People focus a lot about this power law structure and the fact that you have this and that exponent, but what matters is less so the functional form. What matters is, for a given recipe of scaling up your LLM, as you invest more and more FLOPs into the pre-training run of an LLM, you have to be able to predict what the final test loss of this LLM is going to be.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=999">16:39</a>] And why do we care about this question? Why do we care about predicting what our generalization error is? In the classical machine learning world, say we&#8217;re trying to win ImageNet, we have our test loss, which is our classification error for 1,000 different classes, and you run your VGG or your ResNet proposal to get that classification error. That&#8217;s an estimate of how well that model does at classifying among those thousand classes, various different images.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1035">17:15</a>] We can estimate how good our method is going to be by taking a validation set. And then whenever we have an architecture idea for a neural net, we just train it, and then we do a bunch of validation set runs, and we get a cross-validation error that is itself an estimator of our final test error. In this way, you can just iterate on different ideas through this process. But what&#8217;s different in LLM world is every single time you go up for a pre-training run, you&#8217;re about to put in more FLOPs into this run than you&#8217;ve ever done before.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1065">17:45</a>] So it&#8217;s in some sense like a one-shot version of this ImageNet problem. You never get to see the full ImageNet training dataset. You have to practice on MNIST and then CIFAR and then maybe based on those, you try to come up with a method that just works right off the bat on ImageNet. If you were to just do that by itself, as I&#8217;m sure many people have tried, certainly when I was learning how to do all of those different things, you get something, it works really great on MNIST, it maybe even works on CIFAR, and then all of a sudden it breaks on ImageNet.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1096">18:16</a>] You&#8217;ll find out that things don&#8217;t just generalize easily across scale like this. And so much of what we do for LLMs is coming up with recipes, where a recipe is this function that goes from number of FLOPs you&#8217;d like to train on to a training routine for this LLM. And if you can couple this recipe with a prediction rule that can predict accurately what your LLM accuracy is going to be, then you&#8217;re able to make decisions about how to improve your recipe because you can use that prediction.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1129">18:49</a>] That is a ton of context on what LLM research looks like in general. But that&#8217;s an understanding that we got to, that we even thought was feasible thanks to so much initial LLM scaling work that we&#8217;ve seen across the Kaplan paper, across Chinchilla. Since those two papers, there&#8217;s been a lot more work in terms of what other factors are there beyond number of params and number of tokens that you train on that influence your prediction accuracy, like number of unique tokens, for instance.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1168">19:28</a>] But I would say those two foundational papers for LLMs, those are informed by an even longer line of different scaling works going back to, say, the original GPTs. And then Google has had a ton of scaling work across its PaLM papers. This is just a set of works that have informed that viewpoint that I described earlier, that you kind of just need to build up by having gone through that literature review yourself.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1201">20:01</a>] If you were, for instance, if you were trying to pick someone that was going on your team, and the way that you would judge their fitness to help you push the frontier is their understanding of the frontier, including the existing literature, which requires all these prerequisites. I think you called it mathematical maturity in your post.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1223">20:23</a>] Yeah, so I think it&#8217;s easy to read and understand those papers once you have mathematical maturity. So I guess the ones I mentioned in particular, nowadays they&#8217;re table stakes, so I would expect candidates to be familiar with them. I think the general skill set is being able to dive into a paper of that level and then understand it, being able to take a research idea from a paper and implement it yourself. That&#8217;s just a very important skill set to be able to have.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1266">21:06</a>] We get all sorts of different ideas presented. They might not all directly apply to our domain, but if you can deeply understand them, then you can iterate on them and you can improve them inside our domain. And so when we assess people who can work with the mathematical concepts in these machine learning papers, that&#8217;s, I guess, the key skill there. That would be evidence that you can go pick up this arbitrary paper and see to what extent these ideas carry over in the Google setting.</p><h3>21:41 &#8212; Domains that matter for the frontier</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1301">21:41</a>] This probably won&#8217;t be exhaustive, but I&#8217;d be curious to hear other domains that maybe people could dig into to see what kind of matters in frontier AI research. So you&#8217;d mentioned distillation. You also mentioned kernels. It sounds like kernels are helpful everywhere, but are there other areas that come to mind if you were just rattle off areas that are not necessarily exhaustive?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1325">22:05</a>] One thing that I think is quite powerful is actually programming language research. So by looking into how we can create abstractions at the programming language level, we can facilitate kernel development. I think ThunderKittens is a really good example of this. Coming up with an abstraction that allows you to write kernels through four functions instead of arbitrary globs of C code allows you to move really quickly in developing algorithms that fully utilize hardware.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1362">22:42</a>] At that point, it&#8217;s not about the PL research itself, it&#8217;s about having a passion for these kinds of programming language abstractions and working with low-level hardware people who are interested in and will try to work with CuTe DSL. This kind of thing, where there&#8217;s a lot of hardware-specific domain-specific languages. One other thing that comes to mind besides PL and scaling law literature would be reinforcement learning literature.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1397">23:17</a>] So in particular, ever since RLHF, I think we&#8217;ve seen that deep RL algorithms like PPO do have a place in production systems, and there was a time where that was in question. But now it&#8217;s pretty unanimous that we see these kinds of algorithms applied to real production systems. And the theory behind that, you kind of have to start with the basics for reinforcement learning and work your way up to the myriad value-based methods and policy gradient methods that we have today.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1440">24:00</a>] That&#8217;s another domain that I think is just a very rich literature tree to crawl. And then for more of the backend engineer folks, beyond just the kernels themselves, there&#8217;s I think a pretty fun overlap between distributed systems and optimization work, where figuring out how to design neural net training algorithms that allow for training across many GPUs. There&#8217;s all sorts of fun challenges between asynchronicity, how up to date your gradients are, how pipelining affects the staleness.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1483">24:43</a>] All of these system choices that you could make in your training algorithm design will impact convergence and the final quality of your neural net. And those are things that can be analyzed independently of the LLM setting and have been for a while, especially if you&#8217;re more infra inclined, then having a good understanding of how those different algorithms work is a really good place to start.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1509">25:09</a>] Do you see any difference between the demands of the different Frontier Lab? So, for instance, if someone wants to work at <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a>, is there a particular area that you see <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a> cares about more than <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>? For instance, I think in terms of the skill set, it&#8217;s probably pretty similar.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1524">25:24</a>] Yeah, I think there&#8217;s maybe differences in business strategy and the set of offerings that&#8217;s a function of the specialties of the labs and the kind of different customers that the labs could have. But I would say that there&#8217;s quite a lot of overlap between the labs in terms of what people look for. And when I posted my post, you would see people from both OpenAI and <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> saying, yeah, we agree with this advice.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1564">26:04</a>] And so I think that&#8217;s just a little bit of evidence towards that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1569">26:09</a>] I think one reason for the huge demand for wanting to go closer to AI research is because people are thinking, oh, software engineering is not going to be as important in the future. Is there a similar thought when it comes to research where LLMs are also going to handle a lot of that work as well? So there&#8217;s no reason to favor AI research versus software engineering.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1594">26:34</a>] So I think the research skill set is going to become increasingly important. So I would say being able to handle stochastic components in the planning of your work is just going to be a larger and larger part of how we approach our jobs. Figuring out how to leverage AI in whatever thing you work on, which doesn&#8217;t even have to be software related, is just an important muscle to start building right away because these components aren&#8217;t deterministic.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1629">27:09</a>] And thinking about how do I construct systems around these LLMs to do my job more effectively, that&#8217;s going to be the thing that sets you apart in the future. And I think that&#8217;s true no matter what you&#8217;re going to be doing. Look, I think there&#8217;s FUD everywhere, especially with some of the approach to marketing that some people have in terms of AI. It&#8217;s FUD that is being intentionally leveraged. And so I feel like people should really just focus on themselves and trying to be more productive themselves.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1662">27:42</a>] I don&#8217;t think that AI is going to replace all of our roles. And so the reason for that is that one of the important aspects of what we do as humans in an organization, which is really this web of trust from this organization that is this pool of resources and this pool of people that manages these resources. One of the important things that we do is we allocate those resources towards certain goals.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1695">28:15</a>] And even when we can accelerate our execution, there&#8217;s an element of making decisions around how we allocate these resources that will always be something that needs to be attributable to a human making that decision. That&#8217;s simply because you can&#8217;t hand off blame to AI. So we at this point have LLMs that really deeply understand law, and they could review your contract for you or something like that, but they can&#8217;t represent you in court because they can&#8217;t be disbarred.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1733">28:53</a>] And so that&#8217;s, I think, a really sharp way that I might describe, okay, this is why the legal profession will go on, even though LLMs are really good at recalling precedent, is you want to have someone who is responsible, who can validate the output of AI to perform legal work more effectively for you rather than hand off your legal defense to an LLM.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1765">29:25</a>] Yeah, I think the FUD was actually the original motivation for your post.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1770">29:30</a>] Yeah, I really think that the mindset that people should have is a constructive one. And so there was a tweet that I saw, I think by Deedy, that was like some long-form fear mongering about AI permanent underclass or something like that. And it&#8217;s easy to get stuck in that loop. But I think the important thing to think about is that we all have agency over our future and we can start investing in skills that matter for tomorrow, today.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1809">30:09</a>] And that&#8217;s really the only thing you should be doing, right? Worrying about it is not going to help you. And so part of why I wanted to write this post is in response to that, because it was something that I could see echoed. I gave a lecture at Princeton a while back, and a big question that came up is, how do I work at <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a>? And it&#8217;s something that just when people find out what I do, that&#8217;s the top question people ask.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1841">30:41</a>] I figured it would be helpful to add a little bit more constructive direction to the discourse here.</p><h3>30:50 &#8212; Marketing yourself to frontier labs</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1850">30:50</a>] One last thing on the post because if you think about getting a role, there&#8217;s obviously the skills, and we talked a lot about the skills and your fitness for the role, but there&#8217;s also kind of the signaling for that role and what is kind of valued. If you were to be marketing yourself to one of these Frontier Labs, what signals matter most?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1873">31:13</a>] Actual evidence that you&#8217;ve created something of use to other people along the line of kernels. Right. You can take any of the many open-source LLMs that we have and optimize them. You don&#8217;t have to make them better. In every case, you could show that, oh, I have an improvement for this and that setting. It doesn&#8217;t even have to be something that speeds up the model. On GPU, there&#8217;s all sorts of open-source stacks like vLLM.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1903">31:43</a>] There&#8217;s a lot of other things that you can do besides accelerating the LLM inference on device. The serving stack that surrounds LLMs is a very sophisticated distributed system that has to maintain this KV cache memory and deal with all sorts of load balancing and request queuing and very common problems for backend servers. These projects are always looking for help. So contributions to vLLM or SGLang or demonstrations with TensorRT, they have, I think, a distributed system called Dynamo that allows for disaggregated serving where you could show that you made a project using these components, you improved these components.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1954">32:34</a>] That would be an extremely positive signal for any candidate that I&#8217;m looking at and a very welcome contribution to open source, I think.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1964">32:44</a>] Also, a lot of what we said is kind of assuming the path of external hire into Frontier Lab, but a lot of these Frontier Labs have large organizations that aren&#8217;t necessarily doing the cutting-edge Frontier work. So let&#8217;s say, yeah, for instance, I mean, <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a> versus let&#8217;s say there&#8217;s some infrastructure eng that&#8217;s working on Search and they have the backend skill set, maybe not as much domain context, and they try to internal transfer to <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a>.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=1996">33:16</a>] Does any of your advice differ in that kind of case for an internal transfer versus someone who&#8217;s coming from external?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2003">33:23</a>] There&#8217;s someone who I worked with closely on the search side who actually did transfer to my team, Nate Lintz, and he&#8217;s amazing, and now he owns so much of what we do on my team in terms of inference, co-design for Flash and Flash-Lite. And I would say he&#8217;s a really great example of this, where his approach was, how do I help my product area adopt this technology as effectively as possible?</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2040">34:00</a>] So I think there&#8217;s definitely, if you&#8217;re in an organization that isn&#8217;t directly generating these models but in some way trying to leverage them, there&#8217;s a very big gap in terms of applying these LLMs effectively, serving them effectively within your organization, and becoming someone who does that really effectively not only creates a ton of value in terms of the specific business need for your org, which will definitely elevate you in your org, but it&#8217;ll also be the case that you&#8217;re going to just naturally become the partner that we work with on the research side to make sure that our models are effective within your org.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2091">34:51</a>] So at that point, you may or may not want to transfer. Definitely, if you transfer, we&#8217;d be happy to work with you. But at that point, I think you&#8217;re already doing something that is cutting edge, which is integrating this new technology into a real product that people use. And so, yeah, that&#8217;d be my advice there.</p><h3>35:13 &#8212; Concrete steps engineers can take</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2113">35:13</a>] Towards the end of this post, as we leave this topic, you had the concrete invitation because I know you were hiring. Do you want to say what that was?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2123">35:23</a>] Yeah, so I was just trying to think of, how do I put my money where my mouth is? How do I demonstrate, look, this is a good way to show that you have at least some evidence of the skills that I called out as important: intent, mathematical maturity, grit. And so I listed out a couple of exercises that demonstrate some initial knowledge of scaling laws, some willingness to get into the weeds engineering-wise in terms of implementing a real transformer, and sort of willingness to pick up the kind of bread-and-butter math that we use every day to size these LLMs.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2167">36:07</a>] And I won&#8217;t recall the full list of the exercises that I expected here, but if you do the detailed handwritten version of The Scaling Book exercises and send me a video of yourself doing them, along with the transformer exercise on my post, then that&#8217;s something. If you can work in the New York office, I would love to interview you. Quite a few people reached out to me about that. I actually already have had a couple submissions, and we&#8217;re proceeding with the loop with those people.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2201">36:41</a>] So yeah, it&#8217;s quite a bit of work. But impressively, I got a response within, like, I think, a week of posting. So it&#8217;s definitely doable. Yeah, I mean, I don&#8217;t have unlimited headcount, so the offer&#8217;s on the table, but I can only hire so many people. The good thing is, though, that is such a strong sign of self-development that not only is this something that you should be doing for its own sake regardless of whether or not you will get a job at <a href="https://en.wikipedia.org/wiki/Google_DeepMind">Google DeepMind</a> specifically, but I think it&#8217;ll be something that lets you basically prepare for interviews in other places.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2247">37:27</a>] Certainly, if you reach out to me with these exercises completed, even if I do all my hiring, there&#8217;s tons of people who I know who are hiring as well, and I&#8217;d be happy to refer people as well.</p><h3>38:29 &#8212; Overview of pre-training areas</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2309">38:29</a>] On the next topic, I saw you&#8217;re the area lead for pre-training on Gemini, and I just thought it might be interesting to hear you give kind of a high-level overview of what pre-training is in your words and maybe what are the high-level challenges in the area.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2325">38:45</a>] We can talk about that.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2326">38:46</a>] Yeah, so there&#8217;s quite a lot of work that we do in pre-training as an area lead for it. The specific things that my team is responsible for delivering include the Flash model, the Flash-Lite model. These are models that get used for AI Overviews and AI Mode in the search bar, as well as some other 1P models that are used by different orgs like ads and <a href="https://en.wikipedia.org/wiki/YouTube">YouTube</a>. Besides this, we&#8217;re also key technical PoCs for the Google-Apple partnership, and so we do technical work there.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2365">39:25</a>] Those are the actual product-level deliverables from my team. Beyond that, we do research to make sure that these deliverables are SOTA. And also we do general pre-training research that contributes to the pro series model as well. And the nature of the research I would say generally breaks down into three different verticals. There&#8217;s distillation, which I mentioned earlier. There&#8217;s what I like to call inference co-design.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2398">39:58</a>] So creating neural architectures that are efficient to run inference on. So coming up with the network topology, the shapes of the matrices that the matmuls use inside of gating and linear layers for this transformer, as well as the attention shapes, num heads, that kind of thing. So that is effectively utilizing the hardware that you&#8217;re serving on. And then the final pillar here is new quantization methods.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2432">40:32</a>] And so quantization is just something that&#8217;s been near and dear to my heart that I&#8217;ve been working on the research side for ever since I joined Google. And it really changes what&#8217;s feasible for the first two. So that&#8217;s why furthering the state of the art in terms of how you can compress models is also a very important pillar in the research that my team does. Generally, quantization refers to reducing in some sense the size that the neural nets take up in order to represent their weights.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2471">41:11</a>] So typically, a neural net, when you&#8217;re training, it is represented as a series of numbers that make up the matrices inside of the neural net that are stored in FP32, 32-bit floating-point weights. It turns out that when you do these computations, you don&#8217;t need all of that extra precision to still maintain the quality of your neural net. And you can, with pretty simple methods, reduce the precision at which you store these weights down to four bits.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2504">41:44</a>] So all of a sudden, this huge range of numbers that would take this FP32 to represent something that gets you down to seven digits of precision can, with somewhat high fidelity, still be represented well by 4-bit ints, which just cover this tiny range of minus 8 to 7. And it&#8217;s kind of a miracle that you can do this. But what&#8217;s even more of a miracle is that you can apply these kinds of quantization transforms to the runtime activations that the neural net processes.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2547">42:27</a>] And as soon as you do that, the actual math that you&#8217;re performing, because you&#8217;re taking much smaller operands to your matmul than what you were doing before, the amount of electricity that it takes to compute the neural net drops significantly. And what&#8217;s interesting is that 99% of the total cost of operation for AI hardware comes from the power that it takes to run these chips. If you can do these operations, you could just make neural nets run more cheaply, run more efficiently.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2582">43:02</a>] That helps in terms of serving more requests and helps in terms of latency. So the name of the game for quant research is how do we push the frontier beyond this 4-bit range?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2596">43:16</a>] There&#8217;s this take that I see on Twitter all the time, which is just talking about MFU, and someone who&#8217;s not in the space, or model FLOPs utilization, someone who&#8217;s not in the space. They see a number in the low tens and they think, wow, they&#8217;re wasting all of those GPU resources. I was curious if you could just clarify that for people why a low MFU, or I guess naively low, is actually not low at all.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2622">43:42</a>] And maybe also explain what MFU is.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2624">43:44</a>] Yeah, so when we compute MFU, you want to divide the actual number of FLOPs that the neural net is performing here by the total number of FLOPs that the accelerator could have done in the time of your request. And so in some sense, this is giving us the percent of time that we&#8217;re usefully utilizing the FLOPs rate of the accelerator. And to get to 100% MFU, you would just need to be fully utilizing the matmul unit of whatever accelerator you&#8217;re doing here.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2660">44:20</a>] So it would just have to be doing a bunch of matmuls in a loop without reading any memory or doing any other operations. That&#8217;s not a very useful computation in practice. Neural nets have to apply activation functions or do attention or write intermediate outputs back to HBM. And all of those different operations will require utilizing the memory bus or utilizing vector processing units. Or simply they might be mathematical operations that the underlying hardware performs more slowly than it might perform a matmul.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2706">45:06</a>] And so all of those things contribute to not running at the full speed that the processor is rated at. That&#8217;s why you might not see 100% MFU all the time, is because part of the time your neural net was reading and writing to memory, or part of the time it was doing an operation that fundamentally runs slower than certain other units on your device. I think quite a bit of this inference co-design work that I talked about earlier is across all of the different capabilities of the chip.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2744">45:44</a>] So communication to other chips, memory bandwidth, the speed at which we can read parameters from memory, FLOPs, of course this can be matmul FLOPs, this could be FLOPs for processing vectors. So things like doing activations, all of these have different rates in the hardware, and a given computation isn&#8217;t going to match the natural hardware&#8217;s rate of each of those operations. So when you design a neural net, you want to be able to choose shapes for this neural net that fully saturate all of those hardware units to get you as high of an MFU as possible when you are doing inference here.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2792">46:32</a>] What makes this more than just an algebra problem is that those choices translate to different quality outcomes when you actually train this neural net. The process of this kind of inference co-design is how do we come up with neural architectures that scale predictably, have a good prediction, so are high quality and still make the MFU as large as possible during inference. This kind of joint optimization is what makes inference co-design really fun.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2826">47:06</a>] And also this kind of evergreen problem because as the hardware changes, all of those relative constants of FLOPs to memory bandwidth to communication bandwidth change. And those will have different implications for what the optimal neural net shape should be.</p><h3>47:23 &#8212; Jeff Dean spot bonus story</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2843">47:23</a>] On another topic, Google has this idea of a spot bonus where someone can give you a one-off lump sum of money as a thank you for good performance. And I saw on your resume that Jeff Dean, the legend himself, gave you a spot bonus. And if you can tell that story, I&#8217;d love to hear. Why did he give you a spot bonus?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2866">47:46</a>] Yeah, so that one actually was at the very beginning of the Gemini program. He gave out his spot bonus to people who hopped on and launched the first version of Bard. And I had a very small contribution to a very, very large project at the time. I helped with SFT for one of the first versions of supervised fine-tuning for one of the first versions of Bard that got released, like, right? The biggest lesson out of that experience was, at that time I was just doing pure research in Google Brain, and I was super focused on just how do I maximize the number of first-author papers at NeurIPS, ICML, ICLR.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2918">48:38</a>] And I remember distinctly thinking, I had this instinct of like, oh, should I just keep my head down and try to write more papers? Luckily, at the time, my manager, Rohan Anil, really encouraged all of us to get involved in this space, and that was just the right motivation that I needed to roll up sleeves, do a bunch of hyperparameter tuning and engineering work to get this model running on some really old TPUs, to get some extra cycles in for SFT attempts.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=2965">49:25</a>] That very small initial engagement that was recognized by Jeff Dean, I think, blossomed into more and more investments on the LLM side by me and ultimately led me to where I am today. So, yeah, I would say it&#8217;s less so about how much that SFT helped the initial release. And it&#8217;s much more about recognizing that there&#8217;s quite a bit of work, some of it not glamorous, some of it just hyperparameter tuning and golfing the XLA compiler to make your program fit in a certain memory amount that contributes to a wider business goal that is really quite important for getting involved in very high-value projects.</p><h3>50:14 &#8212; Favorite Gemini war story</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3014">50:14</a>] You&#8217;ve been working on Gemini for a while now, and because it&#8217;s a top priority, there have to be some incidents or war stories that you&#8217;ve been involved in. So I&#8217;m curious, what&#8217;s your favorite war story when working on Gemini?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3030">50:30</a>] So I think my all-time favorite would have to be Flash 2.0. So this one was quite a challenge and a very long journey to get there. But one of the main things that we were optimizing for, which Flash 1.5 established, is this category of very fast, low-latency model that&#8217;s still quite good. And in particular it has to be fast because it&#8217;s used by Search to serve responses in AI Mode very quickly. </p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3085">51:25</a>] Because of that, for Flash 1, eventhough at the time we knew about MoE models and how they increase capacity. And so I think one thing that came up was, okay, we sure would like to use this new architecture, but it&#8217;s difficult to just simply switch to an MoE. Because what happens with an MoE is it uses a lot more parameters in general. And because it uses more parameters, it takes up more HBM. These chips that we serve on have a finite amount of HBM.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3115">51:55</a>] So you have to shard the MoE across multiple different chips. So if you have whatever N experts, then you might shard it across N chips or some factor of N. What this causes is a lot of communication in the middle of the model. When you have a token that needs to be routed to an expert, and that token might live on the first TPU, but it needs to go to the last TPU, that&#8217;s a lot of communication that you&#8217;re inducing in the forward pass.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3146">52:26</a>] So the latency of this operation increases dramatically with N. And the challenge with MoEs is they increase N. So that kind of really bottlenecked this approach. And one interesting thing that happened was we definitely knew about pipeline serving for a while. It&#8217;s just in the dense case, it never really ended up mattering. I distinctly remember a very early conversation I had with Sholto about it, and Sholto&#8217;s like, oh yeah, you&#8217;re so FLOP bound and so pipelining is just not gonna change your prefill profile.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3187">53:07</a>] And then he was right. I tested it out and then abandoned the idea. But what&#8217;s interesting is I had a very small team at the time. And one of my reports, Geng Yan, had a very nice idea. He was working with Rahul Arya and a couple folks from the Israel team at Google, and that was to apply pipeline prefill to MoE. Pipelining is a technique where instead of parallelizing those end-machine experts across those end machines, you parallelize layers across those end machines.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3225">53:45</a>] So instead of on a particular layer, you have to route tokens from machine to machine. Now one layer does the computation for one subset of your pre-fill request and then hands off the processed tokens to the next machine to process the second layer, and then the third layer and the fourth layer. And all of the experts can then stay resident to a single machine or a smaller set of machines. So what this does effectively is it changes the communication pattern from something that required a lot of token exchange on every single layer to something that actually can be hidden behind other computation.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3268">54:28</a>] Because you can do this pipeline pre-fill across different parts of your request. So while layer two is working on the first thousand tokens of your request, layer one on the first chip is processing the second thousand tokens of your request. So it was a way of breaking this HBM constraint by moving layers across the machines rather than moving experts across these machines. And because of that, the communication overhead has gone down, and all of a sudden MoE latency looks really attractive.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3085">51:25</a>] Now, the Gemini 2.0 report says it&#8217;s an MoE series of models, and the thing that made that possible is, or one of the things that made that possible, is this serving-time innovation. Dwarkesh and Reiner have an amazing post about exactly this optimization that you can write up in the algebra of The Scaling Book. And it&#8217;s just a wonderful example of how this kind of change can have really dramatic implications on LLM quality.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3341">55:41</a>] What really made Flash 2.0 rewarding is this giant MoE decision. It sounds like a small technical decision at the time, but people were really worried about whether or not the latency of this MoE would actually be reasonable. Luckily, I was able to run a very transparent technical process to get to the bottom of this. And by the end of it, we made the right call, but then we had to train it. This was a bigger model than we&#8217;ve ever trained before at the Flash scale, and we knew this would be the right call.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3378">56:18</a>] But it was just going to be 40 days of grueling work for a really, really small team. We probably had five people on the rotation for training this model. I remember all of us just kind of rotated day by day, handing off all of this SRE-style work of keeping the training job alive, which at the time was a very interactive thing because you had to make sure that everything was moving stably, that you had tuned data iterators that aren&#8217;t slowing down your job, that if there&#8217;s a gap in the data somewhere or an indexing issue, you had to really quickly put up a fix because it&#8217;s wasting all of this GPU time.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3425">57:05</a>] What about at nighttime and on the weekends?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3428">57:08</a>] So, yeah, I think for those 40 days, we did not do a lot of sleeping. We had to do these dual shifts across the Paris office and Mountain View. And the thing that made it so rewarding was when this model came out around the same time DeepSeek-V3 came out, and the Wall Street Journal put out this article that was this giant red scare article about how China&#8217;s going to take over AI with open source models.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3459">57:39</a>] And I remember my friend sent me a screenshot of this table of the LMSYS Chatbot Arena leaderboard, and all the way at the top right you&#8217;ve got <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a> and DeepSeek right behind it. And like, oh, DeepSeek was trained for whatever few million dollars, and they&#8217;re right there. And then my friend was like, oh, Gemini is so behind because they had a version of, I think, Flash 1.5 Pro or something in that table at the very bottom.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3495">58:15</a>] Then I looked at it. It&#8217;s like, oh, that&#8217;s really interesting. I was just looking at this leaderboard because we just released a model, and it definitely doesn&#8217;t look like that when you go to the website. So it turns out there were some elided rows on the Wall Street Journal article. And so now if you go to that article today, you can see what at the time was the state-of-the-art model, Flash 2.0 Thinking up in the top right corner, way far ahead of DeepSeek-V3, might be messing with the open source narrative that they were trying to publish there. But it was a really important accomplishment for the team.</p><h3>58:59 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3539">58:59</a>] Last question for you is if you could go back to yourself when you just graduated college, I guess, undergrad, and give yourself some advice, knowing what you know now, what would you say?</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3551">59:11</a>] You gotta chase the problems that people are facing in the world today. Go after the challenges that people see in everyday life, and don&#8217;t be afraid to tackle a smaller part of this problem or maybe a more menial sounding part of this problem. Even if it&#8217;s not fancy research, math, or something like that. Trust that by working on what&#8217;s important, even if it&#8217;s a smaller part of a larger project for what&#8217;s important, you&#8217;re going to get to see what really matters in terms of moving the frontier forward.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3594">59:54</a>] And it&#8217;s this kind of humility, maybe, in your problem approach that you should really be chasing. That&#8217;s one piece of advice. I think the other bit that I would give, maybe as professional advice, perhaps, would be be the kind of coworker that people would want to see succeed. And so what I mean by that is there&#8217;s this conception of workplace psychopath or Machiavellian leaders or whatever, people who will do anything at all costs to get the results they want, and they might be able to squeeze people to get some short-term gain.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3645">01:00:45</a>] But having interacted with a variety of people professionally for so long, what is interesting to me is there have been a select few. One in particular is probably a very dear friend and mentor of mine, Todd Lipkin, who first got me into computer science, that are just so kind and people that you can learn from, someone that I can follow and be successful by following, that just genuinely inspire me to want to help them succeed.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3693">01:01:33</a>] And so in particular, if you are the kind of person who helps people succeed in their projects, comes up with projects that can leverage other people&#8217;s complementary skills in ways that help them shine, people will notice that people will want to contribute to projects that you come up with in the future and, in general, will want to support you going forward. And so people can get really cynical thinking about the game theory of how to interact at work.</p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3730">01:02:10</a>] But I found that this kind of more amicable approach generally creates this deep sense of collaboration and willingness to help that is so important to get very large projects that require multiple people and multiple skill sets over the line. And so, yeah, I think if I could give any kind of interpersonal feedback or professional feedback or whatever to earlier version of myself, it&#8217;s to be that guy, is to be the kind of person that other people want to see succeed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3771">01:02:51</a>] I love that this advice combats the cynical advice. And I also love that your original post combats the doomer, permanent underclass stuff. So, yeah, thank you so much for your time. It was a lot of fun. Really appreciate it.</p><p><strong>Vlad:</strong></p><p>[<a href="https://youtu.be/cDyi91onoJ8?t=3785">01:03:05</a>] Thanks for having me, Ryan.</p>]]></content:encoded></item><item><title><![CDATA[Co-Creator of Haskell: Functional Programming, Thinking in Types, Useless Languages | Simon Jones]]></title><description><![CDATA[Simon Peyton Jones is the co-creator of Haskell (pure functional programming language) and I interviewed him about functional programming, why it matters, and his thoughts on other programming languages.]]></description><link>https://www.developing.dev/p/co-creator-of-haskell-functional</link><guid isPermaLink="false">https://www.developing.dev/p/co-creator-of-haskell-functional</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 08 Jun 2026 10:02:31 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/200821756/8f5bbc64b28f11f5e661ec5cab7b8e50.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/simonpj/">Simon Peyton Jones</a> is the co-creator of Haskell (pure functional programming language) and I interviewed him about functional programming, why it matters, and his thoughts on other programming languages.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/xcB_LF3cdqw">YouTube</a>, <a href="https://open.spotify.com/episode/5d9VR5MwYDWs9qAzo2H6bv?si=tZEd-XCmS7-4065ifediRw">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-xcB_LF3cdqw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;xcB_LF3cdqw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/xcB_LF3cdqw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/200821756/0039-what-functional-programming-is">00:39 - What functional programming is</a></p><p><a href="https://www.developing.dev/i/200821756/0918-downsides-of-functional-programming">09:18 - Downsides of functional programming</a></p><p><a href="https://www.developing.dev/i/200821756/1053-specialized-hardware-for-functional-programming">10:53 - Specialized hardware for functional programming</a></p><p><a href="https://www.developing.dev/i/200821756/2147-haskell-is-useless">21:47 - Haskell is useless</a></p><p><a href="https://www.developing.dev/i/200821756/2559-rust-vs-c">25:59 - Rust vs C</a></p><p><a href="https://www.developing.dev/i/200821756/2826-haskell-vs-ocaml">28:26 - Haskell vs OCaml</a></p><p><a href="https://www.developing.dev/i/200821756/3526-side-effects-in-haskell">35:26 - Side effects in Haskell</a></p><p><a href="https://www.developing.dev/i/200821756/4426-type-systems">44:26 - Type systems</a></p><p><a href="https://www.developing.dev/i/200821756/5730-how-the-haskell-compiler-works">57:30 - How the Haskell compiler works</a></p><p><a href="https://www.developing.dev/i/200821756/010435-why-haskell-is-talked-about-more-than-used">01:04:35 - Why Haskell is talked about more than used</a></p><p><a href="https://www.developing.dev/i/200821756/010907-avoiding-success-at-all-costs">01:09:07 - Avoiding success at all costs</a></p><p><a href="https://www.developing.dev/i/200821756/011112-llms-and-programming-languages">01:11:12 - LLMs and programming languages</a></p><p><a href="https://www.developing.dev/i/200821756/011357-new-programming-language-design">01:13:57 - New programming language design</a></p><p><a href="https://www.developing.dev/i/200821756/011559-should-students-continue-to-learn-programming">01:15:59 - Should students continue to learn programming</a></p><p><a href="https://www.developing.dev/i/200821756/012233-why-excel-is-is-2nd-favorite-programming-language">01:22:33 - Why Excel is is 2nd favorite programming language</a></p><p><a href="https://www.developing.dev/i/200821756/012504-advice-for-his-younger-self">01:25:04 - Advice for his younger self</a></p><h1>Transcript</h1><h3>00:39 &#8212; What functional programming is</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=39">00:39</a>] I wanted to start by asking you, what is functional programming in your words? And why does it matter to people who might use the other languages, the imperative languages?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=57">00:57</a>] Functional programming is about programming with values instead of mutation. So in a conventional imperative programming, computation proceeds by, if I want to add up the numbers between 1 and 100, I have a mutable cell containing a running total. And then as computation proceeds, I keep adding to my running total. So there&#8217;s a cell that changes its value over time. Computation proceeds step by step.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=83">01:23</a>] There&#8217;s a program counter. Now, that&#8217;s very unlike mathematics. If I say three plus four times seven, you don&#8217;t think of a mutable cell that changes its value over time. You think, well, three plus four is seven, seven times seven is 49. Just the answer is 49. So, in a sense, it&#8217;s more declarative than imperative. Another way to think of it is it&#8217;s more like what a spreadsheet does. In a spreadsheet formula, if you put in a cell, you&#8217;d say, in A1, you put the formula =A2 * A3 + 7.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=115">01:55</a>] Well, then that&#8217;s just the formula that is evaluated. And you expect to compute the values of A2 and A3 automatically before computing A1. It wouldn&#8217;t make any sense to compute A1 from the old values of A2 and A3, right? So, and there&#8217;s no notion of a program counter. There&#8217;s no notion of do this and then do that. It wouldn&#8217;t be sensible to say equals print X plus 3, because who knows when or even whether that cell will ever that formula will ever be evaluated.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=143">02:23</a>] Now, the surprise is that it is possible to do useful things in such a limited language. When I learned programming, I learned it by writing machine code in which the fundamental mode of computation is registers and program counters and mutation of registers as you go along. And memory, like, computation and mutation seem to be inextricably interwoven, right? That&#8217;s how computation gets done. So it was a surprise to me when I learned that, well, it is possible to do useful things with functional programs.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=176">02:56</a>] I still remember that amazing revelation. Now, in fact, one way to look at it is this: right back in the late 1920s, early 1930s, <a href="https://en.wikipedia.org/wiki/Alan_Turing">Alan Turing</a> was doing a PhD in Princeton under the supervision of <a href="https://en.wikipedia.org/wiki/Alonzo_Church">Alonzo Church</a>. Now, <a href="https://en.wikipedia.org/wiki/Alan_Turing">Alan Turing</a>, as we all know, invented the <a href="https://en.wikipedia.org/wiki/Turing_machine">Turing machine</a>, which is fundamentally about mutation, right? It has a little tape, and you can read things from the tape, and then you write back new things onto the tape, and it has an immutable program.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=210">03:30</a>] Now, <a href="https://en.wikipedia.org/wiki/Alonzo_Church">Alonzo Church</a> at the very same time, was working on something he called the <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>, in which there&#8217;s no mutation. The <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> is like the essence of functional programming. Programs execute just by reduction. Now, so these were two different models of computation. You might ask, what can you compute with a <a href="https://en.wikipedia.org/wiki/Turing_machine">Turing machine</a>? And what can you compute with a lambda term? And the astonishing thing that turned out, which is completely nonobvious, is that anything you compute with a lambda term, you can compute with a <a href="https://en.wikipedia.org/wiki/Turing_machine">Turing machine</a>.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=242">04:02</a>] And anything you compute with a <a href="https://en.wikipedia.org/wiki/Turing_machine">Turing machine</a>, you can compute with <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. So that means that the two are identical from the point of view of their computational power. Big surprise to me. Okay, so they&#8217;re equally powerful. And then the question is, which might you want to choose? As I say, I started life from my baseline of imperative programming. Of course, that&#8217;s the way you write programs.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=266">04:26</a>] And functional programming is a kind of weird, anomalous, rather nerdy, academic thing. But I became hooked on the idea that by excluding side effects by default, i.e., imperative programming has side effects by default. By excluding side effects by default, we could make programs that are easier to write, easier to maintain, easier to reason about. And then the question is, what are the practical obstacles?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=293">04:53</a>] Is it slower? Is it less convenient, and so forth? So, in effect, I&#8217;ve devoted my research life to exploring that hypothesis. The wonderful thing about being a researcher is you can take something which, at the time back in 1980, functional programming was very nerdy, geeky, academic. But I was allowed to spend my professional life just saying, what would happen if you took that one idea, programming with values, and just ran with it?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=323">05:23</a>] Where would you get? And I think we got somebody somewhere pretty interesting. I don&#8217;t want to say that it&#8217;s better fundamentally, because after all, <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> on that journey started with functional programming. Then we learned how to mix functional imperative programming together in a way that we might talk about later, right? So it&#8217;s not really an either-or; it&#8217;s more of a both-and in the end. But my position now is, as I often say, when the limestone of imperative programming has worn away, the granite of functional programming will be revealed underneath.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=357">05:57</a>] And why do you say that?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=359">05:59</a>] Well, functional programming is the language of mathematics. That&#8217;s how we fundamentally think about things when we want to be formal. Indeed, if we want to reason about and say, what does this imperative program do? We have to wheel out a lot of mathematics and Hoare triples and so forth. But that&#8217;s all in the world of maths, which is just functional programming. Functional programming is just mathematics brought to life and made executable.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=382">06:22</a>] So you might wonder, I suppose you may be saying, well, isn&#8217;t imperative programming just enough? And I think that there you have to just look at the broad sweep of things. I think what&#8217;s happened is that from the world of functional programming, lots of ideas have got adopted into the imperative arena, right? So they grew up over here in functional programming, and then they have been adopted by the imperative programming thing.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=414">06:54</a>] So you might think garbage collection, you might think lambdas, you might think language integrated query, you might think about lots about type systems, polymorphism and static typing, tons of that. So at least it has been an intellectually interesting experiment that has been a very fertile laboratory in which to grow up ideas that can infect the mainstream. Now my instinct now is that we should start from functional programming and do imperative programming where necessary.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=444">07:24</a>] I think probably the mainstream is we should start with imperative programming and sort of add bits of functional programming where possible. And I&#8217;m not going to say that they&#8217;re wrong to do that, just to say that I think it&#8217;s a pretty interesting dialogue. Now you do have to think very differently about the act of programming, right? You do have to rewire your brain quite a bit. But just let me.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=466">07:46</a>] This is something we might want to come back to. I said easier to maintain. As programs get old. Think of a 10-year-old program, right? If you have a piece of 10-year-old code written by somebody who&#8217;s long since left your company that mutates some more or less global variables, then you might be a little bit afraid about, oh, and there&#8217;s other bits of code scattered about that also read or write those global variables, then you might be a little cautious about making modifications, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=496">08:16</a>] Lest, oh, if I put these two things, if I do B and then A instead of A and then B, maybe the global variable that I read somewhere deep inside the code no longer has the value that it was expected to. So reading and writing shared variables that are visible outside the scope, sort of somewhat global variables, is a form of coupling between bits of code that is very invisible, very invisible.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=524">08:44</a>] So functional programming forces that to become visible. It forces it into the open, and by the act of forcing it into the open, that strongly encourages a style of programming in which you don&#8217;t have such interactions because they are painful to manage. So it encourages a style in which programs execute in modular boxes without interacting. Now, every imperative programmer would want to do that, but I want them to do that in a provably secure way so you can know that these programs don&#8217;t interact, rather than just say I&#8217;ve done my best.</p><h3>09:18 &#8212; Downsides of functional programming</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=558">09:18</a>] Then I guess on the flip side, what are the downsides of a functional programming language?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=565">09:25</a>] Oh, well, it can be very tiresome and inconvenient, right? So you&#8217;re somewhere deep in the, you know, a subroutine of a subroutine of a subroutine of a software stack, and you want to reach out and, I don&#8217;t know, grab the time of day. Oh dear, that&#8217;s a side effect. But we&#8217;re in a pure functional language, the time of day. And indeed there&#8217;s a good reason, because everybody else higher in the stack is assuming that if I call this function, apply to, you know, seven, it will give me the same answer if I call it again, apply to seven.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=596">09:56</a>] But if it could interrogate the time of day, certainly it might give a different answer, right? So a simple thing as I was asking for, the time of day, invalidates a whole set of assumptions and so is not by default allowed at all. And that can be inconvenient because all you say, look guys, all I want to do is get hold of the time of day. Don&#8217;t give me grief. And so every functional programming language, including <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, provides you with a little trapdoor.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=621">10:21</a>] <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s is called unsafePerformIO, and it says, trust me, I&#8217;m going to do something that has side effects, but you can pretend there weren&#8217;t any. All right, so it&#8217;s called, it is literally spelled unsafePerformIO, with a long name that has unsafe in the title, so that you, the programmer, are taking on the obligation to make sure that what you&#8217;re doing is okay. Whereas in imperative language, everything you do could be unsafePerformIO-ish.</p><h3>10:53 &#8212; Specialized hardware for functional programming</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=653">10:53</a>] I was watching one of your lectures, and you did go over a little bit of the history of functional programming languages. And you mentioned somewhere that there was this idea of people building hardware for functional programming languages. This just made me curious because I didn&#8217;t. I only know the von Neumann architecture. Do you have any idea what that might look like? And some high-level thoughts?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=680">11:20</a>] Well, I did. I now think it&#8217;s probably a bad idea. But back in 19, when was John Backus&#8217;s Turing Award lecture? I think it was 1978 or thereabout. What was it called?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=694">11:34</a>] Can Programming Be Liberated from the von Neumann Style?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=698">11:38</a>] That&#8217;s it. Now he was saying two things. Firstly, the designer of Fortran, a quintessentially imperative language, was saying, guys, we should do functional programming. That was what the talk was about. And he introduced a language called FP for functional programming. It was a very particular, somewhat idiosyncratic functional programming language. But there it was. So that was rather amazing, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=726">12:06</a>] This was 1977. He was saying, guys, give up on imperative programming. Just do functional programming. Here&#8217;s how you do it. But he also said in the same lecture, we might imagine designing new hardware to execute this new kind of language. So if you like, we&#8217;ve got Turing machines, they give rise to register machines and microprocessors and so forth. Maybe <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. Perhaps if we designed a machine from the ground up to do <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>, it would look different.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=757">12:37</a>] So that was, of course, very exciting. And similar kind of thinking led to an actual commercial machine called a Lisp machine. That was a piece of hardware that was designed by Symbolics to execute Lisp programs. Lisp is a mostly functional programming language. And so, at the time, I was in my early 20s. We were thinking, oh crumbs. What would a piece of hardware look like that was primarily designed to execute functional programs?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=786">13:06</a>] Would it have a different kind of instruction set and architecture? So the nearest we got to it actually was, although there was a big strand of work then that was called the MIT Dataflow project. So dataflow languages, functional languages, they&#8217;re definitely the same kind of category, right? And so MIT had this great group led by Arvind called the MIT Dataflow Group. And they designed, in the end, a machine called Monsoon that was specifically designed to execute dataflow programs, that is, functional programs.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=819">13:39</a>] It was based on a token store and matching, and tokens flowing around and getting matched up and executing. You might imagine data flowing, going through a graph, when 3 arrives at a plus node and 4 arrives at the other input of the plus node. Bam. You could fire that plus node, right? And the hardware did all of that. At a similar kind of time in England, my colleagues Thomas Clark, Joe Stoy, and others built something called the SKI machine, SKIM.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=847">14:07</a>] There&#8217;s a paper about this in the FPCA conference. And SKIM was designed to execute SK combinators. So at the time, David Turner&#8217;s written a lovely paper in which he described how to take <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> programs and translate them into SK combinators. So you can read about that in my book. It&#8217;s a very simple translation. You can give the whole translation in half a dozen lines, but then you have this mess of S&#8217;s and K&#8217;s, and S and K reduction is very, very easy.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=880">14:40</a>] So it&#8217;s like the machine code of functional programs.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=885">14:45</a>] What is a SK combinator? Just for context, there are three combinators,</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=890">14:50</a>] S, K, and I, and they have simple reduction rules. Here they are: I of X equals X, KXY equals X, S X, Y, Z is X of Z applied to Y of Z. End of story. That&#8217;s all so very simple, right? So now all I&#8217;ve got to do is take my lambda, translate it into a giant tree of SK combinators, right? S applied to K applied to I of 3 and S of Z of just a huge tree of S&#8217;s and K&#8217;s, right? Now simply apply the rewrite rules I gave you, and that is program execution.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=925">15:25</a>] It&#8217;s rather astonishing that that can execute arbitrarily complicated programs, but it can. Don&#8217;t you think that&#8217;s rather amazing? Now, I can take an arbitrary <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program and could translate it into lambda terms, and I could translate those lambda terms into S and K. Very simple transformation. And I can just write those, run those three rules, and it&#8217;ll produce the output of the <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=947">15:47</a>] That&#8217;s pretty amazing. And indeed, there is an implementation of <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> that uses exactly this. It&#8217;s called MicroHs, and my colleague Lennart Augustsson has built it. And its execution mechanism is SK combinator reduction. And if you sit in MicroHs, you can ask it to show you the S&#8217;s and K&#8217;s that it produces. If we had a shared screen, I could probably show you. So, long story short, then the SKI machine was designed to take SK trees, big trees of S&#8217;s and K combinators, and just execute them directly.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=981">16:21</a>] So it was built in hardware and it ran perfectly well, reasonably fast. Now, why didn&#8217;t this catch on? Why don&#8217;t we have functional programming machines? I mean, at the time it was a big idea. We even had a conference called Functional Programming and Computer Architecture. FPCA was right there in the title of the conference. But slowly we became aware that what was really happening is we were doing at runtime what we could do instead at compile time.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1013">16:53</a>] And that is always a bad idea. Imagine an interpreter for, I don&#8217;t know, Pascal or something, or for Java. We can compile Java to bytecode and we could interpret the bytecode, right? So then we&#8217;re doing something at runtime. At runtime, the interpreter is looking at the bytecode, saying, dispatching to say which instruction, blah, blah, blah, right? Now what does JIT do? It takes a sequence of bytecode and says, oh no, no, I&#8217;m going to take that sequence of bytecode and compile it to a sequence of machine instructions that, when executed, will do the same thing much faster because it can take advantage of bytecode sequences, for example, to say if you swap and then swap, it&#8217;s a no-op or something like that.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1064">17:44</a>] So what, in fact, it turned out is this whole SKI business, and any other, and indeed dataflow machines also were simply doing at runtime what you could do better at compile time. We were building an interpreter in hardware. So don&#8217;t build the interpreter in hardware. Instead, build a compiler that translates your lambda term into a sequence of machine instructions that, when executed, will behave as if you had done this reduction business.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1095">18:15</a>] And then you might think, well, machine instructions for what machine? Well, from a practical point of view, it was very hard to compete with <a href="https://en.wikipedia.org/wiki/Intel">Intel</a>, right? They were just spending bazillions of man years on making x86 go faster, and Arm likewise. So we can leverage all of that work by compiling into their instruction set. But even if you say you are God, you are the chief executive in <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> and can tell them to add new instructions or computation mechanisms to support functional programming.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1128">18:48</a>] What would you add? Not very much. Little bits to do with garbage collection, barriers, I think, and such like, but nothing major. So, in other words, I think the whole build hardware to execute functional programming directly thing turned out to be a mistake. Inspiring mistake, fun mistake. I had a great time. But a mistake. It&#8217;s better to build a compiler.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1157">19:17</a>] I would have thought there might be some unique opportunities for parallelism given the immutability in these functional programming languages. So yeah, you described that graph or that SK combinator graph. I imagine you could execute that all in parallel, assuming there&#8217;s no connections across the graph.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1180">19:40</a>] And indeed, that was the dataflow machine. The MIT Dataflow machine was all based on that idea. So you throw the whole graph into the token store and just run all the nodes that are runnable.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1191">19:51</a>] But that&#8217;s incredibly fine grained. You&#8217;ve got to imagine this big tree of a million nodes, and your processor is sort of wandering around over it, doing little things in parallel. There&#8217;s a lot of memory traffic going on there. And if these little threads are operating on the same bit of tree, then there&#8217;s synchronization costs. So very, very fine grained parallel computation like this turns out to be impractical, or not impractical.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1218">20:18</a>] You could do it, but it&#8217;s very slow. So the dataflow people, their history, if you look at the history of the MIT Dataflow project, they started with micro parallelism. Every individual instruction was a separate parallel computation that might take place, and the token store would match it up. And then they built more and more compiler technology that grouped these things together into larger units.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1242">20:42</a>] So they expanded from micro threads of single instruction threads into 10 instruction threads or 100 instruction threads. Right. But even that never really caught on. So you&#8217;re right though, that if I say in a <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program E1 plus E2, then I can do E1 and E2 in parallel. And GHC does support you in doing that. But nowadays, rather than expecting the compiler to do that automatically for every subexpression, you can spark a computation. You say E1 par E2, and that says do E1 and E2 in parallel.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1277">21:17</a>] There&#8217;s a project at Carnegie Mellon for a strict power, a strict language called ML that has a very good variant of Parallel ML going in a similar way. So essentially, we&#8217;ve moved away from the dream of automatically getting parallel, very, very fine grain to the programmer. We give the programmer clues about where it&#8217;s a good idea to do parallelism, and still the compiler will try not to fork two tiny grains.</p><h3>21:47 &#8212; Haskell is useless</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1307">21:47</a>] I was searching on YouTube and I searched your name, and this video came up. It said, &#8220;<a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> is useless.&#8221; Simon Peyton Jones, and it&#8217;s a video from a long time ago, and someone&#8217;s got a low-quality handheld camera. There&#8217;s a group of researchers that are all talking. Butler Lampson is across the table. And it&#8217;s a fun little video. And in the video you lay out this two-dimensional graph where on one dimension there is the useful versus useless axis, and then on the other dimension there is the safe versus dangerous.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1346">22:26</a>] And then you&#8217;re placing programming languages on this two-dimensional graph. And I think you started by putting C as very useful, but it&#8217;s incredibly dangerous. And then you put <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> on the polar opposite side. You said it&#8217;s useless, but it&#8217;s very safe. What are the least safe and most useful languages? And why do you place them there? C, for instance.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1372">22:52</a>] Well, yeah, so let&#8217;s see. In C, you program by mutation. So it&#8217;s unsafe in the sense that any function can mutate any variable at any time. It has a lot. You pass pointers around a lot, and functions mutate the memory pointed to by those pointers. And moreover, typically they can mutate it anywhere. There&#8217;s no array bounds checks or anything. So it&#8217;s kind of like super unsafe. And in fact, and this is demonstrated by the fact, can you imagine all of these exploits that we get every day, right, that Mythos is discovering in huge numbers?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1404">23:24</a>] But why is the Internet so insecure? Primarily because all of our software infrastructure is written in unsafe languages. If we&#8217;d written all of our Internet software and operating systems in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> or maybe in OCaml or ML, 99% of all these exploits would be removed by construction. It&#8217;s like we&#8217;ve built a boat out of paperclips and we&#8217;re surprised that it&#8217;s leaky.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1435">23:55</a>] I mean, you shouldn&#8217;t build boats out of paperclips, right, because they have holes in them. You should build it out of a secure substance, right? But then it&#8217;s too late. So we spend incredible amounts of human ingenuity and effort patching the holes in our boat built of paperclips. It&#8217;s tragic. It&#8217;s tragic how much effort and ingenuity and money has been lost and waste of resources just because we wrote our computational infrastructure for the world in an insecure language.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1465">24:25</a>] That&#8217;s what I mean by unsafe.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1467">24:27</a>] I mean, are there not vulnerabilities? If I wrote something in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, there&#8217;s gotta be some new type of.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1474">24:34</a>] Of course there are vulnerabilities. I just said 99%, I didn&#8217;t say 100. If I write a <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program that says receive message, if the message says tell me everything, then spit out my entire database in reply, no language can stop you doing that, right? But if you look at the program and it doesn&#8217;t have any such things, right? So, I mean, nothing can prevent you against high-level attacks to insecure programs.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1503">25:03</a>] Or, I mean, another example might be deadlock, right? Services, no matter how securely written. If A waits for B and B waits for A, deadlock. Sorry, no. And no language is going to stop you doing that. You might hope for some high-level verification tools, but it&#8217;s like surely if you&#8217;re trying to do something hard, like prove that that doesn&#8217;t happen, you want to have a foundation that is, you know, in which you&#8217;ve got some bedrock to stand on.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1534">25:34</a>] If you&#8217;re standing on sand and trying to prove some advanced property, it&#8217;s very, very difficult. But good point. I&#8217;m not talking about 100% security. Absolutely not. But how many exploits are based on buffer overruns or pointer manipulation that&#8217;s gone wrong? If you couldn&#8217;t have a buffer overrun, you couldn&#8217;t do pointer manipulation gone wrong. Those bars just wouldn&#8217;t exist.</p><h3>25:59 &#8212; Rust vs C</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1559">25:59</a>] So, I mean, C&#8217;s maybe half a century old now. What about more modern versions of those lower-level languages like <a href="https://en.wikipedia.org/wiki/Rust_(programming_language)">Rust</a>?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1569">26:09</a>] Much, much better. Much, much, much better. Right. If we rewrote all of our software infrastructure in <a href="https://en.wikipedia.org/wiki/Rust_(programming_language)">Rust</a>, things would be way, way better. I&#8217;m not actually even sure whether <a href="https://en.wikipedia.org/wiki/Rust_(programming_language)">Rust</a> has array bounds checks built in, but suppose. But it must have the ability to promise that you&#8217;re not indexing out of bounds. I don&#8217;t quite know quite now, but if you compile all your code with that switched on, you are in a way better situation, way better.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1597">26:37</a>] So yes, this is not just functional programming. But you did ask about why I thought C was an insecure, unsafe language.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1604">26:44</a>] If we imagine that graph, there&#8217;s the upper right quadrant, which is useful and safe, and you described that as nirvana in the video. And I guess maybe over time, I mean, C is somewhat really unsafe, really useful, yes.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1621">27:01</a>] So C, <a href="https://en.wikipedia.org/wiki/Rust_(programming_language)">Rust</a> has moved along the axis from useful but very unsafe to stay useful and become safer. Now, <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> started life as being very safe but useless. So it&#8217;s worth just rehearsing that because in the first version of <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> there was no I/O. The only <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program was a function of type string to string. It&#8217;s functional programming, after all; it could do I/O. All it could do was take a string and produce a string.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1651">27:31</a>] So obviously that&#8217;s not very useful. It&#8217;s a little bit more useful. A language that is very, very safe and completely useless is no-op, that does nothing ever, very safe, but very useful. <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> was a bit better, right? At least it lets you apply a function to a string and get back another string. So, of course, we then worked on how could we do I/O in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> in a safe way that will lead to monads and stuff, but just as the, you know, we&#8217;re moving from C horizontally to go safer and safer, but staying useful.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1689">28:09</a>] So from <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, we&#8217;re moving sort of vertically to get more and more useful without stopping being safe. And nirvana is when we sort of meet up, right? But the point is not to say one is better than the other, but just to say that we both seek things that are useful and safe.</p><h3>28:26 &#8212; Haskell vs OCaml</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1706">28:26</a>] When I think of functional languages, I guess in the mainstream, or what people are kind of thinking about, I hear <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. I also hear OCaml. How do these two programming languages compare if someone was trying to pick a functional language to work with?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1721">28:41</a>] So OCaml is a strict language that is called by value. So when you say F applied to 3, 4, it&#8217;ll add 3 plus 4 and then call F to get 7, right? In <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, if you say F of 3, 4, it&#8217;ll build a little suspension that says, well, if you ever need this three, four, you can evaluate it, but maybe you&#8217;ll never need it. And then it calls F. That&#8217;s a lazy language. So because OCaml is strict, it has a defined order of evaluation.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1751">29:11</a>] So if you say F of 3, 4 and a second parameter, 8, 9, it&#8217;ll evaluate 3 plus 4 and then 8 plus 9 in that order. So it makes sense to say you can say F of print hello, print goodbye, right? Because you get hello and then goodbye printed. So once you have a defined order of evaluation, then it&#8217;s pretty easy to incorporate I/O, right? Even though it&#8217;s, quote, a functional language, OCaml allows you to do I/O without saying unsafePerformIO or anything.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1787">29:47</a>] In fact, it has a defined order of evaluation, and it&#8217;s perfectly kosher to do that as a sort of unintended consequence of being called by value. It was also, by default, impure. Now, good OCaml programmers won&#8217;t use many side effects, but OCaml doesn&#8217;t prevent you. <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> prevents you. Why does it prevent you? Because if we&#8217;d allowed you to say F of print hello, comma, print goodbye, those would be thunks, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1818">30:18</a>] Not yet evaluated. Right. So if F happened to evaluate its first argument and then its second, we&#8217;d print hello and then goodbye. If it evaluated its second and then its first, we&#8217;d print goodbye and then hello. If it didn&#8217;t evaluate either, we wouldn&#8217;t print either of them. That doesn&#8217;t sound very good if you want to control what I/O is happening. Laziness forced <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> to stay pure. Strictness allowed OCaml, which grew out of the ML tradition, by the way.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1845">30:45</a>] ML is another functional language that predated <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. It was always strict, and it always had I/O by default. It wasn&#8217;t a goal. But that&#8217;s just the way it worked out at the time. Nobody thought about it. It was just obvious.</p><p>So that&#8217;s a big difference between OCaml and <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, right? Is that now, then, since then, they&#8217;ve both grown up. <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s gained sort of monads. OCaml&#8217;s gained lots of fancy type systems.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1873">31:13</a>] Some of them look at how <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> and OCaml have learned from each other. OCaml has got an interesting effect system recently and a whole lot of new extensions. It&#8217;s an absolute hotbed of innovation at the moment. OCaml. So I view OCaml and <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> as kind of siblings. Brothers and sisters, right? We love each other, we learn from each other, we compete with each other, all of the things that siblings do, but they&#8217;re not the same, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1897">31:37</a>] So siblings don&#8217;t say, I&#8217;m just better than you. You shouldn&#8217;t exist. We say, let&#8217;s enjoy our life together. And that&#8217;s what it is. That&#8217;s the way it is.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1904">31:44</a>] You describe the laziness and the strict. And as a programmer, my initial thought is I would like to control the order of execution, and when I say print X and then print Y, I want it to happen that order. Otherwise it would be a little unintuitive. What is the main benefit of laziness?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1925">32:05</a>] Why is laziness good? Well, John Hughes did this rather well way back, in 1980-something. He wrote a paper called Why Functional Programming Matters. And one of its main theses was that lazy evaluation lets you compose programs in a particularly modular way. So imagine a program that is, I mean, his classic example is a program that plays chess, right? So one thing you could do is you could imagine building a tree of all possible moves, starting from the position we&#8217;re at.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1958">32:38</a>] It would be a very, very big tree. It would be too big, because first of all I compute the tree and then I&#8217;d start deciding what move to do. It couldn&#8217;t possibly do that because the tree of all possible moves has more nodes in it than the number of protons in the universe. So let&#8217;s not do that. Now, if instead you could build the tree, having got this tree, supposing you had it, you could then walk over the tree saying, ah, let me do some minimaxing, and seeing, oh, this looks like a good position.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=1986">33:06</a>] Let me go this way. I could explore the tree right now, then. So in OCaml, then you would have to put the tree generation and the tree exploration in one function, right? In <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, with lazy evaluation, you can generate an infinite tree, and then the explorer that prunes the tree and explores just the bits of it that are necessary is completely modularly separated. I can build a different generator and a different pruner.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2016">33:36</a>] They&#8217;re just completely separate programs. So lazy evaluation is very powerful glue that lets you glue together two programs that you&#8217;d like to be distinct. Strict evaluation forces you to merge them together.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2033">33:53</a>] Yeah, this reminds me, so in Python, for instance, they have the idea of a generator where you could lazily retrieve things.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2041">34:01</a>] Every strict language, every serious strict language has lazy evaluation in it. They&#8217;re called iterators or generators, and so does OCaml. All right, but in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, that&#8217;s the default. Now instead, in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, so this is a bit like, again, the two are converging on the middle, right? <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, lazy by default. But you can make things strict. You can put in exclamation marks to say, please evaluate this before the call, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2069">34:29</a>] Or the IO monad to say value. But that&#8217;s a more brutal strict annotation. So in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, you can make things stricter. In OCaml, you can make things lazier. So it really boils down to what&#8217;s the default? We both want to have a mixture of the two. And then it becomes a bit cultural. As to which you prefer, I don&#8217;t know. Some people sometimes ask me. They say, well, well, if you were designing <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> again, would you make it strict by default with really good support for laziness?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2102">35:02</a>] And I often say, well, I might. Yeah, that does seem attractive because frequently I found myself cursing laziness as an implementer. But I strongly suspect that 10 years after that I&#8217;d be thinking, man, if only it was lazy by default. Oh, when I said strict by default, I would definitely mean strict, but pure, right, no side effects.</p><h3>35:26 &#8212; Side effects in Haskell</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2126">35:26</a>] A few times in our conversation, you&#8217;ve mentioned the word monad, and it seems like it allows us to do the side effects. And what is a monad? And how does it help us do side effects and preserve ordering?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2141">35:41</a>] One way to think about it that&#8217;s very easy to understand and use is just to imagine that the do notation is somehow built into <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. So you can say do print x, semicolon, print y, and that has type IO unit. So a value of type IO unit means I do some input output and return a value of type unit. An expression of type IO int is an expression that when you run it will do some I/O and return an int. Okay, so do lets you combine I/O-performing computations together.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2183">36:23</a>] So print 3 has type IO unit. It does some I/O, namely printing 3. GetChar has type IO Char. It does some I/O, namely reading a character from standard input and returns a character. All right? Notice that&#8217;s different from just Char. So quotes x quote has type Char. It&#8217;s just the character pure, right? GetChar, the I/O-performing operation, has type IO Char. It does some input output and returns a character.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2223">37:03</a>] Okay, the do notation lets you combine together I/O performing computations. So if you could say do and then you say x left arrow getchar; putchar x, then that x left arrow getchar says run the getchar computation and get me the character. Call it x. The putchar x says run the putchar computation to put x, and we combine them together with the do notation that combines two computations to make one I/O performing computation.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2265">37:45</a>] But these things are completely first class. That&#8217;s what&#8217;s new about monads compared to just make it into C. x left arrow, getchar semicolon, putchar x. That&#8217;s a computation whose type is I/O unit. Let me give it a name, so I can say let foo. We type I/O unit equals that do. Now I can pass foo as an argument to something. I could put foo in a data structure. I could return foo as a result; it&#8217;s a value just as much as three or plus in particular.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2301">38:21</a>] For example, I could say do foo, foo that takes foo, uses it twice each time. Do foo; semicolon, foo says do foo and then do foo again, right? So I&#8217;ll read a character and print it, and then read a character and print it again. So these values of type I/O for some type T are first class values, right? In C, I can&#8217;t take X colon equals 3 semicolon, you know, Y plus 4 and pass that as an argument to something and expect it to happen wherever it&#8217;s used, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2341">39:01</a>] It&#8217;s not a first-class value.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2343">39:03</a>] When I first was learning this monad idea, it seems like you&#8217;ve said in a few places it&#8217;s a way to do the side effects, but keeps the language pure. But when I saw it, my first thought was it&#8217;s almost like this little place where we segregate the dirty things we want to do. But then how does that keep the language pure?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2366">39:26</a>] Oh, because. Because we are lying. But look, if your function has type int to int, can it do any I/O? No, no. If it has type int to I/O int, it could do arbitrary I/O. So yes, so it&#8217;s nice. Pure int to int function or dirty int to int function. So you may say, perhaps you&#8217;re about to say, that&#8217;s a bit of a blunt instrument. Either completely pure or completely dirty. Right. It gets you an awfully long way.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2402">40:02</a>] But nevertheless, it would be cool if you could say, oh, I am effectful, doing reading of files only. Right. So, in the type, you&#8217;d like to say, what kind of effects can it have? Can it throw exceptions? Can it spawn new threads? So you&#8217;d like to enumerate the effects this computation could have. Right. That would be cool. That&#8217;s called an effect system. And there&#8217;s a bazillion programming language papers about effect systems.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2444">40:44</a>] And it turns out that you can indeed, in the type system of <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, and indeed OCaml is rapidly becoming the same. You can express not just all or nothing, does it do I/O, game over, but rather which particular effects does it do, including no effects at all? That&#8217;s pure. Right. A good place to start is a library called Bluefin, which my colleague Tom Ellis has designed. It&#8217;s a very nice take on how to do effect systems in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2477">41:17</a>] I guess the purity comes from that. This dirty stuff is signaled through the type system.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2485">41:25</a>] Correct?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2486">41:26</a>] Okay, so we can do it. But here it is. And be aware.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2491">41:31</a>] Yeah, and because you have to thread, you know, because the type system gets in your face, you were trying to map, say, no, map says I apply a function to every element of the list. Well, so it has type A to B to list of A to list of B. If you want to apply an I/O performing function, so the function you&#8217;re applying has type like int to I/O of character, you could map that over a list, but you just get a list of I/O of chars that hasn&#8217;t done any I/O yet.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2523">42:03</a>] Right. You want something that says, take a list of IO chars and perform those actions one at a time. So I want a function that goes from type list of IO char to IO of list, and you might want to perform all those actions top to bottom, or maybe bottom to top, who knows? That&#8217;s what this function of type. Yes. So you could do it in various ways. So sometimes it gets in the way. Right. You say, oh, shit, can&#8217;t I just use Map?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2554">42:34</a>] Well, in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, no, sorry, you&#8217;re going to have to do a little bit more work to tell me in what sequence you want your effects to happen. Because by default, <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> does not specify sequence. So, adding monads to control effects does get in your face a bit. And that&#8217;s the tax we pay. In effect, that&#8217;s part of the big experiment that <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> is doing: to say, suppose we upfront say we&#8217;re willing to pay that tax.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2584">43:04</a>] How many followers can we get? And the more. But I will note that monads have infected quite a lot of other languages. Like F# is definitely a call-by-value impure language. And yet F# had these workflows that were definitely monad. And monads have appeared in Scala and in many other languages. So something that&#8217;s very monad-like has appeared lots elsewhere. It&#8217;s been a very unifying concept.</p><h3>44:26 &#8212; Type systems</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2666">44:26</a>] And so I kind of want to ask you, maybe just even on the highest level, what is a type system, in your words?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2675">44:35</a>] Okay, so what does a type system do? It lets you reject silly programs upfront. So if I have a function that adds one to things and I apply it to a character or to an I/O computation, that&#8217;s going to fail at runtime, right? Because I can&#8217;t add one to a character. Or maybe you overload plus. Maybe I reverse a list or something. If I&#8217;ve got a list reverse function, and I give it something that just isn&#8217;t a list, then it&#8217;s a bit silly to allow that and only fail at runtime.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2713">45:13</a>] So fundamentally, type systems are about rejecting at compile time programs that you do not want to run because they will fail at runtime. Okay, now, we&#8217;ve had static type systems for a long time, going back to Pascal and earlier. But simple type systems are annoying because they get in the way. Imagine a function that reverses a list. In Pascal, you could write a function that reverses a list of integers.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2746">45:46</a>] But if you wanted to reverse a list of characters, sorry, you can&#8217;t. That function that reverses a list of integers, it has type list event to list event, so you can&#8217;t apply it to a list of characters. Game over. You have to write another copy of the code with a different type. That&#8217;s a bit annoying. So what are we going to do? We need polymorphism. We need a more powerful type system. We want to give reverse the type for all a list of A to list of A.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2778">46:18</a>] So that upfront says, I work for any type A. You give me a list of integers. Fine, I&#8217;ll produce a list of integers. You give me a list of characters. Fine, I&#8217;ll produce a list of characters. Notice that&#8217;s much better than just list to list. I want to keep the fact there&#8217;s a list of integers so I know if I apply, when I look inside the list, I know I&#8217;ve got an integer. But also if I just list the list, I might get back a list of some completely different type.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2803">46:43</a>] But I know that reverse returns a list with the same type of things, right? So parametric polymorphism, super valuable, right? If you want a size of type system, it must have parametric polymorphism. The message is, if your type system is too simple, it gets in the way. Very important lesson, because it means that we cannot really say a type system is useful or not unless we say which one, because we know that weak type systems are very inconvenient.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2836">47:16</a>] It must at least have parametric polymorphism. And that is another idea that was born in the world of functional programming. It was born in ML, incidentally, Robin Milner&#8217;s famous dictum: well-typed programs don&#8217;t go wrong. ML was the first, I think, the first parametrically polymorphic functional programming language. But generics in Java and object-oriented programming more generally is exactly the same idea, right?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2862">47:42</a>] into the mainstream and you mentioned polymorphism, and so there&#8217;s this parametric polymorphism in the type system. But in the example you mentioned where maybe you want to do an operation on a list and then the type of the input changes, I was just thinking, what about when people do polymorphism in the classes?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2883">48:03</a>] Yeah, okay, but so now you&#8217;re into a whole more complicated world, right? So as soon as you say class, you&#8217;re talking static type system again, right? So object-oriented programming is another approach to polymorphism. So in an object-oriented language, we say if we have a Ford that is a car and a car is a vehicle, then anything that works on vehicles should also work on cars and should also work on Fords.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2912">48:32</a>] So the code that we write for cars works unchanged for vehicles and for Fords. Okay? Now that&#8217;s a form of polymorphism, not parametric polymorphism. That&#8217;s what you might call object-oriented polymorphism or subtype polymorphism. Okay, so now we have polymorphism in the sense that the same code, the same executable code, actually the same machine instructions work on values of different types. Okay. In both cases, both reversal lists, same machine instructions.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2948">49:08</a>] Code that works on vehicles works on things on Ford. Same machine instructions, right. Okay. That&#8217;s what polymorphism in general means. Parametric polymorphism means this for all. A list of a to list of a stuff. Subtype polymorphism means if it works on vehicles, it works on any subtype of vehicles. Okay? Now the interaction of the two, which you get by adding generics to an object-oriented language, is pretty complicated.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=2976">49:36</a>] Oh, and by the way, adding side effects as well. And so your question involving classes and superclasses and so forth is smack in that complicated world. I&#8217;m not sure it&#8217;ll be very fruitful for us to get a lot more details of the type system sorted out to know what to say. But at this very high-level overview, the big point I&#8217;m trying to make is weak type systems get in the way.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3004">50:04</a>] Type systems are meant to reject programs that will go wrong. But my function that reverses a list of integers, it will also reverse the list of characters. So it&#8217;s tiresome to be told, no, that is a bad program. I want to be able to write that it has type, for all a, list of a to list of a. So the idea of making the type system more complicated is to say programs that you want to run, you can still write in your static type system.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3029">50:29</a>] Now you might say, well, blimey, if that&#8217;s all, why don&#8217;t we just get rid of the static type system altogether? Now we could run all of those programs, but the trouble is, you can run too many programs now. You can run programs that will crash at runtime, and that&#8217;s very, very, very bad. And we all know the costs of programs that crash in deployment, that you could have crashed before you even started to run them, before you even ran your first test, right before you even linked it into an executable.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3060">51:00</a>] That&#8217;s really good. But the biggest benefit of a static type system, in my humble opinion, is maintainability. If you have a program written in, I don&#8217;t know, Perl or Ruby or Lisp in its basic form, and it was written 15 years ago and the original author has left, and all of the people who were involved at the time it was written have left, then that program is very difficult to maintain, and it becomes almost immutable.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3091">51:31</a>] Nobody dares change it anymore. What they do is, it&#8217;s an immutable piece of software, and you do stuff around the edges to impedance match what you really want to do to this now immutable blob. Now of course you have lots of tests, so test-driven development, I&#8217;m totally with it, I&#8217;ll develop all of that. But still, GHC, for example, is itself written in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. It&#8217;s 35 years old, and yet I do large-scale systematic refactorings of it fearlessly because the type system keeps me safe.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3122">52:02</a>] In fact, often what I&#8217;ll do is I&#8217;ll change a few types and then start compiling, and then a sort of wave of changes propagate through, forced by&#8212;I just get type errors. So I know what to do. The thought was that I&#8217;ve changed the representation of this data structure a little bit. I&#8217;ve added a field to this data structure. Where in the entire compiler might that field be read, written, or freshly allocated?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3147">52:27</a>] I can&#8217;t imagine how anybody maintains 30-year-old software and makes large-scale changes like that. Without a type system, it&#8217;s unimaginable. So for me, the benefit of type systems is maintainability. Oh, and designability. Right. So I often write the types of my programs up front. The data types are super perspicuous. Novices say it&#8217;s a bit like writing the classes of an object-oriented language, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3175">52:55</a>] But if you have no types, no classes, nothing, just I don&#8217;t know, S expressions, types are my design language. They&#8217;re the way that I start writing my designs.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3187">53:07</a>] I&#8217;m trying to understand the other mindset. Let&#8217;s say C, for instance. I remember a lot of stuff when I was learning it in college, for instance, many times, right. I&#8217;d add a character to a pointer or something, and it just works because it interprets the character as a number. So is there any value to having that kind of type system, or is that just strictly unredeemable?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3214">53:34</a>] Just use stronger, I think there&#8217;s no benefit to weaker.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3219">53:39</a>] Just off the top of my head, one thing I think of is there is a set of programs that will work but don&#8217;t satisfy the type system.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3229">53:49</a>] Oh, yeah, that&#8217;s why.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3231">53:51</a>] But those are. I mean, I don&#8217;t know if they&#8217;re good. That&#8217;s subjective.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3234">53:54</a>] Oh, no, they may be good. So we started, if you have Pascal and you write a function to reverse a list of integers, then if you apply it to a list of characters, the same machine instructions would work, but it is rejected, right? So we have rejected a perfectly decent program. Bad. Right. Our goal is to expand the collection of programs that satisfy the type system to include as many as possible of the programs we want to run without including any of the bad programs that we don&#8217;t want to run.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3272">54:32</a>] Okay? Parametric polymorphism is a big step in that direction. Other type system innovations are a big step in that direction. But there will always be some programs that would run perfectly well that the type system rejects. Imagine a tree that contains integers at every node. But if you are 17 deep in the tree, or 34 or 51, if you&#8217;re in a multiple of 17 deep, the integers can be characters instead, or the integers all turn out to be characters, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3317">55:17</a>] Now, you could write a program that generated such trees, and you could write a program that consumed such trees, knowing that every 17th layer we switched characters. But most static type systems would make it pretty hard for you to accept that program. And yet it will run. Now you might say, &#8220;Oh, but I really want to write that program, guys. Don&#8217;t get in my way.&#8221; Well, then we should provide a way for you to bail out into dynamic typing, right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3349">55:49</a>] What I&#8217;d like to do is say, okay, so if all else fails, then at least you can, as it were, pair up a value with its type representation. There&#8217;s one reason we don&#8217;t want to interpret an integer as a double precision float, for example, is that they don&#8217;t even have the same representation, which is just nonsense. One possible way untyped languages let you do this is to tag every integer and every double precision float.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3374">56:14</a>] With the fact I&#8217;m an integer, I&#8217;m a double precision float. But that has a lot of overhead. So one merit, I think it&#8217;s not the biggest single merit, but it&#8217;s a major merit of static type systems is you have no tags. You know that if it says it&#8217;s an integer, it&#8217;s going to be an integer. You know that if it&#8217;s a double precision float, it&#8217;s going to be a double precision float, right? But if you&#8217;re not sure, maybe we could make a way to make a pair of a type representation.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3402">56:42</a>] And this value, the type representation, is now like a runtime tag. It&#8217;s like a little runtime data structure that describes the type. And then in your program you could say, now I want to say, oh, I&#8217;ve got this type dynamic. We&#8217;ll call this pair a value of type dynamic. Now when I want to take a value of type dynamic and treat it as a character, I want to say, ah, look, look at the type dynamic.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3424">57:04</a>] See if it says it&#8217;s a character. If it is, return the character. If not, crash.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3428">57:08</a>] Right?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3428">57:08</a>] Runtime failure. That&#8217;s fine, you can do that. So <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> has good support for dynamic typing where necessary. So my story would be static typing should expand to carry as much as possible, and where you absolutely cannot do it, sorry, then use dynamic typing. And we provide facilities to support that.</p><h3>57:30 &#8212; How the Haskell compiler works</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3450">57:30</a>] One thing I wanted to talk with you about is the compiler. How does GHC work on a high level?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3457">57:37</a>] So GHC takes a string like any other compiler. The source code of the program parses it, and then it type-checks, it figures out, is this a type-correct program? Then it converts it to <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. Now <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, the <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> AST, the original source tree, has 50 different data types, 50 different kinds of nodes, some of which have 30 or 40 different variants. So it&#8217;s a really big, diverse, complicated data structure.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3495">58:15</a>] The AST <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> has this variant of the <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. It has eight WANDA data, maybe two or three data with eight constructors. So it&#8217;s like taking a gigantic language and squeezing it down into a tiny one. And that tiny one we can then optimize, right? That&#8217;s what the optimizer works on. So the front end does parse, rename, type check, desugar into <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. That&#8217;s the front end. The <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> particular language is called GHC Core language.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3532">58:52</a>] I&#8217;m quite proud of it because it&#8217;s been very, very stable. It&#8217;s 35 years old, and it has barely changed since birth. That&#8217;s amazing, right? Because <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> has changed a lot. A lot. So almost all of the innovation in <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> has been in the front end. Very little in core. Now the back end, the core optimizer, has changed a lot too. But all of the changes that were useful there would have been useful 30 years ago.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3565">59:25</a>] Right. So they&#8217;re two completely separable things. So I&#8217;m quite proud about that. Core then. We do a lot of core-to-core passes that simply take core programs. We&#8217;ll use core programs, lots and lots, long pipeline. Then we convert it to C, which is a prototypical imperative language. Think of it as a portable assembly code. So that bit is meant to be platform independent. So it&#8217;s simply. That&#8217;s the compiler that takes.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3594">59:54</a>] Remember I said if you take <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> and compile it, that&#8217;s that step. Right. I want to compile the <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> into machine instructions, really. But I don&#8217;t really mean machine, I mean portable machine instructions. That&#8217;s called C. Then I want to convert C into actual machine instructions for various platforms. And there we could either do that directly with the native code backend or go via LLVM.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3618">01:00:18</a>] Interesting. I&#8217;ve never heard of C. Why not just go directly to it? I would have thought maybe assembly or something like that.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3625">01:00:25</a>] Well, because. What happens in assembly? In which processor? Please?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3630">01:00:30</a>] Oh, I see. I guess it&#8217;s the. I would have thought the <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> part was already portable.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3636">01:00:36</a>] So you just. It is, yeah. You could go straight, but, but, but. So there&#8217;s work to go from <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>. You could go all the way to x86, then throw all that away and now go from <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> to PowerPC. Oh dear. I&#8217;ve just duplicated a lot of work by going from <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> to C-- and then from C-- to x86, x C--. We&#8217;ve avoided duplicating the work that went from <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a> to C--.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3665">01:01:05</a>] When you see that, it&#8217;s pretty obvious, isn&#8217;t it? You got to. You want to make the platform-specific bit as small as possible. You would like to have, as it were, a generic architecture, one that can do addition and has a program counter and a stack and so forth. That&#8217;s all C--. It&#8217;s just a portable assembly language. And then you say, oh, the nitty-gritty of whether you have double-precision add and set, the floating-point bit here and there, that.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3691">01:01:31</a>] Well, that&#8217;s platform specific. The core is itself statically typed, but in some ways it&#8217;s surprising because no other compiler has this property. No other production compiler by core is statically typed. I don&#8217;t just mean that the initial program was type correct. I mean that a core program, you could run a type checker on that. I might say, why do you need to? Because after all, if GHC is correct, it started with a type correct core program because it came from a type correct <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> program.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3717">01:01:57</a>] Assuming the desugaring was right, and all the optimizations, if they&#8217;re right, will generate a type-correct Core program. So why do you need to type check Core to discover bugs in GHC? Now, these are serious bugs, right? If you ever take a type-correct Core program and an optimization pass produces a type-incorrect Core program, what will happen if we don&#8217;t have the type checker for Core? We&#8217;ll generate machine code, and we&#8217;ll run it, and we&#8217;ll get a segfault.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3749">01:02:29</a>] Now we have to backtrack from a particular test program that now crashes all the way back up the pipeline. Up the pipeline, up the pipeline, up the pipeline. Oh, it was this pass of GHC that was faulty. That&#8217;s super hard to do because you&#8217;re getting out GDB on some runtime failure. That is an indirect and perhaps distant consequence of the fact you just generated a type. You just made a.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3776">01:02:56</a>] You had a bug in the optimizer. So it is amazing to have a type checker for Core. Now, why does nobody else do this? Well, it&#8217;s because their intermediate language, typically, and this I really am talking typically, because I know of no other compiler that has this property, none. Production compiler, typically. Then they&#8217;re complex syntax trees decorated with all sorts of pragma information and things hanging onto it here and there.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3807">01:03:27</a>] There&#8217;s no type checker for it at all, and there&#8217;s no hope of one. So I&#8217;m very proud of the fact that Core is statically typed. And I also think it&#8217;s a. The most delightful thing is that the way that it is statically typed, it is an implementation of something called System F. So System F, when I said <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>, <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>, as <a href="https://en.wikipedia.org/wiki/Alonzo_Church">Alonzo Church</a> had, it was untyped, had no type system at all.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3835">01:03:55</a>] But Girard defines something called System F, which is a statically typed <a href="https://en.wikipedia.org/wiki/Lambda_calculus">lambda calculus</a>, a rather powerful one. And GHC Core is essentially System F. So he literally adopted something from the nerdy theoretical computer science, logic community, logic of mathematics community, and adopted it directly in a production implementation. So I&#8217;m very proud of that. I think GHC Core is. And the fact that we can do we have 35 years worth of development that has not just not been impeded, but actively aided by statically typed intermediate language is really a big marker in the ground.</p><h3>01:04:35 &#8212; Why Haskell is talked about more than used</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3875">01:04:35</a>] Watching all your talks, reading everything, one of the interesting data points that you brought up was that <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> is talked about more than used. When you compared Stack Overflow volume and actually <a href="https://en.wikipedia.org/wiki/GitHub">GitHub</a> volume, like who&#8217;s actually using the programming language, why do you think that is?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3895">01:04:55</a>] So <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> embodies one key idea: immutability changes everything. There&#8217;s a quote from Pat Helland&#8217;s talk that I think you also looked at or read. So it says programming with values changes everything about the way you think about programming. It&#8217;s just mind changing. It&#8217;s not necessarily better, but it is different. And so <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> takes that idea and runs with it. Everything is driven by that one idea.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3923">01:05:23</a>] Everything else is incidental. In the early days, that meant we just said, well, guys, suck it up. We&#8217;ll keep changing the language, and if it breaks your program, too bad. So it is a bit peculiar and therefore also it felt a bit academic because initially it was really not very powerful. It was taking the key idea, but you couldn&#8217;t do very much with it. We talked about that. Right.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3944">01:05:44</a>] So over time, GHC and <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> in general has become more and more powerful. The type system has become less and less in your way and more and more useful. All the obstacles that make functional programming harder become better. The compiler generates faster code, it can browse faster and so forth. So it has become less peculiar. So we&#8217;ve become more and more taking into account the needs of our users.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3972">01:06:12</a>] But in a way, the sort of cultural heritage is we never give up on the one core principle. We&#8217;re just not going to give you unrestricted side effects. Sorry, you want to say unsafePerformIO and put up with the consequences. So we&#8217;re going to stick to one core principle and then we&#8217;ll do lots of work around the edge to make that better. Right. So that does limit our community somewhat. Right?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=3994">01:06:34</a>] It does mean you really have to think in a different way. Immutability changes everything. That means you&#8217;d have to think a different way about programming. Maybe you don&#8217;t want to think in a different way. That&#8217;s fine. Then don&#8217;t use <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4004">01:06:44</a>] So in a way, we started from a very small user community, a very sort of pure and nerdy one, and growing larger and larger. But all slowly, slowly, while all the time maintaining faithfulness to this core principle.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4018">01:06:58</a>] I thought that was very unique about <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> because I feel a lot of the other programming languages are user-centric. I mean, if people want something, they work on it, they add it, whereas <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> feels more principled. Or it&#8217;s all starting from these ideas. And if you don&#8217;t satisfy these ideas as a user, yeah, well, that&#8217;s fine. For instance, in one of your talks, you mentioned somewhere that there was a release of the compiler where if a file wasn&#8217;t type-correct, then the compiler would delete the file.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4057">01:07:37</a>] Oh, wait, it reported the error message first.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4059">01:07:39</a>] Yeah, it reports there. But I thought that was absurd. I mean, very hostile, I guess, to the. I mean, well, if you&#8217;re type safe, no problems. But.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4069">01:07:49</a>] Oh, it was a mistake, right? It was a bug; it wasn&#8217;t deliberate, and it only happened, you know, how did the bug get out? Because, of course, if it always did that, we&#8217;d have noticed. How did it get into a release? Well, it was because it was only on Windows and only when you compile a module that was not in the current directory. But at that stage, our users were very forgiving. And, you know, somebody wrote to us and said, well, by the way, Simon, you might like to know that, you know, GHC does this, but, hey, don&#8217;t worry about it, you know, I just copy all my files somewhere else before I compile and then I copy them back.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4101">01:08:21</a>] So, of course, those days are long gone. We pay a lot more attention to our users and have much more rigorous CI testing than ever we did. Right. So that&#8217;s a story from a long time ago, but it&#8217;s a good cultural story because it suggests that we care about our users very much, but we care about users who want, in their hearts, want to be principled. We&#8217;re trying to appeal to. One of the things I like best about <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> is people often say, I just enjoy writing <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4128">01:08:48</a>] Right. It&#8217;s fun, right? My boss doesn&#8217;t allow me to because it somehow doesn&#8217;t fit with my production shop. But for me, I would go for it. I love writing this stuff. Every time. Every time. Yeah, that&#8217;s very rewarding to me.</p><h3>01:09:07 &#8212; Avoiding success at all costs</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4147">01:09:07</a>] There&#8217;s also this interesting&#8212;I don&#8217;t know if it&#8217;s a cultural value&#8212;but it&#8217;s a statement that you say often in the context of these older talks, that you say that you avoid success at all costs. Could you explain what you mean by that phrase?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4161">01:09:21</a>] Oh, yeah, this was just a little play on words. It was in a retrospective on <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. I gave an invited talk at POPL a long time ago, probably 20 years ago. So it&#8217;s a little play on words because you can read it as either avoid success at all costs. And that&#8217;s what we&#8217;ve been discussing. Success at all costs means compromise your principles in order to satisfy your users or think that you&#8217;re satisfying users, give them what they say they want, if we build it, they will come kind of deal.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4195">01:09:55</a>] So avoid success at all costs. Or if you parenthesize it the other way, it says, avoid success at all costs, at all costs. Avoid success. And that&#8217;s saying, that&#8217;s a little joke. But it says if you&#8217;re too successful and have too many users, it becomes more difficult to make changes. And we experience that right now. So I devote many, many more of my personal cycles to backward compatibility issues than ever I did.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4228">01:10:28</a>] I&#8217;ve devoted hundreds of hours and hours, well, days and days, weeks and weeks in the last year or two to the following: what seems to be a very simple property. If you can compile a program, a package, a whole program with GHC 10.0 and we release GHC 10.2, you should be able to compile that same package unchanged with GHC 10.2. Seems reasonable, right? After all, 10.2 should just be better. But no, GHC has never had that property, and making it have that property has turned out to be very, very time consuming.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4265">01:11:05</a>] Previously we just never cared. Then we started to care, but thought it was a lot of work, and now we&#8217;re investing the work.</p><h3>01:11:12 &#8212; LLMs and programming languages</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4272">01:11:12</a>] I guess that was from a long time ago. I think nowadays software engineering, there&#8217;s been a major shift in the last year where a lot of code is being generated by these models or these LLMs. How do you see programming language design shifting to accommodate a world where a lot of the code is no longer written by humans?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4294">01:11:34</a>] I think it may be the best thing that&#8217;s happened to statically typed languages for a long time, because as we&#8217;ve been discussing with a static type system, you cut down the space of programs that the <a href="https://en.wikipedia.org/wiki/Large_language_model">LLM</a> can generate because it is perfectly capable of running the compiler and saying, &#8220;Oh darn, that was a bad program, better fix it,&#8221; right? So a zillion iterations get done behind the scenes, whereas in an untyped language, the first one it coughed up, you&#8217;d have had to run it against a test suite or who knows what, but it dramatically tightens up that cycle.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4331">01:12:11</a>] Right. So I think that statically typed languages are a huge boon for LLMs because it&#8217;s too easy to. And programs are just strings, right? It could just generate the next plausible word.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4351">01:12:31</a>] Away if there&#8217;s a slider on, I guess, the strength of a type system. And the other side of the week, what do you see as the absolute strongest type systems among programming languages?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4363">01:12:43</a>] Oh, <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. I think <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> is exploring the bleeding edge. There is an exception, which is that module systems are a... And in particular, sort of functor-style module systems are explored much more deeply in the OCaml world and in the <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> world. We&#8217;ve essentially never gone there. But otherwise, I think <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s right up there. Now, of course, a language like Scala is also as... So Scala has almost everything <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> has, I think, not quite, but it also has subtyping and object orientation.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4404">01:13:24</a>] So that&#8217;s a lot more complicated. A lot more complicated, I think. And they pay a price for it. I think Martin Odersky would agree that they pay a price for it. So in complexity, it&#8217;s probably more complicated than <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s, and maybe in terms of power. So maybe I should have said <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> and Scala are the two leading <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>, Scala, OCaml, perhaps I&#8217;ll just put them in an equivalence class for now.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4432">01:13:52</a>] They&#8217;re not strictly comparable. They all have things which they&#8217;re more powerful than the others, probably.</p><h3>01:13:57 &#8212; New programming language design</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4437">01:13:57</a>] What do you think are the important problems to solve in the future of programming languages today? What are the unsolved problems in the domain that are top of mind for you?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4448">01:14:08</a>] I think it&#8217;s actually hard to identify, to say, here&#8217;s a problem we want to solve, let&#8217;s try to solve it. Well, I think another way to tackle it is to say, what are interesting new languages out there that are exploring very different parts of the design space? And there I think I do have a candidate. So the language that is my day job. I work for Epic, and we&#8217;re designing a programming language called Verse.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4471">01:14:31</a>] Now, Verse is a very exotic language. It&#8217;s really a functional logic language, so it&#8217;s yet more expressive than <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>. It has a path-? type system, but a very different one to <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s. So if you like, the way I think of it is like this. If you look at C and Fortran, they look pretty different if you&#8217;re an imperative programmer. But if you look at them, if you zoom out still, you can see functional languages.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4496">01:14:56</a>] Then C and Fortran are pretty close together, along with object-oriented languages. They&#8217;re all in a clump, right? And then there&#8217;s some functional languages, <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a> and ML. I know OCaml and Scala out here. If you zoom out still further, then the imperative languages and functional languages altogether, and Verse is way out here, right. So Verse is exploring a very new point in the design space.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4519">01:15:19</a>] But just like functional programming back in 1980, it seems sufficiently interesting and cool and unusual and weird that it&#8217;s worth exploring. Right? So back in 1980, nobody would have said, we&#8217;re definitely going to do functional programming and it&#8217;s going to be useful for practical applications. They said, that&#8217;s pretty weird. By all means, give it a try, guys. And that&#8217;s kind of where I am with Verse.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4540">01:15:40</a>] One difference is that back in 1980 we were purely academics, and now, Verse is being developed by Epic Games, so we got much more muscle behind it than was behind functional programming to begin with. So we&#8217;ll see. It&#8217;s a very interesting intellectual endeavor, adventure, I should say.</p><h3>01:15:59 &#8212; Should students continue to learn programming</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4559">01:15:59</a>] Yeah. I think a lot of people, like students, are maybe worried about AI or studying computer science. Would you recommend people learn how to program today, given that AI is starting to write reasonable code now?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4573">01:16:13</a>] Oh, yeah, yeah. So I think people are right to be worried in the sense that I think there&#8217;s going to be considerable dislocation, right? Because if you were in the Industrial Revolution, then lots of people lost their jobs as spinners and weavers, and it wasn&#8217;t easy for them to get a new job in the new economy. Now, the new economy had, in the end, had more jobs, but there was.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4596">01:16:36</a>] If you were one of the people who just lost their job, that was not a happy place to be. From our perspective, you know, from an Olympian perspective of a few hundred years later, we think, well, it&#8217;s just a blip, right? If you&#8217;re part of the blip, not a problem. So I think people are right to be worried. We don&#8217;t know how things will shake out. I&#8217;m actually optimistic that in the medium term, if we don&#8217;t destroy ourselves with some truly existential thing.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4620">01:17:00</a>] But from an employment market point of view, I&#8217;m optimistic that in the end we&#8217;ll just be in a higher place, that AIs will just be a bigger power tool. I mean, everyone likes using compilers, right? We don&#8217;t like machine code anymore. Compilers make us more productive. Maybe LLMs can make us more productive. That&#8217;s what I hope. I sort of believe, modulo dislocation effects. Now, should we teach children or even undergraduates how to program?</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4651">01:17:31</a>] So I think still yes. And the reason is because one way to say it is copilots who need pilots. I think Copilot was quite a good title that <a href="https://en.wikipedia.org/wiki/Microsoft">Microsoft</a> gave their tools because it encourages you to believe it&#8217;s your partner, not your boss. If LLMs spit out a pile of goop and we literally do not understand what it does, we just try it and it kind of works. That might be okay if we&#8217;re just throwing up a quick visualization.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4680">01:18:00</a>] It might be less okay if the quick visualization is going to drive our policy choices about us a nation, whether to go into lockdown because of COVID or if this program is going to run my aeroplane or train signaling system. So now those are extreme ends of the spectrum from quick and dirty things. It really doesn&#8217;t matter if it doesn&#8217;t work, but it kind of does a lot of the time, absolutely fine too.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4707">01:18:27</a>] This is a 30-year code base. It&#8217;s going to last a long time. I really want to make sure that it&#8217;s like putting new stuff into GHC. If somebody sends me a pile of AI-generated code to put into GHC, I&#8217;m not going to put it in unless I&#8217;ve reviewed it or somebody&#8217;s reviewed it. Because in 10 years&#8217; time, I&#8217;m going to want to change that code. How do I even know what it does? If it&#8217;s simply a magic incantation that somebody&#8217;s done that kind of worked on the test they did, but maybe won&#8217;t work in deployment, that&#8217;s no good.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4731">01:18:51</a>] So I really want long-lived, maintainable code to be well reviewed. Sorry. So. And to do that, I need reviewers who can write code, who know. Let me mention one other perspective. If you think about what every child should know. When I think about what every child should know about computing, I would include binary and bits, not, just as for physics, I would include atoms and molecules. Now it&#8217;s not that in real life anybody manipulates atoms or molecules or takes decisions which are based directly on their knowledge about knowledge, but somehow knowledge that all matter is made up of atoms, you know, constituted of a finite number of elements that atoms.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4778">01:19:38</a>] That knowledge underpins everything we understand about the natural world. If you literally have never been told that, you are sort of emasculated as a citizen, let alone as a scientist. So if you literally do not know that everything is composed of bits, that words and music and text and LLMs and everything is all just bits, I think you&#8217;re crippled. So I want every child to learn, you know, it&#8217;s like I want you to learn the bottom.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4807">01:20:07</a>] It&#8217;s all bits, nothing else. It&#8217;s all just bits. I even have a talk. The talk is called Bits with Soul. It&#8217;s so easy to grab for. It&#8217;s a talk I gave to an audience that was not a computer science audience at all. It was a completely lay audience ranging from 14-year-olds to professors of quantum mechanics. Pretty difficult audience to address. And it was meant to be about codes and coding and bits.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4833">01:20:33</a>] So it tries to get at the essence of why I think it&#8217;s important that every child, every person, every human being should understand something about the computational universe that surrounds them and that is founded in bits. Now, just to develop the analogy a bit further, I would then say, and it also, I think, I want them to also know about programming programs. I want them to know that computers fundamentally execute by following machine instructions blindly.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4860">01:21:00</a>] Right. That is not magic, it&#8217;s not hocus pocus, it&#8217;s just remorseless and very dumb. Right. It&#8217;s incredibly empowering then. And also to learn the basics about how neural networks work. In the same talk I explain how a one neuron neural network works, and that&#8217;s enough. Then it&#8217;s actually true to say, not distorting, but that&#8217;s to say <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a> is just a trillion of those wired together, and astonishingly, that very simple.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4892">01:21:32</a>] So all <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a> is, is a trillion floating-point numbers and a lot of floating-point arithmetic. That helps you make sense of a question like, can <a href="https://en.wikipedia.org/wiki/ChatGPT">ChatGPT</a> have feelings? Well, it gives you, I mean, of course that&#8217;s a philosophical question, but it informs your discussion about if you know that all it is is trillion floats and a lot of floating-point arithmetic. That&#8217;s all, nothing more.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4922">01:22:02</a>] Anyway, sorry, long answer to your question, but so I... So yes, basic programming, absolutely, yeah. Becoming fairly skilled in how to use, you know, Django framework, maybe not so much maybe. I think one thing that LLMs are very good at is knowing the arcane and complicated APIs that many of these frameworks present when there&#8217;s just 10,000 functions you&#8217;ve got to know.</p><h3>01:22:33 &#8212; Why Excel is is 2nd favorite programming language</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4953">01:22:33</a>] You mentioned somewhere someone asked you, you know, what&#8217;s your favorite programming language? And obviously <a href="https://en.wikipedia.org/wiki/Haskell">Haskell</a>&#8216;s got to be number one. Yeah, but for your second favorite programming language, you said it was Excel at that time. Why was Excel your favorite programming language?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4969">01:22:49</a>] Oh, because it&#8217;s the world&#8217;s most widely used functional programming language. The formula language is a functional language, isn&#8217;t it? Doesn&#8217;t have any side effects. You program entirely with values. So the formula language of Excel is a functional programming language, a very weak one. It doesn&#8217;t even let you define new functions. And it has a very limited collection of data types, namely just flat arrays and numbers and strings.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=4997">01:23:17</a>] So when I was working for <a href="https://en.wikipedia.org/wiki/Microsoft">Microsoft</a>, I took it as my war cry to say, let&#8217;s take that idea. Excel is the world&#8217;s most widely used functional language by three orders of magnitude. And oh, it&#8217;s the most widely used programming language by three orders of magnitude, not just functional language. Excel is used by many, many more programmers, users, domain experts than any imperative language.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5031">01:23:51</a>] Imperative language is a few million users. Excel, hundreds of millions of users. Even if you just restricted users who are using formulae. So how can we delight those users? Answer: take ideas from functional programming and use them to make Excel&#8217;s formula language more powerful. It took me 20 years, but Excel did finally add Lambda to Excel. Look it up, there are blog posts about it and many YouTube videos about it.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5065">01:24:25</a>] So you can now program in Excel using full lambda as <a href="https://en.wikipedia.org/wiki/Alonzo_Church">Alonzo Church</a> originally defined it. You can write anything, because lambda is computationally complete. You can write any computation in Excel now. It would be a bit slow, but you can. And much more practically, you can take formulae that previously you just copy-pasted here and there and wanted to make reusable, wrap them up in a lambda, and now you can just call the lambda.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5093">01:24:53</a>] And now it is a proper grown-up functional language that is Turing complete. Just search for Excel LAMBDA, those two keywords will get you lots of raw material.</p><h3>01:25:04 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5104">01:25:04</a>] Last question for you is, knowing everything that you know now from your career, if you could go back to when you just graduated from college and give yourself some advice, what would you say?</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5115">01:25:15</a>] All of these people who you see very successful wandering around, looking as if they&#8217;ve made it, I guess you might class me among them. Now, they are all just making it up as they go along. They feel insecure, uncertain, not sure what to do next, not sure what their next steps are, not sure what next big problem they&#8217;re going to tackle, unsure about whether what they&#8217;re doing is going to be successful or not.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5140">01:25:40</a>] And so all of their confidence is, I mean, they project confidence. Maybe that&#8217;s partly a life skill, but often they&#8217;re not. And so the fact that, in those days, of course, I felt very not confident. I would say, since all of these successful people are making it up as they go along, it&#8217;s fine for you to be as well. And they&#8217;ve been lucky. Moreover, they&#8217;ve been lucky. They&#8217;ve had.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5168">01:26:08</a>] But if you want to be lucky, you do need to put yourself in a position where accidents can happen to you, and that means taking risks. So if you want to be lucky, you need to put yourself in positions where lucky things could happen.</p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5188">01:26:28</a>] And that means taking some kind of risk. So if you&#8217;re very conservative and never take any risk, then it&#8217;s very unlikely that the accident that is life transforming will happen. That&#8217;s a balance, of course. But it means that accidents are not so bad. And of course, dying is bad. But lots of accidents may change your life in a way that might be surprising to you, but turns out to be not so bad in retrospect.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5210">01:26:50</a>] Awesome. Well, yeah. Thank you so much for your time. I really appreciate it, Professor Jones. We will.</p><p><strong>Simon:</strong></p><p>[<a href="https://youtu.be/I3DpfUhN5es?t=5214">01:26:54</a>] Nice talking to you. Thanks, Ryan.</p>]]></content:encoded></item><item><title><![CDATA[Turing Award Winner: P vs NP, Zero-Knowledge Proofs, Quantum Computation | Avi Wigderson]]></title><description><![CDATA[His field of study & life's work]]></description><link>https://www.developing.dev/p/turing-award-winner-p-vs-np-zero</link><guid isPermaLink="false">https://www.developing.dev/p/turing-award-winner-p-vs-np-zero</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 01 Jun 2026 10:02:09 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/199877077/0a9ebc30a395270c6d299fb4bcfc44b3.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Avi Wigderson is the only person in history to have won both a Turing Award (computer science) and Abel Prize (math). I interviewed him all about his field.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/5GUcvSAJcJw">YouTube</a>, <a href="https://open.spotify.com/episode/4JZrdA7ciMKLSarhx7m9Lt?si=EXZof0ERRgS3_wsRmjYEzg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-5GUcvSAJcJw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5GUcvSAJcJw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5GUcvSAJcJw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/199877077/0108-p-vs-np">01:08 - P vs NP</a></p><p><a href="https://www.developing.dev/i/199877077/1451-what-if-you-relaxed-correctness">14:51 - What if you relaxed correctness</a></p><p><a href="https://www.developing.dev/i/199877077/2538-why-np-complete-problems-are-equivalent">25:38 - Why NP complete problems are equivalent</a></p><p><a href="https://www.developing.dev/i/199877077/3033-space-vs-time-complexity">30:33 - Space vs time complexity</a></p><p><a href="https://www.developing.dev/i/199877077/4306-why-people-use-sat-solvers">43:06 - Why people use SAT solvers</a></p><p><a href="https://www.developing.dev/i/199877077/4553-randomness-is-a-resource">45:53 - Randomness is a resource</a></p><p><a href="https://www.developing.dev/i/199877077/5548-randomness-depends-on-computational-power">55:48 - Randomness depends on computational power</a></p><p><a href="https://www.developing.dev/i/199877077/012120-zero-knowledge-proofs-and-their-significance">01:21:20 - Zero knowledge proofs and their significance</a></p><p><a href="https://www.developing.dev/i/199877077/013830-quantum-computation-and-why-it-matters">01:38:30 - Quantum computation and why it matters</a></p><p><a href="https://www.developing.dev/i/199877077/015624-math-vs-computer-science">01:56:24 - Math vs computer science</a></p><p><a href="https://www.developing.dev/i/199877077/020816-major-breakthroughs-and-his-experience">02:08:16 - Major breakthroughs and his experience</a></p><p><a href="https://www.developing.dev/i/199877077/021231-advice-for-his-younger-self">02:12:31 - Advice for his younger self</a></p><h1>Transcript</h1><h3>01:08 &#8212; P vs NP</h3><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=68">01:08</a>] Well, in the simplest way, and I will talk at a high level, we are interested in various problems in life, in all sorts of problems: mathematical problems, scientific problems, medical problems, personal problems, intellectual problems. We are in the business, in computer science, of solving problems by computers. Now we would like to somehow classify those problems that we can solve, to know what we can know.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=103">01:43</a>] The problems we can solve are the things we will know, we will understand. There are all these problems that we want to understand. There are many of them, and there are problems that we can understand but can&#8217;t solve. In a very basic way, the first class is NP. All the problems we want to solve are really NP problems. It&#8217;s a class of problems, and a subset of it is the problems we can solve.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=133">02:13</a>] They are the problems we can currently solve, but we are interested in all the problems we can solve in principle. The question of <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> is whether we can solve all the problems we want to solve, whether we can know everything we want to know. This is basically about the limits of our knowledge. Of course, I have to justify why these classes, which are mathematically defined, are actually captured by the intuitive meanings I&#8217;m giving them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=172">02:52</a>] P is very easy to justify. What we can solve is what we can solve in our lifetime using whatever we have, let&#8217;s say all computers that we have using efficient algorithms. Problems we can solve are problems you see in our apps, in our phones. If there&#8217;s a navigation app, it means somebody invented an algorithm that&#8217;s efficient enough to solve any shortest path problem between any two places in any kind of map. All the problems we want to solve are NP problems.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=212">03:32</a>] What is NP? NP is a class of problems in mathematical definition, problems such that we don&#8217;t know whether they are easy or hard to solve. But if somebody hands us a solution, then we can easily check that indeed it is a good solution. Now, why are all problems humanity is interested in of this nature? You can think there&#8217;s no limit to what we may want to know. But I claim that essentially, in any human endeavor, if you are embarking on any problem, if you really seriously want to solve a problem, the least you want to know is that when you hit upon a solution, you will recognize it as a solution.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=259">04:19</a>] Why would you ever start looking for something if you will never recognize it as what you looked for? So you think mathematicians are looking for proofs of theorems? Certainly, if somebody gives a proof and writes a paper about it, we can verify it. If a scientist wants to explain data about the planets or about bacteria or whatever, they want to develop a theory. So that&#8217;s what you are searching for in this case.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=294">04:54</a>] And yeah, again, scientists write papers about their theories. They have to be consistent with the data. We have a way of recognizing that this is a good solution. Engineers are usually tasked with creating something, like a bridge or a phone, under some constraints. It can be monetary constraints, physical constraints, all sorts of limitations, such as what you can use but not this, and so on.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=324">05:24</a>] But anyway, if an engineer designs something, whoever gave them the task can look at it and say, &#8220;Well, you violated some constraints,&#8221; or &#8220;No, it&#8217;s good, you really delivered what we were looking for.&#8221; So I&#8217;m giving you examples. I mean, I can give many more. Detectives are supposed to solve crimes, and we know from detective books they eventually provide solutions. So these are examples of very fundamental human endeavors that encompass almost everything I can think of in all of them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=358">05:58</a>] What people are searching for has the property that when a solution is found, it is easily recognized. Repeating in the beginning, NP are all problems that we can honestly say we really want to solve. NP are those that we can solve. If they are equal, then everything we would want to know, a cure for cancer, anything you can imagine, can be just the very fact that a solution can be easily recognized, which would imply that it can be efficiently found.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=398">06:38</a>] That means that we can know everything we ever want to know. It&#8217;s a fundamental question about human knowledge.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=404">06:44</a>] So I know this is one of those Millennium Prize problems. There&#8217;s a million-dollar prize if you can provide a proof for this. I can imagine you could prove that they&#8217;re equivalent. You could also prove that they&#8217;re not. If you had to guess when and if a proof comes, which direction would your intuition say it goes? And why?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=427">07:07</a>] I think that my intuition, and that probably holds for almost all members of the theoretical computer science community, is that they are different. I think there are several reasons for that, and I will already tell you now that they are not that convincing. One is the intuition that most of us have that finding something is really hard and just checking that it&#8217;s there.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=457">07:37</a>] If you lost your keys or your phone and have no idea where to look, but if somebody points out, &#8220;Oh, you left it on the windowsill,&#8221; you might say, &#8220;Ah, yeah, I did.&#8221; We have this experience from life that finding something is typically a harder task than just checking that it is what we look for. Another reason is that real NP problems occur all over the place. They occur in optimization, in logic, in other aspects, in mathematics, and in verifying that algorithms, protocols, and security systems actually work.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=502">08:22</a>] All sorts of. In this generality, these optimization problems that are NP problems, maybe <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a> problems, people try to solve for completely egocentric reasons to make their company richer. There are thousands of these, and they manifest themselves in many different forms. Over the last 50, 60, or 70 years, people have seriously looked for algorithms for these very different ones and couldn&#8217;t find any.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=542">09:02</a>] This may seem like there&#8217;s no such solution. Getting a bit more technical, NP problems, <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems in particular, seem to require search over an exponentially large space, like all the possible solutions to some system of equations. Somehow, a solution is about having a general method for cutting down this exponentially large search space into searching over something much smaller in a clever way.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=588">09:48</a>] And this should apply to all these problems. So that would be equal. And people don&#8217;t think that there is such a way. But as I said, these are intuitions that are maybe pretty strong. And I think that&#8217;s the only, tomorrow somebody finds an efficient algorithm for <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems using some new idea that nobody ever thought of.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=622">10:22</a>] That would change the world. In software engineering, we use algorithms all the time, and they often have these very satisfying properties where the solution we use takes something where the brute force is much larger down to something that&#8217;s polynomial time or something very reasonable. But in all these NP problems, it&#8217;s pretty unsatisfying that the best we could do is brute force.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=655">10:55</a>] Am I understanding that correctly?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=657">10:57</a>] Yeah, you&#8217;re understanding it perfectly correctly. I think that maybe what you are getting at is that <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems are, a particular <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problem is really a family of instances. I mean, you want to find a <a href="https://en.wikipedia.org/wiki/Hamiltonian_path">Hamiltonian cycle</a>, some <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a> through some map. There&#8217;s this map and that map and this network and other network. And there are many, many instances when we talk about an algorithm for them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=692">11:32</a>] Typically, we say an algorithm is efficient if it&#8217;s efficient and correct on all of them. In software engineering and many aspects of optimization in real-world problems, the instances are a much more restricted set. They come from trying to verify a particular protocol or, I know, protein folding. If you want to think about the biological example, the way it is phrased as an NP problem is that you want to minimize the energy of a system under various constraints that have to relate to the chemistry of the various molecules and atoms that appear in the protein.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=745">12:25</a>] And this kind of optimization problem one can easily prove is <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a>. Nevertheless, our body folds proteins all the time, extremely efficiently. Now why is that? Of course, I don&#8217;t really know why that is. But it would seem that evolution designed proteins in such a way that this, you know, folding them, only them, we don&#8217;t have that many proteins. There are not exponentially many proteins in the body.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=777">12:57</a>] There are only a few that may be more prone to efficient energy minimization. More generally, I think in software engineering, at least in verification and testing, you often want to check that something meets the specifications. You usually translate it into a satisfiability question of Boolean formulas. The formulas that you get, the instances that you get out of these have some structure, and maybe for them, various greedy approaches or clever methods can be applied.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=818">13:38</a>] Not just greedy, but clever and efficient methods work. In other words, making it short. Instances that come up are not the worst-case instances. We have many examples of this that we can actually prove. For example, early on it was recognized that the <a href="https://en.wikipedia.org/wiki/Simplex_algorithm">Simplex method</a>, which is an algorithm to solve linear programming problems, on the one hand, requires exponential time for the worst-case system of inequalities.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=856">14:16</a>] But in practice, it seems to run in linear time. Evidently, the set of linear systems that we tackle in real life has some extra structure that allows this particular algorithm to solve them efficiently. Of course, today we know an algorithm that&#8217;s efficient for all linear programs, but it&#8217;s much more complicated, and I don&#8217;t know how much it&#8217;s used in practice. But the <a href="https://en.wikipedia.org/wiki/Simplex_algorithm">Simplex method</a> is very simple to run, and usually, it works.</p><h3>14:51 &#8212; What if you relaxed correctness</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=891">14:51</a>] You mentioned protein folding, and that&#8217;s an interesting one because it&#8217;s an NP problem, but it was solved to an acceptable extent. When I was reading about it, I found that it wasn&#8217;t solved for 100% of cases. It provides a solution and a confidence score, but it doesn&#8217;t give you the answer that is 100% accurate and fully computed. So when you think about these NP problems, what if you relaxed the criteria of being correct?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=927">15:27</a>] Could the asymptotic complexity be less than the exponential that we&#8217;re talking about?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=933">15:33</a>] Yeah, it&#8217;s a very natural and good question. So in the protein folding, you are talking about <a href="https://en.wikipedia.org/wiki/AlphaFold">AlphaFold</a>, which is an algorithm or heuristic that was learned from existing proteins and works very well on instances it didn&#8217;t see but are still proteins coming from the body. Even if it did perfectly on these, that would still be following the discussion I mentioned before because the instances it is trying to solve come from real life and probably have extra structure that allows efficient solutions.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=973">16:13</a>] But you are right that one of the things that people naturally do when they cannot solve an optimization problem exactly because it&#8217;s <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> or <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a>. Maybe I should say <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> and <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a> problems are those problems in NP that are as hard as all problems in NP. If you solved one efficiently, you solved all efficiently. If you prove one hard, then you prove all of them hard. So they somehow capture the complexity of the class.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1006">16:46</a>] The <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a> is an example. Portraying folding in this framework of energy minimization is one, and Boolean satisfiability and so on. Okay, so you want to optimize something, and it&#8217;s <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> or <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a>. That&#8217;s a sign that you shouldn&#8217;t try to find an algorithm that works always. But one thing you can do is an approximation. That&#8217;s a very natural relaxation. You don&#8217;t want the optimum.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1036">17:16</a>] You are willing to live with a factor of 2 from optimum or 10% from optimum or some factor away from optimum. NP-completeness was discovered in the early 70s. In the early 90s, there was a breakthrough called the <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a>. The <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a> allows you to argue hardness of even approximation problems. For example, you can show that satisfiability of Boolean formulas&#8212;it&#8217;s not just, so you want to, let&#8217;s say, find an assignment to a bunch of constraints, let&#8217;s say Boolean constraints, let&#8217;s say on three variables, for simplicity.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1086">18:06</a>] So you want to satisfy all of them. The relaxation would be okay, I want to satisfy 90% of them, or as good a fraction as I can if I cannot get 100% of them. The <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a>, one way to phrase it, and suddenly one of the major consequences is that we know that you cannot even achieve a very good approximation. To illustrate how tight this understanding is for some problems, when there are only three variables per constraint, if you just guess at random the values to the variables in the assignment, you will satisfy 7 or 8 of the constraints.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1135">18:55</a>] This is because each particular constraint, usually the disjunction of three variables, will be true with probability seven-eighths. So if you randomly assign values to the variables without thinking or looking at the formula, you can satisfy seven-eighths of them. The <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a>, along with a significant strengthening by Johan H&#229;stad, tells you that if you want to satisfy seven-eighths plus epsilon, that&#8217;s already <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a>.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1170">19:30</a>] So we know today to argue the difficulty not only of finding the perfect solution, the optimal solution, but even how close to optimal you can get for lots and lots of problems, large classes of problems, in particular constraint optimization problems, you know, of the type of satisfiability.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1196">19:56</a>] So even if you solve just the slightest epsilon for any trivially small epsilon, it&#8217;s just as hard as doing.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1206">20:06</a>] It&#8217;s as hard as finding the optimum source.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1209">20:09</a>] You mentioned <a href="https://en.wikipedia.org/wiki/NP-hardness">NP-hard</a>, <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a>, and P. I know there are other complexity classes. Maybe you could give some more context on the whole space of complexity classes.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1222">20:22</a>] Yeah, so I&#8217;ll just remark that we have a zoo of complexity classes. In fact, there&#8217;s a website called the <a href="https://en.wikipedia.org/wiki/Scott_Aaronson">Complexity Zoo</a>, which has in it hundreds and hundreds of complexity classes of all types. The website was started by <a href="https://en.wikipedia.org/wiki/Scott_Aaronson">Scott Aaronson</a>. Once we realize that there&#8217;s a whole zoo out there, we want to know the relationship between these classes. Okay, so what is the point about these classes like P and NP?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1262">21:02</a>] We want to classify problems by the amount of resources they take. So P (polynomial time) is a class of problems that can be solved in an amount of time that relates polynomially to the size of the data. Of course, you expect that as the data grows, the instance grows, and you&#8217;ll spend more time. But if it&#8217;s linear time, quadratic time, or cubic time, you say, let me call it at least theoretically efficient.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1295">21:35</a>] And NP, you classify not the time to solve, but the time to verify a solution if it was given. You want this to be polynomial time. So it&#8217;s very different. But first of all, there are many more time functions than just polynomial, exponential, and doubly exponential. And you can go to finite, and one fundamental result, starting with Turing&#8217;s paper, is that there are problems that are simply unsolvable by computers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1336">22:16</a>] So really, the first complexity class is a class of solvable problems, decidable problems. Turing shows that these are not all problems. There are natural problems that are not decidable at all. But time is only one resource, and we are interested, both from practical and theoretical reasons, in lots of other resources. Memory is a costly resource, and we want to minimize the space invested in a computation.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1367">22:47</a>] In problems with several parties, you are interested in the amount of communication exchange, not the internal computational effort that they exert themselves, but how much. If you talk to a satellite, you want to minimize the communication. You can think about energy; there are plenty of other resources, like parallel time. How much can you speed up the computation of a problem if you work on it not with one computer, but with n computers?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1401">23:21</a>] If your data is of size, then does it allow it to solve it immediately? Or maybe it doesn&#8217;t help at all? Complexity classes classify problems according to how much resources they need. There are many of these, and one important part of that is a major theme in the methodology of complexity theory. Coming with it are attempts to understand complexity classes as problems that relate to each other, even if you don&#8217;t know their complexity.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1442">24:02</a>] So, for example, NP is a class, and we don&#8217;t know the complexity of all problems in this or <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems. We don&#8217;t know whether they are all easy or hard, but we do know we have efficient algorithms to translate one into the other. It&#8217;s not obvious that if you want to solve a satisfiability problem, you don&#8217;t know how. But you have an algorithm to solve Sudoku problems, not just three by three, but also four by four. If you had an efficient method to solve Sudoku problems, P equals NP, because you can efficiently take an instance of satisfiability, the <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a>, or factoring integers, and from the solution of the Sudoku problem, you can translate back and get the best solution for your original <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a>.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1504">25:04</a>] So we are interested in efficient reductions between problems whose complexity we don&#8217;t know, but we want to relate them to each other. We build some kind of partial order of the hardness of different problems. There are really, you don&#8217;t want to know how many problems we have summarized for constraints that real-world technology may impose or cause for. Some are naturally arising mathematically because they are interesting.</p><h3>25:38 &#8212; Why NP complete problems are equivalent</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1538">25:38</a>] You mentioned that these <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems are all equivalent in some way. What&#8217;s the proof to equate a SAT solving problem to a graph coloring problem or something like that? Are they all similar in nature? Are they kind of bespoke to the problem you&#8217;re going to and from?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1559">25:59</a>] That&#8217;s a very good question. You will find that in most of these, they call them reduction algorithms to translate one problem to another. Most of them are relatively simple, but not all. I talked about the <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a> before, showing that approximation is hardest is very highly non-trivial. But most <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems prove equivalence, like between graph coloring and satisfiability. This goes both ways; they are both complete.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1592">26:32</a>] So it goes both ways. It is simple. The first one was the satisfiability problem. The original papers of Steve Cook and Leonid Levin, defining NP-completeness and proving that <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems exist, both focused on this. The first problem they proved complete was satisfiability. Why is this so? Why can all these other problems be reduced to satisfiability? The reason is really that computation is local. Computation is local.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1630">27:10</a>] I mean that if you think about your computer, your laptop, or any computational device, usually it manipulates bits by some local operations. You look at these two registers and you add them, or you take these two bits and exclusive OR them or AND them or something, and you do a sequence of simple operations and you arrive at a conclusion. This locality can be translated to a bunch of constraints or what would be a consistent sequence of states of your machine.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1668">27:48</a>] If you look at the time evolution of any computation, memories in some states and in another state change locally. If you could guess, and that&#8217;s the NP part, if somebody provided you the table of computation, you could check that it&#8217;s consistent by writing a bunch of local constraints or some conjunction of very few variables all over this table of computation and write down a satisfiability problem.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1705">28:25</a>] I want this to be true, this to be true, and this to be true for some assignment. What the assignment has to satisfy should be a table like this. All these constraints should be satisfied, should be true, and it has to be consistent with the input given, right? Some of the inputs have to match the first time zero content of the machine. This locality allows for many of these translations.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1736">28:56</a>] If you realize this, you will quickly find a reduction from coloring to satisfiability. It&#8217;s also not clear why coloring is an evolution of computation, but it starts already with local constraints. You want a few colors and that every constraint is satisfied. In some cases, it&#8217;s even easier. But the reduction from any NP computation&#8212;that&#8217;s what <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> means. Anything in NP can be translated to satisfiability.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1769">29:29</a>] NP is just a computation of a Turing machine, say only that you don&#8217;t know what the computation is. If somebody guessed it or showed it to you, then all you need is to verify this verification is local. And that&#8217;s the source of many simple NP-completeness reductions. Let me give you a really simple example. Why is factoring integers reducible to satisfiability? Because if I gave you factors of an integer, you could just multiply them and check that they equal the input.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1815">30:15</a>] This multiplication is a simple algorithm, right? I mean, we know how to do it from second grade or something.</p><h3>30:33 &#8212; Space vs time complexity</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1833">30:33</a>] You really, the factors you mentioned, time complexity and space complexity. Intuitively, in software engineering, we often trade those off for each other. Is there general theory on that, on how these two trade off?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1850">30:50</a>] Yeah, we have all these questions. They are natural questions for complexity theorists. How do two resources relate to each other, and can they trade off? Or maybe you can minimize them both at the same time. It depends on the problem. There are problems for which we know results, like multiplying time and space. It has to be at least the square of the length of the input. So you can either do it with very small space, like logarithmic space, but you need to pay maybe quadratic time.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1892">31:32</a>] And if you want linear time, you have to pay linear space or close to linear space. There are results of this type in various models of computation. That&#8217;s another thing you play with. Not all of them are equivalent in principle, but if you care about exact time or exact space, they are not equivalent. So there are problems for which there is a trade-off. There are problems where you can solve them in linear time and logarithmic space.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1925">32:05</a>] One very important general result that is the breakthrough from last year is a result of Ryan Williams. I&#8217;ll tell you what it is, but first, I need to describe what we knew about time and space. There&#8217;s a clear inequality, right? If you run in time T, you will never use more than space T, because you will never visit so many cells or registers in your machine. So space is at most time. Fifty years ago, Valiant, Paul, and Pipin improved this a little bit.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1961">32:41</a>] They said that if you run in time T, there&#8217;s an equivalent computation that will do this in slightly less space than T over log T. That was a big result. The running time, of course, will blow up. It will be very, very expensive time-wise, but at least you can save on space a little bit. This was believed for almost 50 years to be the best you can do. In fact, we had arguments that in certain models, in stylized models, models of computation, you cannot improve that.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=1999">33:19</a>] I will not describe this restricted model, but in some way, natural restricted models, you cannot beat this. What Ryan Williams found out last year is that, in fact, much better can be done. Any computation that runs in time T can be simulated by another algorithm that uses only the square root of T space, far less than before. The algorithm is quite sophisticated and uses an earlier result of James Cook, the son of Steve Cook, known for defining NP-completeness, and Ian Meltz, which was an essential technical ingredient.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2047">34:07</a>] But anyway, there&#8217;s a really interesting way to save space. A very non-trivial way to save space. Again, this algorithm will run in a lot of time, but at least you have a sense that the two parameters are related in a highly non-trivial way.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2069">34:29</a>] Is this for a general problem?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2071">34:31</a>] General problem for Turing machines? Okay, it may not work on random access machines or.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2077">34:37</a>] Yeah, what&#8217;s the trick? If you could explain intuitively, or is it too deep to explain?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2085">34:45</a>] It is really technically explaining it. It&#8217;s too deep. But I&#8217;ll give you a much earlier result which indicates that really mysterious things can be done in small space. Again, it&#8217;s probably over 40 years ago, people discovered the following. Dave Barrington discovered the following really interesting phenomenon. So suppose all you want is to count. I give you a sequence of bits, and you want to count.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2120">35:20</a>] Say you want to know whether there are more zeros than ones, like the majority problem. Who wants the vote, zeros or ones? Okay, well, you can count. You can just add them up, and this takes logarithmic space; it seems the best you can do. You want to count up to N. To represent N takes log N bits, right? Any number between 1 and N takes log N bits. It seems essential. It turns out that there&#8217;s an algorithm I have to define exactly how it works.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2158">35:58</a>] I will not define exactly how it works, but it solves this problem. In constant space, it seems like you can count arbitrarily high. All you need is access to whatever bit you like, whenever you like. I want the 17th one, I want the 81st one. If you can have random access to the bits, even if you have constant space, you can tell whether there are more zeros than ones or ones and zeros, regardless of how long the input is.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2193">36:33</a>] The trick is that somehow you use noncommutative algebra. Noncommutative algebra is not something everybody&#8217;s familiar with. I mean, we usually put the cars behind the horse and not in front of the horse. It matters. In software engineering, we do this before that; it&#8217;s very important that the order of things matters. So you can think about applying permutations one after the other.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2230">37:10</a>] If I rotate a circle and then take a mirror image, it will give me the... So these are two permutations, right? I mean, say I have endpoints in a circle. I can shift it by one, or I can, let&#8217;s say, flip it around some diameter. There are two permutations, and they don&#8217;t commute. If I first flip and then rotate, I&#8217;ll get something else than if I first rotate. Think of all the points as colored, so they don&#8217;t commute.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2265">37:45</a>] Barrington is using permutations to encode the bits in the input, specifically permutations of size five. When he sees a one, he may rotate, and when he sees a zero, he flips. The non-commutativity of these operations allows him to carry out a general formula, a formula of ANDs, ORs, and NOTs. Any formula of some size S enables him to simulate the ANDs, ORs, and NOTs in this formula through rotations and flips of these five-sided pentagons.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2325">38:45</a>] And so just to remember the configuration of this pentagon, you need five bits, right? So you need constant space. How this captures the computation, I can tell you one more sentence. Maybe it will not be clear to everybody, but these non-commuting things allow you to do the following type of operation: do a rotation, do a flip, then rotate back and flip back. This is called the commutator, and it can simulate an AND gate.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2373">39:33</a>] I can. It&#8217;s too much to describe in words without how it does it. But there&#8217;s a very famous analogy, which is a riddle that captures this <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problem. You want to hang a painting in the following way: you have a painting, it has a string connecting the two sides, but you want to hang it not on one nail, but on two nails.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2408">40:08</a>] Okay. There are two nails on the wall, and you can do with the string whatever you want. The property you want is that if the two nails are there, it&#8217;s hanging, and everybody&#8217;s happy. If you pull out one nail, it doesn&#8217;t matter which, the picture falls to the floor. How would you loop the string around these nails so that this happens? Clearly, this is an AND gate or an OR gate. Right? If one... Yeah. And this has to do.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2443">40:43</a>] I mean, if anybody finds a solution, they realize what non-commutativity I&#8217;m talking about. It&#8217;s a nice fiddle. Anyway, this type of trick, where you can really do this type of result in small space, is something you wouldn&#8217;t imagine possible. It&#8217;s a striking example. When I heard this result for the first time, I was a postdoc, and somebody told me, and I just didn&#8217;t believe it was possible.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2478">41:18</a>] You cannot count arbitrarily high with. The trick that Cook and Mertz have in their algorithm is a solution to a natural problem in much less space than you would think. It uses some sense tricks of this nature.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2504">41:44</a>] If you&#8217;re counting arbitrarily large, that is information. You&#8217;re saying maybe the intuition is that information is encoded in the sequence of operations.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2515">41:55</a>] It is encoded. Of course, you are not delivering the final count. You just say whether there are more zeros than ones. If you have to write it down, and if the answer takes some number of bits, then you need this number to write it down. If you have a decision problem, say like yes or no, like are there more zeros than ones? Then you can do it for any size input in constant space.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2546">42:26</a>] That&#8217;s incredible.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2548">42:28</a>] Pretty incredible, yeah. And by the way, this highly applicable result used in cryptography in fundamental ways is used in various places, and it&#8217;s an extremely useful result here. Unlike the Ryan Williams result, if you have a formula of size S, you can do it in constant space. The time of this algorithm does not blow up really a lot; it&#8217;s just quadratic. So it&#8217;s quadratic time.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2580">43:00</a>] So this is something actually doable, useful, efficient, and magical.</p><h3>43:06 &#8212; Why people use SAT solvers</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2586">43:06</a>] We talked about the equivalence between these <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems. In practice, a lot of people solve other problems using SAT solvers. I would have thought that would be less efficient, though, because you have to translate it and then you&#8217;re doing it in almost like a different problem space.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2605">43:25</a>] Yeah, you usually also enlarge the instance. In translation, you enlarge the instance. So what your question is, why do people use SAT solvers? I would say that there&#8217;s no problem other than satisfiability that people thought so hard about really optimizing the heuristics that efficiently work on many instances. There are very clever ways in which you can try to start. I mean, you are not going to guess all the N bits, but you want to start by guessing some bits that seem more pivotal. If you like in dominoes, maybe when you set them to one value, they force many other values to be set.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2662">44:22</a>] And if that happens, you cut down your subspace. There are many heuristics of this type and also more clever than this that allow solving such satisfiability problems if they have such structure. Again, in the worst case, it will not help you. There is, in fact, a conjecture that is much stronger somehow than NP-completeness, you know, <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems taking exponential time. It says that we don&#8217;t expect any savings.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2696">44:56</a>] We expect satisfiability problems to require really true to some constant times N. It&#8217;s not going to be two to the square or N to the log N or something like this. The worst case may be hard, but people optimize attacks on many such formulas that somehow have all sorts of structural properties arising maybe in practice. I think that&#8217;s the main reason they are used. It&#8217;s also convenient.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2733">45:33</a>] I mean, it&#8217;s very common. People do it for testing specifications that programs or protocols meet. The specifications are usually easily translated into simple constraints, and that&#8217;s almost a satisfiability problem.</p><h3>45:53 &#8212; Randomness is a resource</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2753">45:53</a>] We talked about all the different types of resources in an algorithm, and in one of your talks, you said something where you consider randomness another resource for an algorithm. What do you mean when you say randomness is a resource for an algorithm?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2770">46:10</a>] In the early 70s, of course, randomized algorithms existed since antiquity, and everybody was tossing coins for a lot of reasons. Statisticians do sampling and use randomness. But when algorithmics, designing efficient algorithms, became big&#8212;once we had computers&#8212;people realized that it can really enhance all sorts of computation. For example, we had no idea how to test the primality of a number.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2806">46:46</a>] And in the 70s, both Michael Rabin and Sulwin Slauson found probabilistic algorithms, which are fast to test primality. Okay, so what does it mean, a probabilistic algorithm? It&#8217;s an algorithm that&#8217;s allowed to make random choices. You can think that there is an internal little person or device inside that tosses coins. Of course, that&#8217;s not what happens. There&#8217;s no little person sitting in your laptop.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2835">47:15</a>] So the question is, where do you get these random bits? It&#8217;s very important to stress that in all these probabilistic algorithms, the underlying assumption is that the bits you get are perfect. They are half, half each one and independent of each other. It&#8217;s like a uniform distribution on all possibilities. Now, where do you get this? I mean, where seriously do you get this? If you run a public algorithm on your laptop, since it doesn&#8217;t have this person inside tossing coins, it does something well.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2872">47:52</a>] High-quality randomness of this type costs money, time, and memory. What do I mean by &#8220;cost money&#8221;? You can have a very cheap solution. You can just measure the thermal noise in your computer or use one of the Intel chips. You can measure Internet traffic and sample it, believing it&#8217;s random, and all of these things are used. You can do something mathematical, have some simple procedure that is actually deterministic, like a linear congruence generator or something like this, and believe it&#8217;s random.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2912">48:32</a>] Or you can take the digits of &#960;. It also looks random in some sense. There are many things you can use, but you don&#8217;t know they are random. If you want them to be random, one source&#8212; even this is not a perfect source&#8212; but we have quantum mechanics. People believe it works in life. It seems like a true theory of nature. There are all sorts of arguments about it, but let&#8217;s put them aside.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2941">49:01</a>] They predict. If you measure photons coming out from some source and you measure their spin, whether it&#8217;s up or down, the prediction is that each one is half, half, and it&#8217;s independent of each other. That&#8217;s very nice. But if you are going to build this device, it&#8217;s going to cost you a lot of money, and let alone that it will not be perfect. But let&#8217;s leave this aside. Anyway, since you want high-quality randomness, you have some way of guaranteeing, I mean, you are running it really on your laptop.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=2974">49:34</a>] It&#8217;s not okay. The paper was written; we can test primality with a probabilistic algorithm. Now we want to run it. What do you use? It makes sense to ask lots of questions about randomness. What guarantees the quality of the randomness? Maybe you can minimize the use of randomness. By the way, another issue with randomness is that the outcome is a random variable, and it&#8217;s not always correct. Right?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3002">50:02</a>] The whole point, in most of these algorithms, is that you just have some small probability of error. If you have a deterministic algorithm, there&#8217;s no error, so you also don&#8217;t like the error. Treating it as a resource is simply a convenient, complex theoretic way of saying how do we understand the amount of this resource we have to invest, or the number of bits we have to invest, their quality, and so on.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3030">50:30</a>] So it&#8217;s almost automatic. If you think like a complexity theorist, how do we minimize for a particular algorithm, like primarily testing? Do you really need this? Do you need them to be independent? Maybe you can generate them in a, you know, maybe you need n bits, but you can start from squared n bits or from log N bits that are truly random and make from them in some deterministic way, a sort of pseudo-random sequence you can call it, which is a good name actually, that will, for the purpose of this algorithm, look as if it was random.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3073">51:13</a>] If you could do that, you don&#8217;t need all these bits; you can use much fewer. This whole theory of reducing randomness, derandomization, and removing randomness is a huge field, and there are many problems and many ways of doing it. The story with primality is actually fascinating. I mentioned these two algorithms that were invented, and people were wondering about the deterministic algorithm for primality.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3105">51:45</a>] It&#8217;s a problem that has been articulated already by Gauss in the most complex theoretic way you can. Back in his days, hundreds of years ago, and maybe 150, I don&#8217;t know, he was asking for an indefatigable calculator. Calculator, of course, is a person. But he wanted this to be efficient. That&#8217;s basically what he was looking for, a primality test for large numbers. Because they really wanted to know about numbers, maybe only a few tens of digits.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3143">52:23</a>] They wanted to know whether the numbers are prime or not. There was no efficient algorithm. You can wonder about the Miller-Rabin or Solovay-Strassen algorithms; maybe you can use less randomness or structures, randomness that you can generate from fewer bits, but nobody has any good idea. Then, in the early 2000s, Agarwal, Kayal, and Saxena devised a different probabilistic primality test, and it&#8217;s different.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3175">52:55</a>] So the analysis of randomness in it is different. Once you understand how randomness is used, you can say, &#8220;Ah, we don&#8217;t need to have totally independent bits. It&#8217;s okay if it has some structure.&#8221; They found, using number-theoretic methods, a tool to generate pseudo-randomness from very few bits. That&#8217;s how it works. Sometimes, deterministic versions of probabilistic algorithms are discovered simply by understanding how the algorithm uses randomness, analyzing the algorithm, and realizing it doesn&#8217;t use all that much.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3229">53:49</a>] You mentioned a few times the quality of the randomness. How do you quantify the quality of random bits?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3236">53:56</a>] This is basically a question about what pseudorandomness is. The general answer is that it depends. You want to fool a particular algorithm, let&#8217;s say, for primality. You want the randomness to fool this algorithm; you want it not to notice that you switched from perfect randomness to something that&#8217;s really far less random, has far less entropy, but the analysis works nonetheless. For another algorithm, maybe you need a completely different type of pseudorandomness for it.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3275">54:35</a>] There are a set of examples of algorithmic problems for which, when you look at the analysis, they are much simpler than this primality testing. It seems to use not the independence of all the n bits. You really need only every pair of them to be independent, or every triple. Spaces of random variables that have this property are much smaller. You can generate them from log N bits, and so you can fool them with this.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3308">55:08</a>] So the quality of randomness is a phrase. I like that. The quality of randomness is in the eye of the beholder or in the computational power of the beholder. You really need it to be just as good as the tests that are applied to it. You just want to be completely pragmatic. You don&#8217;t care whether it&#8217;s random or not. You just care that an observer of a particular structure or particular computational power will not distinguish the pseudo-random distribution, which may have much less randomness in it than the perfect one.</p><h3>55:48 &#8212; Randomness depends on computational power</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3348">55:48</a>] In one of your lectures, actually, you mentioned that, and you gave a good example. I was wondering if you could give the intuitive example of why randomness is a function of the observer&#8217;s computational power.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3360">56:00</a>] Yeah. So if you watch this talk, the example I give is taken from this fundamental paper of Manuel Blum and <a href="https://en.wikipedia.org/wiki/Silvio_Micali">Silvio Micali</a>. They tell you to consider three experiments, and the experiments go as follows. They are always between you and me. You are the observer, I am the cause. I have a coin on my finger, and I toss it. Just as it leaves my finger, you are supposed to predict what the value will be when it falls on the floor, like in two seconds.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3398">56:38</a>] You have to immediately say, heads or tails, okay? Before you have to predict, before it falls. Well, this is the first experiment. And you know what I ask and what they ask? What do you think is a success, your success? What are your chances of predicting it? And the obvious answer is, you know, one half. You know what, how can it help? You know, what can you do? The second is when you sit there, but you have a laptop like you have now.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3430">57:10</a>] And, yeah, what can you do? I don&#8217;t know how fast you type, but the coin will be on the floor in a second. The third experiment is where your laptop is connected to a Cray supercomputer, and the Cray supercomputer is connected to a bunch of sensors and cameras and whatever devices you want. They are all trained on my finger, okay? And so as it leaves my finger, the coin. This apparatus certainly is more than enough to calculate all the angular momentum of the coin and the distance to the floor and the humidity in the air and whatever parameter that completely determines its motion.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3477">57:57</a>] In particular, how it will land in a split second, far less than the time needed. So the real point in this example is that the experiment, the random coin toss, did not change in all of them. It is the same what masses of people in math, physics, and philosophy have defined randomness in various ways. There are many definitions of randomness, and they all focus on this event, on the chronos or sequence of coin tosses.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3517">58:37</a>] We in computational complexity theory, starting from this paper, don&#8217;t care about this event. The event stays the same, and all it focuses on is the observer. The only thing that changed in these three experiments is the computational power. We want to know how much entropy is in the coin. If you don&#8217;t have enough computational power, it seems to be full entropy. It&#8217;s half. Half. You don&#8217;t know what it will be; if you have enough computational power, you can predict it completely, and then it has zero entropy.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3552">59:12</a>] So what changed is the observer. The observer is this algorithm we talked about before. Any test that&#8217;s applied to a distribution to test the quality of randomness, depending on the test, you want to make sure to use as little true randomness in a way that the observer will not notice. This is the source of many of the important theorems that show you can remove randomness or reduce randomness in probabilistic algorithms with or without assumptions.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3595">59:55</a>] And really, they are behind this understanding that we have today. Randomness in algorithms is not as powerful as we thought. It&#8217;s like if you know that P is different from NP or you have a <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a> that is exponentially hard or sustainable. If you know that, then in fact, I can give you a pseudorandom generator that will derandomize any. We know that under this assumption, we have what we call P equals BPP.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3628">01:00:28</a>] Anything that has an efficient probabilistic algorithm also has an efficient deterministic algorithm. You just need to know that there is a hard function somewhere. Under what assumptions do we know P equals BPP? The very natural assumption is that one of these <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problems, or in fact even problems in higher classes, requires a lot of hardware and requires exponential size circuits to solve. We believe that it&#8217;s like the <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> problem in strength, and you cannot count down the exponential subspace.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3674">01:01:14</a>] And so maybe people are happier with this assumption, but this assumption about time complexity has no randomness in it. At first, it&#8217;s shocking that it&#8217;s actually related to the problem of removing randomness from algorithms. So that&#8217;s one fundamental connection in this hardness versus randomness paradigm. But we know even more. We know that it goes both ways. We know that if you are trying to remove randomness from some algorithms, if you can do that, then you&#8217;ve found the hard function.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3722">01:02:02</a>] So there&#8217;s really an almost. It&#8217;s not an exact if and only if, but again, there are many variants of this. In some variants, if and only if it&#8217;s stronger, but we really know that it&#8217;s one or the other. Since it&#8217;s hard to believe that all these hard problems are easy, or the seemingly hard problem like <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a>, we tend much more to believe the other alternative, which is that randomness in algorithms is weak and you can remove it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3783">01:03:20</a>] Is there an intuitive explanation for the relationship between problem hardness and randomness?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3805">01:03:25</a>] Yeah, there&#8217;s almost always an intuitive explanation for this. So what&#8217;s the hard problem? I give you the input. You are a limited computer. You are a polynomial time algorithm. I give you the input. Let&#8217;s say it&#8217;s a <a href="https://en.wikipedia.org/wiki/Travelling_salesman_problem">Traveling Salesman Problem</a>, or you don&#8217;t know the answer. So there&#8217;s some entropy in the answer, right? In the same sense we discussed before, there&#8217;s some entropy, a tiny entropy, maybe exponentially small entropy, because if the input is random, maybe only a few instances are hard, and all the others you can solve. But at least you see a hint of a relationship between hardness and entropy.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3852">01:04:12</a>] This, of course, is not satisfactory because we want half. We want the answer to be like a coin toss, not like a very biased coin toss. So we need to find ways to amplify this uncertainty for you. Amplify the hardness. We want not just that some instances will be hard, but since they are random instances, whether the answer is yes or no, for polynomial time, the observer will be half and half.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3886">01:04:46</a>] It will not be able to gain an epsilon advantage over a random guess. For this, we developed all sorts of methods that amplify randomness. Now, this doesn&#8217;t solve the problem at all because I managed to take n bits of a random instance of a problem and manufacture one bit. That&#8217;s hard for you to guess. But I could have taken the first random bit I picked.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3920">01:05:20</a>] What you want to do is the opposite. You want a method that will take a few and generate many. That&#8217;s a pseudorandom generator. The pseudorandom generator starts with a few truly random bits and generates many. For this, we need other tools. There&#8217;s something that I developed with Noam Nisan called the NW generator. I like to say that people say making the first million dollars is easy, and afterwards, you can make many more millions.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3957">01:05:57</a>] You can think of it this way. The hardness in the way I describe it allows you to make $1 million. But once you can make one, there are methods to make many more, in fact, more than the number of bits you started with. But still, their quality depends on the hardness of the function you use. We still use this and your limitation of being in polynomial time. So you can also solve this hard instance.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=3990">01:06:30</a>] You build many instances from a short seed. You start from maybe a logarithmically long seed. You take many subsets of this, like N. You will need n of them. I ask you for the value of the solution to each one. Each one is hard, but we want to say that they are simultaneously hard. You cannot tell them apart. There is no correlation that you can detect between them, even though they are extremely correlated.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4020">01:07:00</a>] So, we have methods to do that as well. The connection you embed, having a hard problem to a limited observer, means there&#8217;s some uncertainty about the answer. That&#8217;s the key. That&#8217;s a starting point.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4034">01:07:14</a>] So when you say that the quality of the randomness is a function of the computational power of the observer, one thought that comes to mind is, imagine I have maybe infinite compute. Then nothing is truly random. Is that accurate to say? Like, you know, if I had infinite compute, the weather is not random. Nothing is random.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4062">01:07:42</a>] Okay. So it points to the fact that you have to really be careful about what question you are asking about randomness. If you want to apply this to probabilistic algorithms, then it&#8217;s really important that you are limited computationally. If you have infinite compute time, you will be able to distinguish a distribution with full entropy from one coming from a generator with low entropy.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4099">01:08:19</a>] So you could do that. But one thing you cannot do with infinite compute is something that&#8217;s information theoretic. There are many other things we need to do with random bits than just run probabilistic algorithms. For example, we want to generate passwords for our cryptography for security systems. The very demand is that they are random, right? I mean, Shannon&#8217;s theorem tells you the quality of your password is as good as the entropy in it.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4138">01:08:58</a>] So the notion of a secret rests on the true randomness of the bits. If I ask you to produce a random bit, or you ask me to produce a random bit, whether you have infinite computational power or not doesn&#8217;t matter. To generate this random bit, it has to be random. Your demand is not that it will succeed in some test; you really want the probability of heads to be half and tails to be half.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4174">01:09:34</a>] Then this definition or this task of producing randomness is not related to your computational power. Lucky or unlucky, I don&#8217;t know. For us, many times we do have implicit or explicit tests for this randomness. You say, &#8220;I can break this system or something.&#8221; Then you can maybe use pseudorandomness. But if your very demand is to have a random event, then you need to produce a random event.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4208">01:10:08</a>] I saw in my research that there are ways to create higher quality randomness by aggregating weaker sources. How does that work?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4222">01:10:22</a>] Well, it&#8217;s another theory I talked about before, the theory of pseudorandomness. There&#8217;s a whole different theory. They are related in nontrivial ways, but it&#8217;s called the theory of randomness extraction or randomness purification, maybe. And this is exactly what you asked about. We imagine that the world may not give us perfect randomness, but we have all these events we cannot predict, like the weather, various quantum phenomena, or sunspots. We cannot predict the stock prices, right? Otherwise, the market will be.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4264">01:11:04</a>] But these are not, even though these are unpredictable events, they are somewhat predictable. I mean, the weather tomorrow is more likely to be similar to the weather today than the opposite. So there are correlations. If you sample this weather, for example, there are also biases. I mean, sometimes in the spring it&#8217;s more, or in the summer it&#8217;s more likely to be hot than cold.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4293">01:11:33</a>] So there are biases and correlations. These are called weak random sources, and there&#8217;s a mathematical quantification of how weak they are. Basically, we talk about the amount of entropy in them. You have N bits; potentially, you can have entropy N, but maybe they were generated from a generator or they came from some source. It has some entropy in it, but you have no idea where. Maybe half of them are fixed and half of them are chaotic, and you don&#8217;t know which half.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4330">01:12:10</a>] Maybe they are each half, three-quarters zero and one-quarter heads or tails, have a virus that also has lots of entropy, but not full. The correlations may be even more complicated. You can imagine all sorts of situations like this, and you mathematically model it by saying, okay, there are N bits. It&#8217;s a probability distribution of N bits which has some entropy, let&#8217;s say the square root of N.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4359">01:12:39</a>] You don&#8217;t know where, and you want to use it in a probabilistic algorithm. So it&#8217;s a basic question: can we use physical events, weak sources coming from nature, maybe in probabilistic algorithms? It&#8217;s certainly not obvious. I mean, just feeding them to the algorithm as is is not going to work. There are very simple algorithms that will succeed on perfect randomness and will fail, making an error.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4394">01:13:14</a>] But the theory of randomness purification wants to take this sample of n bits with some amount of entropy, which is not full, and massage it to create from it maybe a shorter string that is of higher quality. Ideally, it would be perfect random bits. So, ideally, you can imagine that if I have n random bits and I have entropy root N, maybe I can massage out of there square root N, or maybe just the fourth root of n bits that will be essentially uniformly distributed.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4439">01:13:59</a>] Then if I have an algorithm that just needs this amount, that&#8217;s what I feel. Turns out you cannot do that; it&#8217;s impossible. But what you can do is not produce one string of length, let&#8217;s say, the square root of N, but you can produce N to the hundred of them, okay? And what you are guaranteed is that this randomness purification device is an official algorithm that takes any distribution with entropy in it and produces many blocks, polynomially many blocks, of the length roughly the entropy.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4481">01:14:41</a>] What you would hope to be perfect and what you are guaranteed is that 99% of them are truly perfect. So you have many, 99% of them are as good as perfect. And there&#8217;s 1% that is bad. You don&#8217;t know which it is. But this already solves your problem because you can run your algorithm on each one of them and take a majority vote. Doing this is extremely complicated. There are many ways of doing it. There are variants of various assumptions you can make about the entropy.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4521">01:15:21</a>] Maybe you have several weak sources that are not related to each other. There are various variants, and this theory is very well developed. But we know that the statement I made essentially is accurate. If you have one source with some entropy in it, you can generate polynomially many samples of the length of almost the entropy, and most of them will be perfect. So that&#8217;s as useful as having one.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4555">01:15:55</a>] Yeah, I vaguely remember skimming a bunch of papers, and one of the papers had this figure with a two-dimensional array of numbers. I think they were drawing rectangles or something like that. Is that one of the ideas?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4571">01:16:11</a>] So it&#8217;s a notion of pseudorandomness, which is related to one particular variant of this extraction problem where you have several weak sources, not just one. One is the hardest because that&#8217;s... Yeah, but maybe you have a few that are independent, and each of them has some entropy in it. I think what you are referring to is this sum product theorem. I cannot draw the pictures of that, but you don&#8217;t really need to.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4605">01:16:45</a>] So let&#8217;s just think about this problem. You have three sources of randomness; they are each weak, so each has just a little bit of entropy in it. All you want to produce is one source with somewhat larger entropy. You are losing the number of sources, but you are gaining entropy. If you can repeat this, you eventually get to full entropy. So that&#8217;s the idea. What combination of which sources would you take?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4637">01:17:17</a>] It&#8217;s not obvious. And the solution? Yeah, there certainly is one. Maybe you mean this one. There&#8217;s a paper I have with Barack and Impagliazzo about this particular problem, and we are using a result in what&#8217;s called arithmetic combinatorics. I can explain it very simply. I think that we learn about addition in second grade and about multiplication in third grade, maybe. The only motivation for multiplication is that it&#8217;s a way to shortcut repeated addition.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4676">01:17:56</a>] And that&#8217;s all we talk about. Then if you go to college, you learn that sums and products are the basic operations of fields. You can do it with integers, rational numbers, complex numbers, and finite fields. And so you remember these are the basic operations. But how do they relate to each other? It&#8217;s not clear. Maybe there are many ways to ask this question. But one way to ask this question that was suggested by Erd&#337;s and was solved for the first time by Erd&#337;s and Semer&#233;di, and then in a much more applicable form we really need by Bourgain and Katzentau, said the following.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4718">01:18:38</a>] Just think of a set of integers and add all pairs. So you start with, let&#8217;s say, K integers and add all pairs. How many new integers do you get? This you should think of as entropy increase. If you get many more somehow. So you have K integers. Well, it depends on what they are. I mean, if they are the first K integers and you add all pairs up, you get the numbers between 1 and 2K. That&#8217;s not much larger; that&#8217;s a factor of 2 larger.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4751">01:19:11</a>] If you want to increase entropy, you want to get some power of K bigger than one. Of course, if you take a random set of integers, you&#8217;ll get K squared. They&#8217;ll all be distinct because K squared integers. So somehow you want to say, it would be great if, for any set, when you take all pairs, you group. But it&#8217;s not true. By this example of the interval, you cannot take sums; you can take products, but also products don&#8217;t always increase.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4788">01:19:48</a>] If you take 1, 2, 4, 8, you take a geometric progression, you multiply all pairs, you just get twice as many as you started with. So that&#8217;s also not good. The magic in this sum-product theorem is that sums and products are orthogonal to each other. Somehow, if one of them fails to grow a set, then the other will grow this set. This is an amazing result. In this context, you think of the outcome of your sources as numbers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4831">01:20:31</a>] You multiply the first pair and add them to the third. You mix the sum and product. This guarantees growth. This is not obvious, but if you do this, it guarantees that you grow. The number of different values you get is significantly more than K. Maybe it&#8217;s K to the 1.5 or something. Entropy does increase, and that&#8217;s the source. You need to prove it. Distributionally, it&#8217;s not just the size.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4863">01:21:03</a>] The entropy is not just size; it&#8217;s more than that. Anyway, this is behind one solution to one of the problems about purification. It does not solve the one source problem. For this, you need other tools.</p><h3>01:21:20 &#8212; Zero knowledge proofs and their significance</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4880">01:21:20</a>] I saw that you had done some work in zero-knowledge proofs. I was wondering if you could explain what a zero-knowledge proof is. Maybe we can talk about its significance.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4890">01:21:30</a>] Sure, zero-knowledge proof. It&#8217;s pretty amazing that it became a household word. I mean, it was in the domain of theorists for a long time. The original definition of zero-knowledge proof came in a seminar paper by Goldwasser, Micali, and Rakoff, which also contains the definition of interactive proof. So even before zero-knowledge, they conceived of the notion that proofs don&#8217;t have to be written down like in mathematical papers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4926">01:22:06</a>] People can have a discussion, a randomized discussion, in which I try to convince you of something I want to prove to you. Something like, I know the proof of the Riemann Hypothesis, or I can solve this Sudoku puzzle, and we can do it interactively with the random. We are randomized. In particular, you, the verifier of my claim, are allowed to be randomized. So it&#8217;s like in randomized algorithms. But now there&#8217;s a prover.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4957">01:22:37</a>] It&#8217;s not just an algorithm; it&#8217;s a prover, like in NP. Someone who knows the solution. But the convincing has this interactive form. This model was suggested in this paper and also in a paper of Babai, parallel at the same time from different motivations. Basically, it offers a generalization of the notion of NP. NP is just when I send you a message, and it convinces you; you verify it and it convinces you.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=4989">01:23:09</a>] Now, we allow an interaction and we allow a small chance because you are randomized. We allow a small chance that I will convince you of a false claim. But this error can be reduced, like in probabilistic algorithms. You want one in a billion? You can get one in a billion, whatever. Anyway, so there&#8217;s a notion of an interactive proof. In the Goldwasser-Micali-Rothblum paper, they suggested another notion of interactive proof, which is more restricted.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5026">01:23:46</a>] You want to prove something, but the zero-knowledge proof means that you want the verifier to learn nothing, absolutely nothing about the proof except that it&#8217;s true. Now, this sounds really totally ridiculous. I mean, if you think about the last time you convinced somebody to change their mind about anything without providing them any knowledge, any new things they didn&#8217;t know, say, what are you talking about?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5057">01:24:17</a>] Such proofs don&#8217;t exist for anything. There&#8217;s no zero-knowledge proof for anything. They were not thinking about convincing someone of a political opinion; they were thinking about cryptography. Of course, they&#8217;ve created the foundation of cryptography in many other papers, but they are imagining cryptographic protocols in which you have secrets. You use these secrets in your computation, and the protocol tells you to do things that, if you don&#8217;t, then you&#8217;re violating the protocol.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5100">01:25:00</a>] So the others don&#8217;t want you to cheat. They want to make sure you perform the right operations on your secrets. For example, you are supposed to pick a public key by multiplying two prime numbers. If you multiply three or something else, then you are violating the protocol, and maybe security is not guaranteed. So I would like to convince you that the number I give you is actually a product of two primes.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5130">01:25:30</a>] I certainly don&#8217;t want to give you the two primes. What I really want to convince you of is that I computed this number by multiplying two primes. You learn from this interaction absolutely nothing except that you are convinced with very high probability that I did multiply two primes and not any other number, and I didn&#8217;t do anything else that I shouldn&#8217;t. You can think about lots of other cryptographic protocols where people are doing things like multiparty computation, very complex things.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5166">01:26:06</a>] They are confused with their secrets and don&#8217;t want to reveal them, whereas the others want to make sure that they did what they should. There are many applications for this idea. The only problem is that it sounds ridiculous because it seems impossible in mathematical terms. If someone came here and said that they proved <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a>, I would want to see the proof and then tell me, okay, I&#8217;ll prove it to you in zero-knowledge proof.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5200">01:26:40</a>] How can anybody convince me that they have a proof of <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a>? I come out congratulating them and amazed by them, and nevertheless, I know absolutely nothing about the way they proved it. I just know that they did this. So it really sounds ridiculous. And yeah, certainly one of my favorite papers, maybe my favorite, is the zero-knowledge proof paper, which came a year later, co-authored with Omer Goldreich and <a href="https://en.wikipedia.org/wiki/Silvio_Micali">Silvio Micali</a>.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5237">01:27:17</a>] Well, we showed that it&#8217;s not only not ridiculous, it&#8217;s universal. Namely, anything that has a mathematical proof also has a zero-knowledge interactive proof. Anything like <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> or that I multiply two primes, or anything that you can prove revealing your secret, you can prove without revealing your secret and be convinced beyond any reasonable doubt. So that&#8217;s possible.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5266">01:27:46</a>] What&#8217;s the intuition behind that? Let&#8217;s say I have a proof for <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a>, and we want to convert it to a zero-knowledge proof. How does that work?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5275">01:27:55</a>] How does this work? Well, first of all, it assumes cryptography. We assume we have some one-way functions. We assume that some problem, like factoring integers or discrete logarithms, or any number of one-way functions believed to be one-way, exists. One-way functions, if people don&#8217;t know, are functions that are easy to compute in one direction but hard to invert. For example, multiplying numbers is easy. Finding the factors of a number, which is the inverse problem, is believed to be hard.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5312">01:28:32</a>] Of course, we can never know any hard problems. That&#8217;s the pre-first entry question. We don&#8217;t know, but we believe many problems are hard, and this belief is actually shared by the whole world because the entire world is using cryptographic systems that rest on this assumption. All electronic commerce assumes this. So, we assume we have one-way functions. One-way functions allow you to create basically garbled message commitments.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5346">01:29:06</a>] I have a number in my head. I don&#8217;t want to tell you what the number is. It&#8217;s a number between 1 and 100. I can write down another number that looks like a random number to you. On the one hand, you have no idea. I claim this encodes my secret, okay? This commitment scheme, which is built very simply for one-way functions, guarantees two properties. A) You cannot tell what my secret is, even though you can see this other number I gave you.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5388">01:29:48</a>] And on the other hand, I cannot change my mind about my secret. It really commits me to it. So I can later provide you with a certificate that I was thinking about 17, and I can only do it for 17. I cannot do it for any other number. In fact, a very simple example is really using factoring in some sense the product of two primes. This number commits to its factors.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5427">01:30:27</a>] It uniquely defines its factors, right? The only problem is this is not exactly random products of primes. Yeah, but you can do it, so you can do this. Okay? So now that we have to take commitments, that are possible, that&#8217;s very important. I&#8217;m describing very high level; I&#8217;m hiding lots of things. Even the definition of zero-knowledge proof, the formal definition, is quite intricate. The intuition is obvious, and I said it.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5458">01:30:58</a>] But actually, formally defining it is non-trivial. At a high level, I&#8217;ll give you some idea about how a zero-knowledge proof looks. Now we have to think about what theorems I am proving to you. It was very important for us to figure out what problem to think about. Even though today you can do it for others, we were thinking about theorems of the type where, in the graph, we can color it in three colors.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5487">01:31:27</a>] It&#8217;s also a formal mathematical statement; either it&#8217;s doable or not, right? We simply focus on proving zero-knowledge claims of this type. The graph we both know, I claim I can color it with three colors. You don&#8217;t believe me. You want to be convinced of this fact, and you don&#8217;t want me to cheat you. You don&#8217;t want me to be able to provide an interactive proof. Moreover, I want it to be zero-knowledge.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5518">01:31:58</a>] I don&#8217;t want you to have the slightest idea of what my coloring is or what anything you didn&#8217;t know before is. Okay, so roughly the way it works. The honor is proof for this particular set of claims that are certainly not all claims. They are very structured claims. We look like the following. We iteratively repeat the following procedure. I put commitments for the coloring on every vertex. I tell you, you know, I don&#8217;t tell you it&#8217;s green, red, or blue, but I put a commitment to one of them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5558">01:32:38</a>] On each vertex, you will choose at random an edge of the graph and ask me to open these two envelopes to decommit. You will check that the colors you see are in this set. They are not gray or orange; they are either red, green, or blue, and they are different. This should give you some slight advantage or support to the belief that maybe I&#8217;m not cheating you.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5593">01:33:13</a>] Of course, it&#8217;s a very limited one because I can cheat you in some corner. We have to establish two things. We will repeat this again and again. We have to establish that it&#8217;s a proof, an interactive proof that I cannot cheat you. And we have to argue that zero-knowledge part. Okay, so let&#8217;s do one at a time. The correctness that I cannot fool you is very simple, is the following.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5629">01:33:49</a>] If the graph is not three-colorable, whatever I commit for there is one place which is an arrow. Either it has the wrong color and not allowed color, or two colors are equal. You have some non-trivial chance of catching this with your random guess. It&#8217;s one over the number of edges. It&#8217;s not so small. It&#8217;s a graph with a hundred edges. It&#8217;s one in a hundred if thousands.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5656">01:34:16</a>] If we repeat it, not a thousand, but 10,000 times, the probability that you don&#8217;t catch me in any of them drops exponentially to zero. So, of course, this establishes that I cannot fool you. The zero-knowledge proof is a more serious problem because if I keep using the same coloring, you will ask me about this edge and this edge, and then eventually you&#8217;ll know the coloring of all colors.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5690">01:34:50</a>] Here&#8217;s something that&#8217;s really special to the coloring that is being used. If I have one coloring of the graph, I really have six because I can permute the names; I can replace red and green, and it&#8217;s another valid coloring. So what I do in each one of these iterations is not use the same coloring, but use a random one of these possible six. They are all legal, and I just use one of these six.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5721">01:35:21</a>] What&#8217;s the advantage of this? When you open a pair of vertices under this distribution, one of the six, what you will see, if I do have a coloring&#8212;and that&#8217;s the only thing I want to stress&#8212;is that we only have to establish zero-knowledge proof if I really can prove the theorem. So if I do know the solution, the proof, I want to protect my knowledge. What happens when I reveal to you two of these colors on the adjacent vertices of the graph?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5752">01:35:52</a>] It will be true in this distribution. When you are given two vertices, it will simply be two different random colors. This is what you get by all permutations in this experimentation. What did you learn from this? Nothing. You could have picked two random different colors. You didn&#8217;t need me for that. So you didn&#8217;t learn anything. In each iteration, you learn nothing. This is the idea of the proof.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5780">01:36:20</a>] Well, this is the end of the proof that I can show there are zero-knowledge proofs for all graph coloring, three-coloring problems. What about all the Riemann Hypotheses and the <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> problem and factoring and all these other things? I want to claim to you and prove with zero-knowledge. Here it&#8217;s very simple. You use NP-completeness. This is an <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> problem. Anything that can be proved is really something in NP, right?</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5812">01:36:52</a>] This is the definition. So you reduce it to three-coloring, proving it in zero-knowledge. The main point about reductions in NP is that they don&#8217;t only convert the yes/no answer in a consistent way&#8212;if this is solvable, then this is solvable. But actually, if you have a witness, if you have a proof for, I don&#8217;t know, whatever, it will become a legal three-coloring of the graph you generate. So if you have a proof, you can also have the three-coloring.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5846">01:37:26</a>] It&#8217;s very important that the reduction also provides translation not just between the instances, but also between the proofs. NP-completeness theory gives you for free that if you solve the problem for <a href="https://en.wikipedia.org/wiki/Hamiltonian_path">Hamiltonian cycle</a> instances, you solve it for any. You can prove anything in zero-knowledge proof.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5867">01:37:47</a>] So with the zero-knowledge proofs, you can never be 100%.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5872">01:37:52</a>] No.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5873">01:37:53</a>] Okay, but practically, I mean exponentially approaching.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5878">01:37:58</a>] You can make the completeness 100%. Namely, if I do have a proof, you will be. Yeah, there&#8217;ll be no error in your proof. But there is a slight possibility that the graph is not three-colorable, and you will not catch me. You can make this exponentially small. You cannot make it zero.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5898">01:38:18</a>] When I think of cryptography and one-way functions, there&#8217;s this idea of <a href="https://en.wikipedia.org/wiki/Quantum_computing">quantum computing</a> that&#8217;s kind of changed complexity theory.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5907">01:38:27</a>] bit and big time.</p><h3>01:38:30 &#8212; Quantum computation and why it matters</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5910">01:38:30</a>] And I wanted to ask your thoughts on how <a href="https://en.wikipedia.org/wiki/Quantum_computing">quantum computing</a>, or that model of computation, is changing complexity theory. What are the big takeaways?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5922">01:38:42</a>] Okay, let me say what it is. First of all, quantum mechanics is a theory of nature that seems to be universally accepted. So, like with randomness, we can ask, why not enhance computers with this physical knowledge? We allow our computers to operate in the ways that quantum mechanics dictates, namely, manipulate bits in superposition using unitary operations, whatever this means.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5956">01:39:16</a>] But according to the rules of quantum mechanics, really strange things happen. I&#8217;m sure many people know about the interference patterns. Basically, you are working with probability theory. With negative numbers, events cannot aggregate; they can only cancel each other. Okay, so it&#8217;s a model of computation.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=5991">01:39:51</a>] It&#8217;s a generalization of Turing machines. In fact, it&#8217;s a generalization of randomized Turing machines. It&#8217;s very easy to see that a quantum computer is at least as strong as a probabilistic computer. How do you see it? You just measure the quantum bit before you start. You measure them. This is what I mentioned about the photons in the beginning. If you have some basic superposition on a quantum bit, if you measure it, you get a random bit.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6023">01:40:23</a>] So because of that, quantum computers are at least as strong as probabilistic computers. Okay, so it&#8217;s a model of computation. It was suggested in the 1980s. Richard Feynman, along with others, suggested letting algorithms use these quantum mechanical operations, mainly originally for simulating quantum systems rather than building bigger parallel systems, like we simulate other things, such as turbulence. I don&#8217;t know what the power of quantum algorithms is, again, let&#8217;s say running into polynomial time or efficient ones.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6068">01:41:08</a>] Can they do more than efficient classical ones, deterministic or probabilistic? It wasn&#8217;t clear for a while. There were a few examples that were very stylized but were not for concrete natural problems we care about. Then in 1994, <a href="https://en.wikipedia.org/wiki/Peter_Shor">Peter Shor</a> sort of created an earthquake or avalanche. He found quantum algorithms that are efficient, that factor integers and also compute discrete logarithms, the two most basic underpinnings of all security systems that exist.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6108">01:41:48</a>] And this sets the world on fire. Right. So lots of people try to do a lot of things. Of course, people want to have these algorithms they implement; maybe they want to break other people&#8217;s cryptosystems. Since then, billions were invested by companies, by governments, by lots of people trying to build the technological infrastructure. This is extremely complicated. Holding bits in superposition is extremely complicated.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6145">01:42:25</a>] These things are called noise. When people built classical computers like von Neumann, noise was one of the serious problems because the bits were really in vacuum tubes, and they had to contend with it. In fact, it built a nice theory of classical computers that have errors. They have to cope with errors; some of their components can be faulty. But today&#8217;s hardware has no errors to speak of.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6182">01:43:02</a>] We don&#8217;t, there are error-correcting mechanisms, but we don&#8217;t need them for classical computing. Intel chips don&#8217;t have error correction in them. Hardware is very reliable. When you move to quantum, there is this decoherence noise in quantum mechanics. Everything depends on everything. The world can influence the computation in your laptop, and protecting from this is very hard.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6213">01:43:33</a>] There are quantum error correcting codes, and that&#8217;s part of the solution. That&#8217;s only one of the problems. Just holding bits in superposition is hard, and there are many hard technological issues, and there&#8217;s progress on that. So this is one line of huge investment. The other line comes from cryptography because, of course, everybody should be worried. Forget quantum computers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6239">01:43:59</a>] If tomorrow, I mean really tomorrow, somebody finds even a classical factorization algorithm in polynomial time, I think there&#8217;ll be chaos in the world because nobody can do any transactions since most security systems still rely on factoring. Of course, we want to change this. We want to change the underlying assumptions of security. The whole revolution in cryptography was that we risk cryptography on computationally hard assumptions.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6278">01:44:38</a>] And with this, we build all the wonderful public key systems and all the magical things you can do by assuming that players are computationally limited. But now, if the adversary is a quantum computer, suddenly they can break this assumption. You want to find other mathematical problems, computational problems that are somehow hard even for quantum computers. So that&#8217;s a whole field, this change in complexity theory and cryptography.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6319">01:45:19</a>] There is an army of people who are just trying to invent problems. I maybe should stress one-way functions are easy to find. Most problems, most processes in nature are not easily reversible. You make an omelet from an egg. Reversing this doesn&#8217;t take the same amount of time; that&#8217;s the usual example of a physical one-way function. But trapdoor functions, like factoring problems, are what you can use to build public key systems.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6355">01:45:55</a>] And that&#8217;s the most really basic thing for electronic commerce. We don&#8217;t know many. So we know this. I mentioned factoring and discrete logarithm. In the 90s, Itay and Roark created another type of trapdoor function from problems on lattices in high dimensions. I will not describe it, but it&#8217;s another problem. Then people refined and got a few more of similar nature. It turned out when <a href="https://en.wikipedia.org/wiki/Peter_Shor">Peter Shor</a> discovered this, people immediately tried to solve other problems with quantum computers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6391">01:46:31</a>] In fact, even today we cannot solve too many other problems with quantum computers. These are special. The special thing about them is that somehow you can reduce them to finding periods in a signal. Periods are like fluorescence form. A Fourier transform turns out to be in exponential space. But you can somehow perform a Fourier transform with quantum computers using interference with lattice problems and problems related to it.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6422">01:47:02</a>] Some call it learning with errors. Similar to this, to this day, nobody has found an efficient quantum algorithm for them. So there&#8217;s not even a theory. Forget building quantum computers; we don&#8217;t know how to solve them. The world is not just complexity theory. The whole physical world of security systems and also governments like the <a href="https://en.wikipedia.org/wiki/National_Security_Agency">NSA</a> is supporting or asking the world to produce assumptions that may be resilient to quantum attacks.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6464">01:47:44</a>] And so this is a huge change. There is a major change in the interaction between computer scientists and physicists, which grew tremendously since its discovery. It&#8217;s rich in many ways that are not the influence of algorithmic thinking on physical theories, including today&#8217;s theories in quantum gravity, black holes, and so on. The impact is immense. We have new sources of problems, new models, and new complexity classes, et cetera, et cetera.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6499">01:48:19</a>] And there&#8217;s a fantastic interaction. It&#8217;s also enormous complexity theory in that it turns out you can discover quantum algorithms and maybe then dequantize them in special cases. Like I mentioned factoring before, de-randomizing. Sometimes it&#8217;s a way, it&#8217;s a road to discover new algorithms. There are connections between problems. It&#8217;s extremely rich. But one of the most amazing consequences is that we, being what we are, make up models and study them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6545">01:49:05</a>] And the interactive proofs I mentioned before, it turns out that even before quantum, they were generalized to interactive proofs not with one prover, but with many provers, which seems weird but turns out to be very important in itself because it led to this <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a>. Then people said, okay, let&#8217;s allow quantum verifiers, quantum provers, and see what they can do. So one amazing result, which is about five years ago, is the acronym MIP* = RE.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6580">01:49:40</a>] Of course, we have acronyms for all these complexity classes: MIP*, RE. You may have seen them, maybe you didn&#8217;t, and they look weird, but I&#8217;ll tell you what they say. It says that there is a really weird proof system with quantum provers trying to convince efficient verifiers. And it turns out, you know, you ask for which problems can they do it? Forget their own knowledge, just convince.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6612">01:50:12</a>] And it turns out they can do it for the halting problem, for problems that are not computable. But things that are not computable are verifiable by efficient verifiers if the provers are quantum and they are entangled. It&#8217;s a weird proof system, very weird, but it does something that looks totally ridiculous. Things that are uncomputable by any classifiable component are verifiable in this interactive probabilistic sense by efficient verifiers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6651">01:50:51</a>] So what I like to say in talking about this is that it seems the best hypothetical reaction of anybody who hears this is, &#8220;Okay, you complexity theorists, you play in your sandbox and you build all these sandcastles and make up all these models that have nothing to do with anything just because you can.&#8221; Once you get weird enough, you get weird enough consequences.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6680">01:51:20</a>] So one message which I think is very powerful is that this result has absolutely fundamental impact on mathematics and physics. It turns out that it implies, and this was already done in the initial paper, resolution of well-known conjectures in mathematics and physics. It turns out that, weird as it is, it&#8217;s a new mathematical technique to solve problems that nobody had any idea about. Famous problems, important problems that fields were dedicated to, are resolved by this result, by the techniques of this result.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6725">01:52:05</a>] So you ask about the impact of <a href="https://en.wikipedia.org/wiki/Quantum_computing">quantum computing</a>. The impact spans many generations with different motivations and developments, but all of them following the methodology of complexity theory, understanding the power of computational models, proof systems, and so on, has magically led to such a consequence. We are just at the beginning of this, right? This type of proof technique is now being explored and used, and there are more results using this type of technique to address long-standing problems.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6764">01:52:44</a>] Pretty amazing.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6765">01:52:45</a>] That&#8217;s incredible. Yeah, I mean, in complexity theory there are these Venn diagrams of all possible problems. The implicit assumption in this picture is that these are decidable problems.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6782">01:53:02</a>] Yeah, yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6783">01:53:03</a>] In some of these diagrams, there&#8217;s a little in the corner. Undecidable.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6789">01:53:09</a>] Yes, unreachable.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6790">01:53:10</a>] Yeah, right. And it&#8217;s just, I can&#8217;t, I don&#8217;t even understand how you could verify an undecidable problem.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6799">01:53:19</a>] It&#8217;s a 200-page paper. It builds on ten years of understanding, which involves a lot of development in quantum algorithms and quantum proof systems, but also relies on techniques from classical proof systems related to coding theory and various algebraic methods used to prove, for example, the <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a>, as I mentioned. It&#8217;s a huge body of work, and on top of this, they wrote this 200-page paper in which they had to develop many more tools.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6836">01:53:56</a>] And yeah, there you have it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6838">01:53:58</a>] Wow. I mean, when someone produces 200 pages of such complicated work that makes such an outrageous claim, how does that get verified by people?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6852">01:54:12</a>] I mean, it&#8217;s a very good question. This tool is for any mathematical result that is complicated. Of course, proofs of <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> and the Riemann hypothesis are generated very frequently, but some of them were generated by very serious people and took a long time to refute in these two cases. Others, like the Perelman proof of the Poincar&#233; conjecture, one of the Clay Millennium Prize problems, was another one like that.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6893">01:54:53</a>] And this took several years to verify and fix small bugs in. Books expositing this proof were written, and eventually, this particular paper is for some reason still in the referring stage of the Annals of Mathematics now, probably almost six years; it will be done. It&#8217;s robust in the sense that already new papers were written with this type of tools, in fact giving somewhat alternative proofs of the same statements and stronger statements in order to resolve other mathematical conjectures.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6929">01:55:29</a>] But yeah, it&#8217;s a serious thing. In fact, the original paper had a bug, like the original paper with Fermat by Wiles. Fermat&#8217;s Last Theorem also had a bug, and it was fixed. This was very early on, just when it came out, and it was fixed very quickly. So any mathematical claim&#8212;your question is very important. In general, mathematics is a mathematical truth that is a social construct.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6964">01:56:04</a>] That&#8217;s what mathematicians believe. Maybe we are moving into a world in which we have lean proofs. We have formal ways of writing proofs, and you can maybe formally verify them. But anyway, I think that this community is quite certain about.</p><h3>01:56:24 &#8212; Math vs computer science</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6984">01:56:24</a>] Yeah, in complexity theory, it&#8217;s very mathematical in nature when I read the papers, but obviously it has implications in computer science. What&#8217;s the relationship between mathematics and computer science to you?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=6999">01:56:39</a>] Well, the theory of computer science, my field, complexity theory and algorithms, is really lucky to have these two parents. It lives in both spaces. It is a mathematical field in the sense that most of what we generate are theorems and proofs in the mathematical sense. On the other hand, it relates to computation because the kind of questions we ask ourselves, the problems we are trying to solve, are partly motivated by trying to understand computation.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7039">01:57:19</a>] So you can think of theoretical computer science very much like you can think about analysis, algebra, and topology. There are notions in algebra where you are trying to understand equations or systems of equations. In analysis, you are trying to understand, very generally speaking, inequalities and continuous spaces. Right? So there are these notions you are trying to understand. Computation is one of these notions where you are trying to understand what you can produce by an evolution of simple local steps through a given environment.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7079">01:57:59</a>] The input can be a sequence of bits, or it can be the DNA of a person, from which you produce proteins using computation, or you produce a new baby using the evolution of a fertilized egg, or the weather. Computation is sort of everywhere, and we are trying to understand it in ways that are different as we care about resources&#8212;how much resources are exerted in any such computation. We are trying to model it, understand the algorithms, and let the underlying natural phenomena inform our understanding.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7128">01:58:48</a>] So we live in both spaces. We have a lot of input of problems and models from industry technology systems and so on. And we have all the mathematics, so we have the rigorous side coming from and the aesthetic side, I would say, like creating models of proof systems because they are nice, because they are interesting, not because they are implementable necessarily.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7161">01:59:21</a>] So this living in these two spaces is extremely beneficial to this theory of computation.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7172">01:59:32</a>] You mentioned earlier that some people might have the perspective of complexity theory that it&#8217;s kind of solving problems for problem&#8217;s sake. I read in another interview you did that you were never motivated by application. My natural thought was, what motivates you to solve all these really tricky, complicated problems?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7194">01:59:54</a>] Okay, so let me say one thing about the beginning of your question. It says about solving problems. I think that what we do more, much more, is modeling things. Modeling things then gives rise to problems that you want to understand. So I think the modeling part in the theory of computation, modeling various forms of computation, is a large part of theoretical computer science and manifests itself in particular in cryptography.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7223">02:00:23</a>] Well, there are models of adversaries in distributed computation. There are many. I think modeling, just asking questions, and creating definitions is a very central part of the field. It happens more than in mathematics, maybe because they are more ancient and definitions happened earlier. But I want to stress this; it&#8217;s a very important thing. For example, we discussed randomness, and our definition of randomness is totally different, which gives rise to very surprising and interesting practical theorems.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7266">02:01:06</a>] Okay, now that we said that, what motivates me or what motivates other people? Of course, everybody has their own motivation. I am a product of computer science education. All my undergrad, grad school, postdoc, and everything were in computer science. I certainly regret not taking many more math courses because then I wouldn&#8217;t have to learn these later in life. But that aside, I think that the focus on computation, which is what computer science education gives you, is different. Different systems and different issues are, of course, algorithms in databases and algorithms in programming languages, and in fact, famous algorithms in other fields.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7322">02:02:02</a>] There are models and algorithms. Computation is something that I&#8217;m interested in intrinsically. Over the years, I&#8217;ve realized that computation, you know, the field expanded to interact with all the sciences, and that computation happened everywhere in various forms, in physics, biology, and so on. It makes it a fundamental object of study. For me, given the way I was brought up, I guess it is particularly fascinating in itself.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7354">02:02:34</a>] I know from experience that lots of the proofs we develop, not many, but probably with much higher proportions than in other areas of math, are impactful in the world, in real systems, in security systems, in blockchains, in the real world. But this particular side of seeing exactly how you translate your algorithm into a system or your zero-knowledge proof, I predicted when we came up with the zero-knowledge proof that this would never be implemented, because the protocol that I described more or less to you is very costly.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7400">02:03:20</a>] I mean, you want to prove that I&#8217;m producing a product of two primes, and I convert it to a map, and then I do all this complicated procedure many times. It doesn&#8217;t seem that anybody would ever use such a thing. But yeah, I was wrong. I didn&#8217;t realize how motivated people can be and how important applications of zero-knowledge proofs can be in the real world, not in the theory of protocol design. They did simplify, maybe using more assumptions, different assumptions.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7434">02:03:54</a>] There&#8217;s a recent breakthrough of a postdoc here, Ilango, you may have heard, because it gets some publicity in Quanta and so on, where he introduced into the assumptions of cryptography not just hard computational problems, but also things of the nature of G&#246;del&#8217;s theorem, that something is unprovable in some mathematical proof system like formal frameworks, whatever mathematicians use. Somehow using that, they can get zero-knowledge that&#8217;s non-interactive.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7468">02:04:28</a>] You can get. Yeah, it&#8217;s an amazing thing. It gets richer. What I want to say is that, again from experience, I have 45 years of experience in this field, which has been fascinating. I mentioned to you in the email that there are extra benefits to this field. I think the community is amazing. Not just that it has so many brilliant minds, but young brilliant minds are entering the field all the time. It&#8217;s also very lively, interactive, and collaborative.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7514">02:05:14</a>] Yeah, many of my best friends are also colleagues, so it&#8217;s fantastic. The theoretical understanding of many things that we may ask ourselves for aesthetic reasons that are mathematical, we generalize something for generalization&#8217;s sake, which may not have a counterpart in industry. In fact, <a href="https://en.wikipedia.org/wiki/Peter_Shor">Peter Shor</a>&#8216;s work on quantum algorithms raises the question, why do that? There are no quantum computers.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7549">02:05:49</a>] Maybe let&#8217;s wait until somebody builds one and we understand why. So, no, the answer is no. You should try to understand whatever is natural for you to understand. This field has already, maybe it was not clear 40 years ago, but it is clear now that lots of theoretical understanding is not just a product of mathematical results, but because it is about computation, it is meaningful.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7585">02:06:25</a>] I don&#8217;t know often enough in the real world. There are many other examples besides cryptography and <a href="https://en.wikipedia.org/wiki/Quantum_computing">quantum computing</a>, including theory. The field has revolutionized coding theory mainly because of the <a href="https://en.wikipedia.org/wiki/PCP_theorem">PCP theorem</a> tools needed for it. There are lots of examples, and this eventually went into systems, like memory schemes. Look, I&#8217;m interested in <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a>.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7617">02:06:57</a>] This is a question about impossibility, right? This is a problem. We still haven&#8217;t gotten closer to this, at least in the years I&#8217;ve been at it. We are not much closer to showing hardness for anything. We don&#8217;t know that multiplication is harder than addition. It&#8217;s a basic question. So this kind of question by itself is certainly not practical. I mean, it&#8217;s not clear. Well, it is a little clear because if you find R functions, maybe you can substantiate cryptography on a theorem rather than on an assumption.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7655">02:07:35</a>] But anyway, there are questions in the field that don&#8217;t directly relate and will not directly relate to implementations of any system, but they are all part of trying to understand what efficient computation can do and what it cannot. Or when can you minimize resources of some type and when you cannot. And how do different problems relate to each other. This basic methodology of the field has created some wonderful edifice.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7690">02:08:10</a>] And I would say that we are still in the embryonic stage of understanding computation.</p><h3>02:08:16 &#8212; Major breakthroughs and his experience</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7696">02:08:16</a>] You mentioned a few times in this conversation that some people made a major discovery, which kind of broke some assumptions that were maybe 50 years old. Researchers make these advances, and I know in your career you&#8217;ve made a few as well. The time span between breakthroughs is sometimes decades.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7718">02:08:38</a>] a major discovery sometimes because a lot of</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7721">02:08:41</a>] people&#8217;s careers, like maybe if you&#8217;re just an engineer, you got your shipping projects every year. But in your case, let&#8217;s say you discovered <a href="https://en.wikipedia.org/wiki/P_versus_NP_problem">P versus NP</a> or something like that.</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7736">02:08:56</a>] Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7736">02:08:56</a>] How do you feel when you make those discoveries, those sparse discoveries in your career?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7742">02:09:02</a>] It&#8217;s great. Yeah, it happens. The big ones happen. The biggest, rarely I think of science&#8212;I mean realistically, science and math is a community of ants working. Most progress comes from these conferences, theory conferences, and folks, and then we have all sorts of satellite conferences. There are hundreds of papers in them, satellite, I mean, focused on concrete areas and learning theory equipped.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7773">02:09:33</a>] So I know online algorithms. There are many, many thousand researchers working in this, I don&#8217;t know, maybe. And they are producing many papers a year. Most of my papers are, you know, we make a little progress in understanding something. Usually, it&#8217;s not a big problem; we discover a variant of a technique or we can strengthen a bound on some. I find this essential.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7805">02:10:05</a>] I think this is true in science in general. Bigger understandings come more early, and they send a shockwave to everybody who learns them and uses them. But of course, those bigger understandings don&#8217;t have to be a resolution of 50 years of papers. It can be something that, when Barrington discovered his algorithm for counting that I mentioned, he was trying to actually prove that it&#8217;s impossible.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7839">02:10:39</a>] And it was not like somebody asked this question; the answer was obvious. It&#8217;s impossible. Right? So when you have an insight of this type, it&#8217;s amazing. Whenever you realize something really new for you and for the community, it&#8217;s a phenomenal result. It&#8217;s rare. I tell all my students and postdocs that, yeah, it&#8217;s not, you know, you don&#8217;t work for the million dollar. The million dollar is nothing, of course, for some of these problems.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7872">02:11:12</a>] But yeah, you work because you enjoy the practice of it. You enjoy thinking about the problems. In fact, most days in the life of a mathematician, unlike a systems person, you go in the morning, you come back in the evening, and you fail to do what you wanted to do. You just couldn&#8217;t. You just thought more. And I think that it&#8217;s really important. It&#8217;s not for everybody.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7905">02:11:45</a>] Clearly, it&#8217;s not satisfying if this happens every day and every week. It&#8217;s not for everybody, but it&#8217;s for the people for whom this activity of trying to think, of throwing ideas, of failing, but learning from the failures, is meaningful. People don&#8217;t realize that often you learn. The fact that you failed is not just a wasted day. There is something that you gain from it, maybe subconsciously, that will help you later.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7937">02:12:17</a>] So unless you enjoy this activity, maybe this field is not for you. Yeah, and then you see something, and even if it&#8217;s small, sometimes it&#8217;s very satisfactory.</p><h3>02:12:31 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7951">02:12:31</a>] Last question for you. I think people in the field might be curious because you have so much experience. If you could go back to the beginning of your career when you just started becoming a researcher, is there any advice that you&#8217;d give yourself?</p><p><strong>Avi:</strong></p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=7968">02:12:48</a>] I think I was lucky enough not to think that I would change anything. I think I was lucky enough, and many people are lucky in this field because the field is so accommodating to young people. It&#8217;s completely leveled; there are no hierarchies. People would go to conferences and meet. I remember my first conference meeting Dick Harp, who was in the field and invented, you know, did all these <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a> things.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=8003">02:13:23</a>] And I was a first-year graduate student, and somebody introduced me to him. He immediately asked me what I was doing, and I said I proved that some problem was <a href="https://en.wikipedia.org/wiki/NP-completeness">NP-complete</a>, and it was actually interesting. It amazed me. I was lucky to have phenomenal mentors as an undergrad. As a graduate student, I was lucky to have unbelievable collaborators over the years. You learn from all this.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=8032">02:13:52</a>] So it&#8217;s not very informative advice to tell people to be lucky. I think this environment exists. One piece of advice I give everybody is that in the early stages, they should work on things they enjoy the most. When you are discovering your talents as a researcher, one direction may be these are all problems. If you solve them, then you. But maybe those aren&#8217;t your problems; maybe it&#8217;s not your affinity to them.</p><p>[<a href="https://youtu.be/5GUcvSAJcJw?t=8068">02:14:28</a>] You have to experiment with the type of problems. It&#8217;s good to have a variety of areas and papers to read in these areas or problems to understand. It&#8217;s always good to start with a few to better understand your capabilities. Often, what you like most is what you are better at.</p>]]></content:encoded></item><item><title><![CDATA[3 Top Takeaways From Dropbox’s Former Most Senior Engineer | James Cowling]]></title><description><![CDATA[Advice for the AI era, fixing broken incentives, simplicity vs complexity]]></description><link>https://www.developing.dev/p/3-top-takeaways-from-dropboxs-former</link><guid isPermaLink="false">https://www.developing.dev/p/3-top-takeaways-from-dropboxs-former</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 25 May 2026 13:15:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/41eb2847-5e92-44ac-a4b0-77bbe11a71a2_2560x1440.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/jcowling/">James Cowling</a> was the most senior engineer at Dropbox (Senior Principal) before he left to start his own company, <a href="https://www.convex.dev/">Convex</a>. I interviewed him recently digging into all the technical work of his career, learnings he had, and advice tailored towards the new &#8220;AI era&#8221; we&#8217;re in now.</p><p>The conversation was a lot of fun and I&#8217;m glad a lot of his takes are refreshingly not &#8220;doomer&#8221; takes. Below are the top three takeaways I had from the conversation to save you time.</p><p>You can find the full conversation on <a href="https://youtu.be/3XkmNSuHFmY">YouTube</a>, <a href="https://open.spotify.com/episode/1Xt2SfdlM8huFQERnsknqP?si=qU-6R9T4QTiwYlpMeppmzg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>. The <a href="https://www.developing.dev/p/dropboxs-former-most-senior-eng-building">transcript is on Substack</a> if you prefer to skim.</p><div id="youtube2-3XkmNSuHFmY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;3XkmNSuHFmY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/3XkmNSuHFmY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="callout-block" data-callout="true"><h3><strong>Brought to you by:</strong></h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fNjA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fNjA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 424w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 848w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 1272w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fNjA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png" width="212" height="40.47802197802198" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:278,&quot;width&quot;:1456,&quot;resizeWidth&quot;:212,&quot;bytes&quot;:128843,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.developing.dev/i/198909775?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fNjA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 424w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 848w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 1272w, https://substackcdn.com/image/fetch/$s_!fNjA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bf5c97b-aa65-43c6-b11a-3e02c94acad9_3520x672.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><em><a href="https://workos.com/">WorkOS</a>: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code</em></p></div><h1><strong>Takeaways from the conversation:</strong></h1><p><strong>1) Career advice for the AI era</strong> - His take was that software isn&#8217;t about syntax or algorithms. It&#8217;s all about conceptualizing problems and coming up with clean solutions for them. And to build that muscle takes experience. He urged that people shouldn&#8217;t stop exercising that muscle or you&#8217;ll atrophy be left behind. Use AI but also make sure you aren&#8217;t being passive in your learning.</p><p>The other major point he had was that using Claude Code isn&#8217;t that hard if you are a good engineer. The value isn&#8217;t in memorizing the details and learning all the latest AI tools. The important part is building things and solving problems that matter. He said you should just ignore Twitter for the most part and focus on what actually matters.</p><p><strong>2) Fixing broken team incentives</strong> - The problem we discussed is when a team&#8217;s identity, mission and name all revolve around a system they own. What happens is these teams end up trying to protect the system rather than doing what is best for the company.</p><p>The example fix James gave is when he was at Dropbox, he worked on a huge migration to move off of AWS. The resulting team was named after the system they built. He went out of his way to rename the team the &#8220;Storage team&#8221; instead.</p><p>The reason this was so important is he felt that the direction of the team should be oriented around the problem they are solving for the company. Otherwise, imagine if moving back to AWS turns out to be better for the business. The team named after the existing system would have natural incentive to battle doing the right thing. He called this phenomenon &#8220;system bias&#8221;</p><p><strong>3) Simple systems are the goal</strong> - To the untrained eye, simple systems can seem obvious but actually designing simple systems is much harder than building complex ones. And the key James mentioned is that simplicity reduces operational burden. Simple systems are easier to keep running and debug when they break.</p><p>I asked him for a concrete example and he shared how Dropbox managed the metadata for where files are actually stored. All they did was have a cluster of 1000 MySQL nodes that stored the block ID and its location. Many people would say it wasn&#8217;t sophisticated but all the alternative proposals would ruin observability and simplicity of querying this data.</p><p>The idea of complexity being incentivized in larger tech companies frustrated him. To him, the goal is to solve the problem not to check off the box for complexity.</p><div><hr></div><p>I am really excited about some of the future guests that are coming on the show. I love the interviews with the people who have a piece of computer science history in them.</p><p>Also, I&#8217;ll have some AI research folks on soon too. Many people have requested I bring them on so people can learn how to future proof their careers.</p><p>Lastly, I interviewed an award-winning complexity theory professor from MIT and asked him some Leetcode questions. His optimal solution was asymptotically faster than what Leetcode says is &#8220;optimal&#8221;</p><p>Thanks for reading,<br>Ryan Peterman</p><p></p>]]></content:encoded></item><item><title><![CDATA[Dropbox’s Former Most Senior Eng: Building Great Systems and Advice for the AI Era | James Cowling]]></title><description><![CDATA[Transcript & Audio]]></description><link>https://www.developing.dev/p/dropboxs-former-most-senior-eng-building</link><guid isPermaLink="false">https://www.developing.dev/p/dropboxs-former-most-senior-eng-building</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 25 May 2026 10:00:57 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/198900303/7228c016160f5e0b2e81e93a2ea30d39.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/jcowling/">James Cowling</a> is the CTO at Convex and was previously the most senior engineer at Dropbox. We discussed technical details of his past projects, simplicity vs complexity, and career advice given where AI is today.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/3XkmNSuHFmY">YouTube</a>, <a href="https://open.spotify.com/episode/1Xt2SfdlM8huFQERnsknqP?si=qU-6R9T4QTiwYlpMeppmzg">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-3XkmNSuHFmY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;3XkmNSuHFmY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/3XkmNSuHFmY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/198900303/053-systems-work-during-his-phd">0:53 - Systems work during his PhD</a></p><p><a href="https://www.developing.dev/i/198900303/1305-dropbox-technical-deep-dive">13:05 - Dropbox technical deep dive</a></p><p><a href="https://www.developing.dev/i/198900303/2157-why-dropbox-migrated-from-aws">21:57 - Why Dropbox migrated from AWS</a></p><p><a href="https://www.developing.dev/i/198900303/3640-how-to-do-massive-migrations">36:40 - How to do massive migrations</a></p><p><a href="https://www.developing.dev/i/198900303/4431-simplicity-vs-complexity-in-promos">44:31 - Simplicity vs complexity in promos</a></p><p><a href="https://www.developing.dev/i/198900303/4923-what-technical-teams-should-be-focused-on">49:23 - What technical teams should be focused on</a></p><p><a href="https://www.developing.dev/i/198900303/10025-doing-the-right-thing-vs-promo-hypothetical">1:00:25 - Doing the right thing vs promo hypothetical</a></p><p><a href="https://www.developing.dev/i/198900303/10813-why-he-dipped-into-management-sometimes">1:08:13 - Why he dipped into management sometimes</a></p><p><a href="https://www.developing.dev/i/198900303/11136-why-you-shouldn-t-lead-by-example">1:11:36 - Why you shouldn&#8217;t lead by example</a></p><p><a href="https://www.developing.dev/i/198900303/12323-how-to-mentor-senior-staff-engineers">1:23:23 - How to mentor Senior Staff+ engineers</a></p><p><a href="https://www.developing.dev/i/198900303/12730-career-advice-for-the-ai-era">1:27:30 - Career advice for the AI era</a></p><p><a href="https://www.developing.dev/i/198900303/13721-why-he-started-his-own-company">1:37:21 - Why he started his own company</a></p><p><a href="https://www.developing.dev/i/198900303/14605-the-most-technically-challenging-work-of-his-career">1:46:05 - The most technically challenging work of his career</a></p><p><a href="https://www.developing.dev/i/198900303/14810-how-he-got-involved-in-silicon-valley">1:48:10 - How he got involved in Silicon Valley</a></p><p><a href="https://www.developing.dev/i/198900303/15216-career-regrets">1:52:16 - Career regrets</a></p><p><a href="https://www.developing.dev/i/198900303/15554-top-technical-book-recommendation">1:55:54 - Top technical book recommendation</a></p><p><a href="https://www.developing.dev/i/198900303/15636-younger-self-permanent-underclass-advice">1:56:36 - Younger self &amp; permanent underclass advice</a></p><h1>Transcript</h1><h3>0:53 &#8212; Systems work during his PhD</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=53">0:53</a>] Here&#8217;s the full episode. First off, your PhD thesis was huge. It&#8217;s 156 pages. I didn&#8217;t know theses were that long; I thought papers were maybe 10 pages or something like that. It&#8217;s its own book.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=72">1:12</a>] It&#8217;s very similar to <a href="https://en.wikipedia.org/wiki/Spanner_(database)">Spanner</a> ultimately. What&#8217;s funny is I did that work; it came out a little bit before <a href="https://en.wikipedia.org/wiki/Spanner_(database)">Spanner</a>, so I got a citation in the <a href="https://en.wikipedia.org/wiki/Spanner_(database)">Spanner</a> paper. But then <a href="https://en.wikipedia.org/wiki/Spanner_(database)">Spanner</a> came out, and everyone kind of forgot about my paper.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=87">1:27</a>] You know, about the research. If you could kind of describe the problem that <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a> was solving and how it solved it, those types of things. We can talk about that.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=98">1:38</a>] Yeah, absolutely. So my interest in my career has been two parallel threads. One has been abstractions in general, the idea about how to build simple models for complex problems. I find that a very difficult and intellectually stimulating design exercise&#8212;how to design APIs, basically. The other has been large-scale transactional systems. I&#8217;m a big fan of transactions. A transaction is doing a bunch of things at once, and I think transactions are just one of the most incredible abstractions we&#8217;ve invented because they allow us to manage probably the most difficult problem in computer science, which is concurrency.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=142">2:22</a>] <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a> was an algorithm for how to do distributed transaction coordination, particularly for transactions that I described at the time as one-shot transactions. I&#8217;m not actually sure if that was a standard term at the time or whether I made it up. A one-shot transaction means you send some code to the server and say, &#8220;Run this function, all these reads, all these writes on two, three, whatever nodes at once.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=175">2:55</a>] And how to make sure this commits atomically across multiple shards in a distributed system. That was my PhD thesis, and my master&#8217;s thesis was on Byzantine fault tolerance, a consensus protocol in the presence of malicious nodes. It focused on how to achieve agreement in state across multiple parties when there are malicious entities. I guess the most well-known work I did as part of grad school was on a paper called View Stamp Replication Revisited.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=207">3:27</a>] And that was a paper on redefining a protocol called View Stamp Replication Revisited, which predated <a href="https://en.wikipedia.org/wiki/Paxos_(computer_science)">Paxos</a>. It&#8217;s a very similar algorithm. <a href="https://en.wikipedia.org/wiki/Paxos_(computer_science)">Paxos</a>, <a href="https://en.wikipedia.org/wiki/Raft_(algorithm)">Raft</a>, and View Stamp Replication are all based on virtual synchrony; they&#8217;re all basically the same thing. That was a paper I wrote at the time, which ended up being influential to a few really great companies like Tiger Beetle.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=233">3:53</a>] You mentioned transactions, and I saw in the paper there&#8217;s this idea of an independent transaction. If I&#8217;m understanding correctly, a lot of the efficiency in a distributed system is lost by needing to get consensus and to vote for consensus. In your research, you found a way to avoid needing to vote to reach consensus. How did you do that?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=260">4:20</a>] Yes, I mean a lot of people think about performance maybe from the wrong angle because I think of performance as a factor of just this raw horsepower, how fast are your disks, how fast is your network, etc. No one has particularly faster disks or memory than anybody else. What really matters to performance in a large-scale system is eliminating points of coordination. So it&#8217;s how to allow systems to progress without having contention between parties.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=293">4:53</a>] Where you&#8217;re basically reducing parallel throughput to serial throughput across large numbers of transactions. In <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a>, we had this idea called an independent transaction. Again, that was the terminology at the time where there were two entities basically processing pure functions. They have independent state and they&#8217;ll both come to the same conclusion as a result.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=316">5:16</a>] And all they need to do then is serialize those transactions. They have to decide if it was to happen atomically across multiple nodes, what timestamp should it get. <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a>, the paper, was mostly about how to exchange these timestamps very efficiently. So, how to have multiple parties each propose a timestamp and then choose the maximum of these, basically so you could safely serialize the transaction.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=342">5:42</a>] And there&#8217;s a lot more complexity that goes into this. But really the focus there was about how to maximize throughput in a large distributed system without resulting in the alternative, which is two-phase commit and two-phase locking. Two-phase commit with two-phase locking is basically the standard approach for when you want to have multiple nodes agreeing on the same thing, where they both agree to lock their state, not process any other data, and then commit a transaction.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=369">6:09</a>] Two-phase commit can be quite low performance because you&#8217;re blocking basically the systems for the duration of the transaction. It can be high risk too because you are taking a dependency on another node; you&#8217;re basically blocked waiting for another node to return. And so that was what <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a> was. Now, it&#8217;s funny you asked about <a href="https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf">Granola</a> because I forgot about it.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=393">6:33</a>] That&#8217;s so far in my history, I kind of forgot about that work. I even forgot the phrase &#8220;independent transactions,&#8221; to be honest. It&#8217;s great to hear you bring it up again. But I guess all this stuff does feed into all the work you do later on in interesting ways.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=408">6:48</a>] When I think about distributed systems, I think about you&#8217;re at a big company and you got a bunch of machines. But when you were at MIT building this thing, how did you build and test it? Were there spare machines that the college had, or was this in the cloud?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=425">7:05</a>] Yeah, this was a long time ago. Now I&#8217;m showing my age. This is pretty early in the days of Amazon Web Services. We had a rack of servers in our office that we could use. I was fortunate enough to be at MIT where we could afford a rack. There was also a service called Planet Lab, which was a big communal set of nodes that academics could use to run tests. But Planet Lab was a communal system.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=454">7:34</a>] And so it was continually having problems. It&#8217;s just a free for all. For the longest time, I would just sleep next to my desk and wake up whenever. Planet Lab was free. Because you need to run some benchmarks these days, you just spin something up on Amazon Web Services. But then I would sleep next to my desk, and I would wake up in the middle of the night or at random times and check Planet Lab status and then kick off a benchmarking job.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=478">7:58</a>] Now this was right at the cusp of when, in some respects, Google ruined systems research. I say that with affection and respect to Google because before then, most of the research was coming out of academia. A lot of distributed systems research was done on smaller scales, and a lot of the value proposition was the ideas. So, like, hey, I wrote a paper, here&#8217;s an interesting new idea, and yeah, it&#8217;s all in theory.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=508">8:28</a>] And so the value that came out of it was the idea. Around this time, you started seeing papers out of Google, Amazon, and these various companies where it wasn&#8217;t necessarily a paper about an idea; it was a paper about a system. Hey, I built this giant system, it has a whole bunch of features, some of which are interesting, some of which are not, and by the way, it powers Gmail or whatever. That kicked off a pretty interesting transition because at that point, program committees reviewing papers started to expect to see realistic benchmarks that, frankly, grad students were not able to produce at least at the time.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=548">9:08</a>] I mean, I wasn&#8217;t running Gmail on my system, and I think it did, in some respects, obscure intellectual ideas. I&#8217;m more for papers about systems. But I think there&#8217;s also value in a paper about an idea. It&#8217;s not just, &#8220;We built this big thing, and it works, and it&#8217;s up to you to figure out what&#8217;s interesting about it.&#8221; Instead, it&#8217;s, &#8220;Here&#8217;s a new thing.&#8221; There&#8217;s a paper, an old paper just called Leases, where someone invented the idea of a time-based lock.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=577">9:37</a>] And it&#8217;s just a paper called Leases. That was a cool era of systems research. I think a lot of it now has shifted towards industrial research, where people are building stuff for practical purposes. Whereas I think it&#8217;s hard for academia to compete on pragmatism. Academia is really a great place to do impractical work.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=599">9:59</a>] When you look back on getting the PhD in academia, being enabled to completely explore an idea versus going into industry, maybe working on <a href="https://en.wikipedia.org/wiki/Spanner_(database)">Spanner</a> at Google or something like that, which path do you think would be better and why?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=618">10:18</a>] I think that&#8217;s a bit of a misconception a lot of folks have about PhD programs. A lot of students are high-achieving students in college, and they think that the PhD program is college. It&#8217;s like, hey, I like learning things, so I&#8217;m going to go do a PhD, but it&#8217;s not. A PhD is training to be a researcher, and for most people, they shouldn&#8217;t do that. If someone wants to be a professional software engineer, they probably shouldn&#8217;t spend their time training to be a researcher.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=646">10:46</a>] But with a really important caveat, I think there&#8217;s a really interesting and challenging developmental experience you go through in the PhD program at a top university, at least. You reach a certain point in time when you have a problem that you&#8217;re facing that no one else in the world knows the answer to. You have a problem that you are the world expert on, and you can&#8217;t chat about it with folks, but you can&#8217;t ask your advisor because your advisor doesn&#8217;t know either.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=672">11:12</a>] Right. So you face these difficult challenges where you can&#8217;t read a book, you can&#8217;t ask Claude. And so you&#8217;re forced to go through a quite difficult and, frankly, emotionally challenging experience of being unsure, being uncertain, and learning to think for yourself. I think that is really valuable for all engineers. I think that was really, really valuable in my career.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=699">11:39</a>] I do see a lot of folks early in their career think that all knowledge comes from reading or absorbing it from someone else, or that there&#8217;s a right way to do things. But I think being in a PhD program or being faced with really demanding, open questions does train your mind to be comfortable with that discomfort. I feel a little bit lucky that I went through grad school without large language models existing because I had to deal with this uncertainty.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=735">12:15</a>] There was no crutch to help me out, which I think was very valuable.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=738">12:18</a>] So, I mean, a lot of, if you want to be a researcher, go do a PhD. If you want to build stuff, go build stuff. Since leaving academia, I&#8217;ve done a lot of work that I guess one could call innovative. It advanced the industry in certain ways and did novel stuff. But our goal was never research. Our goal was just to solve problems.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=742">12:22</a>] I mean, a lot of, if you want to be a researcher, go do a PhD. If you want to build stuff, go build stuff. Since leaving academia, I&#8217;ve done a lot of work that one could call innovative. It advanced the industry in certain ways and did novel stuff. But our goal was never research. Our goal was just to solve problems.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=767">12:47</a>] And I think that&#8217;s the difference. In academia, your goal is to advance knowledge. In industry, your goal is to solve problems. I gravitate more towards just solving problems, and I find that a more comfortable environment within which to work.</p><h3>13:05 &#8212; Dropbox technical deep dive</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=785">13:05</a>] Going to your time in industry, at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, you became the most senior engineer at the company. Looking through all the projects that you did, I had a series of just technical curiosities. I saw this idea early in your career about multi-homing, and I wasn&#8217;t familiar with that concept. What is multi-homing? What&#8217;s the problem it solves?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=807">13:27</a>] Yeah, I mean, multi-homing is the ability to have data in two locations, two homes. Multi-homing can be valuable for a variety of reasons. One is called primary-secondary or lazily replicated multi-homing, whereby all writes commit authoritatively in one region and get replicated to a secondary region. This is normally used for business continuity. One thing we did at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> was to ensure that if the entire West Coast blew up, which hopefully wouldn&#8217;t happen, <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> would keep running because the data would be replicated in other regions, but with a window of vulnerability, with a window of time where there may be some data lost.</p><p>So that is, you know, that&#8217;s kind of primary-secondary replication or multi-homing. There&#8217;s something else called active-active multi-homing where there is truly an authoritative copy of the data in multiple locations. Basically, when you, for example, write to a system, you don&#8217;t externalize that write as having succeeded until it has landed in all the regions. For example, in the storage system at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, the block storage system, we truly had a multi-region replicated system where we could take down an entire region, say take down the region in Ashburn, Virginia, where everyone&#8217;s data centers are, and there&#8217;d be zero downtime for the company and the data was still safe in multiple regions.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=853">14:13</a>] So that is kind of primary-secondary replication or multi-homing. There&#8217;s something else called active-active multi-homing, where there is truly an authoritative copy of the data in multiple locations. Basically, when you write to a system, you don&#8217;t externalize that write as having succeeded until it has landed in all the regions. For example, in the storage system at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, the block storage system, we truly had a multi-region replicated system where we could take down an entire region, say take down the region in Ashburn, Virginia, where everyone&#8217;s data centers are, and there&#8217;d be zero downtime for the company, and the data was still safe in multiple regions.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=894">14:54</a>] I think multi-homing is something a lot of engineers aspire to work on because it seems like the right thing to do. But the reality is for most companies, it is not. Not because there&#8217;s very high cost, but because there are very high latency costs for this. The speed of light is fixed. If you have to synchronously write data across multiple regions in the United States, you&#8217;re going to have, say, 60 milliseconds in the commit path of your protocol, which for most applications is not tenable.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=932">15:32</a>] So there is a lot of desire; a lot of engineering teams reach out for advice on how to adopt active-active multi-homing for their company. I would normally say don&#8217;t do it. Frankly, if US East is down, if Amazon is down that day, that&#8217;s okay. If Amazon&#8217;s down that day, your company will be down. That&#8217;s a shame, right? But by avoiding that complexity, you&#8217;re going to be able to move much faster and build a much better product.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=960">16:00</a>] And so in my mind, systems are all about trade-offs and making the right ones. I would generally recommend most people do not make the trade-off to have partition tolerance or multi-region availability. Even though for a company like <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, yes, that does make sense. When your job is storing data and you have several exabytes of it and hundreds of millions of customers, I think that&#8217;s when it starts to make sense.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=986">16:26</a>] So it&#8217;s just another term for, I guess, data replication and having it available in other regions. I imagine also, I mean, the cost of storage is a concern. How many replicas would you keep for something like, let&#8217;s just say I stored something in <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>? It&#8217;s my document. Is that on the order of one or two, or are there multiple?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1007">16:47</a>] No, it&#8217;s on the order of many. Many. If you were to store a file in <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, I modeled the storage to be, as we advertise, at least 12 nines of durability internally. The models look around 24 nines of durability. That means the data is secure with 99.9999% where there&#8217;s 24 nines. Right? Which means, at least according to the model, the universe will be extinct before any data is lost.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1040">17:20</a>] And the way that is done is by a combination of what&#8217;s called erasure coding. Erasure coding is how you take several blocks of data and combine them together with an encoding scheme and spread them around in different locations. At that scale, at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> scale, you&#8217;re taking things into consideration, like putting data in different racks, different rows of a data center because they&#8217;re on different power feeds.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1061">17:41</a>] You&#8217;re taking into consideration different eras of hard drives and different manufacturers of drives because they could have correlated failure patterns and then replication across regions. So you could be looking at 27 fragments, for example. That doesn&#8217;t mean you&#8217;re storing 27 times the data. It&#8217;s kind of encoded in many regions. But the replication schemes get really quite sophisticated at that point.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1086">18:06</a>] And we actually had our own custom encoding matrix we had developed. It&#8217;s called a <a href="https://en.wikipedia.org/wiki/Vandermonde_matrix">Vandermonde matrix</a>, where you take a bunch of data and combine it together to produce outputs. We would plug in some variables like how much do disks cost? How much does network bandwidth cost? Because there&#8217;s a trade-off. You can either store, if you don&#8217;t want to lose your data, more copies on more disks, or you can store fewer copies.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1113">18:33</a>] Anytime a disk fails, you re-replicate it really fast. Re-replicating really fast costs disk cross network bandwidth. These kinds of variables go into this equation. It ends up being an extremely complex field of endeavor, but it&#8217;s kind of abstracted away into a part of the system that doesn&#8217;t leak into anywhere else.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1133">18:53</a>] So with erasure encoding, my document in <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> is fragmented into a bunch of different chunks of data and loaded potentially from many different machines.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1143">19:03</a>] Yes, absolutely.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1145">19:05</a>] I mean, the first thought I have is now I might be waiting there, and one of the 27 machines is slow, and I can&#8217;t look at the whole document. So how do you prevent against that?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1157">19:17</a>] It&#8217;s actually faster than not replicated. Because if you imagine, I&#8217;ll pick a simplified example. Imagine this is not the encoding scheme <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> uses, but to reconstruct a file, you have to read six out of nine fragments. So if you read six out of nine fragments, you can just ask all nine and reconstruct and return the data as soon as you&#8217;ve heard from the first six. It&#8217;s actually faster.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1186">19:46</a>] And you can construct these encoding matrices such that it&#8217;s actually faster to have a ratio-coded data than not.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1193">19:53</a>] I see. Okay. So you can oversubscribe, over-request, and then you complete on a portion of them being received.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1201">20:01</a>] Yes. Now in practice, it was a bit more complex than that. We&#8217;d often have a copy and a single disk to provide fast access. We&#8217;d often try to make sure that you could serve your data out of a region close to your home region. So we&#8217;d make sure that your data was mostly served with low latency. But if that region had failed, you could reconstruct it from the remaining regions. So there&#8217;s a lot of talk about building your own infrastructure.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1228">20:28</a>] And you can save money by moving off the cloud. Almost definitely you can&#8217;t unless you have very small requirements, very fixed requirements, or very heavy investment. Because if you want to compete with Amazon, if you want to build a more efficient storage system than Amazon, you have to have a supply chain team that&#8217;s working with Western Digital and Seagate, constantly negotiating on prices of disks and buying shipments at certain times, along with capacity teams and data center teams.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1260">21:00</a>] There&#8217;s a lot of work that goes into optimizing this because ultimately our desire was to use the disks to 90, 95% of their disk size to maximize storage efficiency. To put it this way, that was, you know, I guess a ballpark figure of a $1 billion project. And it was, as far as I know, the largest ever data migration in history at that time.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1288">21:28</a>] And so, an extremely large engineering project with very high technical investment. Yes, if you have that scale and you have the engineering team to do it, and you&#8217;re willing to keep innovating, if you&#8217;re willing to keep optimizing and investing effort in it, then yes, you can do it. But I think the real&#8212; I mean, the cloud has been an incredible innovation. Most people are not experts at this, and most people should not be experts at this.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1315">21:55</a>] Most people should focus on their applications.</p><h3>21:57 &#8212; Why Dropbox migrated from AWS</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1317">21:57</a>] When you say that most people shouldn&#8217;t, my immediate thought was, one of the big projects you worked on at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> was migrating away from Amazon Simple Storage Service. So why did <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> migrate away from S3?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1332">22:12</a>] Yeah, that was a desire at the company for a long time, from even before I was there. I started at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> in 2010, and I spoke to Drew, the <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> founder, about this project. There was a desire to control the destiny of the company from a strategic perspective. At the time, this was before <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> reshaped itself as being more about collaboration.</p><p>At that time, it was in the file sync and share category. Owning the file system was really valuable to the company. Ultimately, we saved a huge amount of money. And this was before the company went public. We really drove massive cost efficiencies through the project. But it was hard, in a way that I think would be very difficult to emulate without a huge investment.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1356">22:36</a>] At the time, it was a file sync and share category. That was the market sector, and owning the file system was really valuable to the company. Ultimately, we saved a huge amount of money. This was before the company went public. We really drove massive cost efficiencies through the project. But it was hard, in a way that I think would be very difficult to emulate without a huge investment.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1383">23:03</a>] And I do think there is a benefit to an organization from having hard problems to solve. If you have a company with extremely hard technical challenges, you can attract engineers who like working on those hard technical problems. When they&#8217;ve solved those problems, they cycle off and work on different parts of the system. After we all worked, we had such a great team.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1408">23:28</a>] It was a very, very small engineering team. After we shipped the storage system reliably, we all went off, and Jamie redesigned the sync protocol, the desktop client, and I worked on the file system and the distributed databases. There&#8217;s value to a business in having that level of technical investment. But it&#8217;s like having a baby; you have to raise the baby. You can&#8217;t just build a system like this and say, &#8220;That&#8217;s it, we&#8217;re done.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1439">23:59</a>] You own it, and you have to keep investing in it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1443">24:03</a>] Did S3 do any counter negotiation before you set out to leave them? Did they say, &#8220;Oh, we&#8217;ll cut you a deal?&#8221;</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1452">24:12</a>] I guess I&#8217;m allowed to talk about this now. It was a long time ago. For the longest time, I don&#8217;t think they were particularly aware that this was happening. But the data center folks talk, and certainly it was noticed that <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> was buying up a lot of data center space. So yeah, we had, obviously at the scales that we were at, we were negotiating very good rates with Amazon Web Services.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1478">24:38</a>] We weren&#8217;t paying sticker price. We were paying very, very, very good discounted rates. But at a certain point, they weren&#8217;t able to meet our cost efficiency because when we launched the system, it really was more efficient than Amazon Simple Storage Service. And that&#8217;s for a variety of reasons. One was that we were using kind of new experimental disks called Shingled Magnetic Recording. We were the first ones, I think, to use these disks at scale.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1502">25:02</a>] And two, we had a very tight understanding of our workload. We were able to design the system specifically optimized for our workloads, whereas S3 has to design the system for everybody. It got to the point where Amazon would not have been able to offer us a more competitive deal because we had a more efficient system. I wouldn&#8217;t recommend another company do this right now, but I think at the time it certainly made sense for us as a company.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1528">25:28</a>] Can you give an example of a tight understanding of your workload leading to something you could do that Amazon Simple Storage Service couldn&#8217;t?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1536">25:36</a>] Yeah, absolutely. So, for example, I know that, well, I&#8217;ll try not to leak any confidential data, right? But when you upload a file to <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, there is a pattern of access. Typically, people access the file very quickly shortly afterwards because you&#8217;re sharing it with someone, or maybe <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> is processing that file to generate an image preview, and then it decays at a certain rate. We understand in general the average block size, and we also understand the access pattern.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1571">26:11</a>] So we could do things like, at a certain point, we had these two clusters. One was designed for temporary storage that was kind of storage inefficient but access efficient. It was very cheap to read and write to, but it was inefficient to store. Data would get written to there first, and then in the background, it would get moved in bulk to this colder storage system. This cold storage was far more static.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1600">26:40</a>] And so it was able to have more efficient algorithms and be written to in bulk. If it went down for writes, that was no problem because it wasn&#8217;t in the live path. We were able to trade off again; like systems are all about trade-offs. We were able to trade off the live data write path from the long-term read path. That was one of many examples where knowing the size of your data, where it&#8217;s accessed from, how frequently it gets accessed, and how long it takes to delete that data, you can really tune a system to your workload.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1633">27:13</a>] You know, stuff like looking at even things down to knowing how much power to put in a rack. You have a rack of hardware; there&#8217;s a power distribution unit, a PDU. At the top of that rack, it has a circuit breaker which can handle a certain number of amps. We would have to figure out how many amps are required for that rack based on access patterns. There were times where we got it slightly wrong.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1657">27:37</a>] There was a time when we got maybe a bad batch of hardware, and disks were failing too frequently. As a result, we were re-replicating the data more regularly. I was getting messages from the data center team saying, &#8220;Hey, we&#8217;re running the racks really hot right now.&#8221; That&#8217;s the level of optimization you can make when you get to that scale. But again, that&#8217;s multi-exabyte scale&#8212;million hard drive scale.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1683">28:03</a>] Yeah. When I was reading about this migration, I saw somewhere in the migration you initially started with Go, and then the racks or something that you&#8217;re running on, the hardware itself was out of memory too much or requesting too much memory. And then you migrated to Rust.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1700">28:20</a>] So initially, the prototype was in Python, if you can believe that. To be fair, Python is actually pretty efficient for I/O. I think people give Python a bad rap for I/O-bound workloads. It&#8217;s pretty good at I/O, but obviously not great for concurrency, not great for memory management, and very hard to refactor. It was critical the system was correct. So we migrated everything to Go, and we built most of the storage system in Go.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1727">28:47</a>] This is before Go was in general availability. We built the system in Go. Go&#8217;s a great language for concurrency, a great language for proxies. It&#8217;s really well designed for servers that move from one place to another. At a certain point, though, we would have, let&#8217;s just pick a number, let&#8217;s say a million. Let&#8217;s say we have a million nodes in the system, and every node has some amount of memory, some amount of disk, some amount of sheet metal in the chassis.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1758">29:18</a>] Right. We would itemize all these things. You&#8217;d have a pie chart of how much money is spent on all the things. It would include items like sheet metal and screws. So you&#8217;re trying to optimize the storage. A big problem for us was the amount of memory these nodes were using. Not just the memory that we were using, but the unpredictability of it. With Go having a runtime, in a storage system, an out of memory error is pretty bad because if a node runs out of memory and restarts, that looks like a disk failure.</p><p>So that looks a lot like a disk has failed and has to be re-replicated. A batch of nodes running out of memory can lead to cascading failures throughout the system. I remember a time it was a band; it might have been De La Soul, I can&#8217;t remember. There was a band that released an album on <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, which caused a big spike in load to a few files. It was very high bandwidth and it caused the machines to run out of memory.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1794">29:54</a>] So that looks a lot like a disk has failed and has to be re-replicated. A batch of nodes running out of memory can lead to cascading failures throughout the system. I remember a time it was a band; it might have been De La Soul, I can&#8217;t remember. There was a band that released an album on <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, which caused a big spike in load to a few files. It was very high bandwidth, and it caused the machines to run out of memory.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1824">30:24</a>] So those machines, those disks oomed. As a result, no problems; the system went to try to recover that data from a whole bunch of other replicas. Now all of a sudden, you&#8217;ve taken one amount, one fire hose worth of load coming in, and you&#8217;ve turned this into seven fire hoses worth of load coming. Because now you have to do a more expensive reconstruction operation, right? So now you&#8217;ve 7x the load.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1847">30:47</a>] And these are the kind of cyclical behaviors that can lead to something called congestion collapse. Congestion collapse is when workload to a system crosses a threshold where it all kind of collapses. Designing against congestion collapse is really one of the hardest parts of Magic Pocket, which is the name of the storage system. Ultimately, we switched to Rust for the storage nodes themselves.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1873">31:13</a>] And this, again, this is before Rust was in GA. So that was a bit of a risky move. But we rewrote the switch to Rust, which coincided with getting rid of the file system entirely on the disks and directly addressing the disk heads. There&#8217;s an instruction set, I think it&#8217;s called ZBC, Zone Based Block Control, something like that. There&#8217;s an instruction set for accessing disks that we were using. The disk manufacturers gave us the draft specs of these new disks, and we were operating off the draft specs and directly controlling the disks.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1905">31:45</a>] And so all that product was tied up together into a product called Discotech, which was the disk technology project. Ultimately, if you look at that pie chart, congestion collapse stopped and reliability improved. But if you looked at the pie chart of where all the money was going, it really shifted to be almost all disks.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1936">32:16</a>] Metal, et cetera, the cascading failure you mentioned was there. That happened, and then there was a postmortem.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1943">32:23</a>] I don&#8217;t think we needed a postmortem. I think we knew as it was happening. When I was getting paged in the middle of the night on these things, it was pretty evident. This is, again, an argument in favor of the cloud. Imagine you&#8217;ve spent hundreds of millions of dollars on a storage system, and it&#8217;s out in production on physical hardware.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1967">32:47</a>] You can&#8217;t just go and put new memory chips in every one of them. I mean, you can, but you have to pay people to come in and swap them out. That&#8217;s a tricky place. This is what I love about the industry. Some of the more challenging moments on Magic Pocket were stuff like, in one week, just a weird coincidence, two trucks crashed that were delivering servers.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=1993">33:13</a>] So there are two trucks showing up to deliver racks, and they both crashed. I don&#8217;t know, the drivers were okay. So we lost capacity for two weeks. We lost capacity for more than two weeks, probably six weeks. What do you do? What happens is the equivalent of your disk filling up on your laptop, except it&#8217;s a million disks, and you can&#8217;t tell the customers to go away; you can&#8217;t delete their files.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2023">33:43</a>] A lot of tricky capacity work. Ultimately, what that led to was trying to build in all these protections against the unknowns, making sure we had the right amount of buffer planned out for anything bad that could happen, making sure we designed the system so there couldn&#8217;t be congestion collapse and so there wouldn&#8217;t be memory spikes. At one point, there&#8217;s a process called FMEA, which is a threat modeling process where you have a big spreadsheet and you kind of write down every bad thing that could possibly happen, and then you know how bad it would be if it happened. Existential risk does someone die? If there&#8217;s a fire in the data center, all these kinds of things get put into this spreadsheet, and you kind of do a bit of a pre-mortem to figure out all the potential failure modes and then design around them. I love that work. I really, I don&#8217;t know, I do like the firefight. I can&#8217;t say I like getting paged because I&#8217;ve spent my whole life on call, but I do like that rubber-hits-the-road stuff.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2061">34:21</a>] Existential risk does someone die? If there&#8217;s a fire in the data center, all these kinds of things get put into this spreadsheet, and you kind of do a bit of a pre-mortem to figure out all the potential failure modes and then design around them. I love that work. I really do like the firefight. I can&#8217;t say I like getting paged because I&#8217;ve spent my whole life on call, but I do like that rubber-hits-the-road stuff.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2092">34:52</a>] I like that. Wow. There&#8217;s congestion collapse and there&#8217;s no one that can help you. You&#8217;ve got to think through this problem.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2101">35:01</a>] When I worked on infrastructure at Instagram, we had this concept of defcon knobs, which are basically these configs that you could flip to gracefully degrade your system while still operating it. But maybe in the case of <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, you store fewer replicas, so you take on a temporary increase of risk in losing data because you need to. Did you have something like that, and did you flip it in that type of case?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2129">35:29</a>] Never. When it came to durability, we had absolute zero, non-negotiable standards; there was no room for negotiation on user durability. We had those knobs for background processes, for CPU and memory, for example. If there was a spike in load, you could turn off background processes and then turn off the test load on the system. Eventually, we built this system called Trampoline, and what Trampoline did was save us a ton of money.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2165">36:05</a>] If we ever got too close to the threshold, we would just start writing data to Amazon Simple Storage Service because it&#8217;s right there. You can run your capacity way closer to the edge if you&#8217;re willing, under worst-case scenarios, just to dump 30 petabytes on S3 and then move it back when it&#8217;s done. Now, that didn&#8217;t happen very often. We would do it to test it. We would do that just to make sure the system worked. But yeah, being able to have an escape hatch for worst-case scenarios was really nice.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2195">36:35</a>] I see. So S3 is kind of like elastic storage.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2199">36:39</a>] Exactly.</p><h3>36:40 &#8212; How to do massive migrations</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2200">36:40</a>] When I looked at this project, it&#8217;s such a massive migration, and my first thought is how do you coordinate this whole project without breaking the system as it&#8217;s running? What are your thoughts on doing such a large migration without breaking things?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2218">36:58</a>] Yeah. Now, in terms of the engineers that built the initial version of the system, it was a handful&#8212;three, four, five, six engineers. It wasn&#8217;t a team of a thousand; it was a very small team. So how do you build such a large system with a small number of people? And then how do you do a high-risk migration? One of the things you do is keep things simple. You try very hard to build cleanly abstracted, simple systems so that their failure modes are very understandable.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2253">37:33</a>] And so they&#8217;re decoupled, so they don&#8217;t have congestion, collapse, et cetera. Focusing on simplicity gets tricky, by the way, with people doing agentic development. They&#8217;re not the best at building simple systems. Simplicity is still the domain of human beings for now, but there&#8217;s a big focus on simplicity. The other was this very thick layer of validation checks during this migration.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2278">37:58</a>] And in fact, when we did the migration off of Amazon Simple Storage Service, we had something called the dark launch where we would be moving data off of S3, but we&#8217;d keep it in both locations. We had to demonstrate to the <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> founders that we would have to keep this system running with no incidents, no downtime, and no data loss for six months before we would delete any of the data from S3.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2306">38:26</a>] So we have double routed, and there&#8217;s one point in time, halfway through this process, where a bug got through to production. Nothing bad happened with the bug, but it was like bugs slipped through our multiple layers of the release process, etc. I went to the VP and I said, &#8220;Hey, a bug made it through to production, and we&#8217;re going to reset the launch clock. As a result, it&#8217;s going to launch later, and it&#8217;s going to cost us some amount of money, let&#8217;s say double-digit millions.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2345">39:05</a>] And they were like, great, thank you, that&#8217;s good, I trust you. That was the cool thing. I mean, <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> had a lot of incredible cultural values, but that was like, okay, cool. If you&#8217;re prioritizing user safety, that&#8217;s the right thing to do. No one was mad about that. It was almost like they were proud, you know? It was just like, yes, you&#8217;re operating in accordance with the principles of this company.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2371">39:31</a>] Okay. So it was double writing both systems for a while and then some subset of data.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2374">39:34</a>] for a while and then some subset of data.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2376">39:36</a>] Yeah, okay. And then you switched over reads once.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2379">39:39</a>] We were sure that the system was durable and we had all the validators running in production 24/7. We were just migrating as fast as we could. I think at some point we got to 764 gigabits per second of peering bandwidth between Amazon servers and ours. Certainly, someone on the network team over there noticed that there was that much data moving out. At one point, I got a slightly nasty email from someone saying it was super weird.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2408">40:08</a>] I think the phrase &#8220;super weird&#8221; is like, it&#8217;s super weird that you&#8217;re doing so many reads and not that many writes. I didn&#8217;t respond to that email. But to be fair, Amazon was a great partner for <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>. So <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> still uses Amazon Web Services. There was no concern that they would do the wrong thing by us as a company. I mean, we only had an excellent experience with Amazon Web Services. But yeah, it was a moment for us.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2436">40:36</a>] It probably strained the relationship somewhat.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2441">40:41</a>] You mentioned simplicity and intuition; it makes sense. Do you have a concrete example, though?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2448">40:48</a>] Yeah, I&#8217;ve got a concrete example for you. Maybe it shows the difference between academia and industry. The storage system is a giant distributed system with files stored in various locations. You need a mapping from the file to where it lives on these disks. All we did was have a cluster of 1,000 MySQL nodes, a big giant database. It was indexed by the block ID and indicated that this block is on these disks.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2481">41:21</a>] And that&#8217;s pretty simple. It&#8217;s not sophisticated. Every time we&#8217;d hire someone out of academia or maybe from other companies, they would say, &#8220;Oh, this is not very sophisticated because you could use a Patricia trie or you could use a distributed hash table, and that would map a block to a set of locations.&#8221; I think that&#8217;s optimizing for the wrong thing because the really nice thing about dumping a list of files and the locations in a giant database is that it is written in one location.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2517">41:57</a>] If I want to validate what happened, if I want to check all the data is where it&#8217;s meant to be, I just walk over the table and check. We had services constantly walking over the table and checking. Whereas if it was a distributed hash table or some giant complex data structure, it&#8217;s very hard to validate. Designing for validation is very important. Designing for understanding is very important.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2539">42:19</a>] It&#8217;s not about getting a system to work. It&#8217;s what do you do when it doesn&#8217;t work. Right. And so having a very simple boundary, that&#8217;s a very basic example. There are more sophisticated examples that take more time to explain. But something like that is a lot of engineers will feed, a lot of engineers will want to do interesting work, will want to advance in their career. They want to be seen as an intellectual problem solver.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2569">42:49</a>] And so the tendency can be to design complex systems. My argument is always that simple systems are way harder to design than complex systems. Simplicity is so hard. And I think to the untrained eye, a simple system can seem obvious. The best compliment you could ever get about anything you design is when people say, &#8220;Oh, isn&#8217;t that the obvious way of doing it?&#8221; It&#8217;s the same as <a href="https://www.convex.dev/">Convex</a>; people say, &#8220;Oh, isn&#8217;t that just the obvious way of structuring?&#8221; Great, because it wasn&#8217;t obvious when we did it. No one else was doing it. Everyone thought we were idiots. If after the fact people think it&#8217;s obvious, then you really nailed it. But I think it requires an understanding that simplicity is the hardest thing in systems. And because simplicity is scalable. Yes, simplicity is scalable in terms of numbers of queries per second.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2600">43:20</a>] What&#8217;s that? That&#8217;s just like the obvious way of structuring. Great. Because it wasn&#8217;t obvious when we did it. No one else was doing it. Right, right. Everyone thought we were idiots. If after the fact people think it&#8217;s obvious, then you really nailed it. But I think it requires an understanding that simplicity is the hardest thing in systems. And because simplicity is scalable. And yes, simplicity is scalable in terms of numbers of queries per second.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2628">43:48</a>] But what I really mean about scalability is you can take a simple system and have it run for five years, have people work in it for five years, have all sorts of features added to it, and have requirements changed because the company realized the product didn&#8217;t work the way it wanted to work and it wants to change things. It still stands the test of time, whereas a complex, over-optimized system will not.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2651">44:11</a>] I think that&#8217;s the tough thing about distributed systems design, especially LLM-augmented distributed systems design. Just because something works doesn&#8217;t mean it&#8217;s maintainable over a long period of time, doesn&#8217;t mean it&#8217;s understandable, and doesn&#8217;t mean it&#8217;s cleanly architected and abstracted. That stuff&#8217;s really very hard.</p><h3>44:31 &#8212; Simplicity vs complexity in promos</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2671">44:31</a>] Absolutely. And I agree with you. I think it&#8217;s the long-term beneficial thing to do. One unusual thing, though, in the industry that I&#8217;ve seen is the incentive system for engineers. I mean, you mentioned the desire for an engineer to want to be seen as someone who can do something difficult. There&#8217;s that, but there&#8217;s also the incentive system of promotions. I&#8217;ve had many friends whose promotions were rejected because their work wasn&#8217;t complex enough.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2701">45:01</a>] And so that kind of forces complexity, which is kind of unusual. I wanted to know what you thought about that.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2708">45:08</a>] Yeah, I mean it almost angers me. I just like it so much. Partly why I started my own company. I think the ideal for anyone is to be doing work where you&#8217;re being appreciated for solving the problem. If we get philosophical, this is what it was like going back to the farming days, right? There was no incentive to make it really complicated to milk a cow because the goal is to milk the cow, and then the reward is you got milk.</p><p>At a startup, that&#8217;s the same thing. The goal is to build the system, have it work, have the users like it, have it grow, and everyone gets rewarded and celebrated for solving the problem. It gets hard to scale that. So at large companies, you end up with so many layers of organization that people end up building alternative incentive structures.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2742">45:42</a>] Right. I think it sounds so silly, but at a startup, the goal is to build the system, have it work, have the users like it, have it grow, and everyone gets rewarded and celebrated for solving the problem. It gets hard to scale that. So at large companies, you end up with so many layers of organization that people build alternative incentive structures.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2768">46:08</a>] Right. It&#8217;s like I&#8217;m so far away from whatever the hell we&#8217;re trying to do over here that my goal now is to get all green check marks on my OKR plan. But who cares about your OKR plan unless it solves the problem? The thing that really drives me insane is when people try to chase artificial goals. I understand that if you&#8217;re in a company like this, you may have no choice in the matter.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2799">46:39</a>] But what I want to tell people is there is a better way. That better way may not be available to you; you may not have job opportunities near where you are, for example. But if you do have the ability to work at a company where you are being appreciated for problem solving, that will make you so much better as an engineer. I see this when I interview people.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2827">47:07</a>] If I do a deep dive with them, they&#8217;ll say they built a system, and I&#8217;ll ask, &#8220;Why did you build it?&#8221; and they&#8217;re like, &#8220;I don&#8217;t know, the VP told me to.&#8221; I&#8217;ll ask, &#8220;How&#8217;s the system used?&#8221; and they&#8217;re like, &#8220;I don&#8217;t really, I think ads use it.&#8221; I&#8217;m not sure. This is a caricature, but I think it&#8217;s very, very hard to do good engineering in that environment. You can do competent engineering, but the best engineering comes from a deep understanding of why. This is something we just drill into the team here at <a href="https://www.convex.dev/">Convex</a>; the team embodies this so strongly at <a href="https://www.convex.dev/">Convex</a>: everything exists for the why. Don&#8217;t build a fancy load balancer unless it&#8217;s needed.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2852">47:32</a>] I&#8217;m not sure. This is a caricature, but I think it&#8217;s very hard to do good engineering in that environment. You can do competent engineering, but the best engineering comes from a deep understanding of why. This is something we just drill into the team here at <a href="https://www.convex.dev/">Convex</a>. The team embodies this so strongly: everything exists for the why. Don&#8217;t build a fancy load balancer unless it&#8217;s needed.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2877">47:57</a>] Turns out we do need a fancy load balancer. We&#8217;re building it right now. But you should always start with, why are we doing this? What&#8217;s the point? I feel for people stuck in environments that are not like this. But you know what? Try to fight the system a little bit. I do see a lot of nihilism, a lot of defeatedness sometimes amongst junior engineers, a lot of this cynicism, like, what does it matter? Who cares? It&#8217;s just a big organization and nothing matters. But I think it does matter. If I think of the happiest times in my life, it&#8217;s been dedicating myself to a cause, trying really hard, and trying to do the right thing. I felt good when I went home, not trying to get promoted, just trying to do the right thing and then assuming I&#8217;m going to get promoted. If not, go somewhere else.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2906">48:26</a>] Who cares? It&#8217;s just a big organization and nothing matters. But I think it does matter. If I think of the happiest times in my life, it&#8217;s been dedicating myself to a cause and trying really hard to do the right thing. I felt good when I went home, not trying to get promoted, just trying to do the right thing and assuming I&#8217;m going to get promoted. If not, I&#8217;ll go somewhere else.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2933">48:53</a>] I know it does sound quaint when I&#8217;m saying this, but I think it&#8217;s possible to do this, and especially possible if you surround yourself with people like this. If someone is in a big company and they&#8217;re feeling frustrated by politics, look around and see if there&#8217;s a team of folks who just seem to want to do the right thing, just seem to want to do good stuff. I don&#8217;t think that&#8217;s selling out. I think that&#8217;s being true to yourself. That&#8217;s what real engineering is. Not trying to make a complicated fancy thing to get promoted. Just build the coolest thing that solves the problem.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2953">49:13</a>] I think that&#8217;s being true to yourself. That&#8217;s what real engineering is. Not trying to make a complicated, fancy thing to get promoted. Just build the coolest thing that solves the problem.</p><h3>49:23 &#8212; What technical teams should be focused on</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2963">49:23</a>] This really reminds me of something you had written. I thought it was really good writing. In the writing, there was this idea of system bias. You have this quote you&#8217;re writing. It says, here are some examples: the team is spending six months to improve performance by 10% when it was completely fine to begin with, or the team is trying desperately to force their tooling on clients who don&#8217;t need it.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=2989">49:49</a>] Or the team is riding their outdated system to the grave like the captain going down on the Titanic. I&#8217;ve definitely seen examples of all those types of things in industry. And so, yeah, I think it was in the context of your writing about what you should orient your team around. Not systems, but actually missions. And maybe that&#8217;s a way to fight system bias.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3020">50:20</a>] Yeah, I mean, one of my jobs at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> wasn&#8217;t the most fun job, but it might have been one of the most impactful jobs. It was shutting down projects. You know, looking around and being like, huh, that thing over there that has had 60 people working on it for two years doesn&#8217;t seem to make a lot of sense to me. And then I, you know, it wasn&#8217;t a hostile thing, but I&#8217;d go and chat with a team and I&#8217;d say, hey, what are you all doing?</p><p>Do you believe in what you&#8217;re doing? Does this make sense? And the team in private would say, I don&#8217;t really know. But inertia is so strong. You know, this whole desire to not get in trouble, to just keep doing what you were previously doing is so strong, and talented people can end up doing things that don&#8217;t make a lot of sense. One of the things I said in that article is it&#8217;s kind of a cheesy story, but when we started building the storage system at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, the team was called the Magic Pocket Team because that was a silly code name for the system we built.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3044">50:44</a>] Do you believe in what you&#8217;re doing? Does this make sense? The team, in private, would say, &#8220;I don&#8217;t really know.&#8221; But inertia is so strong. This whole desire to not get in trouble, to just keep doing what you were previously doing, is so strong, and talented people can end up doing things that don&#8217;t make a lot of sense. One of the things I said in that article is it&#8217;s kind of a cheesy story, but when we started building the storage system at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, the team was called the Magic Pocket Team because that was a silly code name for the system we built.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3077">51:17</a>] The team was oriented around building that system, but as soon as we shipped it, I renamed the team to the Storage team. That actually took a bit of work because you had to rename all the email addresses, the channels, the repos, and the things. It seems like a waste of time. But my argument to the team was that the responsibility of the storage team is not to advocate for Magic Pocket, the storage system. It&#8217;s to solve the needs of storage for the organization. Because who else in the company knows more about storage than the storage team? If there was a point in time where S3 was a better idea, it would make sense to move back. Or maybe there was a different kind of storage system that meant to use. Right. It&#8217;s the job of the storage team to advocate for moving back. Right. And so I&#8217;ve seen this before.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3106">51:46</a>] It&#8217;s to solve the needs of storage for the organization. Because who else in the company knows more about storage than the storage team? If there was a point in time where Amazon Simple Storage Service was a better idea, it would make sense to move back. Or maybe there was a different kind of storage system that meant to use. Right. It&#8217;s the job of the storage team to advocate for moving back. Right. And so I&#8217;ve seen this before.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3132">52:12</a>] You have a team called the Puppet team. People don&#8217;t really use Puppet that much anymore, but for the job manager, Puppet versus Chef. They&#8217;d be kind of advocating for their team&#8217;s thing when really a team should be oriented around what problem they solve. They should not care about the system that survives. Because if you are on the Magic Pocket team and someone says we should move back to S3, that&#8217;s pretty threatening to your identity and your career.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3164">52:44</a>] But if you&#8217;re on the storage team and it turns out it makes sense to move back, I don&#8217;t think that&#8217;s the case. But then that&#8217;s an exciting new product for you to own. I think it seems like such silly management philosophy, but I think it&#8217;s really, really important to orient a team and an identity around solving a problem and not owning and defending a system. You just see this in big companies. Inertia is so strong.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3187">53:07</a>] Inertia is so strong.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3190">53:10</a>] Yeah.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3190">53:10</a>] And you see people doing things they don&#8217;t believe in because that&#8217;s just what they do.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3194">53:14</a>] If inertia is so strong, how did you fight it and close down all those projects?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3200">53:20</a>] I think I got lucky insofar as I was there pretty early on and I worked hard enough. I was putting in probably 16-hour days at the start. I&#8217;m not advocating for that, but I was dedicating my life to the company, and I think it became pretty obvious to people that I cared. This is a guy over here that really wants to do the right thing and cares about the company.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3222">53:42</a>] At that point, you build up enough confidence, capital, that you feel comfortable saying things. I wasn&#8217;t afraid for my career; I was afraid of the wrong decisions getting made. So I felt psychologically comfortable making observations. Then at a certain point, there were a lot of engineers. I would mentor many of the staff plus engineers at the company.</p><p>People would sometimes get grumpy because they felt we weren&#8217;t doing the right thing or that something was inefficient. Every engineer listening to this has a story like this; they&#8217;re annoyed about some inefficiency at the company. My response was generally, do you think we should solve this problem right now? Because if we should, let me know the team to take some engineers off and the product to shut down, and I can redirect resources so we can solve this problem right now.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3256">54:16</a>] And people would sometimes get grumpy because they were like, &#8220;Oh, we&#8217;re not doing the right thing over here,&#8221; or &#8220;This is inefficient.&#8221; Every engineer listening to this has a story like this. They&#8217;re annoyed about some inefficiency at the company. My response was generally, &#8220;Do you think we should solve this problem right now? Because if we should, let me know the team to take some engineers off and the product to shut down, and I can redirect resources, and we could solve this problem right now.&#8221;</p><p>And they&#8217;d be like, &#8220;Oh, well, we shouldn&#8217;t shut down any other stuff.&#8221; I&#8217;m like, &#8220;Cool, well, we just have this many engineers right now. If there&#8217;s anything lower priority, let&#8217;s stop doing the lower priority thing and do the higher priority thing.&#8221; Oftentimes the answer was, &#8220;Oh no, nothing else is lower priority.&#8221; Then the answer is we just have to accept it. Just have to accept it, right?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3282">54:42</a>] And they&#8217;d be like, oh, well, we shouldn&#8217;t shut down any other stuff. I&#8217;m like, cool, well we just have this many engineers right now. If there&#8217;s anything higher or lower priority, let&#8217;s stop doing the lower priority thing and do the higher priority thing. This thing. Oftentimes the answer was, oh no, nothing else is lower priority. Then the answer is we just have to accept. Just have to accept, right?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3306">55:06</a>] There&#8217;s no point in being angry or upset that we&#8217;re not doing the right thing all the time. I think there&#8217;s a dimension to &#8220;you break it, you bought it.&#8221; I don&#8217;t think, when I said part of my job was shutting products down, it was just going around causing problems. Ideally, it&#8217;s about solving problems, like, &#8220;Oh, this product is not going in the right direction, let&#8217;s redirect it and do this alternative thing.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3328">55:28</a>] And so I think the thing that helped me, I guess, was having a sense of ownership that instead of complaining, I just wanted to go fix problems. That was part of the culture. I mean, when I started at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, the infrastructure team was, I don&#8217;t know, seven, eight, nine people. We&#8217;d have someone join from Google, for example, and they&#8217;d say, &#8220;Well, someone should go build.&#8221; I can&#8217;t do anything without this logging framework. I couldn&#8217;t possibly do anything with this logging framework. I&#8217;m like, &#8220;Well, we&#8217;re going okay without it.&#8221; And then someone needs to build this thing. Everyone would be like, &#8220;Well, who is someone?&#8221; Right? Because it&#8217;s just us. It&#8217;s just us. We build it or we don&#8217;t build it. They&#8217;d pretty quickly come to understand, &#8220;Oh wait, it&#8217;s just us.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3355">55:55</a>] I can&#8217;t do anything without this logging framework. Couldn&#8217;t possibly do anything with this logging framework. I&#8217;m like, well, we&#8217;re going okay without it. And then someone needs to build this thing. And everyone will be like, well, who is someone? Right? Because it&#8217;s just us. It&#8217;s just us. We build it or we don&#8217;t build it. And they&#8217;d pretty quickly come to understand, oh wait, it&#8217;s just us.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3377">56:17</a>] There&#8217;s no other idiots out there. We&#8217;re the idiots, right? So that was, I don&#8217;t know, I loved that time because it was just a time of accountability. Life gets easier and harder when you realize that everyone else is not an idiot. When you realize that everyone else is just dealing with their own stuff, right? I do not think someone will have good luck going around complaining about stuff and just saying this is a dumb idea and being negative.</p><p>I think people will have&#8212;everyone wants problems solved, though. So if you&#8217;re someone in an organization who is willing to put their head up and say, you know what, I think this thing over here is a bad idea, but here&#8217;s a different idea and I&#8217;m willing to own it and put the effort behind it, I think that&#8217;s a recipe for success.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3410">56:50</a>] I think people will have&#8212;everyone wants problems solved, though. So if you&#8217;re someone in an organization who is willing to put their head up and say, &#8220;You know what, I think this thing over here is a bad idea, but here&#8217;s a different idea, and I&#8217;m willing to own it and put the effort behind it,&#8221; I think that&#8217;s a recipe for success.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3431">57:11</a>] From that article, you had a great quote or a great question to think through this. It said you went to everyone or a bunch of people, and you would say, if we could be spending these resources working on any project at the company right now, would this still be the best use of time? I feel like it frames exactly what you just described really cleanly.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3452">57:32</a>] Yeah, I mean, I guess this is like maybe a trick for being a tech lead or a manager. Nothing&#8217;s a yes or no question. It&#8217;s a prioritization question. It&#8217;s not like, should we redesign the database? I don&#8217;t know, maybe, I guess. Is it the most important thing to do right now? No. Cool, let&#8217;s not do it. I think it&#8217;s much easier to have those conversations than to think about it as a yes or no.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3474">57:54</a>] I often hear, &#8220;Oh, my VP won&#8217;t let me do blah.&#8221; Okay, well, it could be that your VP is uninformed, but it&#8217;s probably not. It could be your VP has a different set of priorities. Maybe their VP knows that you need to ship these features, and if you don&#8217;t, then the company is going to struggle. I don&#8217;t know. But to really frame it around prioritization, the way I think about the career ladder for engineers is that some people&#8212;maybe it&#8217;s not common&#8212;think that becoming a senior engineer means you get better at programming.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3507">58:27</a>] But I don&#8217;t know, I think my programming abilities went down from level four onwards. I think once I got to level four, that was peak programmer for me. Then I probably went downhill from there, and I got wiser, whatever. But I think the real thing that happened is the scope that I cared about increased. At a certain point, you&#8217;re just thinking about what matters most at the company for the next five years.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3532">58:52</a>] And I don&#8217;t think a junior engineer in their first year on the job should try to do this because you probably don&#8217;t yet have the wisdom, insight, or knowledge to make a good assessment. I think it would be a mistake to try to come up with redirecting company strategy. But as you grow, the real key part about growing is that the IC 6, 7, and 8 engineers are not necessarily the best programmers. They&#8217;re just getting better at having a broad perspective in decisions within a company.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3562">59:22</a>] They&#8217;re just getting better at having a broad perspective in decisions within a company.</p><h3>1:00:15 &#8212; Doing the right thing vs promo hypothetical</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3615">1:00:15</a>] You mentioned this idea of doing the right thing and then, you know, the byproduct, you also get promoted. I think that&#8217;s the dream, you know, do the right thing, get promoted.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3625">1:00:25</a>] But in reality, oftentimes people would have to make the trade-off. Imagine a two-by-two matrix of doing the right thing, doing the wrong thing, getting promoted, and not getting promoted. Obviously, do the right thing, get promoted&#8212;great. Do the wrong thing, don&#8217;t get promoted&#8212;obviously bad. But I&#8217;m curious about the other two quadrants. Which one would you have picked when you were earlier in your career?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3650">1:00:50</a>] Let&#8217;s say I came to you, I said, hey, you can do the wrong thing, but you&#8217;re going to get promoted, or you can do the right thing and I guarantee you&#8217;re not going to get promoted. Which one?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3660">1:01:00</a>] Absolutely the second one. I&#8217;m going to say a really tacky thing, right? There are a set of engineers who at a certain point just make infinity money. The amount of money they can make, you know, going to work at <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> is a huge amount of money. Right? And so there&#8217;s a certain point where it just doesn&#8217;t matter anymore. The money is whatever. But if you really want to maximize long-term income, I don&#8217;t think you should try to do that.</p><p>If you get to a certain level of experience, skill, and seniority, you&#8217;ve made it; that&#8217;s it, you&#8217;re done, the money&#8217;s fine. I think there&#8217;s this desire amongst junior engineers early in their career to be kind of over-optimizing for promotion and salary, etc., as opposed to investing in themselves.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3691">1:01:31</a>] But if you wanted to maximize long-term income, if you get to a certain level of experience, skill, and seniority, you&#8217;ve made it. That&#8217;s it, you&#8217;re done; the money&#8217;s fine. I think there&#8217;s this desire among junior engineers early in their careers to be kind of over-optimizing for promotion, income, and salary, etc., as opposed to investing in themselves.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3719">1:01:59</a>] They probably don&#8217;t want to hear me say this because it&#8217;s easier for me as an old guy to say this, right? But I think, for the longest time at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, there was a time at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> where we just didn&#8217;t have our level system. I was tech leading the team and I was making the least amount of money on the team because I was at some point a technical manager. I saw everyone&#8217;s salary, so I was making the least money, but whatever, right?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3740">1:02:20</a>] It all worked out in the end, right? Now that was a happy story. But I do really think I benefited so much from working with the best people. If you have two choices&#8212;making 20% more money now or working with the best people in the world&#8212;just be around the best people because that&#8217;s going to set you on the ship to success. Now, not everyone wants to do that. Not everyone wants to go on that. There&#8217;s no shame in just wanting to have a regular job and just be chilling and getting paid. Fine, there&#8217;s nothing wrong with that. But if you do want to maximize your career growth, then the way to maximize that is to maximize your skills. Hopefully, you&#8217;re not so cynical to think that there is no correlation between talent and compensation.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3769">1:02:49</a>] Not everyone wants to go on that. There&#8217;s no shame in just wanting to have a regular job and just be chilling and getting paid. Fine, there&#8217;s nothing wrong with that. But if you do want to maximize your career growth, then the way to maximize that is to maximize your skills. Hopefully, you&#8217;re not so cynical to think that there is no correlation between talent and compensation. I don&#8217;t. If you think that&#8217;s the case, okay, I don&#8217;t know what to say, but there is a correlation between talent and compensation and growth. So my advice to people early in their career is to land at the best company with the best people doing the most important problems. Your life will be great because we are so lucky as engineers. Where else can you work solving problems for a job, solving puzzles for a job, and getting paid so well? Engineers get paid so well compared to most jobs.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3799">1:03:19</a>] I don&#8217;t. If you think that&#8217;s the case, okay, I don&#8217;t know what to say, but there is a correlation between talent, compensation, and growth. So my advice to people early in their career is to land at the best company with the best people doing the most important problems. Your life will be great because we are so lucky as engineers. Where else can you work solving problems for a job, like solving puzzles for a job and getting paid so well? Engineers get paid so well compared to most jobs. The real privilege that we get is to work on something cool. That&#8217;s what I advocate for.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3833">1:03:53</a>] The privilege, the real privilege that we get is to work on something cool. And that&#8217;s what I advocate for.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3842">1:04:02</a>] I love the idea you mentioned earlier. That&#8217;s kind of counter to the common opinion I often hear. The common opinion I hear is that someone working at a big company is in this machine, and it&#8217;s kind of defeatist. They&#8217;re going, &#8220;I ship this thing, but I hate this&#8221; or &#8220;I don&#8217;t believe in this.&#8221; I really liked your perspective on doing the right thing and basically giving a damn that someone outside of you is doing the right thing too and expanding your sense of ownership.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3875">1:04:35</a>] And I want to know what is your motivation for that? Because there are so many other people who are faced with the same inputs, and they come to a very different conclusion. So what motivates you to actually do the right thing when the machine doesn&#8217;t necessarily incentivize you to do so?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3891">1:04:51</a>] Yeah, well, I think the machine does incentivize you to do so, but I don&#8217;t think it&#8217;s visible immediately. I think people who do less job hopping really do grow the most. Because I would say very strongly, if you&#8217;re not in a job for three years, you&#8217;re not going to see whether your decisions were good. You can get more money as a junior engineer, but you cannot become a very talented senior engineer without being around for long enough to own the consequences of your decisions.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3923">1:05:23</a>] It&#8217;s like playing a basketball game and leaving before the game&#8217;s over. You&#8217;re just not learning. I do think there&#8217;s an actual structural incentive towards staying in a job now. If you&#8217;re in a bad job, leave the bad job. Right? But there&#8217;s a structural incentive to be able to stay long enough to have an impact. But also, I don&#8217;t know, I guess it&#8217;s easy for me to say. I joined the tech industry when it wasn&#8217;t a very lucrative field.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3950">1:05:50</a>] It wasn&#8217;t like it is now, but I just like it. The reason I left academia was because I wasn&#8217;t confident I was making the right decision. So what would happen, how I was in academia&#8212;I&#8217;m not saying everyone was like this, but you don&#8217;t know what to do. So you have to make up a problem to solve. You make up a problem, and then you make up a solution, hopefully a good one.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=3978">1:06:18</a>] And then you write a paper where you try to convince everyone that it was a really good idea. I hated it because I didn&#8217;t want to convince people. I just wanted to build it and see if it was good. I just wanted to build it and ship it and see; I didn&#8217;t want to play pretend and argue about whether it was good. That&#8217;s what makes me feel good. I don&#8217;t know. Obviously, if you&#8217;re an engineer who&#8217;s living paycheck to paycheck, you should maybe ignore what I&#8217;m saying.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4010">1:06:50</a>] I don&#8217;t want to come across as unempathetic to anyone who&#8217;s really struggling financially. But if you are not struggling financially, the best thing you can do for your quality of life is to enjoy what you do every day. I don&#8217;t know, you could have a fancier car or you could enjoy what you do every day. I love cars. I&#8217;m a motorhead. But I promise that enjoying what you do every day is going to have a much bigger impact on your life.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4042">1:07:22</a>] Personally, I don&#8217;t enjoy going to work and working on products that I believe in. Jamie, my co-founder, is the same way. We really get along well in this respect. We can only put up with doing things we actually care about and believe in. I think that&#8217;s the luxury&#8212;it&#8217;s more luxury than taking a first-class flight.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4069">1:07:49</a>] That&#8217;s more luxury than going to a three-million-star restaurant. The luxury is you don&#8217;t get up and have to do a terrible job. You get up and go to a place where you like your coworkers, and you solve cool problems, and you go home and feel proud of yourself. If you can construct your career that way&#8212;and it&#8217;s not easy, it requires very active effort&#8212;then you&#8217;re going to have a good life.</p><h3>1:08:13 &#8212; Why he dipped into management sometimes</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4093">1:08:13</a>] I saw in your career journey it said that you&#8217;re an occasional manager, and you seem like someone who really enjoys the technical aspects of things. So how did you decide to occasionally dip into management?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4107">1:08:27</a>] I feel like most people should not want to be managers, and I sometimes think it&#8217;s a bit of a red flag if someone wants to be a manager too much. Right. Because being a manager is a hard job. I think most people should go into management because they need to, because it&#8217;s necessary. What happened every time I went into management&#8212;I&#8217;m in a management role now as well&#8212;was because you have to.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4130">1:08:50</a>] Right. But someone needed to manage the team, and so I do genuinely enjoy accountability. I like responsibility; I like having weight on my shoulders, I suppose. I went into management several times, but as soon as I had the opportunity to get out, as soon as someone else would come in and manage the team, I would bounce out of that role back into engineering. Now, going into management is tremendously educational.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4162">1:09:22</a>] Being a manager really lets you see the world in a different way. You realize companies are more complicated than you thought. You realize the engineers are more complicated than you thought. You realize everyone on the team is going through something, and you realize, oh wow, now I understand why that thing happened. So being a manager is very educational. One thing I would say, I would strongly caution people against is going into management too early.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4184">1:09:44</a>] In their career. I see this happen a lot with well-intentioned people who want to push others into management as a means of career advancement. I think it&#8217;s doing people a disservice. You should let folks take their time in an organization, in a career journey. I don&#8217;t think you should be going into management under most circumstances in the first three years of your career.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4208">1:10:08</a>] I think you should ideally get to staff engineer before you do that. That&#8217;s not going to happen for everybody. But ideally, you take the time because if you go into management before you&#8217;re an excellent technician, it will limit your ability to influence strategy later in your career.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4224">1:10:24</a>] How does that play out?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4225">1:10:25</a>] It plays out with people who are stuck or have roles as people managers where they see their job as making a team happy or maybe coordinating a team or dealing with all the day-to-day challenges of people on a team. But they&#8217;re not organizational leaders. They&#8217;re not pushing the team towards excellence. They&#8217;re not asking, &#8220;How can we reframe what this team is about? How can we help influence technical strategy?&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4259">1:10:59</a>] And again, people management is a fine job. But I think the best companies are ones where all the managers are very technical, so they&#8217;re able to make sure the company&#8217;s doing well at a certain point. If you&#8217;re not a particularly technical manager, you&#8217;re not going to be able to evaluate the work of your team. You&#8217;re not going to know whether your team is even doing well. I guess maybe you might contribute to some of the stuff you were saying about cynical organizational attitudes where my manager doesn&#8217;t understand me.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4289">1:11:29</a>] Maybe your manager doesn&#8217;t understand you if they haven&#8217;t spent enough time developing technical skills and struggling.</p><h3>1:11:36 &#8212; Why you shouldn&#8217;t lead by example</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4296">1:11:36</a>] Yeah, there&#8217;s another piece that you wrote I thought was really good about leading by example, and you actually say it&#8217;s bad to lead by example. But there&#8217;s a quote in there that says modern tech workers can be an anti-authoritarian bunch at the best of times. Let&#8217;s say you&#8217;re a tech lead or manager. How do you strike that balance so you don&#8217;t lose credibility as a tech lead?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4323">1:12:03</a>] Yeah, there are command and control companies where the manager just tells people to do things, and they just do them. Those typically are not excellent companies, and not every company needs to be excellent. But if you want to have a company where your team is really innovating, where the engineers feel very personally responsible for making high-quality decisions, you have to have them believe in what you&#8217;re doing.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4350">1:12:30</a>] And so, at <a href="https://www.convex.dev/">Convex</a>, I&#8217;m the founder and CTO. I can&#8217;t really go to someone&#8217;s desk and say, &#8220;Do this thing.&#8221; Now, they might do it just because they like me. There&#8217;s a good chance they would do it, but not because I&#8217;m the boss. They&#8217;d be like, &#8220;Well, why?&#8221; And that&#8217;s because that&#8217;s our culture. We don&#8217;t just do things because we&#8217;re told. If I want someone to do something, I&#8217;ll spend time with the team talking about why we&#8217;re doing it, where the market&#8217;s trending, where the gap in our product, what, you know.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4377">1:12:57</a>] And then once we&#8217;ve really well articulated why something matters, people are just going to do it anyway. Sometimes you have to have hard conversations, but I think about it in terms of conflict in an organization. There&#8217;s this kind of hierarchy of the values you have and then why, what, and how. Engineers are very often debating the how. They&#8217;re often debating what algorithm we should use for this and whether we should use this container service or that container service. These are kind of like the implementation details. Most times when I see organizational conflict, it&#8217;s because well-intentioned people are debating as best they can about how to do something, but they don&#8217;t agree on why we&#8217;re doing it. If one team thinks the most important thing we can do right now is get more features out to expand our customer base, and the other team thinks the most important thing we can do is increase reliability because there&#8217;s a risk to the business.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4404">1:13:24</a>] And should we use this container service or that container service? These are kind of like the implementation details. Most times when I see organizational conflict, it&#8217;s because well-intentioned people are debating as best they can about how to do something, but they don&#8217;t agree on why we&#8217;re doing it. If one team thinks the most important thing we can do right now is get more features out to expand our customer base, and the other team thinks the most important thing we can do is increase reliability because there&#8217;s a risk to the business, they&#8217;re both very valid perspectives. However, it&#8217;s going to lead to them doing very different things. Within an organization, I&#8217;m a strong believer that everyone needs to have 100% alignment on why. The stuff we argue about or debate or talk about ad nauseam is why. Largely, I just trust the team to do the right thing. I think that&#8217;s the case for any tech lead. If you want to have credibility on a team, if you think that you can just tell someone to do something and they&#8217;ll listen to you because you&#8217;re the senior guy, no, they won&#8217;t.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4434">1:13:54</a>] They&#8217;re both very valid perspectives, but it&#8217;s going to lead to them doing very different things. Within an organization, I&#8217;m a strong believer that everyone needs to have 100% alignment on why. The stuff we argue about or debate or talk about ad nauseam is why. Then largely, I just trust the team to do the right thing. I think that&#8217;s the case for any tech lead. If you want to have credibility on a team, if you think that you can just tell someone to do something and they&#8217;ll listen to you because you&#8217;re the senior guy, no, they won&#8217;t. They won&#8217;t listen to you. If they listen to me, it&#8217;s because I&#8217;ve come in and explained it in a way that resonates with them. Again, one of my jobs at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> was to resolve situations. Someone would say, &#8220;Oh, that team&#8217;s an idiot; they won&#8217;t do blah.&#8221; Then I&#8217;d go talk to that team. That team was not an idiot. We talked through it, and they would decide to do the project. It wasn&#8217;t because they were scared of me.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4470">1:14:30</a>] If they listen to me, it&#8217;s because I&#8217;ve come in and explained it in a way that resonates with them. Again, one of my jobs at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> was to resolve situations. Someone would say, &#8220;Oh, that team&#8217;s an idiot, they won&#8217;t do blah.&#8221; Then I&#8217;d go talk to that team. That team was not an idiot. We talked through it, and they would decide to do the project. It wasn&#8217;t because they were scared of me.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4491">1:14:51</a>] I hope they weren&#8217;t. At least it was because I took the time to figure out what their motivations are. That&#8217;s something you really have to learn. If you&#8217;re a tech lead, you don&#8217;t really have that much authority over people, and you have to encourage them and get them to believe in what you&#8217;re doing. If you are leading through authority, you&#8217;re not going to have a culture where good ideas arise from within the organization.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4517">1:15:17</a>] People are just going to do what they&#8217;re told.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4519">1:15:19</a>] Influence without authority is huge. Even big companies like Meta technically don&#8217;t have titles; everyone&#8217;s just a software engineer. Is that also how <a href="https://www.convex.dev/">Convex</a> is run?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4532">1:15:32</a>] I mean, we&#8217;re a very flat organization. There are people in tech lead roles. I don&#8217;t completely buy that everyone&#8217;s a software engineer thing. It&#8217;s like, you know, <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a>. Everyone&#8217;s a member of technical stuff because it&#8217;s kind of like a wink-wink thing. You kind of know, it&#8217;s like no one&#8217;s mentioned the title, but you kind of know if the most senior person in the company comes to your desk, you&#8217;re probably going to notice. I think there&#8217;s no point in playing pretend. People know. But I will say that you should have an organizational culture where you don&#8217;t do something just because a senior person says something. Now, I do think that reputation matters. I certainly think that if a very experienced person at a company comes and tries to explain something to you, you probably should listen to them because they probably have some wisdom.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4563">1:16:03</a>] People know. But I will say that you should have an organizational culture where you don&#8217;t do something just because a senior person says something. Now, I do think that reputation matters. Like, I certainly think that if a very experienced person at a company comes and tries to explain something to you, you probably should listen to them because they probably have some wisdom. There might be something to be learned there. You should be open to it. But ultimately, even the most senior person, I don&#8217;t think should be leading through authority. They should use the benefit of their experience to be able to articulate, to win the hearts and minds. And that&#8217;s how your engineering culture is so important. And if you want a culture of ownership and innovation and drive and enthusiasm, and people trying to do the right thing, not get promoted, you have to have a culture where everyone believes in what they&#8217;re doing, and no one&#8217;s going to believe in what they&#8217;re doing.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4587">1:16:27</a>] There might be something to be learned there. You should be open to it. But ultimately, even the most senior person, I don&#8217;t think should be leading through authority. They should use the benefit of their experience to be able to articulate, to win the hearts and minds. And that&#8217;s how your engineering culture is so important. If you want a culture of ownership, innovation, drive, and enthusiasm, where people are trying to do the right thing and not just get promoted, you have to have a culture where everyone believes in what they&#8217;re doing, and no one&#8217;s going to believe in what they&#8217;re doing.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4620">1:17:00</a>] If the senior principal engineer comes to the desk and says, &#8220;I&#8217;m not going to tell you why, but you have to delete this database and do this other thing,&#8221; that&#8217;s not an empowering statement. What the empowering statement is, &#8220;Hey, let&#8217;s spend some time together to talk about where this product&#8217;s trending and how this is probably not going to work out and how there might be a different way we can solve this problem.&#8221;</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4640">1:17:20</a>] On the topic of tech leadership, in that article I mentioned, the title was &#8220;Don&#8217;t Lead by Example.&#8221; I think that might be confusing for people. Can you explain why you think you shouldn&#8217;t lead by example?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4653">1:17:33</a>] Yeah, leadership by example is a very passive thing to do. Engineers are passive people at the best of times. I think that&#8217;s stereotypically a little bit part of that personality. The very concrete example is when I first started becoming an engineering leader, I was trying to lead by example. I wanted to demonstrate the behaviors I wanted everyone else to have. Very specifically, with regards to on-call and people getting paged, I wanted people to have high ownership. I wanted people to jump on issues as soon as they happened. So I would do it. I&#8217;d be the first one to respond to a page. I would always be writing up the reports. I would always be jumping on all the bugs and stuff, really falling over myself to show how I wanted people to be. But from their perspective, all they see is that the lead is just doing all these jobs, and they don&#8217;t know. They&#8217;re like, &#8220;Oh, maybe that&#8217;s James&#8217;s job,&#8221; or &#8220;Maybe James knows how to do it, and I don&#8217;t know how to do it.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4685">1:18:05</a>] I want people to jump on issues as soon as they happen. I would do it. I&#8217;d be the first one to respond to a page. I would always be writing up the reports. I would always be jumping on all the bugs and stuff, really falling over myself to show how I want people to be. But from their perspective, all they see is that the lead is just doing all these jobs, and they don&#8217;t know. They&#8217;re like, &#8220;Oh, maybe that&#8217;s James&#8217;s job,&#8221; or &#8220;Maybe James knows how to do it, and I don&#8217;t know how to do it.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4717">1:18:37</a>] Turns out I didn&#8217;t know. I was just kind of figuring out, or maybe he likes doing those things. At a certain point, being a leader is about understanding human psychology. I don&#8217;t think you can just act a certain way in front of people and wait for them to copy you. I think you should act with integrity and values, and you should own the values of the team.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4741">1:19:01</a>] But you sometimes have to explain stuff; sometimes you have to tell people very, very specifically. There&#8217;s a trajectory almost every high-achieving leader goes through. They become a tech lead and care so much that they become a micromanager. They review every line of code and are involved in every decision, and at a certain point, they become the bottleneck for the team.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4766">1:19:26</a>] So this person is so overwhelmed; you might have gone through this yourself. They&#8217;re so overwhelmed, and they&#8217;re like, &#8220;Wait, my team doesn&#8217;t even seem busy right now. And I&#8217;m so busy. I&#8217;m reviewing all this code. I&#8217;m doing the strategy, what&#8217;s going on?&#8221; Then a manager will come along and say, &#8220;Hey, you&#8217;re micromanaging. You got to let your team have more ownership.&#8221; So the tech lead says, &#8220;Okay, sure, whatever.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4789">1:19:49</a>] I&#8217;ll let them own stuff. And they just take their hands off the wheel. Then the team falls apart because you can&#8217;t just stop doing the things you&#8217;re doing. You have to go and have a conversation with people. I see leadership as this kind of slider between oversight and accountability. When someone&#8217;s new to the team, when they&#8217;re very junior, they&#8217;re in a mode of oversight, like you&#8217;re checking their work. But at a certain point, you have to dial down the oversight and, very importantly, dial up the accountability. Instead of saying, I&#8217;m not going to look at what you&#8217;re doing anymore, you say, okay, cool, you&#8217;ve got this project. Let me know when it&#8217;s going to get done. Next Thursday.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4817">1:20:17</a>] Right. But at a certain point, you have to dial down the oversight and, very importantly, dial up the accountability. So instead of saying, &#8220;I&#8217;m not going to look at what you&#8217;re doing anymore,&#8221; you say, &#8220;Okay, cool, you&#8217;ve got this project. Let me know when it&#8217;s going to get done. Next Thursday.&#8221;</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4834">1:20:34</a>] Cool.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4835">1:20:35</a>] All right. What&#8217;s the plan? This is going to happen. How are you going to know it&#8217;s correct? Great, great, great. It&#8217;s on you. I expect you to do that. Let me know if there are any issues. The ownership relationship is explicitly on them. What you want to do is basically encourage ownership within teams. You can&#8217;t go from owning something to not owning something and expect that to develop.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4860">1:21:00</a>] You have to go have those conversations. You have to go and give people accountability. Now, what I have found, even though that can feel like an awkward conversation, most people genuinely like accountability. Most people like to own their work. Most people like to say, &#8220;Hey, this is on you. We&#8217;re all going down with the ship, right? I&#8217;m not going to leave you high and dry.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4883">1:21:23</a>] I&#8217;m the tech lead. I still take responsibility for this project. That&#8217;s how you can develop people within your team. A tech lead shouldn&#8217;t be about you as the boss and the team as the people who do the work. It&#8217;s about this kind of flow where you&#8217;re the more experienced person, typically maybe the more organized or more strategic person, and you&#8217;re working on developing your team members so they can take your job and then you can do something else.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4915">1:21:55</a>] That slider you mentioned, what if you give someone accountability and they blow it? Do they go back down the slider to where you start micromanaging?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4926">1:22:06</a>] You have to recognize it. Right. So this is growth. I mean, this is growth. Right. It doesn&#8217;t always work. I had another article about paper cuts or something. So we&#8217;d call these paper cuts. Some decisions don&#8217;t matter that much. You get them wrong, and maybe it sets you back a week. You get them wrong, and maybe something&#8217;s a bit suboptimal. These are great candidates to give people accountability for because they do it.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4953">1:22:33</a>] And if it doesn&#8217;t work, they get to experience it, and it will feel bad. I guess that&#8217;s good. It&#8217;s good to feel bad. You shouldn&#8217;t be demoralized, but it&#8217;s good to try something. It doesn&#8217;t work, and you&#8217;re like, oh, wow, that didn&#8217;t work. Then you feed that back into your little LLM in your brain, and you get better next time. Right? That&#8217;s a growth experience.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=4976">1:22:56</a>] I think what you don&#8217;t want to let people do is lose their arm. Paper cuts are fine, but I would not let someone, a junior person especially, design the replication system at <a href="https://www.convex.dev/">Convex</a>. You have to have safeguards. One of the arts as a CTO, a leader, or tech lead is figuring out the right level of altitude for how to know whether you can trust someone to take on ownership for something.</p><h3>1:23:23 &#8212; How to mentor Senior Staff+ engineers</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5003">1:23:23</a>] You mentioned that, because you were the most senior engineer at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>, at some point you had to mentor other very senior engineers. When I think about mentoring a junior engineer, it&#8217;s relatively straightforward patterns. But when I think about, let&#8217;s say, needing to mentor a senior staff engineer, maybe even a principal engineer, how do you mentor someone like that who&#8217;s already so polished and knows how to take ownership?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5031">1:23:51</a>] Yeah. How do you mentor someone so senior?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5033">1:23:53</a>] Yeah. I mean, there&#8217;s in two ways. One is there is still just a whole bunch of commonality. Everyone goes through the same problems. They don&#8217;t know how to deliver harsh feedback. There is a bunch of just standard stuff people are working on. But I do think that at a certain point everyone needs to become the best version of themselves. Sounds so cheesy, right? But there is not an archetype for what. According to me, there&#8217;s not an archetype for a senior principal engineer. You&#8217;ve got to be your own brand of engineer, right? So you might be. For me, I&#8217;m like the strategic kind of collaboration, simplicity, abstraction engineer. And my coding just fell off a cliff. You know, some engineers are like the deep science fiction, you know, the super hardcore science, you know, problem-solving engineer.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5058">1:24:18</a>] According to me, there&#8217;s not an archetype for a senior principal engineer. You&#8217;ve got to be your own brand of engineer, right? For me, I&#8217;m like the strategic kind of collaboration, simplicity, abstraction engineer. My coding just fell off a cliff. Some engineers are like the deep science fiction, the super hardcore science, problem-solving engineer.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5082">1:24:42</a>] So everyone&#8217;s going to find their own brand. Oftentimes, what I&#8217;ll be working on is a combination of two things. One, finding their strengths and helping them be more spiky&#8212;helping them really excel in the area that makes them special. At the same time, I&#8217;m almost always working on the personal side of being an engineer, like the organizational and personal understanding, people understanding the why. I don&#8217;t know, we&#8217;re so mathy as an industry. We think that somehow, even silly things, like for example, I have to tell so many senior engineers that you can&#8217;t change someone&#8217;s mind in a meeting.</p><p>You just can&#8217;t do it. You can disagree with someone, and they&#8217;re going to be mad or whatever, right? But to change someone&#8217;s mind involves going through a complex series of neurological processes that don&#8217;t happen live in front of ten people in a meeting. If you force someone to agree with you about them being wrong, they&#8217;re just going to go along with it, and their ego is going to get bruised and whatever.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5123">1:25:23</a>] You just can&#8217;t do it. You can disagree with someone, and they&#8217;re going to be mad or whatever, right? But to change someone&#8217;s mind involves going through a complex series of neurological processes that don&#8217;t happen live in front of ten people in a meeting. If you force someone to agree with you about them being wrong, they&#8217;re just going to go along with it, and their ego is going to get bruised.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5152">1:25:52</a>] So, one thing, you know, firstly, meetings generally aren&#8217;t for decision making. Most of the time when you identify a disagreement, I would point out the disagreement and why I don&#8217;t agree or where the problems are, provide enough information, and then let it sit. Let them go and reflect on it and come back and have a conversation a week later after they&#8217;ve thought about the why.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5176">1:26:16</a>] Right. And so this is a lot of very psychological, but it makes no sense to be like, well, they should be able to change their mind. I mean, well, they won&#8217;t, right? That team should do this. Well, they didn&#8217;t, right? You know, people should like my API? Well, they don&#8217;t, right? I think it&#8217;s like that whole, it&#8217;s almost like a capitalist attitude towards interpersonal behaviors.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5201">1:26:41</a>] Right. It&#8217;s like capitalism rewards success or impact. It doesn&#8217;t matter how well-intentioned you were. If you didn&#8217;t manage to convince, you could be the smartest person in the world. But if you can&#8217;t change someone&#8217;s mind, then you are kind of useless. Right? And so I think a lot of seniors still struggle with this. It doesn&#8217;t matter how smart you are, it doesn&#8217;t matter how right you are; are you effective?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5227">1:27:07</a>] And a lot of being effective at that level is about uncertainty, project management, and simplicity, but also kind of interpersonal dynamics because most hard problems happen with a team. Not always, but generally when you&#8217;re at that level, you&#8217;re going to have to have 10 or 20 people working with you to get something done. And that&#8217;s a whole different ball game.</p><h3>1:27:30 &#8212; Career advice for the AI era</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5250">1:27:30</a>] On the topic of career advice, the industry&#8217;s changed a lot in the last five years because of all these agentic tools. I wanted to know if you thought there is any career advice that has majorly changed in the last five years. Is there something that you used to say five years ago that you don&#8217;t say anymore, or vice versa?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5272">1:27:52</a>] No, I don&#8217;t know if I&#8217;ve changed my perspective much, but I think the industry has changed this perspective dramatically. I mean, let&#8217;s just be honest about it. There is incredible demand in <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> for senior engineers, and it is getting harder for junior engineers to succeed and grow for a variety of reasons. Why is there demand for senior engineers? Well, because large language models can&#8217;t do everything.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5297">1:28:17</a>] Well, the architecture and simplicity in design are still the domain of human beings, despite what you might hear on Twitter. Every company, including the labs, is desperately hiring senior engineers. But junior tasks are getting a little bit commoditized, and that worries me because I do think that learning&#8212;wisdom is kind of facts put into practice and then synthesized. People would argue that it&#8217;s easy to learn now because of ChatGPT; you can just go ask it a question about how two-phase commit works or what&#8217;s the difference between snapshot isolation and serializability, and it will give you probably a pretty good answer. But I think growth as an engineer does require wisdom, and wisdom only really happens when you synthesize it, in my opinion. I&#8217;m still very bullish on young people. We&#8217;re hiring junior engineers at <a href="https://www.convex.dev/">Convex</a>, and I&#8217;m very excited about it. I love working with junior engineers who are really hungry to grow.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5336">1:28:56</a>] Or what&#8217;s the difference between snapshot isolation and serializability? It will give you probably a pretty good answer. But I think growth as an engineer does require wisdom. Wisdom only really happens when you synthesize it, in my opinion. I think I&#8217;m still very bullish on young people. We&#8217;re hiring junior engineers at <a href="https://www.convex.dev/">Convex</a>. I&#8217;m very excited about it. I love working with junior engineers who are really hungry to grow.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5368">1:29:28</a>] But what I would say is train your mind. Do not listen to anyone who tells you that there is an advantage to having less knowledge. I&#8217;m not sure if you&#8217;ve seen people say these ludicrous things, like, oh, maybe in the future, not knowing engineering will be an advantage because you won&#8217;t have biases and you&#8217;ll just use Claude. I think these are ludicrous statements. Software engineering is an intellectual discipline that helps you think. The best software engineering is not about knowing syntax and it&#8217;s not about knowing an algorithm; it&#8217;s about being really good at conceptualizing problems, breaking them down into building blocks, and coming up with clean solutions. That requires experience. It&#8217;s like doing weights with your mind. Just like if you went to the gym and it never hurt.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5396">1:29:56</a>] The best software engineering is not about knowing syntax, and it&#8217;s not about knowing an algorithm. It&#8217;s about being really good at conceptualizing problems, breaking them down into building blocks, and coming up with clean solutions. That requires experience. It&#8217;s like doing weights with your mind. If you went to the gym and just picked up really light weights, you&#8217;re not growing. Also, if you went to the gym, picked up a heavy weight, and let the robot pick it up for you, you&#8217;re also not really growing. You have to do the reps. It&#8217;s tough, but find a way to stress your brain every day.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5417">1:30:17</a>] If you went to the gym and just picked up really light weights, you&#8217;re not growing. Also, if you went to the gym and picked up a heavy weight and thought, &#8220;Okay, cool, I think I can do it,&#8221; but let the robot pick up the weight for you, you&#8217;re also not really growing, right? You have to do the reps. It&#8217;s tough, but find a way to stress your brain every day. Find a way to.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5447">1:30:47</a>] To avoid. Now, obviously, agentic coding is here, right? Obviously, that&#8217;s. I could never tell someone to never use a coding agent, because that would be silly. But I can say that it is easy to fall into a passivity trap, where you&#8217;re just being passive about your learning. I do think you need to spend some time in the intellectual wilderness of not being able to solve a problem and struggling.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5476">1:31:16</a>] I would say if you&#8217;re running into a new problem, try to think of a solution yourself and then go check with a large language model. It&#8217;s going to be hard for you. It&#8217;s almost like, for example, I&#8217;m not very good at reading anymore. I got to be honest, I find it hard to sit down and read a book because my brain has been fried by the stimulation economy, right? I find that I go home and I&#8217;m tired, and I watch a YouTube video. I don&#8217;t tend to go home and read a novel. I should be better at that, but it takes discipline to do that. I would say similarly, it&#8217;s getting harder to solve a difficult problem without reaching for help. It&#8217;s getting harder and harder every day to be faced with a very difficult intellectual problem and not be like, well, I&#8217;ll just do a Google search or I&#8217;ll just ask Claude.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5505">1:31:45</a>] I find it hard. I go home and I&#8217;m tired, and I watch a YouTube video. I don&#8217;t tend to go home and read a novel. I should be better at that, but it takes discipline to do that. I would say similarly, it&#8217;s getting harder to solve a difficult problem without reaching for help. It&#8217;s getting harder every day to be faced with a very difficult intellectual problem and not think, &#8220;Well, I&#8217;ll just do a Google search or I&#8217;ll just ask Claude.&#8221;</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5530">1:32:10</a>] You&#8217;re whatever, I&#8217;m still doing the work. I don&#8217;t know. I would really encourage people to practice using your brain every day.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5539">1:32:19</a>] If I was just playing devil&#8217;s advocate or thinking from the junior engineer perspective, I might think, well, the proof that Claude can do this work today means that I don&#8217;t need to know it today or tomorrow or in the future. So why even build that skill in the first place? Also, these agentic tools have positive trajectories too.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5563">1:32:43</a>] Yes. So firstly, the agents are not particularly good at a lot of parts of engineering currently. Probably everyone should agree Claude is not good at designing distributed systems protocols right now, for example, or managing a 3 million line code base. There are parts of engineering where engineers are valuable. And like I said, you know this because the labs are all hiring engineers.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5590">1:33:10</a>] No matter what they say, they&#8217;re still hiring engineers, right? Desperately hiring engineers. Really, really aggressively hiring engineers. So engineers still have value. And maybe they won&#8217;t in the future, but I doubt it. I really do think there&#8217;s a role for human ingenuity in engineering. And so here&#8217;s the trade-off I would pose to people. I mean, there&#8217;s kind of three paths. One, get out, get out of engineering, go mow lawns and do whatever if you want.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5621">1:33:41</a>] That&#8217;s the most nihilistic attitude. I don&#8217;t believe in that. I really love engineering. I believe that there&#8217;s promising roles for human beings in engineering. So the second is to realize that there&#8217;s value right now for humans in engineering. And maybe one day AGI will be here. Let&#8217;s say in a year&#8217;s time AGI will be here. And there&#8217;s two people. One person just gave up on problem solving right now and they&#8217;re just feeding the machine. They&#8217;re running 17 coding agents in parallel, accepting everything Claude says, and they&#8217;ve given up.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5655">1:34:15</a>] Right. And one person said, you know what, I still want to have an active role in my learning. I want to understand what it&#8217;s doing. I want to think hard about problem solving, right? Played out one year, AGI arrives. Who is going to be better placed for the future, right? The person who&#8217;s been training their mind, right? You don&#8217;t&#8212;engineering is not a means to an end. It&#8217;s a mechanism for improving your mental processes.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5681">1:34:41</a>] It&#8217;s like, I still do whiteboard coding interviews, and by the way, <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> still does whiteboard coding interviews, just in case you were wondering. They don&#8217;t say, just use Claude. So I still do whiteboard coding interviews with candidates. Candidates are getting worse at coding. That&#8217;s absolutely true. But I don&#8217;t know any better vehicle for evaluating someone&#8217;s intellectual capacity for problem solving than seeing them solve an engineering problem.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5708">1:35:08</a>] It&#8217;s a great mechanism for doing so. If you were doing a math degree, a big part of doing advanced mathematics is proving theorems. By the way, the theorems are already all proven. Part of the exams is proving a theorem that has already been proven. You might say, well, what&#8217;s the point of proving that theorem? It&#8217;s already been done. The point is that the act of proving that theorem improves your mind, and then you can go and do more innovative stuff. I would say, sure, you may not. I still think that there is a very promising role for human beings in engineering. Maybe not in coding, but coding and engineering are very different things. Even if you don&#8217;t believe me, you can either give up now, but.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5731">1:35:31</a>] It&#8217;s already been done. Well, the point is that the act of proving that theorem improves your mind, right? And then you can go and do more innovative stuff. So I would say, sure, you may not. I still do think that there is a very, very promising role for human beings in engineering. Maybe not in coding, but coding and engineering are very different things, right? But even if you don&#8217;t believe me, you can either give up now, but.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5756">1:35:56</a>] Or you could just keep trying to improve your brain, and you&#8217;ll be better off anyway. Imagine if you&#8217;re wrong. Imagine if you take the defeatist view. Imagine if you say, &#8220;Well, AGI is coming next week, why bother doing anything?&#8221; And imagine you&#8217;re wrong. Oh my God, you just got off the ride. You get off the ride, and there&#8217;s so much cool stuff. I mean, this is a cool time for engineering, right? This is a really cool time. There&#8217;s so much cool stuff happening. And I hate this talk about how AI means humans don&#8217;t have a role anymore. AI means no one&#8217;s going to have a job. AI means we&#8217;re not going to do anything new. What I like is, hey, look at this. All this new cool stuff we can build. Look at the ways we can make people&#8217;s lives better.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5779">1:36:19</a>] This is a really cool time. There&#8217;s so much cool stuff happening. I hate this talk about how AI means humans don&#8217;t have a role anymore, that AI means no one&#8217;s going to have a job, or that AI means we&#8217;re going to do nothing else new. What I like is, hey, look at this&#8212;all this new cool stuff we can build. Look at the ways we can make people&#8217;s lives better.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5802">1:36:42</a>] And to be honest, I really wish my peers and my cohort would stop it with the real doomer, human elimination kind of narrative. I think there is a little bit of a psychological anchoring. As part of <a href="https://www.convex.dev/">Convex</a>, we didn&#8217;t start <a href="https://www.convex.dev/">Convex</a> to eliminate jobs; we started <a href="https://www.convex.dev/">Convex</a> to make it easy for people to build cool stuff.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5830">1:37:10</a>] I think the more we as an industry get behind, let&#8217;s do everything we can to make it possible for people to do more cool things. I think it&#8217;s a really exciting future we have ahead of us.</p><h3>1:37:21 &#8212; Why he started his own company</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5841">1:37:21</a>] I wanted to talk about what you&#8217;re working on now at <a href="https://www.convex.dev/">Convex</a> and why you quit <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> to build <a href="https://www.convex.dev/">Convex</a>.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5847">1:37:27</a>] <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> is a great place. At a certain point, I was there eight years. I don&#8217;t know if I outgrew the company, but I&#8217;d been the most senior engineer for a while. At a certain point, I wanted to grow and do my own stuff, so I left the company very amicably and started <a href="https://www.convex.dev/">Convex</a>. Why <a href="https://www.convex.dev/">Convex</a>? <a href="https://www.convex.dev/">Convex</a> started pre-agentic era because my observation was that the real differentiator in the success of a project, especially a large project, is the quality of the abstractions, the quality of the design, the quality of the architecture. Systems that thrive over time and are extensible are systems that are architected well. In my mind, the most difficult challenge in engineering is distributed state management. How do we store state reliably and reason about it, modifying it concurrently with other users? We designed a platform for application building based on our experiences building large-scale systems.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5888">1:38:08</a>] Systems that thrive over time and are extensible are systems that are architected well. In my mind, the most difficult challenge in engineering is distributed state management. How do we store state reliably and reason about it while modifying it concurrently with other users? We designed a platform for application building based on our experiences building large-scale systems.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5914">1:38:34</a>] So <a href="https://www.convex.dev/">Convex</a> is a transactional database where the transactions are written in TypeScript; they run as stored procedures. TypeScript stored procedures are serializable. There&#8217;s automatic reactivity, so what the client sees is a consistent view of what&#8217;s on the server. I would say <a href="https://www.convex.dev/">Convex</a> is a very designed platform because <a href="https://www.convex.dev/">Convex</a> is designed to be very composable and fit together well. We designed <a href="https://www.convex.dev/">Convex</a> for developers to use.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5939">1:38:59</a>] In particular, we wanted to make it so that application developers were able to build complex full-stack applications. That was the goal of <a href="https://www.convex.dev/">Convex</a>. Now all of a sudden, a coding agent development came along. That&#8217;s been really interesting for us because it turns out that what humans find hard is also what agents find hard. The coding agents are not particularly good with large code bases.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5964">1:39:24</a>] They&#8217;re not particularly good at reasoning about action at a distance, race conditions across services. They&#8217;re not particularly good at simple architectures over time. And these are the things that <a href="https://www.convex.dev/">Convex</a> gives you as a developer. So the idea now, and almost everyone using <a href="https://www.convex.dev/">Convex</a> is using <a href="https://www.convex.dev/">Convex</a> because they have the coding agent doing their front end. But they need a backend abstraction, a higher-level abstraction than something like Amazon Web Services or something like hosted PostgreSQL, which makes their problems go away.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=5996">1:39:56</a>] So this is a layer of abstraction on top of those types of primitives that makes it easier for an application developer.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6011">1:40:11</a>] It ties into a lot of the stuff I said earlier about making problems go away. Frankly, I watched the interview you did with <a href="https://en.wikipedia.org/wiki/Barbara_Liskov">Barbara Liskov</a>. Barbara was my advisor in grad school, and we worked a lot together on abstraction and the value in clean designs that minimize complexity. AWS is a fine tool. PostgreSQL is a fine tool, although none of the mainstream databases are that great, frankly.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6036">1:40:36</a>] But they&#8217;re fine tools, but they don&#8217;t make problems go away. Right. And so the idea of <a href="https://www.convex.dev/">Convex</a> is a higher level set of abstractions that you can use and not reason about state management, not reason about concurrency, not reason about scheduling, not reason about transactions, not reason about polling and data sync and type safety and all those things. So <a href="https://www.convex.dev/">Convex</a> is a, if you think about, you know, the history of engineering over time, the abstraction floor raises.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6065">1:41:05</a>] You know, when <a href="https://en.wikipedia.org/wiki/Barbara_Liskov">Barbara Liskov</a> was first starting, she was using punch cards. I don&#8217;t know if she mentioned it to you, but when she started as a programmer, she never heard the word &#8220;programmer&#8221; before. That was the first time she heard the word. And then you went from punch cards to having proper operating systems, and then languages like C, and then higher-level languages, and then you had cloud computing.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6091">1:41:31</a>] And over time, the abstraction floor raises, and you largely forget about what&#8217;s going on beneath the surfaces. Most people don&#8217;t think about how Amazon Simple Storage Service is implemented. I do, but that&#8217;s what I used to work on. Most people just use it, and it stores your data and gives it back. That&#8217;s great; that&#8217;s a successful abstraction. But I do strongly believe that the world is and has been overdue for a new abstraction one level up the stack.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6115">1:41:55</a>] And especially now that people are doing agent development, they don&#8217;t want to own a PostgreSQL instance. They don&#8217;t want to think about Kafka versus RabbitMQ. They don&#8217;t want to think about what set of tools to use. They want it just to work so they can focus on building their application.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6131">1:42:11</a>] When you talk about the abstraction, there&#8217;s obviously a lot of stuff going on behind the scenes in <a href="https://www.convex.dev/">Convex</a> and the technical side. What is it that <a href="https://www.convex.dev/">Convex</a> is building behind the scenes that you&#8217;re most excited about and why?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6147">1:42:27</a>] Basically, <a href="https://www.convex.dev/">Convex</a> is a new operating system in some respects. We have the primitives: queries, mutations, actions, subscriptions. What I think is cool is how we built this. We have our own distributed database that tracks read ranges and write ranges and does very efficient subscriptions over WebSockets, et cetera. So that&#8217;s the current operating system set of primitives.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6176">1:42:56</a>] But <a href="https://www.convex.dev/">Convex</a> is getting much larger workloads now and much more interesting workloads and more high-performance workloads. So we&#8217;re in the process of developing a slightly lower-level API for doing very efficient background processes, singletons, and APIs like fork, like operating system primitives. I&#8217;m pretty excited about launching these and how much faster it&#8217;s going to make various <a href="https://www.convex.dev/">Convex</a> components like the workflow system.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6208">1:43:28</a>] And to be honest, the thing I find exciting every day, challenging every day. I still find <a href="https://www.convex.dev/">Convex</a> very hard. I struggle every day. I don&#8217;t find my job easy. I feel confident at my job, but it&#8217;s not easy. Designing the new API for this is super hard. I can&#8217;t just go ask Claude; it&#8217;s not going to give a good answer, right? Because it&#8217;s innovation, it&#8217;s new ideas. I really enjoy it.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6236">1:43:56</a>] I find it stressful sometimes. I find it challenging and tiring, but I also find it exciting. I would encourage engineers to try to find this kind of stuff to work on, where it&#8217;s like you&#8217;re on that edge of, &#8220;I&#8217;m really liking this, but also it&#8217;s a bit tricky.&#8221;</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6256">1:44:16</a>] You mentioned fork, and in operating systems, I&#8217;m familiar with it. You take the existing process and kind of split it. What&#8217;s the idea of fork in a distributed system?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6267">1:44:27</a>] So <a href="https://www.convex.dev/">Convex</a> almost never has scale issues with regards to live traffic. Live traffic is typically bound by user-facing interactions, people clicking on stuff, running a website, acting on a website. Every now and then, someone will come to <a href="https://www.convex.dev/">Convex</a> and want to kick off a million background jobs to do some background processing. It&#8217;s a big workload. You can programmatically trigger huge workloads, right?</p><p>And so one of the things we have to scale is these background workloads, and a lot of them involve things like scheduling. There are a lot of workloads in <a href="https://www.convex.dev/">Convex</a> that would be very efficient if you had a background singleton process to perform things like aggregates. I&#8217;ll give a very silly example, right? If you&#8217;re building an election on <a href="https://www.convex.dev/">Convex</a>, a voting system, and every vote is a new row in the...</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6290">1:44:50</a>] And so one of the things we have to scale is kind of these background workloads, and a lot of them involve things like scheduling. There are a lot of workloads in <a href="https://www.convex.dev/">Convex</a> that would be very efficient if you had a background singleton process to perform things like aggregates. I&#8217;ll give a very silly example. Let&#8217;s say you&#8217;re building an election on <a href="https://www.convex.dev/">Convex</a>, a voting system, and every vote is a new row in the table. You want to show a tally of the votes. One way of doing this is having a bunch of background processes or cron jobs adding these things up. One way is doing a table scan, which is the obvious way to use PostgreSQL, which doesn&#8217;t scale. The other is to have a background job, which, if there are new votes, adds them all up and keeps a tally. If there are no new votes, it goes to sleep and waits on a condition variable to wake up again when there is a new job to perform.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6319">1:45:19</a>] In the table, you want to show a tally of the votes. One way of doing this is having a bunch of background processes or crons adding these things up. One way is doing a table scan, which is the obvious way to use PostgreSQL, which doesn&#8217;t scale. The other is to have a background job that, if there are new votes, adds them all up and keeps a tally. If there are no new votes, it goes to sleep and waits on a condition variable to wake up again when there is a new job to perform.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6345">1:45:45</a>] And so these are the kind of primitives that we&#8217;re working on right now. Most people won&#8217;t even know they exist, but they allow us to build this very high-performance primitive for scheduling, aggregates, background aggregations, et cetera. I&#8217;m pretty excited about the next generation of workloads we can support.</p><h3>1:46:05 &#8212; The most technically challenging work of his career</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6365">1:46:05</a>] As a result, when you reflect on your career, it sounds like you&#8217;ve done a lot of gnarly technical work across your PhD. <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> seemed like pretty intense systems work, and <a href="https://www.convex.dev/">Convex</a> is also doing a lot of cool stuff. When you look back on your career, what was the most technically stimulating work you&#8217;ve ever done? Why was it hard, and what did you learn from it?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6389">1:46:29</a>] There were certainly times in grad school where we were formally modeling consensus protocols and stuff. I&#8217;d be on the phone with <a href="https://en.wikipedia.org/wiki/Barbara_Liskov">Barbara Liskov</a> on weekends, talking through and trying to reason about this in our heads. That was pretty intellectually stimulating and fun. But I think the stuff I found most stimulating was working on a very large storage system with a team where things were going wrong.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6413">1:46:53</a>] Where the rubber hits the road, that&#8217;s where I find, and this is every day at <a href="https://www.convex.dev/">Convex</a>. The rubber hits the road. We have a compaction process that runs in the background, but it&#8217;s running into issues. We might have to redesign it using partitioning, etc. I feel most intellectually stimulated where there&#8217;s a really clear constraint in front of me.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6436">1:47:16</a>] And that to me is engineering. If I don&#8217;t actually know what the definition of engineering is, I&#8217;m just going to make it up in my mind. Engineering is science with constraints. It&#8217;s like, how do you solve problems in the presence of resource constraints? I&#8217;m not particularly interested in constraint-free environments. That&#8217;s art. I like craft and engineering. The more visceral and difficult the constraints, the more fun that is for me.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6467">1:47:47</a>] And I&#8217;ve been lucky enough to, whether it&#8217;s luck or intention, I don&#8217;t know, but I&#8217;ve always placed myself in those environments. You know, let&#8217;s go get on the hardest team and own the hardest problem and then put the effort in to survive.</p><h3>1:48:10 &#8212; How he got involved in Silicon Valley</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6490">1:48:10</a>] This question might be a little bit off topic, but I know you were a consultant for the TV show <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a>. I love that show, and I gotta hear, how&#8217;d you get involved with that?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6501">1:48:21</a>] Yeah, that was a lot of fun. A lot of folks might not know this: I had nothing to do with season one. Many TV shows don&#8217;t know whether they&#8217;re going to survive as a series. Mike Judge, who wrote <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a>, is also known for Beavis and Butt-Head and Office Space. He started his career as a software engineer at, I think, Lockheed or something. So he actually was a software engineer, which a lot of people don&#8217;t realize.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6528">1:48:48</a>] And so <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> was like a throwback to the kind of work he did. If anyone&#8217;s seen the movie Office Space, you would get this. That&#8217;s really a dystopian cubicle era tech industry film. They did season one of <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a>, and then it was very popular, so they got picked up. They had to figure out what to do for season two, but they didn&#8217;t know what to do because they had written a storyline that gets to the point where there&#8217;s a compression algorithm, and then what happens?</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6558">1:49:18</a>] And so they needed to find an expert on compression. I guess ostensibly that was me. I don&#8217;t know whether I was an expert on compression; I guess I was an expert on storage at least. They came to the office, and we just chatted. It was so much fun. I was pretty heavily involved in the show. A lot of it was storyline. So first, like, yeah, sure, what would you do with the compression algorithm?</p><p>What would you design for a storage system? Coming up with story ideas that were technically accurate, you&#8217;d be surprised to know how much they care about accuracy. A lot of people I know can&#8217;t watch that show because it&#8217;s just so creepily accurate. They find it so cringy. Partly why <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> can be so cringy is because it&#8217;s real. Those stories are almost.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6587">1:49:47</a>] What would you do? Could you design a storage system? Coming up with story ideas that are technically accurate, you&#8217;d be surprised to know how much they care about accuracy. A lot of people I know can&#8217;t watch that show because it&#8217;s just so creepily accurate. They find it so cringy. Partly why <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> can be so cringy is because it&#8217;s real. Those stories are almost...</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6614">1:50:14</a>] Almost. Maybe not everyone. So many of the stories in <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> are just real stories. They went to a bunch of companies, just farmed everyone for stories of crazy things that happened in the tech industry, and they wove them into the series. All the characters are based on real people and real archetypes. But they also cared very much about technical accuracy. So I would also do technical consulting.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6637">1:50:37</a>] And then they&#8217;d say stuff like, &#8220;Oh, we&#8217;re building a data center in our house. What should the rack look like? And what should the diagram on the wall look like?&#8221; Part of me wanted to say, &#8220;Oh, well, it doesn&#8217;t really matter. No one&#8217;s going to care.&#8221; But they were like, &#8220;No, no, no, it matters.&#8221; They really cared to get it right. So, yeah, I love the show, but it can be hard to watch just because of, oh my God, how real it can feel. Yeah.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6661">1:51:01</a>] How real it can feel. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6665">1:51:05</a>] Are there any Easter eggs where you look at that and go, that&#8217;s unusually accurate? Or, you know, that system diagram actually is very spot on.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6675">1:51:15</a>] I can&#8217;t. I don&#8217;t think I can even say them because there were stories of early <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a>. There were stories of having ideas ripped off by other companies and being tricked into having meetings with folks, only to find it may be the competitor&#8217;s team there to steal the information. A lot of those stories are real. And so there are people who watch <a href="https://en.wikipedia.org/wiki/Silicon_Valley_(TV_series)">Silicon Valley</a> and think, oh, wow, that was something I went through.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6703">1:51:43</a>] And probably it was because it was about that situation.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6708">1:51:48</a>] That&#8217;s such a cool experience. Did you get paid for that, or was it just for fun?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6712">1:51:52</a>] Yeah, that&#8217;s a complicated question. I got paid because I had to get paid; it was like a Hollywood union thing. I didn&#8217;t want to get paid because it made my visa more complicated. So I&#8217;ve got a green card now; I&#8217;m all good. But there was something where they had to pay me $400 anyway. I made a grand sum of $400 off that show.</p><h3>1:52:16 &#8212; Career regrets</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6736">1:52:16</a>] Looking back on your career, is there any regret that comes to mind that maybe other people can learn from?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6743">1:52:23</a>] Yeah, I mean, I think I underinvested in my personal life, to be honest. People don&#8217;t probably say that much because I think you can be all about growth, and sure, I could have grown more. I could have dropped out of grad school, say, three years in or four years in. I probably would have learned just as much. I could have taken a job at <a href="https://en.wikipedia.org/wiki/Dropbox">Dropbox</a> two or three years earlier and made a lot more money.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6768">1:52:48</a>] Everyone who has been in the industry long enough has been offered the chance to co-found several billion-dollar companies. Everyone has a story about the times they could have been a billionaire several times over. I don&#8217;t really regret those. Yes, it&#8217;s been a lot of sacrifice, to be blunt. I&#8217;ve been on call my whole career. I&#8217;ve carried a laptop almost every day. There are many dinners, parties, and events I&#8217;ve had to skip, and there are people in my personal life who have suffered as a result.</p><p>I really appreciate these people; they love and care about me, and they know that I have a passion for this, so they accept me for who I am. I think this is old person talk, but everyone has to decide how much they want to really drive their career because there&#8217;s a trade-off. Absolutely. I made a tremendous amount of sacrifices in my career, and I really prioritized building as probably the number one.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6801">1:53:21</a>] And I really appreciate these people. They love and care about me, and they know that I have a passion for this, so they accept me for who I am. I think this is old person talk, but yeah, I think everyone has to decide how much they want to really drive their career because there&#8217;s a trade-off. Absolutely. I made a tremendous amount of sacrifices in my career, and I really prioritized building as probably the number one.</p><p>I mean values first and then building. But if I went back in time, I would have had a good life, but I probably would have done more vacations and just had a bit more of a balanced life. I really do think. I see that with all the 996 stuff and this kind of performative photos of being in a bar with a laptop, and I&#8217;m like that&#8217;s not real.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6832">1:53:52</a>] I mean values first and then building. But if I went back in time, I would have had a good life, but I probably would have done more vacations and just had a bit more of a balanced life. I really do think&#8212;I mean I see that with all the 996 stuff and this kind of performative photos of being in a bar with a laptop and stuff, and I&#8217;m like that&#8217;s not real.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6856">1:54:16</a>] That&#8217;s not real. I mean, sure, I was working more than I do now back then, but I probably still work more than I did then. But I don&#8217;t do it as a checkbox; I do it because I really want to be doing stuff. I would caution people against that hustle culture. That&#8217;s not, firstly, like, you&#8217;re only young once; you should have fun. You know, I&#8217;ve got quite a few grays in here.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6886">1:54:46</a>] But that&#8217;s acting. I mean, focus on solving problems, work hard, be passionate. Yeah.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6895">1:54:55</a>] When I studied your career, there are mentions of you working 16 hours a week in various places. But you like being on call, you like fires, you like ownership; all those things are a recipe for working obscene hours.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6913">1:55:13</a>] Yeah. I go home and I&#8217;m tired, and I wind down by building stuff. I&#8217;m lucky enough to have a little workshop at home, so I go home and make things with my hands. You become an infra person; you become an engineer in all aspects of your life. But yeah, I would just say the cool thing is doing cool stuff that humans use.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6941">1:55:41</a>] The cool thing is not working long hours, the cool thing is not showing off that you were running a coding agent all night. Who cares? The cool stuff is enjoying doing important things.</p><h3>1:55:54 &#8212; Top technical book recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6954">1:55:54</a>] Do you have a best technical book recommendation for people?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6960">1:56:00</a>] I have to be honest, I haven&#8217;t read almost any technical book. I was in academia for a long time, so I read a lot of papers. Learning is awesome, reading is great, but balance it, right? Read something and then go put it into practice. Most of my career has been about doing. Sure, I did a PhD, so I guess that is like the academic side. But after that, most of my learning has been by doing because there&#8217;s nothing like being faced with a real problem to really develop as an engineer.</p><h3>1:56:36 &#8212; Younger self &amp; permanent underclass advice</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=6996">1:56:36</a>] Yeah, I think if I answered the question, I&#8217;d probably say the same. I think learning by doing is really where it matters most. And then the last question for you is, if you could go back to the beginning of your career and give yourself some advice, what would you say?</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7011">1:56:51</a>] I&#8217;d say, it&#8217;ll be okay. Don&#8217;t sweat the small stuff as much. It&#8217;s hard because my whole engineering brand is about caring about details, and I love design, so being obsessive is a little bit part of my DNA. But I think I would go back and say careers are long. I mean, there&#8217;s just any story anyone reads about a 22-year-old billionaire, blah, blah, blah. Just ignore that story. That&#8217;s not real. That&#8217;s not repeatable. That&#8217;s not normal. And it&#8217;s not that healthy. And it&#8217;s not that good for the people either. You probably won&#8217;t, ideally, max out your growth for 20 plus years as an engineer. I&#8217;m still learning all the time. I&#8217;ve been an engineer for several decades. So my advice would be don&#8217;t sweat it. There&#8217;s time to grow. And I do think that, again, this is a little bit of a modern phenomenon, but there is a feeling right now, oh my God, AGI has come and better max out my growth in the next three months.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7046">1:57:26</a>] That&#8217;s not real. That&#8217;s not repeatable. That&#8217;s not normal. And it&#8217;s not that healthy. And it&#8217;s not that good for the people either. You probably won&#8217;t, ideally, max out your growth for 20 plus years as an engineer. I&#8217;m still learning all the time. I&#8217;ve been an engineer for several decades. So my advice would be don&#8217;t sweat it. There&#8217;s time to grow. And I do think that, again, this is a little bit of a modern phenomenon, but there is a feeling right now, oh my God, AGI has come and better max out my growth in the next three months.</p><p>Well, guess what? You ain&#8217;t going to do it. It&#8217;s not going to happen. You can&#8217;t max your growth out in the next three months. It won&#8217;t happen. You don&#8217;t have to be running 17 agents at the same time. All you gotta do is orient your career around learning every day, getting better at what you&#8217;re doing, and trying to solve things in the most simple ways.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7084">1:58:04</a>] Well, guess what? You ain&#8217;t going to do it. It&#8217;s not going to happen. You can&#8217;t do it. You can&#8217;t max your growth out in the next three months. It won&#8217;t happen. You don&#8217;t have to be running 17 agents at the same time. All you gotta do is orient your career around learning every day, getting better at what you&#8217;re doing, and trying to solve things in the most simple ways.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7106">1:58:26</a>] I mean, on Twitter, I see this take on the idea of a permanent underclass, where if you don&#8217;t make it in time for AGI, then you&#8217;re going to be part of this permanent underclass.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7119">1:58:39</a>] Yeah. I mean, look, it&#8217;s a bit challenging economic times for a lot of folks. I don&#8217;t want to be unsympathetic to people who are having financial difficulties. At the same time, it&#8217;s just not an instructive attitude. There&#8217;s not much you can do with that information other than feel bad about yourself, and I don&#8217;t. Someone will argue back.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7140">1:59:00</a>] No, what you can do is get really good at using Claude. Well, guess what? It&#8217;s not very hard to use Claude. Sometimes people tell me, &#8220;Oh my God, I&#8217;m finding it hard to keep up with all the new models.&#8221; The new model drops, and I don&#8217;t even know what the new models are. I use them, but I forget the latest model because it doesn&#8217;t matter. Right? Like there was this thing called Ralph.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7163">1:59:23</a>] Right. I guess it&#8217;s still RALPH. It&#8217;s like a loop thing. I don&#8217;t really know what RALPH is, and I haven&#8217;t heard anyone mention RALPH in the past few weeks, but it was the biggest thing on Twitter for like a month. And it just doesn&#8217;t. This is noise somehow. Sometimes I feel like it&#8217;s like tech tabloids. It&#8217;s like people think that they&#8217;re learning by somehow knowing, like listening to what Jensen said today or like, oh my God, Boris said that Claude writes itself.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7193">1:59:53</a>] I don&#8217;t think people&#8212;look, that&#8217;s like&#8212;it&#8217;s just like reading about Beyonc&#233;, but you&#8217;re a nerd, and so you&#8217;re reading about Jensen, right? But it doesn&#8217;t matter, you know? Growing&#8212;like, you don&#8217;t have to know. It&#8217;s okay. The new coding agent could come out, and you could miss it. And then next year, if it turns out it&#8217;s the big one, you&#8217;ll just use it. It&#8217;s not hard. I haven&#8217;t seen any skill so far.</p><p>I guess there&#8217;s some skill, but it&#8217;s not a hard skill. If you&#8217;re good at engineering, you can figure out how to use Claude code, right? Or open code. So just watch out for tech tabloidism. It doesn&#8217;t matter. Just be building stuff. Just do real work.</p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7219">2:00:19</a>] I guess there&#8217;s some skill, but it&#8217;s not a hard skill. If you&#8217;re good at engineering, you can figure out how to use Claude code or open code. Just watch out for tech tabloidism. It doesn&#8217;t matter. Just be building stuff. Just do real work.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7237">2:00:37</a>] I love that mindset. Yeah, well, thank you for your time.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7242">2:00:42</a>] I&#8217;m on Twitter, too. I&#8217;m part of the thing, but just ignore me.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7251">2:00:51</a>] Oh, God. All right, well, thank you so much for your time, James. I really appreciate it. This was a lot of fun.</p><p><strong>James:</strong></p><p>[<a href="https://youtu.be/3XkmNSuHFmY?t=7256">2:00:56</a>] Ryan, it was great. Thank you.</p>]]></content:encoded></item><item><title><![CDATA[Takeaways from Bjarne Stroustrup (Creator of C++)]]></title><description><![CDATA[What made Bell Labs special, negative overhead abstraction, C++ garbage collection]]></description><link>https://www.developing.dev/p/takeaways-from-bjarne-stroustrup</link><guid isPermaLink="false">https://www.developing.dev/p/takeaways-from-bjarne-stroustrup</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 18 May 2026 13:15:40 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ee791b3e-12fd-4342-88c7-32bb5da4d1ac_2560x1440.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://en.wikipedia.org/wiki/Bjarne_Stroustrup">Bjarne Stroustrup</a> is the creator of the C++ programming language and a former researcher at Bell Labs. I got to interview him about his career and programming language design.<br><br>This conversation was a lot of fun. I brought my equipment to his office and we shot it right there. That&#8217;s why the background of his shot is almost entirely C++ books (many of which he wrote).</p><p>In addition to language design questions, I asked him for his personal anecdotes since I felt only he could tell those stories. Below are my top three takeaways to save you time.</p><p><em>You can find the full conversation on <a href="https://youtu.be/U46fJ2bJ-co">YouTube</a>, <a href="https://open.spotify.com/episode/52pEgoAglP6iPSNCxNeEvi?si=20L-_ZwnQyK_PzqaK0zNZQ">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>. The <a href="https://www.developing.dev/p/creator-of-c-bell-labs-negative-overhead">transcript is on Substack</a> if you prefer to skim.</em></p><div id="youtube2-U46fJ2bJ-co" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;U46fJ2bJ-co&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/U46fJ2bJ-co?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="callout-block" data-callout="true"><h3><strong>Brought to you by:</strong></h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!32cf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!32cf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 424w, https://substackcdn.com/image/fetch/$s_!32cf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 848w, https://substackcdn.com/image/fetch/$s_!32cf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 1272w, https://substackcdn.com/image/fetch/$s_!32cf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!32cf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png" width="450" height="81.28434065934066" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:263,&quot;width&quot;:1456,&quot;resizeWidth&quot;:450,&quot;bytes&quot;:250860,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.developing.dev/i/197216250?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!32cf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 424w, https://substackcdn.com/image/fetch/$s_!32cf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 848w, https://substackcdn.com/image/fetch/$s_!32cf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 1272w, https://substackcdn.com/image/fetch/$s_!32cf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccff88b0-f9b7-4398-9c75-fb855fc5bbe1_9065x1635.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><em><a href="https://www.developing.dev/p/cursor.com">Cursor3</a>: a unified workspace for building software with agents. I used it to build tools for the podcast and shared the process in this episode&#8217;s ad read.</em></p><p><em><a href="https://workos.com/">WorkOS</a>: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code</em></p></div><h1>Takeaways from the conversation:</h1><p><strong>1) How Bell Labs did research </strong>- When I asked Bjarne about how project selection went, he talked about two ways of doing research.</p><p>In one way, you have a well-designed project that is carefully chosen by management. Then the project is staffed with many researchers. This was not the Bell Labs way.</p><p>The other way was to hire the best people on the planet and don&#8217;t tell them what to do. Each year, the researchers would be asked to explain what they did in 9pt font or larger on a single page. If you can&#8217;t explain what you&#8217;re doing concisely, it probably isn&#8217;t that interesting.</p><p>If the work was interesting enough, then you could continue to work on it. Even though this was unusual, the results speak for themselves as much of the technology we take for granted today was invented at Bell Labs.</p><p><strong>2) Negative overhead abstraction - </strong>Coming into the conversation, I assumed that more abstraction meant you&#8217;d lose performance but Bjarne corrected my naive assumption.</p><p>A simple example is to look at C vs C++. Because C++ encodes more information in the language itself, the compiler has more it can use to optimize the final result. Therefore, the abstractions in C++ can be used to get better performance actually.</p><p><strong>3) There was a garbage collection API in C++</strong> - I had always assumed that where there was C++, there was manual memory management. I was wrong. </p><p>In 1995, Bjarne introduced a standard interface for garbage collection due to user requests. However, C++ developers tended to prefer manual resource management so the garbage collection API never got much use. Later, this garbage collection interface was removed from the standard. This is probably why you don&#8217;t hear too much about this.</p><div><hr></div><p>By the way, thanks for the input last week on where to place ads. It looks like most people prefer the ads grouped together instead of multiple shorter breaks. I would have picked the same personally. I factored in that feedback for this episode.</p><p>Looking at other podcasts, is it seems like some creators like to lump their ads together at the beginning instead of in the middle of the episode. I can see arguments for both sides:</p><ol><li><p>Beginning &#8594; get the interruptions out of the way so the episode can be interruption free</p></li><li><p>Middle &#8594; reading ads before delivering value could make people bounce </p></li></ol><p>Curious your thoughts on what you&#8217;d prefer:</p><div class="poll-embed" data-attrs="{&quot;id&quot;:512104}" data-component-name="PollToDOM"></div><p>Thanks for reading,<br>Ryan Peterman</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.developing.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to receive new posts and support my work:</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Creator of C++: Bell Labs, Negative Overhead Abstraction, Mistakes | Bjarne Stroustrup]]></title><description><![CDATA[Transcript & Audio]]></description><link>https://www.developing.dev/p/creator-of-c-bell-labs-negative-overhead</link><guid isPermaLink="false">https://www.developing.dev/p/creator-of-c-bell-labs-negative-overhead</guid><dc:creator><![CDATA[Ryan Peterman]]></dc:creator><pubDate>Mon, 18 May 2026 10:02:48 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/197420708/c47e40c8b1a8d8fc5fdc3ec277f2e12a.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/bjarnestroustrup/">Bjarne Stroustrup</a> is the creator of the C++ programming language and a former researcher at Bell Labs. We talked about what Bell Labs was like, programming language design, and interesting anecdotes from his experience.</p><p>Check out the episode wherever you get your podcasts: <a href="https://youtu.be/U46fJ2bJ-co">YouTube</a>, <a href="https://open.spotify.com/episode/52pEgoAglP6iPSNCxNeEvi?si=20L-_ZwnQyK_PzqaK0zNZQ">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835">Apple Podcasts</a>.</p><div id="youtube2-U46fJ2bJ-co" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;U46fJ2bJ-co&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/U46fJ2bJ-co?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h1>Timestamps</h1><p><a href="https://www.developing.dev/i/197420708/050-the-origin-of-c">0:50 - The origin of C++</a></p><p><a href="https://www.developing.dev/i/197420708/846-what-bell-labs-was-like">8:46 - What Bell Labs was like</a></p><p><a href="https://www.developing.dev/i/197420708/1724-dennis-ritchie">17:24 - Dennis Ritchie</a></p><p><a href="https://www.developing.dev/i/197420708/2400-when-to-build-a-programming-language">24:00 - When to build a programming language</a></p><p><a href="https://www.developing.dev/i/197420708/3159-bootstrapping-a-language">31:59 - Bootstrapping a language</a></p><p><a href="https://www.developing.dev/i/197420708/3358-c-is-not-object-oriented">33:58 - C++ is not object-oriented</a></p><p><a href="https://www.developing.dev/i/197420708/3732-discussing-type-systems">37:32 - Discussing type systems</a></p><p><a href="https://www.developing.dev/i/197420708/4620-memory-safety">46:20 - Memory safety</a></p><p><a href="https://www.developing.dev/i/197420708/4926-standards-committee-anecdotes">49:26 - Standards committee anecdotes</a></p><p><a href="https://www.developing.dev/i/197420708/10940-adding-automatic-garbage-collection-to-c">1:09:40 - Adding automatic garbage collection to C++</a></p><p><a href="https://www.developing.dev/i/197420708/11825-template-instantiation-is-turing-complete">1:18:25 - Template instantiation is Turing complete</a></p><p><a href="https://www.developing.dev/i/197420708/12157-abstraction-and-performance">1:21:57 - Abstraction and performance</a></p><p><a href="https://www.developing.dev/i/197420708/12851-ai-writing-code">1:28:51 - AI writing code</a></p><p><a href="https://www.developing.dev/i/197420708/13554-his-motivation">1:35:54 - His motivation</a></p><p><a href="https://www.developing.dev/i/197420708/13918-famous-quotes">1:39:18 - Famous quotes</a></p><p><a href="https://www.developing.dev/i/197420708/14648-reflecting-on-building-c">1:46:48 - Reflecting on building C++</a></p><p><a href="https://www.developing.dev/i/197420708/14912-top-c-book-recommendation">1:49:12 - Top C++ book recommendation</a></p><p><a href="https://www.developing.dev/i/197420708/15059-advice-for-his-younger-self">1:50:59 - Advice for his younger self</a></p><h1>Transcript</h1><h3>0:50 &#8212; The origin of C++</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=50">0:50</a>] What is the origin story behind <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=54">0:54</a>] Well, let&#8217;s start from the real beginning. I got a job at <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, which is a really great place in New Jersey. It&#8217;s not like that anymore, but at the time it was the best applied math, applied engineering place in the world. I looked around at the great people who were there; they built <a href="https://en.wikipedia.org/wiki/Unix">Unix</a>. They did a lot of the theory behind it, and I realized I had to do something important; otherwise, it didn&#8217;t belong.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=90">1:30</a>] So I decided I was going to build a distributed <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> because it was clear that computers were getting better, networking was getting better, so we needed one of those. If I had succeeded, we would have had <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> clusters ten years earlier or something like that. But of course, I couldn&#8217;t do it. That&#8217;s not a one-person job. The first thing I realized was there wasn&#8217;t a language in the world that could do what I needed.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=125">2:05</a>] It needed two things: low-level access to hardware, such as memory managers, process implementations, process schedulers, network drivers, and device drivers. Then it needed high-level features that indicate, well, there&#8217;s a module here on this computer and there&#8217;s a module there on that computer, and here&#8217;s the communication protocol they&#8217;re using, things like that. There are lots of languages that could do either, but none that could do both.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=163">2:43</a>] The obvious language for the low-level stuff was C because <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a> and <a href="https://en.wikipedia.org/wiki/Brian_Kernighan">Brian Kernighan</a> were down the hall, and I said distributed <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> because I was in the home where <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> was invented and still being built. For the high-level languages, there were a fair number, but they were all too slow and they couldn&#8217;t manipulate hardware. But I learned to use <a href="https://en.wikipedia.org/wiki/Simula">Simula</a>. I knew Kristen Nygaard and Ole-Johan Dahl, who invented object-oriented programming and <a href="https://en.wikipedia.org/wiki/Simula">Simula</a>.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=202">3:22</a>] And so I decided I had to merge these two. The way that was practical was to take the class concept from <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> and stick it into C so that it could run much faster and be used for systems programming. At the same time, I made the type system a bit more regular. User-defined types and classes were handled the same way as built-in types. That&#8217;s basically the start of what neither C nor <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> could do, which gets us to generic programming.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=243">4:03</a>] Eventually, many years later, I had to add overloading. We have always had overloading. You can add two integers, you can add two floating-point numbers, you can add a floating-point number to an integer with a plus. That&#8217;s a single name, right? So I had to generalize that to be able to have a unique set of rules for both built-in and user-defined types. That&#8217;s where it came from.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=276">4:36</a>] In one of the lectures that I saw that you gave, you talked about rewriting a simulator in <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a>. Is that the distributed <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> work?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=286">4:46</a>] No, that&#8217;s before that. I went to Cambridge, England, to get a PhD, and at some point, I decided I needed a simulator of software on a distributed system to do the PhD work on distributed systems. Of course, the idea of distributed <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> three or four years later came out of the same way of thinking. What I did was I wrote a really nice simulator in <a href="https://en.wikipedia.org/wiki/Simula">Simula</a>. <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> is very good at that. It&#8217;s misnamed because it was a general-purpose programming language, and its bad name didn&#8217;t help it at all.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=337">5:37</a>] But anyway, I wrote this simulator and I wrote little examples, test cases, etc. It all worked nicely. Then I tried the first real run full scale and I took the department&#8217;s mainframe and used it for a very significant time. Well, PhD students can&#8217;t do that. The chemists and the astrophysicists would never accept it. So I was kicked off the machine, and it was clear that <a href="https://en.wikipedia.org/wiki/Simula">Simula</a>.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=379">6:19</a>] I could write the program in <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> very well, but I couldn&#8217;t afford to run it. So I took the ideas and moved them to a little-used experimental computer, which was the <a href="https://en.wikipedia.org/wiki/CAP_computer">CAP computer</a>, which had hardware protection and capabilities and great stuff for hardware. It was somewhat unusual, so the astrophysicists couldn&#8217;t use it. They weren&#8217;t computer scientists as such, but I could. The only problem was I couldn&#8217;t run <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> there because <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> was never ported to that kind of machine; it was being used on mainframes and it was proprietary, and everything was wrong in the context of the <a href="https://en.wikipedia.org/wiki/CAP_computer">CAP computer</a>.</p><p>So I basically rewrote my simulator in <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a>. <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a> is a language that will make C look like a high-level language and has only one data type, the word. It was a very painful exercise, but once I&#8217;d done it, my program ran, I guesstimated about 50 times faster. I got my data and I got my PhD. So that was good. But I was convinced I would never again attempt a problem with tools that inadequate as I had tried on the mainframe in Cambridge.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=433">7:13</a>] So I basically rewrote my simulator in <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a>. <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a> is a language that will make C look like a high-level language and has only one data type, the word. It was a very painful exercise, but once I&#8217;d done it, my program ran, I guesstimated about 50 times faster. I got my data and I got my PhD. So that was good. But I was convinced I would never again attempt a problem with tools that inadequate as I had tried it on the mainframe in Cambridge.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=487">8:07</a>] And so I had a list of things that my ideal language should have, and, well, C didn&#8217;t have all of that, but it came closer than any other language that existed. C++ came out of there.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=505">8:25</a>] Yeah, in that lecture you said something like that. Writing that program in <a href="https://en.wikipedia.org/wiki/BCPL">BCPL</a> was so difficult you lost half your hair debugging.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=515">8:35</a>] That&#8217;s almost exactly true. And I lost the other half getting <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> going over the years. But anyway, it worked.</p><h3>8:46 &#8212; What Bell Labs was like</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=526">8:46</a>] You mentioned <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, and I think there&#8217;s a lot of curiosity about that topic just because it&#8217;s such a legendary place. When you graduated from your PhD and were thinking about where to work, what was <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> known as at that time?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=544">9:04</a>] <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> was the place to go if you wanted to do practical engineering at a large scale, at sort of world class. I think it was easily the best. We probably had twice as many computer scientists as MIT at the time, things like that. The Computer Science Research Center had great people, and some of them had come from Cambridge. One day, in my last year in Cambridge, one of the people from <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> came along to give a talk.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=594">9:54</a>] And the tradition in England and in the computer lab is, after a day&#8217;s work, you go to the pub and you chat with other people to see what has been going on. He said, &#8220;Well, when you need a job, give us a buzz.&#8221; So I did, and I flew over to New Jersey on my own tab, actually. My later boss, Sandy Fraser, a great guy working with networking, told me that I&#8217;d come at a wrong time. They didn&#8217;t have any jobs. This is not what you want to hear when you&#8217;ve just flown over the Atlantic. Anyway, the next day I gave a talk to a development group, not the research group, and then they changed their minds and took me up to the research group, and I worked there for the next couple of decades.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=634">10:34</a>] This is not what you want to hear when you&#8217;ve just flown over the Atlantic. Anyway, the next day I gave a talk to a development group, not the research group, and then they changed their minds and took me up to the research group. I worked there for the next couple of decades.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=653">10:53</a>] What was the interview process like?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=656">10:56</a>] You just talk to some people. I remember having a long chat with <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a>, for instance, and I talked to people doing networking mostly. There wasn&#8217;t an interview process as such. They hadn&#8217;t actually hired anybody new for five years. So no, they just did it by the seat of the pants.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=687">11:27</a>] So it&#8217;s kind of like the belief and credibility that other people say that you have. Like <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a> talked to you and he knew that you knew what you were talking about?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=698">11:38</a>] Sandy Fraser and such. They just talk to you, see what you know and don&#8217;t know. At the end, they go to the director and say, in this case, we&#8217;ve got a good guy; can you let us have him? I, of course, didn&#8217;t know anything about that. I wasn&#8217;t there. I was out talking to somebody in California, and I get a phone call from the director saying, would you like to come and work here a week later?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=730">12:10</a>] What gave you the conviction to fly on your own tab to go? And there wasn&#8217;t even a promise of a job yet?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=737">12:17</a>] No. Well, it was the best place in the world, right? Do you need any more? If it worked, it was the best. And if it didn&#8217;t work, so what? You can&#8217;t succeed at everything.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=755">12:35</a>] Today, there are more industries that have a stronger pull than at that time. Would it have been IBM or something? That would have been the best industry.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=763">12:43</a>] I talked to IBM; they weren&#8217;t as good as the <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> Computer Science Research Center. I was up at Yorktown Heights, and I talked to the researchers and the young researchers. I just didn&#8217;t think they were doing the right stuff and not in the right way. They were much more controlled and directed than the researchers at <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=792">13:12</a>] At a place like <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, how does project selection go?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=799">13:19</a>] I think still there are two philosophies about how to get good research. The one is that you have a well designed project chosen carefully by management and higher management, seriously funded and maybe you do put 20 or 30 people at the problem and you solve it and you have something great. The other philosophy is you hire the best people you can find and don&#8217;t tell them what to do.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=841">14:01</a>] My job was described as do something interesting in a year&#8217;s time, tell us what you did, and if we like it, we&#8217;ll extend. We will give you the same deal next year. By the way, the way you tell us is you write one sheet of paper using more than a nine-point font or more. Because if you can&#8217;t say what you did fairly briefly, you probably haven&#8217;t done something interesting enough, very unusual. So there they were actually worrying when they built <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> because it eventually involved five or seven people, and it was getting too big for that model of the world of individuals doing interesting things.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=903">15:03</a>] Very different. I would say that on average, this fairly anarchic organization did better than the well-organized one. Most of the things you&#8217;ve heard of from <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> came out of there. In another part of the building, they were doing hardware things. So fibers, as we use them today, came out of there. A lot of the wireless technology came out of there. The charge-coupled devices that are our cameras came out of there.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=943">15:43</a>] They tried to do videophones and couldn&#8217;t get it to work because hardware hadn&#8217;t grown up to it, but they were trying to do it. The system of cells for cell phones came out of not that building, but another building. For <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, it was just a great place. The computer science people tended to talk to people doing other things. I remember when I was doing simulations, I was helping somebody build a simulator for some networking stuff.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=978">16:18</a>] A lot of early <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> had to do with doing things like what happens when a network get overloaded, how do we handle the overload protocols. And in this particular case, they did a good job. And they called me back and they had a slightly bigger problem. They wanted to simulate the computer traffic of Manhattan. Even then, my answer was no, we don&#8217;t have the compute power to do that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1022">17:02</a>] Back then, though.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1024">17:04</a>] Now they probably could, but of course the computer traffic has become much more. So maybe they can&#8217;t; I don&#8217;t know. I don&#8217;t have the numbers now. Then they gave me the numbers. That was why I declined to help them because it was impossible.</p><h3>17:24 &#8212; Dennis Ritchie</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1044">17:24</a>] I saw somewhere when I was doing research that you said you had gotten lunch with <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a> once a week for like 16 years or something like that. And you know, he&#8217;s also a very legendary name. And I was curious, you know, if there&#8217;s anything that you, you learned from him or anything that impressed you about him that maybe influenced you or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1069">17:49</a>] He was a great guy and we talked about a lot of things. He never said anything rude or negative about <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. Actually, in his Hubble paper, he points to <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> as the obvious successor to C. So all of these C versus <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> language wars are ridiculous. They should never have happened, and they certainly didn&#8217;t happen because, well, I knew Dennis. We were not fighting. I still know <a href="https://en.wikipedia.org/wiki/Brian_Kernighan">Brian Kernighan</a>. I was talking to him this Friday.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1113">18:33</a>] We&#8217;re good friends. And yeah, language wars are silly. Dennis helped me design Const, for instance, for <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. It used to be called Read Only and Write Only, but the C guys couldn&#8217;t handle two words and they were too long. So we got what we got. But that&#8217;s one specific thing I remember Dennis being helpful with. He was a bit worried. Worried about overloading because you had to look at the declarations of functions before you knew what the meaning of a call was.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1161">19:21</a>] But that&#8217;s a very reasonable way of thinking. It just happens that it works. Anybody who writes see today are using my handiwork essentially all of the time, because the modern syntax for function definitions and function declarations and the call semantics came out of the early <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, early serial classes work of mine. So when people start ranting, they should remember that they&#8217;re actually using my handiwork every day.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1204">20:04</a>] You had written these almost like historical accountings of the history of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, like maybe three really long papers. I think for some conference I forgot the exact. And so in there, there was one anecdote about <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a>. It was about this concept where he proposed to the C Standard committee this idea of a fat pointer where it also has. It stores its size as well. You mentioned the C Committee didn&#8217;t approve.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1243">20:43</a>] Dennis did see, but he didn&#8217;t take part in the Standards Committee. I&#8217;ve even heard people from the C Standards Committee say, no, Dennis isn&#8217;t a C expert. He&#8217;s never come to meetings. Very strange attitude. But anyway, we knew the problem about buffer overflow and range errors. The obvious solution is to use what Dennis called a fat pointer, which is a pointer with its number of elements it points to attached to it.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1284">21:24</a>] But that&#8217;s two words. And that&#8217;s probably&#8212; no, that is why it wasn&#8217;t used in early <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, because then they had 48 kilobytes of memory. When I had this discussion with Dennis, I think we had a whole megabyte. When I started with <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, we had 256 kilobytes. I knew that we were going to get a megabyte. We were talking about it, and he called them fat pointers. Today in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, they&#8217;re called span.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1324">22:04</a>] And the span came out of my work with others on the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> core guidelines. We needed something like that. We couldn&#8217;t provide the degree of control and safety that we needed. So we built span, and it came into the standard a bit later. But some of these ideas are very old.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1350">22:30</a>] When you put together really impressive people, like the world&#8217;s greatest people, you look around and see how great the others are. Even though each individual is great, just the greatness of others can give people this feeling of imposter syndrome that&#8217;s not necessarily founded. Is that something that you ever felt or saw at <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1372">22:52</a>] Definitely, maybe I still have a bit of it, but certainly when I came to <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> and saw the names on the doors and I had read the papers, they had created the fields I like to work in. Yeah, I thought I have to up my game. I have to do something bigger and better than what I had imagined. Also, you talk to them and you learn things. I learned a lot over lunch where there were a bunch of them talking about what they&#8217;re doing and why they&#8217;re doing it.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1410">23:30</a>] There are some places that have that effect on people and have the density of talent. Cambridge University was one of those places. I learned a lot there. The doors were always open. It was almost a policy that we kept the doors open because of Howells.</p><h3>24:00 &#8212; When to build a programming language</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1440">24:00</a>] And when we talk about a programming language and just generally about language design, if I wanted to build a programming language today, what are all the pieces that I&#8217;d need to create to make a programming language?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1457">24:17</a>] Well, that&#8217;s a relatively easy question to answer. Everybody asks that question, and I think it&#8217;s the wrong question. What you need is a problem that needs a solution. A lot of people just want to build a language that is better at what they are doing now and what they particularly are doing. Most of the time, that can be done reasonably well with existing languages. If you build a very specialized language, that&#8217;s fine.</p><p>But if we&#8217;re talking about more general-purpose languages, you&#8217;re then building something that, when you want to work with somebody else, it&#8217;s not ideal for them. I&#8217;m sometimes asked why <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is so big and complicated. There are two reasons. One is history. I could not build the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> I wanted back in the 80s for a variety of reasons&#8212;partly technology, partly computers, and partly because I didn&#8217;t know enough and I had to learn.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1495">24:55</a>] But if we&#8217;re talking about more general-purpose languages, you&#8217;re then building something that, when you want to work with somebody else, it&#8217;s not ideal for them. I&#8217;m sometimes asked why <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is so big and complicated. There are two reasons. One is history. I could not build the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> I wanted back in the 80s for a variety of reasons&#8212;partly technology, partly computers, and partly because I didn&#8217;t know enough and I had to learn.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1537">25:37</a>] So you do the standard engineering thing, you do the best you can, then you see what works and what doesn&#8217;t work and try to fix the problems. Then you repeat. That&#8217;s how <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> grew. There are some leftover things that just get in people&#8217;s way now. You have spans, you very rarely use pointers, and you should certainly not use pointers for resource handling. That was already built in with your classes in &#8216;79, but people didn&#8217;t get it, so there&#8217;s teaching to do.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1579">26:19</a>] But anyway, once you figure out that you have a problem that requires a new language, then you start looking at what there is, and you have lots of help, lots of books about analysis and about code generation. You have frameworks like <a href="https://en.wikipedia.org/wiki/LLVM">LLVM</a> that most of the modern languages use to generate decent code. So most of the languages that compete with <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> do it by using a C infrastructure. It&#8217;s highly amusing.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1622">27:02</a>] But anyway, focus on the problem and don&#8217;t think you&#8217;re the only user. If you think you&#8217;re the only user, you build a special-purpose programming language, and that&#8217;s fine. Domain-specific languages are great when you find the right solution to the right problem, but identify the problem first. In my case, the problem was I needed high- and low-level facilities in the same language. Otherwise, I had to use two languages, and I had to have them communicate properly.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1661">27:41</a>] High-level languages at the time tended to use interfaces that took away performance. Quite often, they required garbage collection, which is not very good for device drivers, for instance, or for building garbage collectors. So yes, identify the problem and try to solve it.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1686">28:06</a>] Back when you were creating <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> and you had identified the problem, you went off to build <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and there were all these pieces, right? There&#8217;s the compiler, there&#8217;s a linker, and in the implementations of those, there&#8217;s a parser and all those things. When you were building the original thing, what was the most technically challenging part to implement?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1713">28:33</a>] I don&#8217;t think any part was particularly challenging. It was more those many parts, as you point out. One of the things I decided that caused trouble later was that I wasn&#8217;t going to touch the linker. It came simply because I asked around and realized people were using about 25 different linkers at <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, just in <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>. If I wanted to serve my obvious initial uses, I would have to write interfaces or modifications to 25 linkers.</p><p>And nobody wants you to touch their linker because if you make a mistake, everything breaks. So I decided on a rule: don&#8217;t mess with a linker. Later, people have messed with the linkers and made them better for <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. But that was after C became a major issue. The other thing was that there were many different optimizers. Every computer from different sources had a different optimizer, and again, I couldn&#8217;t write a dozen optimizers.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1754">29:14</a>] And nobody wants you to touch their linker because if you make a mistake, everything breaks. So I decided on a rule: don&#8217;t mess with a linker. Later, people have messed with the linkers and made them better for <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. But that was after C became a major issue. The other thing was that there were many different optimizers. Every computer from different sources had a different optimizer, and again, I couldn&#8217;t write a dozen optimizers.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1796">29:56</a>] I mean, I&#8217;m going to write a language here, right? And I can&#8217;t write. If I wanted to be an optimizer specialist for deck computers, I could become that. I have the background, I have the training, but that wasn&#8217;t what I wanted to do. I wanted to build first a distributed system, and then when my friends and colleagues started using <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> classes, I wanted to help them. They were doing things like network simulations, hardware layout, positioning of satellites, all kinds of interesting stuff.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1839">30:39</a>] So that was worth doing. I decided that actually there was a common interface to all these optimizers and code generators. It&#8217;s called C. So let&#8217;s use C as the assembler, and that worked nicely. C was very good at the low level, part of the reason I chose it. So let&#8217;s use it for the low level. I could have hidden C and recreated probably a better interface, but I decided that I&#8217;ll just use C and have C compatibility.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1881">31:21</a>] At the time, what I said was that we can have <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a>&#8216;s mistakes, which we know, and we can have my mistakes, which we don&#8217;t know yet, so we&#8217;ll take Dennis&#8217;s. That&#8217;s much more manageable and understandable, and I don&#8217;t have to teach people how to write a for loop and things like that. So that&#8217;s how C compatibility came in, partly as an implementation technique and partly to get into the culture and tool support and such.</p><h3>31:59 &#8212; Bootstrapping a language</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1919">31:59</a>] I saw somewhere in my research that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> was used to write some part of the language toolchain, and immediately I had this thought of a chicken and egg problem. How do you use this language to build something that it is using itself? How does that work?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=1943">32:23</a>] This is bootstrapping, and it was not an unusual thing. So I started with <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. I wrote a preprocessor that did some of the fundamental things in what became <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> classes, including fairly simple inheritance and overloading. Then, in that, I wrote a simple compiler for a subset of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. Now I can use operator overloading and overloading in general classes, so I can build a scope class that handles lookup and naming and things like that.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2002">33:22</a>] And then you work from there, just writing the next version in the previous version. You keep going, and after a couple of years, you have something that became known to the world as <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. I wrote a book about it, and a compiler came out in the world. But I didn&#8217;t invent this technique. This was known as bootstrapping. I think I was taught it as an undergrad that you could do things like that.</p><h3>33:58 &#8212; C++ is not object-oriented</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2038">33:58</a>] Most people look at <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> and think that&#8217;s an object-oriented language. I&#8217;ve heard you say multiple times that that&#8217;s not the case or that&#8217;s not your immediate thought. Why is that?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2050">34:10</a>] Yeah, I never called it an object-oriented programming language. If you look at the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> programming language, the first edition, the closest I come is to say some people call these techniques object-based. Actually, it&#8217;s more focused on classes; it&#8217;s type-oriented, class-oriented. It actually supports the techniques of object orientation very well. In particular, it follows <a href="https://en.wikipedia.org/wiki/Simula">Simula</a>&#8216;s model of defining types, defining classes, and defining class hierarchies to handle groups of related classes.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2100">35:00</a>] But that was never all it was. For instance, I do not want object-oriented complex numbers. I don&#8217;t want to say 2 dot something to get to some parts of numbers. I really want to say two plus Z, and I want that to end up being roughly the same as Z plus 2&#8212;no dots, no arrows. Math has developed a notation over the last 300 years or so. Descartes was, I think, the first one to use this notation, and it is very good.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2154">35:54</a>] I didn&#8217;t want everything to be object-oriented. Furthermore, I wanted things that did not require inheritance, that did not require runtime resolution, not to use it. So for arithmetic and for complex numbers and such, I wanted Fortran compatibility. I was rather keen on what&#8217;s called reuse in those days. But I saw it slightly different from a lot of researchers. A lot of researchers wanted to build a language, a system that allowed reuse.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2198">36:38</a>] I wanted to reuse things that existed. Fortran was there with some great software, C was there with some great system software, and actually helped with compilers and such. There was a <a href="https://en.wikipedia.org/wiki/Simula">Simula</a> that was used a fair bit too. So I wanted to reuse that, and I wanted to make sure that worked. And that meant I couldn&#8217;t go too far away from the hardware. I couldn&#8217;t build all of the things that were considered ideal, or even the things I would consider ideal.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2239">37:19</a>] This is the real world; this is the real set of problems you are attacking. You have to respect the constraints that come with that view of what you&#8217;re doing.</p><h3>37:32 &#8212; Discussing type systems</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2252">37:32</a>] At the time that you wrote <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, C was already there and it had a weaker type system than what <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> eventually had. Just give your thoughts on the trade-offs behind that. Why did you choose to make the typing system stronger in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2267">37:47</a>] Because we needed it. The weakness in the type system is one of the most obvious sources of errors, and it&#8217;s certainly one of the sources of endless testing and debugging. I hate debugging. I would much rather do design. You can&#8217;t really have either design or debugging. Some people claim they can, but they can&#8217;t. So I want to move the arrow towards more design that helps the debugging and makes fewer mistakes at runtime.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2313">38:33</a>] And the type system is one of them. And actually what you get and see today to a large extent is much stronger. Well, no, it is a much stronger type than it was in those days, and partly because of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. Also, there are things you can&#8217;t express unless you have a strong type system. I mentioned overloading before. Overloading is essential for generic programming. And if you want to write, say, a vector of T where T is a parameter type, you have to have overloading because you can only operate on T&#8217;s, providing all the T&#8217;s have the same interface for what you need.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2357">39:17</a>] So you need the type system to resolve those things, and that can be resolved at compile time. The compiler gets a bit more complicated, probably a bit slower, but you don&#8217;t need to do so much debugging. There was a large-scale experiment done at <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> in Chicago where they had some groups using C switching to <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and they wanted to know whether they were more or less productive. Some people claimed that the slower compilation slowed them down.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2404">40:04</a>] Somebody simply measured how much compile time was used before and after switching to <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and they found that the amount of compilation time on compute power was roughly identical. That is, <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> was slower by about a factor of two at that time. But the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> people compiled twice as often. This is just one experiment. The factor of two is just one experiment, but I wanted to move towards using more compile time resolution.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2451">40:51</a>] Still doing that.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2453">40:53</a>] I mean, for every language there&#8217;s this dichotomy of having it being statically typed versus dynamically typed. And <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is one of the most famous statically typed languages. Why did you choose a statically typed language?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2469">41:09</a>] Because of the problems I wanted to attack. What do you do when you get a runtime error in something like Smalltalk? You go into the debugger, and that makes a lot of sense. If there&#8217;s a programmer sitting at a screen getting the error, it doesn&#8217;t make any sense. If a telephone switch finds a runtime error, then you have to resolve it. Furthermore, you want performance, and you want small programs to fit into memories.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2508">41:48</a>] This is true even today because I think 99% of all computers are embedded systems, and they tend to be memory constrained. Again, if you do runtime resolution, you need to have enough information, enough data to do the runtime resolution. I wanted to fit into small memories&#8212;small meaning 120k, 250k, 1 megabyte, things like that. I think it&#8217;s still relevant for many systems.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2552">42:32</a>] You can build a camera like that; it can still have several megabytes of memory. But if you put in a lot of memory, it gets bigger, costs more, and the battery runs out quicker. So we don&#8217;t do that. Phones, cameras, and things like that are still memory constrained. Statically typed languages, languages optimized for memory consumption, are just better at that. That&#8217;s why we use it; we&#8217;re using it right now.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2590">43:10</a>] I suspect that the microphones have chips in them too. There&#8217;s a lot of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> in that world. The composition there is C and assembler.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2604">43:24</a>] And you mentioned the research that was done on the compile time on</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2612">43:32</a>] If you catch things earlier, you compile less often, but maybe it takes longer in this case. I could see a similar analogy where you catch errors way earlier if you have a statically typed language because the compiler is yelling at you before you put together that final thing. Whereas in a dynamically typed language, the errors may come later.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2612">43:32</a>] You catch things earlier, you compile less often, but maybe it takes longer in this case. I could see a similar analogy where you catch errors way earlier if you have a statically typed language because the compiler is yelling at you before you put together that final thing. Whereas in a dynamically typed language, the errors may come later.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2640">44:00</a>] I don&#8217;t know any solid research on that, but you can look at it. <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a> and Python are very popular, and they are runtime checked and they run much slower. I mean, raw Python runs something like 70 times slower than raw <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. The reason it&#8217;s viable is that a lot of key Python libraries are written in C or <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> to get the performance. So you get the performance by actually getting to the point that I was starting out with.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2683">44:43</a>] You need high-level stuff, and you need the thing that can manipulate hardware. Here they are using two languages, but still the same fundamental needs. It&#8217;s easier to try out things in a dynamically typed language because you don&#8217;t have to know enough about the language; you don&#8217;t have to know about type systems. Your average web developer or astrophysicist is not a computer scientist and doesn&#8217;t want to become one.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2720">45:20</a>] So there are advantages there. But the problem is that errors found by the type system in a statically typed language are discovered at runtime later. As systems grow, performance problems start. Furthermore, it becomes harder to write reliable software. You need much more unit testing, for instance, in a dynamic language because the compiler doesn&#8217;t do it for you. If you want things to guarantee to work&#8212;like the telephone switch mustn&#8217;t crash, your car mustn&#8217;t crash, your plane mustn&#8217;t crash&#8212;you want guarantees, and they&#8217;re harder to provide in a very flexible dynamic type system.</p><h3>46:20 &#8212; Memory safety</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2780">46:20</a>] One thing that I think <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is infamous for is memory safety issues or foot guns that exist there.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2791">46:31</a>] I&#8217;m so tired of that. I haven&#8217;t had those problems for years. Somebody did a study of the obvious problems with buffer overflows and people hacking in using that kind of stuff. Almost all of these cases involve people writing C-style code or in C. Herb Sutter has a talk with actual numbers, and they are quite significant. It&#8217;s sort of that kind of problem. More than 90% are from people who don&#8217;t write modern <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>.</p><p>They use raw pointers to pass things around without the number of elements. No fat pointers, no spans. You have them in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>; you can use them, you can use vectors. We have hardened libraries; everybody has hardened libraries that do the runtime checking. Apple has it, Google has it, Microsoft has it. It&#8217;s just not standard until now. C26 has a hardened option that is standard. The work I&#8217;m doing on profiles will give you a way of guaranteeing that you don&#8217;t do the stupid things.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2844">47:24</a>] They use raw pointers to pass things around without the number of elements. No fat pointers, no spans. You have them in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>; you can use them, you can use vectors. We have hardened libraries; everybody has hardened libraries that do the runtime checking. Apple has it, Google has it, Microsoft has it. It&#8217;s just not standard until now. C26 has a hardened option that is standard. The work I&#8217;m doing on profiles will give you a way of guaranteeing that you don&#8217;t do the stupid things.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2892">48:12</a>] So anyway, fundamentally, theoretically, the problem was solved many years ago, and people just do what they&#8217;ve always done and get the problems they&#8217;ve always had. That makes me sad. It&#8217;s one of the things that makes me work on coding guidelines, enforced profiles, and education.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2920">48:40</a>] I mean, education is one way to solve the problem. Is there a way to get the compiler to just prevent people from doing all those risky things? And is that enabled by default in modern <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> today?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2933">48:53</a>] No, but it should be. I&#8217;m proposing that for C++29, the simpler versions of that should have been in C++26. But there are still a lot of people, even in the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> Standards Committee, that are very devoted to their old code and their old ways of doing things. There are people who say you should only standardize what is common in industry, but when the bugs are common in industry, you should do something else.</p><h3>49:26 &#8212; Standards committee anecdotes</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2966">49:26</a>] The Standards Committee is a topic I want to talk about. It&#8217;s interesting. I mean, the language is now run by a democracy. One question I wanted to ask you is if it was a dictatorship, and you had full say over what language features would be included, would that make it harder to get by?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=2988">49:48</a>] First of all, it never was a dictatorship. I never had full control. Once you have some users, in my opinion, you gain some responsibility for making sure that they&#8217;re helped and their stuff works. You can&#8217;t keep breaking the language. That&#8217;s what academic language development does. They break to improve all the time, and then they can&#8217;t maintain a user population. I didn&#8217;t actually choose to have a standards committee.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3023">50:23</a>] I chose responsibility to the community. But one day, two guys came in representing IBM and HP, and I can&#8217;t remember if it was Sun or DEC that was the third thing they represented. But anyway, the biggest computer and software suppliers in the world at the time came into my office in SH9 and said, &#8220;Well, Bjarne, do you want to help us standardize <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> under <a href="https://en.wikipedia.org/wiki/International_Organization_for_Standardization">ISO</a> rules?&#8221; I said, &#8220;No, I can&#8217;t do that.&#8221;</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3070">51:10</a>] I&#8217;m still doing experiments. It&#8217;s still not complete. So they said, &#8220;No, Bjarne, you don&#8217;t get it. Our organizations cannot use a language that&#8217;s not standardized. They cannot use a language that&#8217;s owned by a corporation that we might compete with. And we do sometimes, okay, we trust you, of course, but not your employer. We compete with them sometimes. And you can get run over by a boss. No, no, no.&#8221;</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3109">51:49</a>] We need a standard, and we need a standards committee. So this goes on for about an hour. They twist my arm. Ow, ow, ow. In the end, they said, okay, I will standardize <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> under ANSI rules just like you suggest. The computer community needs that. By the way, what&#8217;s the ANSI rule for standardization? They told me, and we started a year later. But this was the way it came about.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3145">52:25</a>] Some very important organizations wanted that standardization. <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> was on track to get standardized, and AT&amp;T, being primarily a user of software, was also in favor of standardization. I found out, and so they supported it. The documentation I had written was based on it. Actually, I rewrote the documentation that became the ARM, the Annotated <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> Standards Manual, which provided the definition, the manual of the language, and for every feature, some rationale and some way it could be implemented or was implemented.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3196">53:16</a>] And that became the foundation document for the standardization.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3201">53:21</a>] When they were strong-arming you, what if you had just said no? What would have happened?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3206">53:26</a>] Well, I think <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> would have faded into becoming an academic cube language that was loved by some small community, and it would have disappeared out of the mainstream of computing. There are people who say this stronger than I do. They say that C is spread and its usage is that it has a standard; it&#8217;s not owned by a corporation. It is one of the things that sometimes blocks the wannabe <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> killers.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3257">54:17</a>] I remember the ads for Java and people standing up saying, &#8220;We&#8217;ll kill, absolutely kill <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> in two years.&#8221; I thought that was rude. Anyway, we have 10, 12 times more <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> developers today than we had when they said it, so it didn&#8217;t work.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3283">54:43</a>] How is it that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>&#8212;not exactly that it&#8217;s a war&#8212;but just if we looked at adoption, clearly <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> gained a lot more adoption than Java. Yet I know Java had the backing of a big company that was putting a lot of marketing dollars in, and C was kind of, I think you&#8217;ve said that it had almost zero, next to zero marketing done for it.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3313">55:13</a>] Next to zero was $5,000 to be used over three years. Sun used much more money on advertising and marketing Java than was ever used in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> development. To this day, the standards committee has a problem. It has no funding, and that means it&#8217;s hard to do extra experiments; it&#8217;s hard to deploy things. Other language communities keep sort of stealing <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> compiler and tool developers because they&#8217;re good.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3359">55:59</a>] But it makes it hard to predict how fast we can implement things today. Last I checked, the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> standards committee had 527 members. We work on consensus because if you don&#8217;t have consensus, then you get dialects. We don&#8217;t want to have a feature that&#8217;s voted in say 60 to 40 or even worse, 52 to 48. No percent. We don&#8217;t do that. And that&#8217;s painful and tedious and good.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3405">56:45</a>] When you say consensus, that 100% need to approve is not necessary. We don&#8217;t need unanimity. We need a massive majority. And basically, I would like to see 90%, and we often do 80%. I start to worry.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3407">56:47</a>] To approve 100% is not necessary. We don&#8217;t need unanimity. We need a massive majority. And basically, I would like to see 90%, and we often do 80%. I start to worry.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3428">57:08</a>] What&#8217;s the lower bound that&#8217;s coded into the rules?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3430">57:10</a>] There&#8217;s no lower bound coded into the rules. The rule says that the convener of the <a href="https://en.wikipedia.org/wiki/International_Organization_for_Standardization">ISO</a> committee determines what is consensus. Pure numbers don&#8217;t say it. Could you imagine you had a vote, 95% versus 5%? But the implementers of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> compilers and standard libraries from Google, Apple, Microsoft, and others were all in the 5%. Is that consensus? I can reassure you that no convener would call that consensus.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3482">58:02</a>] And that makes sense intuitively. I kind of wonder, with democratic decisions, there need to be objective rules. So what if the convener made the wrong decision?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3493">58:13</a>] It happens. But you can&#8217;t just have numeric rules. Not everybody cares for the whole language. Not everybody understands what&#8217;s going on. You can vote at your third meeting. So you might have somebody with a vote that has eight months of experience with the standardization and doesn&#8217;t understand standardization and knows only what they know from their development organization that they have been part of, which might be a small one.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3535">58:55</a>] You need some judgment, and you hope that the convener has that judgment. The convener always asks the national representatives. I mean, the other way of getting a consensus is that you have a massive consensus, but you have ten countries where the representatives didn&#8217;t agree. That&#8217;s not consensus. And even when there looks, if there is a look, if everybody is for it, if it&#8217;s massive and all of that, there&#8217;s not a problem.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3575">59:35</a>] But if there&#8217;s a problem, the convener asks the national body heads, he asks the implementers, sort of key people that are necessary for getting the voted change into real use. Sometimes educators also before they make that decision. Not everybody weighs equally once there&#8217;s a disagreement.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3705">1:01:45</a>] I saw in some of your writing you said one of the most negatively received ideas you&#8217;d ever presented was Auto. I know Auto eventually made its way in there, but what&#8217;s the story behind why it was so negatively received?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3722">1:02:02</a>] At that time, it was just unusual. People thought it was weakening the type system. It also opened the door to fairly general generic programming that is not heavily syntax-based. Auto is the beginning of concepts, which is the ability to put constraints on generic code. Auto is just the simplest constraint; it must be a type as opposed to seven. Maybe I didn&#8217;t explain this well enough, and there&#8217;s a variety of backgrounds in the committee. Maybe they didn&#8217;t know languages of the generic types, such as ML and Haskell.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3784">1:03:04</a>] So it was a horrible ache. Anyway, we still got it because we needed something like that, but it wasn&#8217;t enough. I have looked at industrial software and problems with the overuse of Auto. You should only use Auto when you have an idea about what is needed there. It&#8217;s good in generic code where you go and eventually check that the type is correct, that Auto has resolved to something that supports the operations that you&#8217;re going to do on it.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3830">1:03:50</a>] And that&#8217;s what concepts formalize. But it was always checked at the end, and I noticed a group of people that was overusing auto, actually in a framework for networking. They were saying that they were being slowed down, not so much with bugs, but they had to look up the functions being called to see what that auto could possibly bind to. And then they had to put in comments that said what the auto was meant.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3871">1:04:31</a>] So you have auto, and the comment says must be an input channel. Now you simply define input channel, and then instead of saying auto, you say input channel auto. Fine. That&#8217;s what the system is. My design of that simply said that auto is the simplest concept, and you should simply only set input channel <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> equals blah blah blah. But anyway, the committee wanted an indicator that this was going on.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3910">1:05:10</a>] Oh well, I saw another anecdote about the standards committee being heated at times. You mentioned there&#8217;s this thing about shuttle diplomacy between two corners of the room. I think it was IBM and <a href="https://en.wikipedia.org/wiki/Intel">Intel</a>. They both needed different support. What&#8217;s the story behind that?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3930">1:05:30</a>] I was actually talking to Brian McKnight, who was the IBM representative at the time. We were discussing some of the things that were happening then, and it&#8217;s still remembered. So basically, IBM was doing the... What&#8217;s that architecture called?</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3954">1:05:54</a>] PowerPC. And <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> was doing well. They had different models of the underlying hardware, especially coordination with the caches and things like that. The guy representing <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> wasn&#8217;t actually an <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> guy. Anyway, also a good guy. I used some of his slides in my presentations. I knew these guys, but they were totally deadlocked.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3955">1:05:55</a>] PowerPC and the x86. The <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> was doing well. They have different models of the underlying hardware, especially coordination with the caches and things like that. The guy representing <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> wasn&#8217;t actually an <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> guy. Anyway, also a good guy. I used some of his slides in my presentations. I knew these guys, but they were totally deadlocked.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=3996">1:06:36</a>] I mean, these are massive organizations with massive amounts of code out there. Basically, some things could be done. But the IBM guy, Brian, said that they had a lot of software, mostly in the lowest level, even down in the microcode, that was relying on the way they had done it. The people who had done it had left the company a long time ago, and they just couldn&#8217;t rewrite all of that code and get it right, even if the <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> guys were right.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4058">1:07:38</a>] Okay. And the <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> guys had similar arguments and similar points. This is better. We are using it. I was shuffling; they were in different corners of a large room. So I go up to the <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> guy, the guy representing <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> here, and what&#8217;s the problem? Tell me about it. Explain it. I go down, explain. He&#8217;s saying this. They say, well, there&#8217;s this, this, this. I go back again and I spent a couple of hours literally doing shuttle diplomacy, walking from one corner of the room to the other.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4098">1:08:18</a>] And we reached an agreement. That&#8217;s in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. A couple of years later, they both agreed that they were now using a combination of what they had before and what the other guys brought in. So actually, the result was improvement, cross-pollination.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4125">1:08:45</a>] That&#8217;s funny. Why did you have to do shuttle diplomacy? Why not just a group conversation?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4130">1:08:50</a>] Because they&#8217;ve been trying that for days and probably for meetings before that, and it didn&#8217;t work. So I guess I was just translating and asking questions. I mean, those were experts. I don&#8217;t consider myself an expert at that level. I&#8217;ve done hardware, I&#8217;ve done microcode, so I&#8217;m not an amateur, but these guys are really good. So the IBM guy is now the guy doing most of the synchronization on the Linux.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4170">1:09:30</a>] We&#8217;re still using his stuff today. He has a <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> version of it, so he can get it into the bottom of the Linux kernel.</p><h3>1:09:40 &#8212; Adding automatic garbage collection to C++</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4180">1:09:40</a>] When I was reading your writing in these papers, there&#8217;s this part where it seems like in 1995 you had this idea to introduce some form of automatic garbage collection into <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. That kind of surprised me because when I think about <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, one of my immediate thoughts is, no garbage collector. We&#8217;re going to manage the memory ourselves. How would that even work?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4210">1:10:10</a>] There&#8217;s two things there. One, I wanted to automate resource management in general, not just garbage, not just memory. For that, you have constructors, destructors, and the techniques later known as RAII&#8212;Resource Acquisition Is Initialization&#8212;which is probably my worst naming ever. But I was busy at the time. So in the standards committee, those people insisted that we needed to be able to do garbage collection.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4250">1:10:50</a>] And there were garbage collectors out there. Oh, there was Hans Boehm here. He was the one representing the <a href="https://en.wikipedia.org/wiki/Intel">Intel</a> model of stuff. He has a conservative garbage collector still used today. We thought we needed an interface so that it could be standardized how you used such a garbage collector. Basically, I was listening to the users, expert users, and they thought it was necessary. I thought that support for memory management and resource management was important.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4292">1:11:32</a>] I&#8217;ve thought that from the beginning. I didn&#8217;t think garbage collection was appropriate for a lot of what I was doing. But certainly, automating the management was ideal. After a long set of discussions, we found an interface that people agreed on, and we put it into <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. What we found was that over the next 10 years, the amount of usage of garbage collection decreased as RAII, the resource management that had been there all the time, got more and better understood and was used more.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4344">1:12:24</a>] Furthermore, the people who still used garbage collectors didn&#8217;t use the standard interface because they had figured out ways of doing it better. Today, there are still a few people doing garbage collection, but it&#8217;s not part of the standard.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4363">1:12:43</a>] How does that work? Is it kind of like a wrapper around the memory allocation methods?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4368">1:12:48</a>] Yeah, okay. You have a different implementation of new or malloc or operate a new or whatever it is you&#8217;re using at your lowest level, and then delete becomes something slightly different too.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4391">1:13:11</a>] There&#8217;s this cautionary tale in the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> community about this ship, Vasa, and I was kind of curious why that&#8217;s popular.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4401">1:13:21</a>] Oh, we had a meeting in Stockholm at some point, and they have a wonderful ship that, if you ever get to Stockholm, you should see: the Vasa. It&#8217;s a battleship from the 1600s. There&#8217;s a story to that, and that&#8217;s the one I tell people. The King was ordering a battleship that should be the best and most beautiful battleship around. It was going to be a good fighting battleship and used for diplomatic visits. So it should be beautiful. They laid down the keel and started building it. Then they heard that a likely opponent was building battleships with two gun decks. This was an old-fashioned battleship with only one gun deck. If you put a one gun deck battleship next to a two gun deck battleship, the highly predictable result is a lot of holes in the one gun deck battleship, and it&#8217;s gone.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4445">1:14:05</a>] So it should be beautiful. And they laid down the keel and they started building it. Then they heard that a likely opponent was building battleships with two gun decks. This was an old-fashioned battleship with only one gun deck. If you put a one gun deck battleship next to a two gun deck battleship, the highly predictable result is a lot of holes in the one gun deck battleship, and it&#8217;s gone.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4482">1:14:42</a>] So the King orders that this ship should now have two gun decks. And they&#8217;ve already started building it. So they add another gun deck; they add cannons up there. And the King also watched. Now the ship is bigger, they want more statues and beautiful things. So it becomes a bit top-heavy. And rumor has it&#8212;I&#8217;ve never checked this rumor, so it might be wrong&#8212;that the ship designer committed suicide out of horror.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4523">1:15:23</a>] Also, a thing that I don&#8217;t believe is just a rumor was when it was built, they tested it for stability. The way you test a ship like that for stability is you take the whole crew and you run them from one side to the other back and forth, creating harmonic sway. If you can do that 14 times, then it will stand up to the Baltic and the North Sea. Rumor has it&#8212;actually, as I said, I think it&#8217;s a fact&#8212;that they did it seven times and then they stopped because it looked dangerous.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4568">1:16:08</a>] So this was 1624, I think the ship gets finished. It&#8217;s sailing out, the most wonderful ship you&#8217;ve ever seen. It&#8217;s sailing out in Stockholm harbor, trumpets blaring, flags flying, families of crew on board, the whole thing. It gets halfway across the harbor, a gust of wind comes, it kills a roar, and it&#8217;s gone. And it ends down in some place where there&#8217;s not much oxygen. So it was well preserved, and they fished it up again, and you can see it.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4616">1:16:56</a>] And so I tell this story to the standards committee and I point out there&#8217;s something they did wrong. They built more features on top without improving the foundation. Always improve the foundation to make sure that it&#8217;s not just a random set of features that you have added, because that&#8217;s complexity. Furthermore, do not compromise your testing. That&#8217;s really dangerous. And furthermore, you&#8217;ve all noticed when your high bosses say something should be done, and the high bosses don&#8217;t always know what&#8217;s right. Sometimes the professional thing is to say, no, we are not doing this.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4663">1:17:43</a>] We have to take it easy. If they had said, okay, we&#8217;ll build a one-gun deck battleship, we&#8217;ll just not call it the Vasa, call it something neutral. Then next year, you can have a battleship that&#8217;s been designed from the bottom up to be a two-gun battleship. You wouldn&#8217;t have any problems with that. But the high management, meaning the king who was in Poland at the time, so he couldn&#8217;t even see it, says, nope, must be delivered on time.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4697">1:18:17</a>] And so they delivered something on time that just couldn&#8217;t do the job. But go see the ship. It&#8217;s great.</p><h3>1:18:25 &#8212; Template instantiation is Turing complete</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4705">1:18:25</a>] At some point, someone demonstrated that the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> template instantiation mechanism was Turing complete. So what the compiler is going to do to kind of preprocess that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> program can actually be used for computation. I was just trying to understand how that is possible. It was mentioned something about calculating prime numbers at compile time. How does that work?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4737">1:18:57</a>] Well, the prime number thing was just a curiosity. It used the error messages to report the result. But when you build something, you can get Turing completeness. You need some form of iteration or recursion, and you need a comparison. That&#8217;s about it. Then you can achieve Turing completeness. At least some of the theoreticians say, &#8220;Hey, we can&#8217;t do that. It&#8217;ll run forever.&#8221; The guy who came up with the first example of this, Erwin Unruh, actually thought I should prohibit it somehow.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4790">1:19:50</a>] I should ban it. My reaction was, &#8220;This looks useful. Great.&#8221; I think I was right. Furthermore, nothing runs forever. If you have a Turing machine, you have the tape, and the tape has to be infinite. So if you imagine building a real Turing machine the way Turing designed it, you have to have a bunch of navigators building track all the time when it gets out there. Of course, we don&#8217;t do that. The point is that the compiler will run out of resources long before we get into real problems.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4830">1:20:30</a>] Machines are finite, so the problem doesn&#8217;t become real unless there are bugs and the bugs get caught. Guaranteed. So, it&#8217;s not a problem. What happened, though, was that people were misusing templates to do simple calculations like prime numbers or trustworthiness. Sieve or calculating factorials and such. It&#8217;s so awful, and it&#8217;s so expensive, and it uses up so much memory that it becomes a problem. So that was why I and Gabbitas Reyes built constexpr, which basically says you can calculate perfectly ordinary code at compile time. It is much simpler, much more what we&#8217;re used to, and much faster to compile, usually resulting in much faster code.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4897">1:21:37</a>] And you have that today, and you have Constexpr if you want to guarantee that this is done. That takes care of the obvious misuse of the idea of templates being true and complete. It turns them into ordinary functions.</p><h3>1:21:57 &#8212; Abstraction and performance</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4917">1:21:57</a>] Generally with programming languages there&#8217;s this high level intuition that the closer to the machine you are, the higher the performance is. And I tend to see C as closer to the machine than <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, for instance.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4938">1:22:18</a>] No, it&#8217;s not the case, it&#8217;s not the case. It&#8217;s not as good as compile time calculation that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is. And anyway, we have exactly the same machine model because C borrowed the C11 machine model. So if you write the same code in both languages, you get the same result. Except the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> compilers can do more at compile time and so <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> runs as faster, faster than C. In most cases there&#8217;s more information. If you give Optimizer more information, it can do a better job.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4980">1:23:00</a>] Okay. Yeah, because that was what I was going to ask you. You had said somewhere that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> can be more performant than C, but I tend to think that more abstraction costs you something.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=4992">1:23:12</a>] It&#8217;s compiled away. This is why I talk about zero overhead abstraction, and people are beginning to take me to task for that because that&#8217;s understating the ability of the <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> compiler. We can do negative overhead abstraction.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5014">1:23:34</a>] What if I was really good at writing assembly and I had all the time in the world to write it? How does that compare?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5022">1:23:42</a>] If you are very smart and you have infinite time, you can do better. By and large, we are not as smart as the optimizers anymore, and we don&#8217;t have infinite time. So if we are smart enough, we can only do a small piece of code. And now the question is, did we get enough time to use our smarts? This is even starting to affect clever code. I gave a talk to Slack last year, which is a group of very performance-oriented people from the finance industry.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5073">1:24:33</a>] My title was &#8220;Don&#8217;t Be Clever.&#8221; Actually, the written title was &#8220;Don&#8217;t Be (Too) Clever,&#8221; but I can&#8217;t pronounce parentheses. I got out alive. My main point was that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is good enough for more than 98% of your code. So if you want time to be clever, you use these techniques, and I showed modern <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. That way, you get time so you can do all the clever optimizations. The problem is, clever optimizations these days tend to be machine dependent.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5115">1:25:15</a>] That is, if you get a new computer or if you get a new version of the compiler, you might actually have pessimized your code. I&#8217;ve seen this repeatedly ever since the 80s. There are people who do nothing but use different optimizations on the next generation hardware. My standard techniques for improving things actually are to first throw away the clever stuff and then see if you run faster or slower.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5158">1:25:58</a>] Usually, you run faster because clever stuff, at least 1990s store style clever stuff, which there&#8217;s a lot of it still today, because the techniques carry on in people&#8217;s heads and some of the code remains, tends to use a rats nest of pointers, and that gives the compilers and optimizers problems. They also sometimes use more allocations, which is not good. You want to minimize memory access, you want to maximize your cache performance and things like that.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5201">1:26:41</a>] And compilers are getting very good at that. I have seen this kind of thinking. I wrote a paper about it together with a friend of mine in Spain doing fluid dynamics, and we threw away the clever stuff&#8212;actually a performance test suite example. So it was not a toy, and we got only a 20% improvement by reducing the code to about 80% of what it was before. Some people didn&#8217;t think that was significant.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5246">1:27:26</a>] I thought it was a significant proof that the technique was appropriate. You apply optimizations only when you need them. Knuth says don&#8217;t do premature optimization, but he also pointed out that 2 to 3% is where you should optimize, which is exactly the number I&#8217;m using. First, build the stuff using high-level facilities, see if it&#8217;s good enough, and if it isn&#8217;t, and you have to time it, you don&#8217;t guess your time. Then you figure out where the time is spent and then you optimize that.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5290">1:28:10</a>] But a lot of the time, you don&#8217;t need to go to that stage; it&#8217;s fast enough.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5295">1:28:15</a>] I see. So when you say cleverness here, it&#8217;s like human-level manual management to eke out performance.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5303">1:28:23</a>] Yes.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5304">1:28:24</a>] And you&#8217;re saying that actually if you don&#8217;t do that, you&#8217;re giving the compiler more to optimize, and it can do a good job. It&#8217;s much, much better than it used to be. Code that was cleverly and correctly optimized in the 1990s is often pessimized today because machine architectures have changed and the compilers have improved.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5311">1:28:31</a>] A good job, and it&#8217;s much, much better than it used to be. Code that was cleverly and correctly optimized in the 1990s is often pessimized today because machine architectures have changed and the compilers have improved.</p><h3>1:28:51 &#8212; AI writing code</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5331">1:28:51</a>] When I look at the industry today, more and more code is being written by machines than by humans. I feel like a lot of programming language design is thinking about how to make it amenable to humans solving problems and writing the code. I&#8217;m curious if you have any thoughts on whether you think programming language design will change if more and more of the code is written by models and machines.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5363">1:29:23</a>] I think that in the field I&#8217;m mostly interested in, code will still be written by humans, and they will use abstraction. The examples I&#8217;ve seen of attempts for AI to generate code in this domain have not been successful. They generate more bugs, more security holes. They have bloated code, which pessimize again because you use more memory, and it&#8217;s hard to validate. The senior developers that would be needed to validate it have started to retire because they don&#8217;t want to deal with the validation of something that changes every time you make a change in your code or your prompts. Furthermore, a lot of the things I think about involve regulatory bodies for this validation. You have to be able to validate what you changed. When you make a change and the AIs, the tools change, even if you make a slight difference in the prompt, a lot of the code will change, and you have to check it again.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5408">1:30:08</a>] I&#8217;ve seen some of them starting to retire because they don&#8217;t want to deal with the validation of something that changes every time you make a change in your code, in your prompts. Furthermore, a lot of the things I think about involve regulatory bodies for this validation. You have to be able to validate what you changed. When you make a change and the AIs, the tools change, even if you make a slight difference in the prompt, a lot of the code will change, and you have to check it again.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5447">1:30:47</a>] All of the code that was generated knows more code generated than if it was written by humans. When a human makes a change, it will make a change that&#8217;s localized. You can look for the effects of that localized change. If an AI writes it, you don&#8217;t actually know where it&#8217;s changed. You have to try and figure that out. So if you&#8217;re doing something that has been done many times before, like writing a standard web app, what you say is correct.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5482">1:31:22</a>] Also, AI is not useless. That&#8217;s not what I&#8217;m saying. It can be used to write documentation. Again, it has to be humanly validated, but it helps write things. It&#8217;s good at text. It&#8217;s not, at least now, good at safety-critical, performance-critical code. Now, let&#8217;s say that 70 or 80% of the world&#8217;s code doesn&#8217;t fit that pattern. But it&#8217;s that 10 or 20% of the code that I&#8217;m interested in. And there, it&#8217;s not there.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5529">1:32:09</a>] And I don&#8217;t see it coming with the LLM model. Furthermore, when fed with training data, it has to be trained with old code. My job, as I see it, is to make sure people write new things and use new techniques that are improvements over the old code. I find that LLM-based code is imitating old code and getting old performance and old bugs again. Maybe you can improve that. I hear rumors of Bjarne apps being written that fit my writings, but even that is problematic because I&#8217;m not saying exactly the same as I did 20 years ago.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5586">1:33:06</a>] But anyway, we&#8217;ll see. Also, even Dijkstra was looking into the possibility, and he claimed that the idea of having natural languages as the programming language was idiotic. He&#8217;s less polite than I am. I think that a language like English is very flexible, and what we say is often very ambiguous. We need a programming language that&#8217;s precise, that&#8217;s engineering, that&#8217;s math; it&#8217;s not English.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5628">1:33:48</a>] And for that code that is performance or safety critical, I imagine there will be some group of people using LLMs for that. Based on what you&#8217;re saying, the intuition is that you would foresee more breakages and bugs because it&#8217;s not valid.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5650">1:34:10</a>] People that are really good at that kind of stuff tend to not want to spend all their time validating. Another problem is that they want to eliminate junior programmers because there are lots of them. But if you do that, where do you get the senior programmers from? We&#8217;ll see. I mean, you can ask me the same question again in ten years, and there will be more knowledge. Undoubtedly, some of what I said will not be correct. And my guess is some of what I say will be correct in ten years. I&#8217;m always told by AI proponents that either the problem has already been solved or it&#8217;ll be solved in the next release. But I hear <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> 4.7 is having more problems than 4.6 for reasons I don&#8217;t understand. The idea that the next version will solve the problem is always a dangerous assumption. Furthermore, it&#8217;s getting more and more expensive.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5682">1:34:42</a>] And my guess is some of what I say will be correct in 10 years. I&#8217;m always told by AI proponents that either the problem has already been solved or it&#8217;ll be solved in the next release. But I hear <a href="https://en.wikipedia.org/wiki/Anthropic">Anthropic</a> 4.7 is having more problems than 4.6 for reasons I don&#8217;t understand. The idea that the next version will solve the problem is always a dangerous assumption. Furthermore, it&#8217;s getting more and more expensive.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5723">1:35:23</a>] If you have to build a $100 million center and use and run the electricity for it, how many junior developers does it take to be cheaper? They&#8217;re starting to need money. It&#8217;s not unproblematic. And I&#8217;m in a subfield where it is probably more problematic than most.</p><h3>1:35:54 &#8212; His motivation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5754">1:35:54</a>] One thing I saw in a profile that you did is they asked you what keeps you going on <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> or what motivates you, and you said one is the fun of kind of building the future. The second thing was the obligation to make sure <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> moves forward. When you started <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, I can&#8217;t imagine you knew that you were embarking on a journey for decades.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5781">1:36:21</a>] And so not decades, but I knew it was a longer journey because I knew I couldn&#8217;t build the language I wanted. I could build a subset of it. There were two reasons for that. One was, well, I was the team that did it. Secondly, there was a lack of resources and a lack of time. I didn&#8217;t have the input needed to make sure that what I designed was right. So we have the engineering issue: build what you can, see what works, improve it.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5819">1:36:59</a>] And so I knew I was getting into something like that. I knew I was building a language meant to evolve. And meant to evolve means that you make certain decisions in knowing that that is different. For instance, that&#8217;s one reason <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> wasn&#8217;t just an object-oriented programming language. Because I could see in the world that there were things that didn&#8217;t seem to fit that paradigm. And so I knew we would evolve.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5855">1:37:35</a>] The other half of that answer to that question, what keeps me going is applications. It&#8217;s really nice to see interesting uses and such. I was at JPL and I talked to the people who are doing the Mars rovers. That&#8217;s cool stuff. I&#8217;ve been to CERN. I&#8217;m going to CERN this summer to see how you do high-energy physics. I don&#8217;t know anything about high-energy physics. Well, probably more than the average, but nowhere near being a physicist.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5890">1:38:10</a>] And so you can go there and see they do interesting things, and there are things that surprise you. I was talking to a guy at CERN some years ago, and his job was to open and close doors. These doors weigh a couple of tons and are made of lead; they move across to close off an area to protect against radiation or something like that. I don&#8217;t know the details, but the point is he has to start up this door, which is not too hard. You have engines. But then you have to make sure you stop it. Because when you have a couple of tons this wide going into a wall, it will not stop. Normally, you have to write. And the code for that was interesting. I learned something. I still travel around and talk to people and see what <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is being used for, what it can be used for, and what it can&#8217;t be used for. Just learning and learning is fun.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5925">1:38:45</a>] You have engines, but then you have to make sure you stop it. Because when you have a couple of tons this wide going into a wall, it will not stop. Normally, you have to write. The code for that was interesting. I learned something. I still travel around and talk to people and see what <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is being used for, what it can be used for, and what it can&#8217;t be used for. Just learning is fun.</p><h3>1:39:18 &#8212; Famous quotes</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5958">1:39:18</a>] There&#8217;s a few quotes that you have which I thought would be interesting. If you could just give some context behind them. One of the quotes is, &#8220;<a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> makes it easy to shoot yourself in the foot. C makes it harder, but when you do, it blows your whole leg off.&#8221;</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=5978">1:39:38</a>] Somebody asked a question at a talk I was given in Boston back in the 80s, and I shot that one back not thinking, but it&#8217;s a good quote and it&#8217;s correct. Arno Penzias got a Nobel Prize in Physics, so he&#8217;s not a nobody. I was trying to explain to a large group of <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> managers about <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, and he said, &#8220;You can&#8217;t have a power tool without knowing how to use it.&#8221; So if you have a saw, you saw like this.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6022">1:40:22</a>] If you have a power saw and you try to do that, it&#8217;ll bounce, and you will have to be very lucky not to get hurt. Notice it&#8217;s roughly the same story. What is behind that is if you get a power tool and you misuse it, you will encounter more problems. Get a car that accelerates faster, and it can wrap you around a tree in a way an old-fashioned, slow-accelerating car can&#8217;t. It&#8217;s fundamental to having power tools.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6063">1:41:03</a>] You also have this other great quote: nobody should call themselves a professional if they only know one language. Obviously, you&#8217;d recommend people learn <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, but if they had to know a second or a third language for the sake of being a better engineer or programmer, what would you recommend?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6082">1:41:22</a>] Yeah, and if you&#8217;ve heard, if you&#8217;ve seen that interview, you&#8217;ll know I waffle on that deliberately. It is not so much which other languages you know, but that you get a set of ideas that are embedded in those languages. So what you should do is learn languages that are different from yours. I&#8217;m not too fuzzy about which languages they are. I think I said learn a scripting language today that would be Python or <a href="https://en.wikipedia.org/wiki/JavaScript">JavaScript</a>.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6127">1:42:07</a>] Then I guess it was <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> shell or something like that. Have a look at a functional language. ML or Haskell would be obvious solutions. Or just pick something different. The point is that you mustn&#8217;t get stuck with just what&#8217;s in your language. It&#8217;s not good for you to be monoglot. I mean, you know what you call somebody who knows three languages? Trilingual. Who knows two languages? Bilingual. One language? American is a very popular joke, at least outside America. And it&#8217;s the same idea with programming languages. But it&#8217;s more important with programming languages, I think, because you&#8217;re building things, and you should broaden your mind with ideas and techniques.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6163">1:42:43</a>] American is a very popular joke, at least outside America. It&#8217;s the same idea with programming languages. But it&#8217;s more important with programming languages, I think, because you&#8217;re building things, and you shouldn&#8217;t. You should broaden your mind with ideas and techniques.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6188">1:43:08</a>] Another quote is, &#8220;People who think they know everything really annoy those of us who know we don&#8217;t.&#8221; I was curious about the context behind that or your thoughts on it.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6199">1:43:19</a>] Well, that&#8217;s. I mean, that&#8217;s very simple. There&#8217;s so many people who think they&#8217;re simple solutions to just about everything in the world. In this context, they&#8217;ll come and tell me how much simpler <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> could be. And this is true. If you only want to do one thing, usually the one they have in mind, you can make a much simpler language. But this is like if we threw away this part of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, it will be much simpler and nicer.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6237">1:43:57</a>] Usually, they want to throw away things like <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, but then you annoy a few million people, and you don&#8217;t actually succeed because they&#8217;ll stick to the old stuff. So it&#8217;s a way of expressing my frustration with people who oversimplify. People think they can program without learning to program. They think they can be engineers without learning engineering. They think they can be politicians without knowing how to run a company or a country.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6276">1:44:36</a>] Oversimplification annoys me, and I probably shouldn&#8217;t express annoyance. I very rarely do, but in this particular case, my frustration showed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6290">1:44:50</a>] Yeah, I saw somewhere in response to <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> being difficult for some people, that programming should be approachable and anyone can learn programming. I think you expressed the opinion that <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is not necessarily for everyone; it&#8217;s for serious programmers.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6311">1:45:11</a>] Yeah, I mean, the first line of the C programming language, version one, first edition, was, &#8220;C is designed to make life more pleasant for the serious programmer.&#8221; I took away the first version, which was professional, because I saw amateurs, and it was really, really good. A serious programmer is probably programming for somebody else. If you program for yourself, it doesn&#8217;t matter; it&#8217;s you, and that&#8217;s your problem.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6356">1:45:56</a>] If you do it for your friends, you can lose friends. If you build something for a million people, you can do harm in the world. And so that&#8217;s what it&#8217;s for. I mean, Guido van Rossum built Python with the explicit aim of allowing many people, or even everybody, to program. And he succeeded. I designed <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> to be a really good tool for serious programmers, engineers, mathematicians, and such. And I succeeded too.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6398">1:46:38</a>] It&#8217;s just not the same problem. Remember where we started? I said the problem, look at the problem, and then learn from what worked and what doesn&#8217;t.</p><h3>1:46:48 &#8212; Reflecting on building C++</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6408">1:46:48</a>] Looking back on <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> and the whole journey, is there any part where you think, &#8220;Oh, that was a mistake,&#8221; or something that you learned from in the design?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6420">1:47:00</a>] Many, many times I learned something. I think most of the things never made it into <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. That is, this is what you have experiments for, and that&#8217;s what you have initial uses for. I think I got the major part of the language right, and I think I could improve every single detail. But stability and compatibility are essential. If you make an insignificant change, it will annoy a few people, and it wouldn&#8217;t matter. If you make a significant change, you will annoy a lot of people, and it will not work because, say, a million people will stick to the old way. So I try to grow the language without breaking it. I have this thing that happens again and again. I explain it. People come up and say to me, &#8220;<a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is too complicated. You must simplify it.&#8221; And I need these two features. I need them yesterday. You must, when you&#8217;re doing this, give me these two features.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6468">1:47:48</a>] If you make a significant change, you will annoy a lot of people, and it will not work because, say, a million people will stick to the old way. So I try to grow the language without breaking it. I have this thing that happens again and again. I explain it. People come up and say to me, <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> is too complicated. Yeah, you must simplify it. And I need these two features. I need them yesterday. You must, when you&#8217;re doing this, give me these two features.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6514">1:48:34</a>] Yes. And whatever you do, don&#8217;t break my code. I have a million lines of it that doesn&#8217;t work. That&#8217;s impossible. And so that is why I&#8217;m working on coding guidelines and on profiles, which are enforced guidelines. That way, you can design a profile that ensures you can use the libraries that you need and ensures that you don&#8217;t misuse the features that are unnecessary and dangerous in your field.</p><h3>1:49:12 &#8212; Top C++ book recommendation</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6552">1:49:12</a>] For anyone who wants to learn <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>, what is the top technical book recommendation you have?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6562">1:49:22</a>] There&#8217;s a book I wrote when I was teaching undergrads. This is accidental; I didn&#8217;t mean it to be there. Anyway, this is the second edition of Programming Principles and Practice Using <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. This is a big fat book written for undergrads. The third edition is not as thick because the language has improved, and I can actually get the ideas across better with less text.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6599">1:49:59</a>] But use the latest <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. Learn the modern way first. Don&#8217;t start learning all the bad ways of writing C as the starter. A lot of courses still say you learn C first, so you learn to misuse malloc and pointers. Then later, you can learn how to use a vector and a string and not have the problems. But yeah, profiles are there to enable compiler and static analyzer support for that kind of thinking.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6640">1:50:40</a>] And educators are asking for something like that too. A lot of people think profiles are simply to deal with memory safety and performance. No, it has to give people a better tool both for learning and for doing specific kinds of work.</p><h3>1:50:59 &#8212; Advice for his younger self</h3><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6659">1:50:59</a>] And then the last question for you is, if you could go back to the beginning of your career and give yourself some advice, what would you say?</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6666">1:51:06</a>] Oh dear. Yeah, that&#8217;s the time machine question. I sometimes set that for my students. You have a time machine. Go back and give Dennis some advice. And once you&#8217;ve done that, step 10 years forward and give me some advice. It&#8217;s a good exercise. I usually get some really good stories out of it. Some suggestions. I think a lot of&#8212;I tried to avoid the two-way conversions of the building types in <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. I should have fought harder for that.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6718">1:51:58</a>] I tried but was stopped by the people in <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a>, and these were more experienced people than me. I should have gone further there. Furthermore, I should have delayed the release of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> until I could have something template-like so I could do a better standard library. It wouldn&#8217;t have been good enough, but it would have gotten people into the habit of using a standard. Everybody was building standard libraries, and we got saved by Alex Stepanov with the STL.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6761">1:52:41</a>] But that was real luck because I made a mistake in not delaying until I could have built a good vector and class hierarchy. And then finally, if I&#8217;d known what I know now about standards committees and bloated bureaucracies&#8212;and we have more subgroups now than we had members to start out with&#8212;I would have tried very hard to set up some kind of steering group so that people could make suggestions.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6807">1:53:27</a>] But we wouldn&#8217;t have a vote with, say, 500 people. We would have suggestions from a community of 500 people, and we would have maybe a group of five or six people with vast experience who cared for the whole language and made the decisions based on what was proposed. Something like that. But I did not have the experience or the knowledge to make such a suggestion. Notice that I did not mention tools.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6843">1:54:03</a>] <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a> has a weakness in tools, and that was because it grew up early in a time with limited tools, limited compute power, and limited memory. One constraint on the exercise I give to the students for time machines is to try and make sure that it would be possible to follow your advice. If I just said I want this, then a lot of the things couldn&#8217;t be done until 20 years later and therefore would never have happened.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6889">1:54:49</a>] There&#8217;s a lot of languages designed to be perfect for the future computers and the future programmers. Most of them die because by the time ten years later they get the language, the world has changed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6907">1:55:07</a>] On that first one, I imagine that would have been really tough to do because the <a href="https://en.wikipedia.org/wiki/Bell_Labs">Bell Labs</a> people were so senior.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6915">1:55:15</a>] I failed. I tried, but it&#8217;s obvious that you don&#8217;t want narrowing conversions. I even wrote a paper about how to get rid of them last year in a library. But it is a fundamental flaw in the type system of <a href="https://en.wikipedia.org/wiki/C%2B%2B">C++</a>. It came because they needed, being people like <a href="https://en.wikipedia.org/wiki/Dennis_Ritchie">Dennis Ritchie</a> and the <a href="https://en.wikipedia.org/wiki/Unix">Unix</a> team, to handle both integers and floating point. They didn&#8217;t think of explicit type conversion and thought explicit type conversion was too clunky.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6959">1:55:59</a>] They didn&#8217;t actually get casts until about five years after they got floating point integers. And of course, you have to be able to turn a floating point into an integer, right? So that you get implicit conversions whenever you can. Also, things like integers into characters are problematic. But since you are writing fundamental software, you didn&#8217;t want to write something complicated, and you didn&#8217;t want to have runtime checking.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=6993">1:56:33</a>] You couldn&#8217;t afford that. It was established long before I came, and my attempts to deal with that failed.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=7007">1:56:47</a>] It sounds like it wasn&#8217;t for no reason. It saves resources, maybe, or is very limited.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=7013">1:56:53</a>] They didn&#8217;t have the resources to deal with it. They built lint, the static checker, to deal with some of it and small machines. Another thing was that that group of programmers was significantly smarter and significantly more experienced than the average developer. Today we have the law of large numbers. I think the latest estimate I&#8217;ve seen on the number of software developers in the world is 47 million.</p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=7048">1:57:28</a>] And at that time with <a href="https://en.wikipedia.org/wiki/Unix">Unix</a>, the number of programmers was probably a few dozen. The ones that didn&#8217;t have a PhD from a good university were geniuses. It&#8217;s easier to get a PhD than to be a genius. They offered a different set of problems with a different set of machines and a different set of people. Boy, it was a pain and still is.</p><p><strong>Ryan:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=7082">1:58:02</a>] Awesome. Well, thank you so much for your time, <a href="https://en.wikipedia.org/wiki/Bjarne_Stroustrup">Bjarne Stroustrup</a>. I really appreciate it.</p><p><strong>Bjarne:</strong></p><p>[<a href="https://youtu.be/U46fJ2bJ-co?t=7085">1:58:05</a>] Okay, thank you.</p>]]></content:encoded></item></channel></rss>