(Music: “No One Is Perfect” by HoliznaCC0)
Anne Brice (intro): This is Berkeley Talks, a UC Berkeley News podcast from Strategic Communications at Berkeley. You can follow Berkeley Talks wherever you listen to your podcasts. We’re also on YouTube @BerkeleyNews. New episodes come out every other Friday. You can find all of our podcast episodes, with transcripts and photos, on UC Berkeley News at news.berkeley.edu/podcasts.
(Music fades out)
Oliver O’Reilly: Welcome back, everyone. Please join me for, this lecture is going to be given by Professor Pieter Abbeel, who’s going to share anecdotes about some of his most successful students, a group of incredible Berkeley alums that includes co-founders of over a dozen AI companies and a NASA astronaut. He is one of the world’s leading researchers, educators, and entrepreneurs in artificial intelligence, and as a professor in artificial intelligence and robotics here at Berkeley, he researches generative AI, reinforcement learning and robotics.
His pioneering AI contributions with his students include diffusion models, large world models, UniSIM, TRPO, SAC, Ring Attention, MAML, Hindsight Experience Replay, Domain Randomization, Decision Transformer, LLM as Zero Shot Planners for Robotics, and RFM1. His awards and honors include the Presidential Young Investigator Award, an NSF Career Award, an ONR Young Investor Award Investigator Award, the DARPA Young Investigator Award, TR35, IEEE Fellow, and the most prestigious one of the lot, the ACM Prize in Computing. So please join me in welcoming Professor Pieter Abbeel to the brilliance of Berkeley.
Pieter Abbeel: Thanks for coming out, everyone. Pleasure to be talking with you. I’ve actually never had the pleasure to talk after somebody who talks about love. I have no expertise in love, but now I feel like it’s my opportunity to share my non-expertise in love for just one minute. If I reflect, just a reflection, no professorial expertise. If I reflect on my love life, I’ll highlight two things. One is, even though it’s kind of unavoidable when you’re broken up with that you’re sad for a bit, the time is spent being sad and obsessing still over that person, you never get back. So keep it short. And the other one is, I think most of you probably don’t have kids. So right now love is this thing you think about your partner, that’s love. But once you have kids, you experience a whole different level of love. It’s something completely different. Keep that in mind. There’s something yet way beyond out there, and I hope you all get to experience that also.
OK. Leaving that behind, do have three kids, very lucky to have three kids. I want to talk about some success stories of some of my students, but since I and my students work in AI, it’s kind of an interleaved story of what’s happening in AI and what some of my students have done to make this happen in AI. Let’s see. So I guess Oliver, you did plenty long an intro, so I’ll keep this short. I’m professor at Berkeley. I also spend time in industry on all kinds of things, including helped start GradeScope, which some of you might use for some of your classes. Helped start OpenAI, Covariant. Recently I’m at Amazon doing a lot of AI and robotics work there. Also did a lot of investing for a while. Still do some of it now, but a little less time for it.
I think for many of you, not necessarily all of you, but many of you, probably the moment AI became something that felt like you like it or not, it’s going to be there for you, it’s going to be present, is when ChatGPT came onto the scene. Before that AI plays Go. Sure, cool. But how does it affect your life? Not really. But AI starts having conversations, starts taking exams really well, like Wharton MBA exam, the bar exam, which is something that lawyers study for a very, very long time, and somehow it scored in the top 10%. And you can ask it things like, “Hey, can you describe UC Berkeley in a poem where every line starts with the letter R?” And that doesn’t exist on the internet. That’s not just like a Google search. That’s actually a creative effort that AI has to do to do this for us. And sure enough, (slide shows a poem that meets his prompt example): Resplendent jewel on the bay, Berkeley stands proud. Rigorous minds in pursuit of knowledge, loudly avowed, and so forth.
Everything starts with letter R. It rhymes because it believes that that’s the style of poem we might be looking for, and it actually is grounded in facts about Berkeley. So this would take most people a good amount of time to come up with. I’m not saying the best poets in the world can’t do this fast, but many of us would take a long time. Now this tool is available to us. We can prompt AI with the vision of what we want and it can fill in the details and generate it for us.
It can also do math for us. 5X plus 3 equals 33. X equals 6. That’s right. Just as good as the AI on this particular problem. AI is still worse on many other things, but here AI is pretty good, and it even explains how it got there. So when you’re stuck on a homework problem or something you’re just studying on your own, AI cannot just give you an answer, but also show you how it got there. And that’s nice because it allows you to learn beyond just seeing the answer.
With these capabilities, a lot of money started pouring into AI. In 2023, everybody was like, “Oh my God, Microsoft, $10 billion in OpenAI. That’s so much money to invest into a company.” And OpenAI worth $30 billion. That’s kind of crazy. But then OpenAI became worth $80 billion, $300 billion, $500 billion, $850 billion, and this just keeps going up because more and more use cases surface and more and more use happens of AI in all kinds of productive ways.
Other companies, Anthropic, essentially a bunch of folks who were at OpenAI then left OpenAI to start their own company. They started in 2020. They raised at a $380 billion valuation just in February. If you look in the betting markets valuation of Anthropic right now, it crossed a trillion today, so it’s likely the next round will be at a trillion, which is pretty wild from a 3X in two months when you’re already worth that much. It’s pretty unheard of.
Now of course, if you write software, you’ll know why. You use Claude code, you write software at 10 times faster. Well, if every software engineer in the world is 10 times more productive, sure, companies are willing to pay a good amount of money for that productivity tool to make their engineers more productive and the revenues go up quickly for Anthropic.
Elon Musk was actually one of the co-founders of OpenAI back in the early days when I was there. He would show up every other week trying to understand what’s happening, give some suggestions. But he left at some point due to some disagreements with what OpenAI was planning to, what he wanted to do, and then he started a new AI company, XAI, which was valued at $200 billion in his latest fundraiser, and then he kind of said it’s worth $250 billion in merger with SpaceX. But this is the actual market valuation that was the latest.
And then Ilya Sutskever, the chief scientist of OpenAI, left about two years ago and a little after that started a company and immediately raised at a $20 billion valuation. So there’s this strong belief, at least with people with a lot of money, that investing in AI is probably worth a lot of money in the future because somehow it’ll create a lot of value. So you’re willing to pay a lot already now to get into these companies.
Now, the most widely used system today, generally most widely used system is ChatGPT. Actually, quick question. Anybody here not using ChatGPT ever? Never used it? Do you use anything else instead?
Audience 1: Gemini.
Pieter Abbeel: Gemini. OK. Equivalent. That’s kind of like the competitors who popped up a little later to kind of imitate ChatGPT. How does it work? A bunch of words go on the input. For example, what you type, and then it generates a word on the output and he repeats that behind the scenes multiple times feeding this N plus first word back on the input to go through that same neural net.
Neural net, just think of it as something like the brain. We don’t know how the brain works, but we know that it somehow takes inputs and generates outputs, and in between there’s this massive amount of neurons that is very flexible. It can be adjusted as you grow up to internalize knowledge, and we try to imitate that inside a computer, and so it’s learned from a ton of internet textbooks and so forth that after a certain sequence of words, another word is pretty likely. And it actually has a probability distribution over next word so it can generate sample from that distribution to generate a variety of responses as needed.
Now the question is how do you build it? The way it started was what is now called pre-training. You learn this next token/next word prediction on lots of text data, essentially entire internet. Rough amount of training data used these days is probably like roughly 10 trillion tokens. So it’s more text than any human would ever read, but of course not all the text is equal quality, but it’s still, either way, it’s a lot more than any single person would ever read in their life. It’s used to learn to predict the next word from all the previous words. If you left for long enough, you get some kind of thing that can auto complete. Let’s say it was training a lot of emails, you type email, be able to auto-complete your email that you’re typing.
What actually was found though is that it didn’t see a whole lot of use in that format. This was in 2022. People type with it, but it doesn’t really respond in an interesting way. The problem was that it had seen so much data that was not exactly how you want to have conversations. It was other data, how books are written, how internet pages are written. And you would say like, “How do I make a boiled egg?” And it would reply with, “How do I make bacon? How do I make hash browns?” And so forth. And you’re just like, “Wait, that’s not the answer to my question.” But actually if you go on the internet, you’d see that listed one after the other after the other, and so you’re like, “OK, it makes sense. That’s what it saw in sequence.” So something needed to change. It had all this internal knowledge, but it wasn’t willing to expose it the right way.
Then John Schulman, Berkeley graduate, show his picture here, one of my former students, invented the second step named post-training. And the idea there was that instead of just launching the model as is, you make it have a lot of conversation with humans where it has to generate multiple responses to everything you tell it, and then the person rates the responses, and then it is trained to maximize its ratings. Something called reinforced learning where you maximize your score. In this case, the score is your ratings of your responses. And so once you start doing that, because it sees a lot of human ratings where natural responses are rated higher than unnatural responses, it starts generating those more often. And remember, it’s really a probabilistic system, and so it’s shifting its probability distribution from kind of evenly over what all the text on the internet looks like to the text that looks like conversational AI, and now it starts doing that.
That’s ChatGPT. That was launched in November 2022 and immediately took off from there. John’s work here on step two is of course what unlocked it, but not just that. In fact, if you go back, it’s not just that you have to have a good technology, you have to have the conviction to launch it. And John also had the conviction, even though a lot of people at the time would be saying, “Oh, who wants to chat with an AI?” Sure. Sometimes it’s funny. Maybe people will chat with it for five minutes to get some jokes and then stop chatting with it. John said, “No, no, no, no, no. That’s not the right way to think about it. These things have the whole world of knowledge that’s written in text in them. People are going to want to chat with them to get information out.” In fact, the use case today is almost never getting jokes out of the AI. For example, 30% is healthcare cases where you’re wondering like, “I’m feeling this, I’m feeling that. What do you think it could be?” And it brings the entire internet to bear to provide a response.
And John had the conviction, let’s launch this as a product. Because until then, the OpenAI GPT models were just APIs. It was just like an API for auto-completes in various applications that people were building. But he was convinced and it made it happen to make it into product.
(Video clip: A simulated 3D robot learns to walk through repeated attempts)
By the way, at Berkeley, the same techniques that we used for step two of ChatGPT, the same reinforced learning techniques, he developed them at Berkeley during his Ph.D. in the context of robots learning to walk. And these robots at the time, this was the first time it was possible to have a robot like this learn to walk on its own. It shows something about the generality of these techniques.
The principle behind learning to have a better conversation is actually the same principle as the principle in terms of learning to run better. It’s about optimizing a score that you are given after every attempt. Give a score, attempt multiple times, compare the scores, compare the attempts, and then optimize for a higher score.
And the beauty actually was that at the time, the mainstream in robot locomotion was Boston Dynamics, like highly qualified engineers carefully engineering every aspect of their robot systems. But then every new robot required so much new work here. You have a new robot, sure it’s sim but the same has happened later in real. You can run the same algorithm, no change, and the new robot learns to run. So it’s pretty amazing how general these techniques are.
And in fact, the latest advances in LLMs also use reinforcement learning. So the next big wave, ’22 was the ChatGPT edition. 2024 was the reasoning edition where models now can reason, think through multiple steps of kind of getting through, trying to find a solution to a problem, maybe backtracking if it doesn’t look right, try another path until they find a solution. That’s the reasoning.
OpenAI announced this in ’24. They said, “We used reinforcement learning.” Nothing else. Big, big headlines. ChatGPT can now do so many things like solve advanced math problems and so forth. But nobody knew how they did it. The big thing that happened a few months later, OpenAI kept it secret, not telling anybody. Little company out of China, DeepSeek released an open source reasoning model showing, “Hey, our results match the OpenAI results and we will write a paper showing exactly how we did everything.”
And how’d they do it. Well, for OpenAI, it was a lot of speculation, but DeepSeek said, “This is how you do it.” You essentially set up a specific conversation. You tell the bot, this is the structure of the conversation where you’re an assistant and the user will provide a prompt that will be inserted over here, and then you have to first think, and after you’re done thinking, you provide the answer and that’s how you do things. And then you run reinforcement learning against the accuracy of the answer and the formatting of the language. You need to write in a reasonable language, not just provide a number at the end that is correct.
So what does that mean for today’s AI? It’s essentially done this way now. First you pre-train all the data you can get essentially that’s reasonable quality. Next token, next word prediction. Step two, post-training. You learn to optimize human chat quality ratings. That’s ChatGPT ’22.
And then since ’24, there is more post-training. Learn to optimize against verifiable rewards. So have the bot try to solve a math problem. It can think, then provide an answer. If it’s correct, that’s a positive example. That’s a high score. If it’s incorrect, that’s a negative example, that’s a low score. And it tries to keep optimizing that score. It allowed DeepMind and OpenAI to claim gold in the IMO, International Math Olympiad. Very hard math problems. Very few people can solve these problems, yet somehow the bots can do it now. Same for the main college programming level contest. They won gold there also.
Revenue in coding has gone up like crazy. Just a half year ago it was big news that Cursor hit $1 billion in ARR in 24 months of existence. But since then, Anthropic has just kind of taken it through the roof recently hitting a $30 billion run rate and expecting to hit a hundred billion dollars in ARR by the end of the year, which is kind of crazy to think that a company can go from essentially no revenue to a hundred billion in about two years and might keep going up. So pretty exciting. And to be personally pretty excited how much a role reinforced learning and John’s work has played in enabling this.
Now, one of the things that you might’ve noticed with these models is sometimes they’re incorrect. So here’s a great example from a court case actually. So in this case, a passenger was chatting with a chatbot from Air Canada and asking for a refund, and the chatbot said, “Sure enough, you can get a refund.” A little later, passenger does not receive the refund. They say, “What’s going on?” And the people working at the airline say, “Well, actually the chatbot was wrong. You should not be getting a refund. They made a mistake.” The person takes it to court says, “Well, I was told I was going to refund by the official communication channel from this airline.” And the court said, “Yes. As an airline, if you let chatbots have your conversations, you’re responsible for whatever they say they will pay the passengers in refunds,” and they were forced to pay. I mean, for the airline, it’s probably a couple hundred dollars or something like that. But the principle is very interesting here that the court said, “If you roll out a chatbot and it promises things to people on behalf of your company, you are responsible to honor that.” So very, very interesting precedent.
Now, let me tell a little about a company that Aravind Srinivas started, one of my other former students, called Perplexity. Perplexity from day one was essentially a version of ChatGPT where you are always getting answers grounded in context of the sources. So if you ask it a question, it wouldn’t say, “Let me go find sources to look at,” and then instead of answering just from memory like somebody might do like ChatGPT used to do, you answer based on reasoning by looking at these sources, looking at the question and reasoning through how these sources are providing in some sense the relevant content to answer the question and then formulate the answer to the question. Much more grounded that way, gives you more confidence and it gives you the references.
If for example, you say, “Hey, you know what? This NerdWallet thing, I don’t trust it. Answer again without relying on it.” You can just click it away. You get a new answer that’s just relying on, for example, the Air Canada source in the front if that’s what you want.
I’ll tell a little bit about … Actually, I forgot to tell a little bit about John. John, very interesting case. John Schulman, absolute genius. He came to Berkeley, did his undergrad at Caltech, came to Berkeley for Ph.D. Interestingly, he started working in my lab as just a rotation student for three months. At the end of the three months, I’m like, “My God, this guy is an absolute genius. Never seen anything like it. I feel so bad because he’s not officially my student. He’s supposed to be in another program. He’s supposed to be in a neuroscience program and I’m in computer science.” And I felt so guilty. I’m like, I go to my colleagues in neuroscience, I say, “It seems like John wants to keep working with me. I feel so bad because I’m like, this must be like your best student possibly.”
And I think it’s kind of pretty reflective on Berkeley. They said, “No, if that’s what the student wants to do, it’s what the student should do. We’re 100% supportive. They should pursue what they want to do, and we’re very happy for you.” And I’m like, “OK, well, I’m definitely happy.” And work with John from their own words for Ph.D. and the first couple years at OpenAI, which was really, really fun.
I’ll also say one thing that’s very interesting about John that will stick with me, the first time I took him to a lab meeting of another professor, there’s some students saying something and the professor’s commenting, and John goes, after the professor says their thing, John just goes, “I think you’re actually wrong. I think you should do it the other way.” And the person’s like, “What just happened? This John person, I’ve never seen him before. Invite him to my lab meeting because Pieter said he could give interesting comments. First thing he does is telling me I’m wrong, and he’s probably right, but it still felt very uncomfortable.” And I was like, “Yeah, I’ll tell John to phrase it a little differently next time. Compliment sandwich these feedbacks.”
So another interesting thing with John was he essentially did a complete Ph.D. in three years, and then he concluded that all this learning from demonstrations, which he had worked on, which we all do, we learn from watching other people, was never going to be enough. You needed reinforcement learning, but nobody could get reinforcement into work. In fact, if you went to the machine learning AI conference, they’d talk about reinforcement learning, you’re relegated to a small room in the side of everything and just six people, and then the thousand plus people are somewhere else and nobody comes and talk to you. He’s like, “Oh, these people are doing reinforcement learning. That stuff never works.” But John had the conviction that the future of AI had to have reinforcement learning. And then just at some point it needs to be figured out. He might as well try to figure it out now and spend all his time on that.
And he went from publishing a paper every three months in top conference venues to not publishing for a whole year, cranking away at trying to find a way to get reinforcement learning to work, and sure enough, after a year, he got it working. He had that humanoid learning to run and everything fed into ChatGPT and the current reasoning. But it was a very interesting case where somebody essentially just had conviction. And though most students like to write papers every three months or four months, he just was happy to not write any papers for a whole year to then get to the result he was really hoping for.
Aravind, another interesting case, he actually joined Berkeley hoping to work with me, but my lab was pretty full. I’m very happy he ended up working with me in the end. Very, very thankful. But I was like, “My lab is a little full. I don’t know if it’s going to work. A little out of time.” He said, “You know what? I’m happy to take my chances. I’m going to come to Berkeley. Let’s work together, see what you think after a half year.”
And honestly, the only big advice I ever had to give Aravind was the question, “Did you sleep?” Every time he showed up, he got so much work done. I’m like, “Aravind, are you sure you’re sleeping enough?” And he would say, “I cannot afford to sleep until I’m officially in your lab.” And I said, “Aravind, don’t think that way. You really need to sleep. I would really appreciate you sleep while you do your work.” And he’s finally started getting enough sleep and did a really amazing Ph.D.s. Went to OpenAI, then started Perplexity to ground answers in text.
Another interesting story reflection, you saw the DeepSeek story where, actually right now, the open source frontier of AI is largely in China. DeepSeek, Qwen are two of the top model providers in open source. Open weights, really not fully open source, open weights. There’s nothing in the U.S. Misha saw the opportunity, former postdoc of mine. He actually heard through the grapevine that Jensen Huang from NVIDIA would love to see a U.S. strong effort on open source for AI. So there was a kind of part of that here. Why? Typically, when things are open source, the whole ecosystem thrives more. A lot more people can build on it, and so the kind of AI industry could likely blossom more on top of open source, even though maybe for a company like OpenAI or Anthropic, giving everything away is not very productive for their bottom line. For the full industry, it’s good to have strong open source players.
So he went to talk with Jensen and told him, “I think I know how to do it.” And then sure enough, raised $2 billion in October and are currently just a half year later looking to raise another $2.5 billion out at a $25 billion valuation, trying to establish essentially the U.S. counterpart of DeepSeek and Qwen. Misha is another interesting case of somebody with very strong conviction. Actually, he did his Ph.D. at University of Chicago in quantum physics. Very strong Ph.D. He was one of the top Ph.D. students in the field. He decided that actually he should instead spend time on AI. So he applied OpenAI. I was not at OpenAI anymore at the time, but still a lot of my friends there, and I said, “Hey, we have this applicant, Misha Laskin. He has no background in AI. We can’t really incorporate him right now into OpenAI, but maybe we want to talk with him for … Maybe he wants to switch field in his postdoc at Berkeley.”
I talked with him and I said, “This is hard, because I mean, you have no background. I don’t know if you’re going to do well, but clearly you’re very smart. Let’s see what happens.” He said, “No problem. I’ll just volunteer.” And so he volunteered for like two or three months, and after three months, he was pretty much the most productive person in my lab. I was like, “OK, I cannot have the most productive person in my lab be a volunteer. I got to hire him as a postdoc.” And then we wrote just a ton of amazing papers together. Also really great influence on more junior students in the lab. And then of course now started this company.
Another story I want to highlight here is one thing you’ve probably seen is that AI cannot just go text to text, even though that tends to be the more work and use case that people use the most for productivity. AI can also entertain in going from text to illustration. For example, you can say, “Hey, a masterful oil painting a Persian exotic cat discovering their astounding crypto losses while checking their phone.” OK, can somebody make a piece of art that reflects this? Well, surely that requires some creativity, and that’s why this is an interesting prompt to use because we want to check, can the AI be creative rather than just parrot things it’s seen before?
(AI-generated image reflecting the prompt)
And sure enough, it can. And this obviously has never happened in the real world, but somehow the AI knows how to put these things together into a new composition.
A Victorian man struggles with his addiction to TikTok. The Victorian era in England is like 200-ish years ago. There was obviously no TikTok at the time, but it somehow still knows how to put this all together.
(AI-generated image reflecting the prompt)
This man is holding a flask, likely loves alcohol a little too much. You see the red cheeks, small eyes. He’s holding his phone and just staring at it, and that’s kind of just all he’s doing. Then you can also do it for whole videos. So this one here, the prompt is a man rides a horse which is on another horse. It’s even talking, but I don’t know if that comes through.
(AI-generated video reflecting the prompt)
(Video)
Man riding the horse: Steadiest pair I’ve ever had under me.
Pieter Abbeel: An ostrich steals dad’s hat and dad chases after it. This is my favorite.
(AI-generated video reflecting the prompt)
(Video)
Dad: Whoa.
Child: Took your hat. (Giggles with delight)
Dad: Hey, hey. Give that back. Come on. That’s mine, you hear me? My hat.
Child: Dad, it’s running too fast.
Dad: Come back here, you feathered feet. You’re not going to catch it. Oh, I’m trying.
Pieter Abbeel: And so these things can now just be made without having to actually put somebody there with a camera. Google Nano Banana is also very interesting because it’s focused on how you can edit imagery and combine different concepts together. So it was a picture of a butterfly, make it into a dress, what could it look like? Picture of flowers, some functional boots, but not the boots the way you want them to look. And it can project for you what that would look like if you combine things.
(Photo of a room, adding different features, including a bookshelf and couch)
Same thing for rooms. You can ask questions like maybe you’re looking for a new apartment and it’s empty. You’re trying to get a sense what would it look like. If you have pictures of your own furniture, you could give that. You can also just ask questions generically, like put a bookshelf there and so forth and see what everything would look like.
(Photo of a person, cycles through three different makeup looks)
Makeup. I imagine some of you sometimes spend a lot of time doing makeup and then deciding did it look good or not look good? Here can do it at least give it a trial run very quickly before having to put in any physical effort, which might save you a bunch of time.
And the person who invented essentially the key idea behind all of this was a Berkeley undergrad, Berkeley Ph.D., Jonathan Ho, and he did the work here at Berkeley. It’s called diffusion models. And the key idea is that everything in AI these days is there’s an input, there’s an output, you get enough data of inputs and output pairs, and at some point it just knows for a new input, what should roughly the output be. But for illustrations, it was proving harder, and the reason it was harder is that nothing is 100% deterministic in text, but it’s at least somewhat close to deterministic. A lot of words are the natural follow most of the previous words. Not a lot of options available.
But in illustrations, I mean, think back to the example of the cat. You could choose out of so many types of cats, so many phones, so many background sceneries. None of that is described. So you have a very open-ended output here that I need to generate, which is harder to represent and actually very hard to learn.
So the trick Jonathan invented and made the mainstream practice is to actually not just have the AI solve this problem, but initially when it’s still warming up and it hasn’t learned much yet, you also take what you want it to put on the output, you put it on the input, but you put a bunch of noise on it. So it’s kind of like a noisy, blurry version as well as a text describing what you want, and all it has to do is de-noise the thing into the correct thing. And now it’s kind of already in the right vicinity because it has a noisy version of what it needs to generate rather than having the very open-ended problem.
If you do that, that’s called diffusion. Things just kind of work. And that’s now widely used for image generation, video generation. And in fact, AlphaFold 3, I’ll say more about Alphafold in a second, is also diffusion, and most robotics uses diffusion.
AlphaFold is this thing where it was shown that you can actually, from the sequence, so the amino acid sequence that makes up a protein, which is very cheap to measure, you can actually compute the 3D structure with AI. And the 3D structure is very difficult to measure, very expensive. Of course, people did this. People paid the money many, many times to measure 3D structure with matching sequences, and then after they had a large enough data set, AI could learn the pattern, and for new sequences generated 3D. And that is now also done with diffusion. There’s competitions there.
And here you see kind of level of accuracy. Green is the x-ray crystallography measurement called experimental result. Blue is the computational prediction. You see they match very, very closely. And AlphaFold is kind of the thing that made it take off. Previous methods were kind of mediocre in performance. AlphaFold kind of shot it up much closer to 100% than anything before.
And actually the creators of AlphaFold won Nobel Prize not too long ago. In fact, just a couple years after AlphaFold. It’s very rare to win your Nobel Prize within a couple years of your invention, but somehow they managed to do it because for the longest time, biologists had said the biggest unlock in biology would be to be able to compute protein 3D structure with a computer around and having to measure it, and they did it, and so the Nobel Prize came very soon thereafter.
A lot of this is digital, then bio. How about robotics? Can robots start doing things? Turns out that a lot of my own time is actually spent in robotics and a lot of my students are there. I want to share a couple stories of those students.
So I co-founded Covariant with Peter Chen (and) Rocky Duan, and a year and a half ago we moved to Amazon as part of a deal we signed with Amazon. Interesting stories there. Peter and Rocky, both undergrads at Berkeley, Ph.D. students at Berkeley.
I’ll start with Rocky. Very interesting story. Never seen anything like him. It’s kind of crazy. He’s a freshman at Berkeley and it’s like January and he applies for a research position. And I look at his transcript. I’m like, “I’m not going to work with a freshman. It’s too early. They need to do other things, learn more, explore more.”
But his transcript is like, he took six classes, which is a lot for your first semester, especially coming from abroad, and had five A pluses and then his English writing was a B, but all his technical classes were A+. And so it’s like, “OK, well, let me see. At least encourage him rather than just ignore his request to do research.”
And I talk with him and I say, “Look, Rocky, it’d be better to take more classes, and then maybe next year we’ll do research together, but I just want to say I really impressed with your record. Let’s reconnect next year.”
And he says, “But I would really like to do research this semester.” I’m like, “OK, well, clearly he’s going to do research with somebody else. I don’t let him do research, and then I’m never going to get to do research with him, so maybe I should do research with him anyway.” That’s what’s going on in my head.
And I’m like, “OK, well, maybe we can do research, but probably you should take less classes then because six classes is just not that compatible with research, and I want you to do research seriously if you’re going to do it.” He says, “What do you mean with doing research seriously?” And I’m like, “Well, why are you avoiding the question on classes?” Well, “I say 20 hours a week would be good.” He says, “Oh, oh. I can keep doing my current classes then.” I said, “What kind of classes are you taking?” “Oh, I’m taking eight classes now.” I’m like, “What? You’re taking eight classes now?” I’m like, “What classes are you taking?” He’s like, “Oh, it’s OK. I’m taking seven technical classes. They don’t take time. But I will admit, I’m taking another English writing class. It’s really killing me.” I’m like, “My God, what’s going on?” I’m like, “He’s definitely not going to sustain that.” But I’m like, “I should do research with him because clearly he’s one-of-a-kind and I want to work with him.”
So we start doing research. He’s in my class actually, 188. We do the midterm and he stands one standard deviation above everybody else in the midterm. I’m like, “OK, this just doesn’t seem right.” And sure enough, by the end of the semester, there’s A+ on everything, including English writing, and he has done amazing research. And so I’m like, this is like 2013 or 2014, and still today I’m working with Rocky. I worked with him through his undergrad, which he did in two and a half years. Worked him through his Ph.D., which he did in two and a half years. Worked with him starting Covariant, which took us six, seven years to get to where we wanted to be, and now we’re working together at Amazon. Absolutely amazing.
And I would say the most interesting thing about Rocky, I think, in the end is that there’s this moment where Covariant is still pretty small. We’re just 12 people, 10 folks from Berkeley, from my lab, and two other people.
And we go on a retreat and we go around, and one of the new people, not from Berkeley, who we didn’t know, who just joined Covariant a month ago, and we were asking, “OK, what has mostly struck you about Covariant? What’s been most good or most bad or whatever?” And he goes, “What’s mostly struck me is the speed at which Rocky answers all my, I would argue, very stupid questions and the patience with which he answers them.” And I think that’s just so amazing. It’s like he’s somebody who can just do so much. I mean, he types his code fast, and back when people still typed code, typed his code faster than I can type emails. But he’s ready to answer anybody’s question anywhere, anytime, just the sweetest person ever. It’s absolutely amazing. Then Peter Chen and Rocky are good friends. The story is Peter Chen. Yeah, go ahead.
Oliver O’Reilly: 20 minutes.
Pieter Abbeel: Perfect. Peter Chen goes to a party at Berkeley right when he arrives here as freshman. He walks around at the party and he sees somebody at the party sitting on a bouncy ball in the middle of the party with a laptop and somehow able to focus doing his homework. And Peter’s like, “That’s the guy I need to get to know.” He goes up to Rocky and gets to know him and sure enough from them they’re best friends and they’ve done literally everything together: Undergrad at Berkeley, their research and OpenAI, Covariant and so forth. And so been working with Peter ever since.
Also the funny thing is Peter didn’t want to do AI. So Rocky was like, “I’m going to do AI.” Peter’s like, “I’m not going to do AI.” This is 2014. He’s like, “AI doesn’t work. I’m going to do something else.” But Rocky was like, “I’m going to do AI.” So they were briefly on a split path, very, very briefly. And then after a couple months, Rocky said, and that’s how I really got to know Peter, Rocky said, “I have this friend, Peter Chen, and he actually wants to do AI, and he’s not in the Ph.D. program. He’s not a student, but he’s in industry now, but he wants to do AI. He wants to come volunteer in your lab.”
And I’m like, “Rocky, do you like working with him?” And Rocky’s like, “I love working with him.” “OK. Rocky, if you’d like to work with him, let’s bring him in.” And sure enough, he came as a volunteer, and it was another one of those scenarios where he was one of the top two, three most productive students in my lab, but he was a volunteer. Again, a broken system.
Quickly got him as a Ph.D. student, but I had to wait a whole cycle. He actually started at OpenAI even before he started his Ph.D. Really amazing. Super nice also. It’s been real pleasure. Those are the people I worked with the most in my career so far.
Two other people I work with a ton of ton are Chelsea Finn and Sergey Levine. Couple of things I want to highlight here. Again, people willing to make bets. Sergei did his Ph.D. at Stanford on graphics, computer graphics. He applied to my lab at Berkeley to switch to robotics saying, “Hey, I think robotics in the end is going to be bigger than graphics, more important than graphics. I want to make the switch,” and made the switch, of course, extremely successfully. Was postdoc here, became professor here, and started his company, Physical Intelligence, that’s doing amazingly well. I’ll show a video in a moment.
And Chelsea Finn, Ph.D. student here at Berkeley, undergrad at MIT. One of the things that I want to highlight about Chelsea, which I think a lot of the folks do, but it’s not always as clearcut, and definitely I like to do it, is she doesn’t just work. She makes her exercise regimen very high priority. Chelsea would literally go swim every day. Like lap swimming, right? She would go lap swimming every day. Unnegotiable. We had a conference somewhere in some city anywhere in the world, major academic conference, and Chelsea’s like, she found the pool that she could get access to and was going swimming every day. I think there’s something to it, honestly. I think good physical exercise, better blood flow, more oxygen to your brain. I do think it really helps. Can also help in other ways, just like being fit is good in many, many ways, but I think it actually also helps in ultimate productivity.
Deepak, another student here at Berkeley. The lesson learned from Deepak in terms of what I experienced with him was, I think the most unique thing about Deepak was he would do this amazing work, and I’d be like, “My God, Deepak, that’s so amazing.” And he wouldn’t come to me to hear that it was amazing. He’d be like, “I know there are things.” He’s like, “Pieter, I know there are things you don’t like about the work. Let’s just get to it. Tell me what you think is still not good enough about this work.” Which in research essentially defines the next directions you’re going to take on and then would push very hard to get to the next phases.
In terms of what’s possible, here are a couple of videos. This is in the hills behind Berkeley. Humanoid’s walking around. This here is some work we did at Amazon recently. Humanoid wants to climb on that table. It’s a little high, so puts the chair there, gets on the chair, gets on the table. Should it jump? I don’t know. Maybe it should. And sure enough, there it goes. Nice roll.
We decided to take it to another level of speed. Got it to do a wall flip. Pretty amazing. And the interesting thing with all of these is that actually this is all trained in simulation. So simulation environments are good enough today that you do reinforcement learning in simulation to train a neural network that understands how to control a robot. You train at a very large scale and then you can make it do the same thing in the real world.
In robotic manipulation, it’s a little bit different. This is Chelsea and Sergei’s startup, PI
(Physical Intelligence).
(Video of a robot doing laundry)
There typically you collect data in teleoperations, so you make the robot do everything under your direct supervision, and then you take all that data, you learn it to, teach it to mimic what you did, and then it can do the same things, but on different items on its own. So pretty amazing. These robots can fold laundry now, do all kinds of practical tasks in the home, some cooking. It’s not all the way there or you’d probably have seen it in your home already, but it’s getting. I would say robotics is in a phase where it’s starting to show the signs that the methods we use today with a couple more years of work can likely solve a lot of the problems we’ve though about for decades in robotics.
Who here has ever used GradeScope for any of their classes? So almost all of you. Arjun (Singh), Sergey (Karayev) and Ibrahim (Awwal) built it. I did some of the product vision with them, but they built it all, ran the company. This started actually in a very interesting way. Arjun was undergrad here, loved being a TA, loved helping students, hated grading. Felt like the grading part was not the thing a TA should be spending that much time on because it’s not fun. Helping a student, that’s fun. Grading helps students in some way, but most grading is more like getting through it in a way that you get a grade back rather than really teaching a student something. He wanted to change that. And we trial ran some of the ideas he had in my class and from there built into a product that is very widely used.
I think one of the things that really stands out about Arjun is that in my own experience working with Arjun, I can tell you as a professor, sometimes you’re just tired and you just want to take the shortcut to get things done. And we have a TA like Arjun, he’s like, “No, no, no. Think about it again. That’s not in the best interest of the whole class. Let’s do it this way.” I’m like, “OK, you’re right, Arjun. We’ve already worked so crazy hard, but let’s work a little more and make it even better for the class. You’re right.” Really amazing person.
Last student I want to highlight is Woody Hoburg. Again, very interesting story of inspiration. So Woody is currently an astronaut. He’s stationed in Texas. But when he came to Berkeley, he had done his undergrad at MIT. He came to the Berkeley visit days. He’s talking with me. He says, “Hey, I want to do my Ph.D. in your lab,” and so forth. And I say, “Woody, OK. Sounds like a great fit.” He says, “Yes, it looks like a great fit based on what I’ve done at MIT where you can see the research I did there, but my real goal is not to keep doing research. It’s to become an astronaut.” And I say, “Woody, what’s the probability of that happening?” He says, “Oh, probability is effectively zero. But even though it’s pretty much zero, I still want to maximize it. If I can make it from 0.1% to 1% by giving it my all, that’s what I’m going to do. That’s my vision for what I want to achieve.”
And I asked him, “Well, Woody, what does that mean for your Ph.D. and what you want to do in your career?” He said, “Look, research is definitely my second most exciting thing, and I’m going to commit myself to that because there’s not that much you can do to improve your chances, but everything I can do to improve my chances, I’m going to do.” And so he’s like, “Often astronauts are like professors at top institutes in the aeronautics department.” He’s like, “If I can do a strong Ph.D., maybe become professor at MIT Aeronautics Department, maybe that improves my chances.” I’m like, “Sure.” And that’s pretty well aligned with what everybody does for their Ph.D. and tries for.
And the other thing he said, which really stood out, he said, “You know what? And sometimes I’m not going to do research.” I’m like, “Oh, sometimes not. In my lab, I like it when people do research pretty much all the time. What’s the point of doing other things?” He said, “If ever I’m going to be an astronaut, I’m going to be in situations that are very unpredictable and I need to react fast. It might be a life or death situation with a couple seconds to react, and I need a way to simulate that on Earth ahead of time.” And he’s like, “OK, I’m going to be a pilot. I fly my own plane and make sure that I’m very good at that because maybe I’ll be a pilot for a rocket later.” But also he said, “I’m going to go spend the whole summer in Yosemite.”
You’ve all been in Yosemite, I hope. If not, you should really go check it out. But he wouldn’t hike there. He would actually be there part of the rescue team, and he would just sit there waiting for a call. Essentially he’d be doing research, and then when a call comes in, like all the other members of the team would be hanging out, having fun, talking to each other. He’d just be doing research. The call comes in, he closes his computer, he runs out, jumps on a helicopter, and he’s the guy hanging off the helicopter, grabbing somebody off the rock who is stranded and couldn’t make it to the top and couldn’t come back down safely. And that’s what he was doing all summer, essentially rescuing people stuck in Yosemite. And it’s what somehow, all of that combined, it made it work.
Three years ago, he was called upon to become an astronaut, did the training. Again there, there’s a lot of selection, but he made it through. And he actually spent six months on the ISS from March 3, 2023 until Sept. 3, 2023. He was the pilot for the rocket that took him there. So pretty amazing. Very, very proud of him. He’s actually this year’s engineering commencement speaker, so you have a chance to actually not just see a picture of him, but actually hear him speak and see him. Absolutely wonderful person.
Again, I think one of the things that has stood out to me at Berkeley so far in my career, and I hope it stays that way, is all the people I’m highlighting here, super accomplished, but super nice at the same time. All of them super nice. And what I mean with that is it’s easy to be nice to your boss. You have to. But super nice to all their peers, everybody around them, new talent that wants to join. Not everybody’s going to be equally good, but always super welcoming to everybody. And Woody is another prime example of that. Very, very proud of him. Thank you.
Oliver O’Reilly: Thank you so much. That was amazing. Thank you.
Pieter Abbeel: Thank you, Oliver.
Oliver O’Reilly: So we have time for one quick question.
Pieter Abbeel: I can also hang out. I’ll just say I can hang out a couple more minutes after. If you can’t get your question in, I’ll stay till 2:15 p.m. or something.
Oliver O’Reilly: Thank you.
Audience 2: Hi, Professor. Thank you so much for the presentation. It was really insightful. So your students have co-founded OpenAI, Perplexity, Skild and Physical intelligence now. So looking across them, what they did, what did they get right early before the PMF, before the funding, and how they chose the problem that they wanted to solve?
Pieter Abbeel: Yeah. Great question, and I think that’s a very important question for anybody starting a company, right? What’s the right way to do it? I think there’s something very interesting to highlight here as I think traditionally the way startups were done was you target a specific market customer, and from day one, you’re iterating with that customer to best understand what they’re looking for and build your product around it.
That’s still often the case, but OpenAI kind of changed the playbook because OpenAI, Sam Altman had enough money at OpenAI to say, “We don’t need to find product market fit. We can just take a very big long-term direction. Namely, if we build highly capable AI, something will emerge.” And it happened to be ChatGPT. It wasn’t the goal to make ChatGPT. It was never the goal. It was just John, one of the many people at OpenAI said, “We should now do this.” It wasn’t like a top down, like, “This is our mission.” It’s just like we do the best AI research, something valuable will emerge.
And so a bunch of companies are founded that way now. Skild is an example. Physical Intelligence is an example. Those two companies are founded in a kind of open AI paradigm of you take some of the world’s best researcher on a certain topic and you assume that if you make even more progress, something will open up and investors are willing to make the bet. You actually delay the product market thing because you assume that today’s technology is not good enough. If you try to do it today, you’ll just be stuck and it will never be good enough for your product. You first do a couple years of research and then you might be in a place where the product can be good enough and then you go start doing the product market fit cycle. Because obviously OpenAI, ChatGPT, there’s a product market fit cycle going on at all times also.
Perplexity was a little different. There it was like, let’s look at limitations in current products and take it to another level at the product level. If you look at Perplexity, the grounded answers was something that was missing from existing products. Of course, it’s been copied by others since, but it’s always still good. If you’re first on something, you’re often getting more adoption in that direction. And since then, Perplexity have of course built computer agents that can do computer work for you. So I would say we need to know which path you’re on. The first we need to do research or we go directly product market fit.
The other lesson learned I would say from my own experience seeing so many students do startups and being part of many startups along the way is there’s a lot you need to do at a startup. There’s a lot of bits and pieces. But the thing you need to avoid, the biggest failure mode I would argue is that you are, essentially there’s like a highest order bit. Are we actually working on the right thing? And then there’s like everything else. Are we hiring the right people? Are we ramping them up quickly? Are we firing people who aren’t performing well? Are we being a friendly, productive work environment? Are we marketing ourselves well? Are we doing this and that and that and that? But if you’re not working on the right thing, all of that, you can nail it everywhere else, but it’s not going to matter. And so you get busy with everything else because that’s your day-to-day thoughts. Like today I need to get all these things right, but you need to frequently think back and like, “Are we going in the right direction?”
Your board is supposed to help you do that every few months of course, but also as a founder, you have to do that because otherwise you might feel like you’re firing in all cylinders, but it’s not on the right thing. And you nail everything, but then the company as a whole doesn’t really move forward the way you want to move forward.
Oliver O’Reilly: So please join me in thanking Professor Abbeel one more time.
(Applause)
(Music: “No One Is Perfect” by HoliznaCC0)
Anne Brice (outro): You’ve been listening to Berkeley Talks, a UC Berkeley News podcast from Strategic Communications at Berkeley. Follow us wherever you listen to your podcasts. You can find all of our podcast episodes, with transcripts and photos, on UC Berkeley News at news.berkeley.edu/podcasts.
(Music fades out)