One Reason for Leaving Google: Small Teams Can Focus Intensely!
Gemini Was Indeed Not Very Good at Coding in Its Early Days
You have to hand it to Jeff Dean—he really is refreshingly candid.
In his first public interview since leaving Google, he made no attempt to hold back, effectively delivering a “former employee’s candid critique of his old company.”
At Stanford’s 2026 Frontier & Pioneer Symposium, Jeff Dean openly discussed a number of topics he had rarely addressed in public before—
including why he left Google, why Gemini fell short, his secrets to conducting research, and what his new company Discovery Loop plans to focus on next. He spoke freely about all of them during the interview.

The interview was packed with information. Jeff Dean’s key conclusions this time can be summarized as follows:
- AI’s capabilities in cybersecurity may already have reached the level of top human attackers, and could even go beyond it.
- TensorFlow made two fairly clear mistakes back in the day—mistakes Jeff Dean himself acknowledged.
- In the future, one iteration of a scientific experiment could be compressed from a day or a week to a minute or an hour.
- Jeff Dean’s vision of true recursive self-improvement will encompass model parameters, training data, Eval, and even the model architecture itself.
- When looking for major research directions, it may be better to skim the abstracts of 10—or even 100—papers than to read one paper in depth first.
Below is a transcript of the key points from the interview, selected and organized around its central ideas. Some passages have been lightly abridged or edited without changing their original meaning.
Gemini Was Indeed Lacking in Coding in Its Early Days
Q: Looking back on Gemini’s development, what surprised you most? What lessons could influence the next generation of AI systems?
Jeff Dean: Gemini was ultimately the result of several early research projects within Google converging—
including the original DeepMind, Google Brain, and work from other teams at Google Research.
Over time, we gradually realized that everyone was moving in very similar directions.
For example, they were all continually scaling up their models, while several teams were separately researching how to give language models multimodal capabilities so they could understand information such as images.
So at one point I wrote a one-page memo and thought: it’s a bit silly for everyone to work separately. We should collaborate directly.
We could bring together everyone’s people, ideas, and computing resources to train a model with multimodal capabilities from the outset, bringing together the best people from multiple research organizations within Google.

Later, Oriol Vinyals and I launched the Gemini project together, serving as its initial co-technical leads, and truly integrated these teams.
Looking back now, giving the model multimodal capabilities from the very beginning was a highly successful decision.
If you ultimately want a model that can be used for a wide range of tasks, it should understand text, language, code, images, video, audio, and other modalities simultaneously.
We even included some LiDAR data in the training data, at least so the model would know that LiDAR is a type of data. It could potentially become an important use case when Gemini was trained further in the future.
But at the time, we wanted the model to perform well across many different tasks. So I think we were a little late in recognizing how important it was to make Gemini’s Coding capabilities truly impressive.
We eventually realized this and have been working hard to catch up. Some very good work is now underway.
I also think Coding is an extremely important capability.
When you specifically improve a model’s Coding abilities, you often end up with a system that is better at reasoning, because writing code requires the model to process problems step by step, break a complex problem down into multiple subproblems, and solve them one at a time.
So once Coding capabilities improve, that ability often transfers to many non-Coding tasks as well.
Admitting That TensorFlow Made Two Clear Mistakes
Q: Looking back at TensorFlow, which design decisions would you reconsider today? If you were building it again from scratch, what would you do differently?
Jeff Dean: I think there were several things we didn’t get quite right.
First, we didn’t include an Eager Execution mode from the beginning.
This approach later became very popular in frameworks such as PyTorch and JAX, and TensorFlow eventually added the feature as well. I think it made the overall abstraction much better.
The other issue was that, when we open-sourced TensorFlow, we created a subdirectory called contrib, allowing many external developers to contribute various utility libraries and alternative implementations.
Over time, this caused a great deal of confusion in the community, because the same thing could gradually have ten different ways of being done. Which one you should use depended on which subdirectory or library in contrib you chose.
Looking back now, we should have kept TensorFlow’s core much simpler and provided these components as external libraries built on top of the core.
If we were doing it again today, we would not design it that way.
Of course, TensorFlow as a whole still helped many people get started with machine learning and gave everyone a common framework for solving the problems they cared about. I think that was extremely valuable.
One Reason for Leaving Google: Small Teams Can Focus Intensely!
Q: Why did you choose to leave Google and start a company? What advantages do small teams have over large companies when conducting frontier AI research?
Jeff Dean: I genuinely loved the years I spent at Google. I worked there for 27 years and met many outstanding colleagues.
So I think Google is in a very good position today. It has its own plans and will continue making the Gemini models excellent. At the same time, I was eager to go out and work on what I’m doing now.

Sometimes, there is something inherently appealing about a very focused small company where everyone works toward the same mission.
And the circumstances today are very different from what they were in the past. The growth of cloud computing and the deployment of large-scale machine learning infrastructure across various cloud platforms mean that even a very small team can raise funding and immediately use this infrastructure, without having to build an entire system from scratch.
So if a small group of people has a dream, a vision, or a particular direction they want to explore, they can now pursue it within a very small organization.
We can rely on Google or other cloud providers to handle much of the heavy infrastructure work.
For us, one of the most exciting things about being a small company is that we can focus intensely on automating science and engineering.
To be frank, this is probably something we could have done within Google as well.
But if there are only around 10 people, all sitting in an office somewhere in Palo Alto and focused exclusively on this one thing, many of the minor distractions that come with a large organization can simply be avoided.
Of course, large companies also have many wonderful aspects. Over the years, I built deep friendships at Google and benefited enormously from the resources and support that a large company can provide.
So leaving that support is certainly a little nerve-racking; at the same time, it is incredibly exciting.
After Leaving Google, He Wants to Turn Scientific Discovery into an Automatically Iterating Loop
Q: Recently, there has been increasing discussion of recursive self-improvement: Can AI systems continuously learn how to improve themselves and even accelerate the progress of AI? What do you think about this?
Jeff Dean: Using machine learning to improve machine learning has actually been explored for many years.
My co-founder Quoc Le, for example, was working on Neural Architecture Search very early on: having a model automatically generate model architectures, then continuously receiving feedback based on metrics such as learning speed and training cost to gradually find better designs.
Later, they developed Evolved Transformer, using evolutionary algorithms to recombine Transformer components. The resulting architecture was around 30% more efficient than the standard Transformer.
So I have always believed that having AI participate in improving AI itself is a very important direction.
The fundamental problem recursive self-improvement needs to solve is how to continuously improve every element required to build a model through automated means.

Today, an entire team will typically investigate what data is most useful for improving model quality, what kind of Eval is needed to evaluate the model, and which model architecture should be selected.
I believe these stages can eventually form a highly effective automated Loop. Each component can be optimized continuously, and the results can then be combined to improve the model’s overall capabilities, quality, and data mixture.
In fact, if you look closely at many modern scientific and engineering problems, you will find that they share a similar structure.
First, you pose a major question and break it down into many subproblems. You propose possible solutions to one subproblem, implement them, run an actual experiment, and evaluate the results.
You then feed the experimental results back into the process and use that feedback to determine what experiment to conduct next.
This is essentially the most basic scientific method. Engineering design works the same way: you continually iterate on your design and compare the various properties of different approaches.
What we want to do is make this complete Loop increasingly automated.

Discovery Loop’s Goal: Give One Model the Research Capabilities of “20 PhDs”
Q: You just discussed recursive self-improvement, which also seems related to your newly founded company. What exactly are you trying to do?
Jeff Dean: The idea behind our new company, Discovery Loop, is to automate machine learning, science, and engineering in order to accelerate the pace of scientific discovery across many different fields.
Initially, we will certainly focus on a small number of fields, because maintaining a certain degree of focus is very important in the early stages.
But we believe there is a great deal of reusable infrastructure and general-purpose technology shared across different fields.
Moreover, if we can build a model that truly understands many different scientific and engineering disciplines, it may be possible for a single model to possess expertise approaching the level of a PhD in 20 different fields.
No human could hold PhDs in 20 different fields simultaneously.
But if a model had this kind of capability, it could identify which subproblems within a major question were truly important, then assign Agents and Multi-Agent systems to solve those subproblems separately.
It could then recombine the results from each subproblem into a solution to the overall question and continuously repeat this cycle.
We also want to make this cycle run faster and faster. By building the right tools, the system could conduct an experiment or evaluate one extremely quickly.
That way, an experimental iteration that used to take a day or even a week might eventually take only a minute or an hour.

What’s more, we could run thousands of experiments simultaneously, gather feedback from them, and use that feedback to determine what the next batch of experiments should be.
So we hope to improve two things at the same time: the speed at which experiments run and the quality of the experiments themselves.
If both of these can continue to improve, I think the results could ultimately be astonishing.
Rather Than Read One Paper Closely, Browse 10 Papers—or Even 100 Abstracts
Q: How do you determine whether a technology has truly foundational value or is simply popular at the moment? Are there any research or engineering principles that have remained unchanged?
Jeff Dean: I often tell students: instead of reading one paper extremely carefully, it can sometimes be better to quickly browse 10 papers.
That way, you will have 10 more ideas in your head about “which things may be starting to become feasible.” You can even quickly browse the abstracts of 100 papers.
The capability you really want is to connect important ideas that had not previously been linked.
Sometimes, when you are facing a difficult problem, having a sense that many things are gradually becoming feasible can help you rethink the entire problem.
At first, the problem may appear to consist of seven parts, none of which can be solved. But when you look at it again, you may find that research in five of those directions has already begun to take shape and may be sufficient to solve part of the problem.
That leaves two problems where you currently have no idea how to proceed. But if you devote yourself to them, you may be able to solve them.

To me, these are the kinds of problems that are truly worth pursuing over the long term—perhaps for five years.
If a problem appears to require 20 years to solve, and you have no idea how to make progress on any part of it, then it is usually too early.
Likewise, a problem that can be completed in two years and has an obvious path forward is more of an engineering project.
What is truly worth seeking out are problems that may require a very different approach but are also just beginning to become feasible.
Another engineering tool I use frequently is making quick order-of-magnitude estimates. For example, if I need to process this much data, how long will it take? If I need to transmit this data over this kind of network, is it actually feasible? Will it take 100 years, or only 10 seconds?
Ten seconds and 100 years are completely different things.
So, you need to be able to use first principles and some engineering experience to quickly estimate the approximate scale of different solutions in your head.
I don’t have any magical answers, and I have tried many things that ultimately failed. So here is another piece of advice:
Try more things that may not succeed. Some of them will inevitably work.
AI Can Already Do What Top Human Hackers Can Do—and May Go Even Further
Q: What is your view of the rapid improvement of AI capabilities in cybersecurity?
Jeff Dean: I think these models can be used for many different purposes, just like many other technologies.
For the vast majority of applications, I believe they will have an extremely positive impact on the world.
For example, AI in healthcare and education can help people solve problems that were previously difficult to address on their own and enable them to accomplish more. All of this is incredibly exciting.
But models can also be used to find security vulnerabilities, so this really is a double-edged sword.
You can use them to discover and patch the many security vulnerabilities that exist in the real world, but malicious users can use the same capabilities to exploit those vulnerabilities.
These models are now genuinely capable of doing some of the things that highly skilled human cyberattackers can do, and may even go beyond them.

At the same time, they may discover vulnerabilities that even highly capable human cybersecurity engineers would not necessarily find.
So cybersecurity has always involved a balance between offense and defense: on one side, people work to protect computer systems, while on the other, people work to attack them. The tools available to both sides are now much more powerful.
I am not a cybersecurity expert, but I do think this is a legitimate cause for concern. For certain things we do not want models to do, we may also need to adopt nontechnical measures, such as laws and regulations.
Ultimately, society as a whole needs to gradually figure out what we want models to do and what we want to prevent them from doing.
Reference: