Next-Generation Agents Will Naturally Move to the Cloud and Larger-Scale Computing Resources
henry from Aofeisi Temple
“As long as I want to reset, I can reset whenever I want.”
That was the confident declaration made moments ago by Tibo—the cyber godfather, AKA the Codex quota-reset deity—in the latest episode of Matthew Berman’s podcast.
What’s even more outrageous is that Tibo had a physical button built specifically for resetting quotas.

As the current head of Codex, Tibo not only explained how he gives users extra benefits in this episode, but also looked back in detail on his career at Google and revealed several important areas of work currently being pursued by himself and OpenAI.
This was, in a sense, one of the few opportunities the outside world has had to gain an in-depth understanding of OpenAI’s newly appointed product lead.
They discussed almost everything, from OpenAI’s product culture and next-generation agents to GPT’s ultimate high-speed mode, Ultra Fast; the future of human-computer interaction; OpenAI’s decision to stop developing frontier-model training; and recursive self-improvement (RSI).
Tibo’s core views are as follows:
- ChatGPT and Codex will eventually converge. In the future, there will not be two separate products—a “coding agent” and a “chat assistant.” Instead, there should be a single, highly personalized AGI that automatically adapts its interface according to each person’s tasks, abilities, and habits.
- The key to next-generation agents is not adding more Skills, Memory, or sub-agents, but making those mechanisms “disappear.” What users really want is a companion that continuously understands them, their goals, daily lives, and team context—not something that requires them to constantly manage skill files, memories, and agent networks.
- Laptops will become the new bottleneck for AI agents. Next-generation agents will naturally move to the cloud and larger-scale computing resources.
- Running 10–15 agents simultaneously today compensates for models’ relatively slow speed through concurrency. Once Ultra Fast compresses response times to something close to the speed of human thought, workflows will return to real-time interaction and “flow,” rather than constantly switching contexts.
- Competition is not the focus. OpenAI’s differentiation from Anthropic is not simply to emphasize that its “models are stronger,” but to “put the most powerful capabilities in the hands of as many people as possible.” OpenAI emphasizes broad distribution, community participation, and lowering the barriers to use. This is also an important reason Codex was integrated into ChatGPT.
- OpenAI is already using “recursive self-improvement” in practice. This does not just mean having models research models. It also includes using the most powerful models to optimize CUDA kernels, inference stacks, and infrastructure, creating a flywheel of “stronger models → greater efficiency → more compute → stronger models.”
(Tibo’s full name is Thibault Sottiaux. He is from Belgium and studied applied mathematics as an undergraduate at the Catholic University of Louvain (UCLouvain). In 2015, he joined Google, initially working on Google Maps-related projects before moving to Google DeepMind, where he was responsible for building AI research infrastructure and contributed to frontier AI projects including AlphaGo. In 2024, he joined OpenAI and began leading the Codex project. As the AI coding wave took off, Codex quickly became one of OpenAI’s fastest-growing products, bringing Tibo into the developer community’s field of view. Because he often personally responds to user feedback and helps developers reset their Codex usage quotas, online users affectionately dubbed him the “cyber godfather.”)
Here is the full, edited transcript:
Experience at Google and DeepMind
Host: Great, I’m really looking forward to talking with you. I’d like to start with your time at Google. You were on the DeepMind team. Before ChatGPT appeared, Google had something called LM Chat.
You once posted that Google was too nervous to release it, and that DeepMind had also been prevented from launching products that might disrupt Google. I think about that a lot. You were working on these products back then, long before ChatGPT truly changed the world. What were you thinking at the time?

Tibo: It was an incredibly exciting period. DeepMind was a very creative place.
My personal specialty was building infrastructure and products to accelerate research. At the time, one team was working on language models and scaling them.
Eventually, they achieved some fairly impressive results, so it was only natural to ask: Could we turn this into a product that people could have conversations with and use for all kinds of things?
That’s how ideas like LM Chat naturally emerged. It began as an internal project, and later people developed a vision of turning it into a publicly available tool.
Host: What year was that?
Tibo: About a year before ChatGPT appeared.

Tibo: We were also working on various other projects, which I won’t go into here. It really was an extremely creative place. But DeepMind was not an organization designed to deliver products.
OpenAI is very different in that respect. Our research and product teams now work extremely closely together.
We brainstorm together and co-design many things. We are highly inclined to ship products and to make them available to people. I love that. It’s one of the reasons I was drawn here: the mission, the people, and the density of talent. There are so many great things about OpenAI.
Host: When you were involved with LM Chat, did you already know that it was special, or that it might become something special?
Tibo: It definitely felt very special. With those models, you realized for the first time that they could generate coherent text and provide some degree of assistance. At first, it was simply fun; over time, it gradually became more and more useful.

Host: You said you think about that a lot, and I completely understand why. I feel that Google managed to trip itself up in many ways. What lessons did you learn there and bring with you to OpenAI?
Tibo: Yes, that’s exactly why I think about it so often. I consider it from the perspective of both team culture and OpenAI’s broader culture: which good aspects should be preserved, and what should be avoided.
OpenAI’s culture is very bottom-up and highly empowering. People can propose all kinds of ideas, come together, and ship products very quickly. There is very little overall resistance to new product ideas. It’s exciting and fun; everything is about making a positive impact on the world. Preserving that is extremely important to me.

Another equally important thing is not letting everything turn into a mess, right? You don’t want to end up with a giant hodgepodge full of new features but lacking an overall direction and consistency. So that energy also needs to be balanced with a sense of simplicity and pride in product quality.
I think the ChatGPT iOS app is one of the best apps on the market. We want to keep it that way. We invest heavily in delight, performance, efficiency, and simplicity. Those are our overall principles, while still empowering everyone to try new things and ship quickly.
Building OpenAI’s Culture
Host: If you were giving advice to entrepreneurs, how would you suggest they cultivate a culture like that? What more specific elements or practices within OpenAI would you recommend that entrepreneurs adopt?
Tibo: I think you need strong convictions. At the same time, you need to find a way to stay close to users and iterate quickly based on their feedback. You also need to be willing to disrupt yourself. That may not be as directly relevant to every entrepreneur, but it is highly relevant to a company like OpenAI.
We are constantly seeing new research and new ideas. Being able to determine when to invest in them—even when that may mean reallocating resources from the core business—is critically important. It’s difficult, but it really matters.
Host: Exactly. That’s what you just described Google as failing to do.

Tibo: To be fair, they did have a plan. It’s just that everything was part of a larger plan. For me, that wasn’t the right environment.
Host: As OpenAI—or any company—matures, does it become harder to preserve a culture of shipping quickly and being willing to disrupt itself? Especially when you already have a cash cow that keeps generating revenue, while on the other side there’s something potentially exciting and innovative.
Tibo: We are very future-oriented. The future of AI, what it will ultimately become, and how humanity will benefit from it will not stop and wait, nor will it care about what you have built over the past month or three months.
So I think it’s extremely important to commit fully, remain open-minded about where things are heading, and then figure out how you should position yourself to make sure you can catch that wave.
That’s true even for OpenAI. We train models and only then discover what they are capable of. Benchmarks cannot tell you everything.
We have to spend a great deal of time using the models ourselves before realizing: Oh, perhaps we hadn’t considered that we could benefit from using it in this particular way; or, Oh, it can actually do that.
That, in turn, changes how we think about products. For example, we’ve now launched a new voice feature. It’s extremely delightful and feels very natural to interact with. It can also use tools.

That changes a lot of things. Now I spend more time talking directly to it. Another thing I’ve been doing consistently is voice dictation, because the dictation quality is extremely, extremely good. It’s far more efficient than typing prompts.
So in the morning, I’ll sit there with my phone and say something like, “I want ChatGPT to do a few things…” Then it goes and does them; it has access to all my tools. That simply wasn’t possible before we had genuinely excellent voice models. So it suddenly changes the way you think about the product entirely.
The Future of AI Agents
Host: All right, let’s continue with new models and new harnesses. A few weeks ago, I wanted to start with another excellent post of yours: “Codex will look primitive in two or three months. We’re about to go through another major evolution. The next generation of models will need more than your laptop.” So let’s start with the harness. As models become more capable, where is there still substantial room for innovation in harnesses?

Tibo: There’s so much—really, so much. Take voice, for example. As I mentioned earlier, if you’re an advanced user of Codex or another coding agent, you’ve probably become somewhat accustomed to that lack of smoothness, right?
You have to manage skill files. That’s one way to teach it things, but I think many people have also realized that maintaining those files over the long term is quite difficult. Memory is another issue: it doesn’t always remember everything. If you have sub-agents, you have to worry about those too, and you have to build some kind of small network. At every stage of the interaction, that illusion that it is a complete partner gets broken.
What you really want is simply something that deeply understands you, understands your goals and your daily life, and understands what your team is working on. Ideally, it can respond, take initiative, help you throughout your day, and preserve the illusion that it is the perfect little companion by your side. That is exactly what we are working toward.
Another thing is that when you have extremely, extremely powerful models, you discover that the laptop itself becomes a limitation. The amount of work a laptop can handle is designed around humans.
△AI-Generated
It is roughly designed around the amount of work you can produce, the speed at which you type and think, and the number of applications you need to have open simultaneously—all of which are human limitations.
Models do not have the same limitations. For example, a model might eventually be able to handle 100 open applications simultaneously without any problem. So from the perspective of resource access, it’s clear that future models will need more resources than a laptop can provide.
Host: You mean cloud-based agents, I assume? And once we have the Ultra Fast we’ll discuss later, with token speeds reportedly 10 to 14 times faster than Fast, bandwidth constraints will change. The CPU will become the bandwidth. Literally, tool calls, the network, any tool, and any overhead in the technology stack will become limiting factors.
Tibo: You can also compensate by doing multiple things concurrently. You can explore while writing tests, compile at the same time, and simultaneously validate a new hypothesis. In this way, you keep moving the bottleneck, because you can process more things in parallel and the model can think and make progress extremely efficiently and quickly.
Host: At current token speeds, I find myself launching 10 to 15 agents in parallel. That creates a fairly significant cognitive burden: constantly switching contexts, constantly launching them, and knowing that a task may not return for another 30 to 45 minutes.
Now, with Ultra Fast, that workflow will change significantly. I don’t think I’ll continue running 10 or 15 agents at once, which might be a good thing. Maybe I’ll only need three or four at a time. How do you think independent developers’ workflows will evolve over time?
Tibo: I think managing your attention and making the experience better aligned with your attention are things we care about very deeply.
How AI Will Change Developer Workflows
Tibo: After all, we are building products for humans. We want to create technology that empowers people as much as possible, which means designing around your ability to multitask: How do you want to manage your attention? Should something be presented to you now, or would it be better to present it 30 minutes from now?
When Ultra Fast’s speed is combined with voice, you suddenly feel: Okay, this thing can operate as quickly as I can, or even faster.
Then you can stay in flow, keep developing ideas, see prototypes, and generate small reports in real time. It feels really good. You suddenly think: Right, I don’t actually want to go back to the way I was multitasking across 10 agents.
So we are working to deliver an experience that feels extremely natural while also seeming tailor-made for you. You shouldn’t have to adapt to it; the technology should adapt to you.
Host: Over the past few months, there has been a lot of discussion about agentic coding techniques. Loops were once very popular and still are; now I’m also hearing about graphs. Are these techniques all intended to help independent developers manage their attention, or, as you put it, to be more attention-friendly? I like that framing.

Tibo: I would divide the question into two categories. The first is building the best personal AGI—or personal agent, if you prefer: something that can stay in flow with you, proactively suggest important new ideas, and execute your intentions extremely efficiently whenever it identifies an opportunity.
Whether it’s a technical problem, research, advice, or something else, it can handle the task while aligning closely with you. This is extremely important. It is deeply rooted in understanding you as a human being and as a unique individual. We are pushing hard on this category of problems.
The second category is comprehensive automation. Here, you are more focused on building intelligent systems that can take over a highly complex process—one that may require, or may once have required, intelligence and appear extremely complicated.
For example, reviewing production logs and automatically optimizing performance, or detecting and automatically fixing regressions. We are seeing the same trend in cybersecurity.
Suppose a scanner discovers a vulnerability. Can the system patch it automatically and reduce the window of exposure to virtually zero, without humans participating in the loop—or with only minimal human involvement? You would approve only high-risk actions, while the system itself would remain largely automated. In that case, you would not need to maintain such direct control over it.
Host: Understood.
The Convergence of ChatGPT and Codex
Host: I’d like to shift topics slightly. Over the past few months, ChatGPT and Codex have been moving toward convergence.
My first question is: How is that going? What does it feel like internally? What feedback have you received from users?

Tibo: It has been extremely beneficial. At first, the feedback we received was: “Why merge them? Is this really necessary?” But the models of the future require us to merge them.
So that’s what we are going to do, because it is the simplest and most correct way to build that highly personalized, incredibly capable agent that can help you in all kinds of ways.
Under the hood, they will be based on the same technology: the same harness and the same overall approach. It will be highly multimodal, voice-first, and extremely efficient. Whether or not you are coding won’t matter. This agent will be able to do anything, and it will do so with very high efficiency.
The interface you use should adapt to your needs. You shouldn’t have to decide in advance, “I’m a programmer, so I need a programmer interface,” or “I’m not technical, so I need a nontechnical interface.”
People exist along a continuous spectrum. Labels such as software engineer and designer are human concepts we created to cope with an overly complex reality.

Ultimately, everyone occupies a different position on that spectrum. So we want to build the perfect interface for each person. Whether or not you are technical, it will adapt to your specific personality. That’s why we are doing this.
Host: But does that mean it will inevitably become a single unified interface? There won’t be a dropdown menu where people can choose different products? It’s truly astonishing to imagine my mother using exactly the same interface as I do.
Of course, it would be customized to my needs. If I were doing more complex work, I might need more information. But what does the final form look like to you?
Tibo: Exactly—it will be the same thing. You and your mother will use the same thing. It will be each of your personal AGIs. You will have very different tasks, derive different value from it, connect it to different tools in your lives, and bring different ideas and needs to it. Then it will continuously adapt itself to benefit you as much as possible. It will serve your friends and everyone else in the same way.
The Future of Human-Computer Interaction
Host: I want to return to something you said earlier. When describing the end state, you used the word “illusion” several times. For ordinary users, what does that perfect “illusion” look like? If you imagine a few years into the future, what will interaction between AI and humans look like?
Tibo: To me, it will be something extremely attuned to humans. That is also why large language models succeeded: they use natural language. Natural language is inherently a human concept, right? We are accustomed to talking to one another. If tomorrow you wrote me a letter