Skip to content
Main Site News Console

[AINews] OpenAI to Meet AGI Standards by the End of 2026

· Latent Space Translated
播客深度访谈

Usually, at Latent Space, we avoid discussing AGI timelines because such claims are ambiguously defined and impossible to hold anyone accountable for. But it now seems that missing this topic may be the more serious failure. The last time we looked at OpenAI’s AGI timeline was 9 months ago. Now, as expected, Chief Scientist Jakub Pachocki says that the unreleased Astra model is the “Automated AI Research Intern” he plans to achieve by September 2026. In an interview with TIME, Sama went even further, estimating that OpenAI will internally announce that it has achieved AGI by December 2026.

TIME:

  • OpenAI leadership believes the company is close to AGI. Sam Altman believes that by the end of 2026, OpenAI will have an internal system that meets the standard for AGI. Mark Chen, meanwhile, believes OpenAI has already completed 80% of the path to AGI.

  • OpenAI already has an Automated AI Research Intern—that is…

Start the clock.

AI Twitter Recap

Open-Source Robotics Breakthrough: Hugging Face and Pollen Launch the $399 Microduck

  • Microduck launch: The most eye-catching hardware launch this time was Microduck, a 25-centimeter-tall open-source bipedal robot jointly launched by Pollen Robotics and Hugging Face. It is priced at $399 and is expected to ship before Christmas. It supports training in simulation and deployment on real robots, and comes equipped with 15 actuators plus a remarkably extensive set of sensors and communications components, including cameras, speakers, lidar, NFC, Bluetooth, and Wi-Fi. Introductions from @pollenrobotics, @Thom_Wolf, and @ClementDelangue highlighted reinforcement-learning customization capabilities and a variety of pretrained policies available out of the box.

  • Technical significance: What makes it interesting is not merely that it is a “cheap and cute robot,” but the overall design of the solution: an open simulator, transfer from simulation to hardware, and a sufficiently affordable form factor and price that enable the community to train robot policies rather than simply watch demonstrations. Hugging Face has already made the simulator available through a Hugging Face Space, which @HuggingApps highlighted. It was precisely this open loop—from community training to real-world deployment—that prompted multiple researchers to purchase the device immediately, including @yacineMTB and @gneubig.

  • Early excitement and community experimentation: The launch received unusually broad attention across the robotics community. Thom Wolf shared several experiments, including rapidly integrating an image detector so that the robot could follow a laser pointer in real time (@Thom_Wolf). He later reported that Microduck was selling at a rate of one unit every 5 seconds, and subsequently announced that sales had reached $1 million (@Thom_Wolf, @Thom_Wolf). The combination of a low price, an open simulator, and embodied reinforcement learning makes this one of the more credible recent launches of “consumer physical AI.”

GLM-5.3-Flash/Ox Alpha Revealed and the Surge in Local Open-Source Models

  • Ox Alpha confirmed as GLM-5.3-Flash: One of the biggest model stories this week was the confirmation that the mysterious Ox Alpha model is actually Z.ai / Zhipu’s GLM-5.3-Flash. @theo, @UnslothAI, and @togethercompute all noted this. The public specifications repeatedly cited across these posts include 320 billion total parameters, 18 billion active parameters, a 1-million-token context window, and hybrid attention. The model performs strongly on coding and agent benchmarks.

  • Open weights + quantization + local deployment: The model attracted attention so quickly because users soon integrated it into local workflows. Unsloth said that the model can run in a 3-bit GGUF format with 128GB of RAM (@UnslothAI), while @danielhanchen claimed that 4-bit quantization retains 93% of its accuracy, making it practical to run on a Mac with 256GB of memory or two DGX Sparks. This is exactly the kind of post-launch ecosystem response that open-source model engineers watch for: quantization recipes, deployment methods, and real-world deployment constraints emerged almost immediately.

  • The price/performance narrative: Multiple posts described GLM-5.3-Flash as a new breakthrough in efficiency. @togethercompute said that it nearly matches Luna on DeepSWE while completing more than twice as much work on the same budget; @theo said its performance was already strong enough to change his model rankings; and @zainhas recommended using high rather than max reasoning effort, since accuracy is essentially the same while token usage doubles. Baseten also emphasized that the model achieved over 122 TPS in serving throughput on its first day (@baseten), while Databricks cited a speed of 270 tokens/s, claiming that it delivers 10% higher quality at one-tenth the cost compared with GLM-5.2.