Skip to content
Main Site News Console

A 3D Heavyweight Enters the Global Generative Model Race, Bringing Scene-Level Generation into Production Pipelines!

· 量子位
国内AI

From One Image to a Fully Generated World: Bringing a Top Conference Best Paper to Life

By Luyu, from Aofeisi

This is huge! 3D generation has just taken a giant leap from “creating objects” to “creating worlds”!

Have you seen Inception? With nothing more than a wave of their hands, dream architects can make streets fold according to the laws of physics, while buildings rise around them.

Now, something similar is happening in 3D generation:

Upload any scene image, wait 2–3 minutes, and a complete 3D scene is generated directly.

The bowl on the table can be picked up independently. The chair can be rotated or moved. Furniture you do not like can be replaced at any time…

Unlike three-dimensional images that can only be viewed, every asset here is “alive.”

And if you are tired of cattle, you can also have horses or tigers!

Taking this step is leading 3D generation company Hyper3D, which has officially released its world-generation model, WorldGen, pushing 3D generation from individual assets toward complete scenes.

Compared with the industry’s broad concept of a “world model,” WorldGen has a much more specific and clearly defined focus:

It does not generate a world that can only be viewed. Instead, it delivers 3D scenes that are interactive, editable, and ready to enter real production environments.

It is like building with LEGO: you can keep adding modules, or remove and replace any piece at any time. Each piece also comes with its own physical properties. If a key support is accidentally removed, nearby objects will shift or fall as well.

Simply put, none of the assets here are permanently locked to the background, and the assets also have clear physical relationships with one another.

That specificity also brings it closer to real-world work.

Games need levels, films need sets, and robots need training environments… Now, all of these can begin with WorldGen. Its generation capabilities are already sufficient for real production scenarios.

This did not come out of nowhere. As early as last year, Hyper3D won the SIGGRAPH 2025 Best Paper Award, one of the top honors in computer graphics, with its self-developed scene-level generation technology, CAST.

The only commercial companies to receive awards that year were Google and Meta.

CAST is also the technological foundation of WorldGen.

So, if the first question AI 3D answers is what an object looks like, WorldGen answers the second:

How can multiple objects, complete with spatial relationships and physical properties, form a scene?

One Image → A Set of Assets → A Complete Scene

At present, the vast majority of holistic scene-generation systems share one common problem:

They look good, but are of little practical use.

There is very little users can do inside the scene beyond walking around and looking. Want to make changes? Sorry—you have to start over.

The root of the problem lies in the generation method. As Wu Di, founder and CEO of Hyper3D, explains:

Traditional approaches generate the entire scene as a single model. All the content can easily be merged into one inseparable whole, making the scene difficult to continue using.

The key change in WorldGen is that it breaks the scene down into individually operable assets while preserving their spatial and physical relationships.

The entire generation process can be broken down as follows:

Open the official website (https://hyper3d.ai) and select WorldGen.

Here, we upload a partial image of a cartoon-style kitchen.

The system automatically identifies the main objects in the image. Users can also manually draw bounding boxes to specify which parts should be generated.

Next, each foreground object is generated individually as an independent asset. A preview takes approximately 2–3 seconds to generate.

The background environment is completed using 3D Gaussian Splatting, preserving visual completeness. The entire generation process can be completed in 2–3 minutes.

After enabling SimReady mode, the system also adds colliders to each asset and estimates the physical properties required for simulation, such as mass, friction coefficient, and coefficient of restitution.

The result is a successful recreation of a 3D version of Overcooked!.

Once the initial scene draft is ready, industry applications are well within reach.

Let us start with embodied intelligence, where the demand for data is most urgent.

Real-world data collection requires the cost of facilities, equipment, and personnel. Dangerous environments such as fires and structural collapses also cannot be safely recreated. Traditional simulation can help fill the gap, but engineers still need to spend substantial time building scenes manually, resulting in high costs and significant deviations from reality.

WorldGen strikes a delicate balance between the two.

Take the closed-loop solution jointly released by Hyper3D, Deep Robotics, and Moxinfei in July this year as an example:

WorldGen generates interactive 3D scenes from real-world footage. All 3D objects in the scene are ready to be imported into simulation environments.

Moxinfei’s MotrixSim engine handles physics and parallel simulation, allowing robots to move, grasp, and interact within the scenes.

Deep Robotics provides the computing power and development toolchain.

Together, the three parties form a complete pipeline spanning scene preparation, simulation execution, data collection, and policy validation.

As a member of NVIDIA Inception, Hyper3D is also using WorldGen to further connect with NVIDIA’s global simulation ecosystem, expanding the use of “individual assets in Isaac Sim” to “entire scenes in Isaac Sim.”

The application in the gaming industry is even more direct.

After receiving a level concept image, artists and designers traditionally still have a long list of tasks to complete.

Image-to-3D technology has already shortened the production time for individual assets, but those assets still have to be assembled one by one by human artists.

WorldGen aims to move this part of the workflow further upstream.

It first generates an object-based scene from the concept image. Designers can check navigation paths and spatial layout, level designers can move or replace props, and artists can determine which key assets are worth refining.

Even more conveniently, the independent assets generated by WorldGen can be directly imported into DCC and real-time engine environments such as Blender, Unity, PlayCanvas, Tuanjie Engine, and Unreal Engine for secondary creation, including material adjustments, animation rigging, level editing, and performance optimization.

In July this year, Hyper3D partnered with Unity China to jointly build an engineering pipeline connecting 3D generation with real-time game applications.

According to the team, during its closed beta, WorldGen had already attracted interest from numerous game industry professionals.

The film and television industry, meanwhile, sees its potential as a “director’s console.”

Wu Di gave a very concrete example:

If a director wants to move an apple from the plate on the left side of the frame to the plate on the right, it is difficult to control the result consistently with pure video generation.

It is much more straightforward to use WorldGen to create a controllable 3D foundation, then use a video-generation model to improve the final image quality.

Some creators are already combining WorldGen with Seedance 2.5:

WorldGen provides the 3D scene, making it easy to adjust camera movement and composition. The result can then be passed to a video-generation model to add character performances, material details, lighting, and the final visual style.

WorldGen’s AI Render module offers AI rendering, free camera movement, and other capabilities. A single scene can be transformed into multiple visual styles, bringing greater efficiency and quality to product concept demonstrations, spatial tours, and other content.

This new approach means film and television generation no longer has to choose between controllability and visual appeal. Every shot can look like a blockbuster.

WorldGen has reportedly already entered the production workflows of actual film and television projects, with the related works expected to be released successively upon completion.

XR and spatial computing are testing its versatility.

If a scene can be viewed from any angle and its primary objects can be replaced and moved, then an ordinary photograph could become an immersive space that users can enter, explore, and modify.

Interior design, immersive education, and spatial showcases could all become potential applications.

Paired with XR devices such as Apple Vision Pro, Ready Player One is becoming reality.

Of course, these represent only a small portion of WorldGen’s possible applications. Its potential is vast and still waiting to be explored.

From a Top Conference Best Paper to a Product: WorldGen Spent an Entire Year in the Making

To understand how WorldGen made its leap “to the heavens,” we need to start with last year’s CAST paper.

What makes scene generation so difficult?

It is not something that can be achieved by running image-to-3D generation a dozen times in succession. A single image inherently lacks information about depth, real-world scale, and the backs of objects. It also contains occlusions and perspective distortions.

For example, if an apple is blocked by a plate, the model first has to complete the hidden portions. If it generates the apple slightly too large, the apple may intersect with the plate when placed back on the table. If its position is even slightly off, the resulting colliders and support relationships may also be incorrect.

Errors in a scene can cascade: one mistake leads to another, and the final product reflects the real cost of rework for users.

Image

The CAST approach can be summarized in four steps: understand first → complete next → place back → perform physical correction last.

Step 1: The model decomposes the input image into multiple objects and estimates their relative depth, turning an RGB image that lacks structural information into a list of objects and spatial cues.

Step 2: It independently generates complete 3D geometry for each object. The occlusion-aware mechanism refers to the visible portions and scene information to reconstruct areas that cannot be seen in the image.

Step 3: The generated objects are placed back into a unified space. The system estimates their rotation, translation, and scale while attempting to preserve the composition and positions in the original image.

Step 4: The system establishes a more detailed object-relation graph describing relationships such as contact, support, and suspension. It then combines this with SDF-based physical correction to reduce interpenetration, floating objects, and implausible stacking.

This framework went on to win last year’s SIGGRAPH Best Paper Award. The committee stated:

CAST supports open-vocabulary reconstruction tasks and demonstrates outstanding performance in handling occlusions, precisely aligning objects, and ensuring consistency between the physical world and the input image. It opens up new possibilities for virtual content creation and embodied intelligence.

Image

But there was still one final hurdle before productization.

The team spent nearly another year specifically addressing the accumulation of errors caused by chaining together multiple modules.

By reducing redundant modules, restructuring the underlying workflow, and improving the stability of each individual stage, WorldGen was ultimately released in its fully realized form.

Image

As for why it is called a “world-generation model,” rather than a broadly defined “world model” or a “scene-generation model,” Zhang Qixuan, co-founder and CTO of Hyper3D, defines it this way:

We believe a world-generation model uses a text prompt or a single image to reconstruct the scene depicted there as a usable 3D world.

It can generate open-ended scenes and is also closer to the expression of a “world.”

It remains an important component of the broader concept of world models. After all, a true “world model” needs some medium to carry the appearance and properties of the world. Hyper3D has simply chosen 3D as that medium.

3D Generation Enters the Era of Scene-Level Creation

Let us return to the original question: What exactly does WorldGen solve?

The answer is already clearly demonstrated in its applications.

It will not directly generate a game ready for release or a complete film, nor will it replace embodied intelligence systems in training and decision-making. Instead, it addresses a preliminary challenge shared by all these industries:

How to combine ideas with vast numbers of assets and quickly turn them into complete scenes.

3D-generation models such as Hyper3D Rodin Gen-2.5 have already reduced the production cycle for individual assets from days to minutes or even seconds. But real-world projects rarely need isolated assets alone.

The assets inevitably need to be assembled despite differences in form and physical properties. This includes identifying, calibrating, and importing them into an engine—steps that still require human involvement.

WorldGen changes the fundamental unit of delivery for AI 3D:

From a single generated object to a scene that has already been preliminarily assembled.

Although there is still room for improvement before these scenes become finished products, WorldGen already provides game designers, filmmakers, and simulation engineers with a starting point they can continue to edit.

Image

For a scene to be truly usable, merely “looking right” is not enough.

Simulating the laws of physics is what determines WorldGen’s upper limit. The more details it can generate, the more realistic and useful the scene becomes.

When generated scenes can carry content, simulate interactions, and connect to existing workflows, they attract not only ordinary users looking to try something new, but also professional creators who genuinely need to improve production efficiency.

3D generation was already one of the AI applications closest to real production workflows. Its value will only be fully realized when generated results can be continuously edited, reused, and put into production.

The scene-level gap that once limited the practical adoption of 3D generation is precisely what WorldGen now fills.

Its arrival could hardly be better timed.

And so, moving from generating objects to assembling worlds, 3D generation has entered the scene-level era.