Skip to content
Main Site News Console

Open-Source Chinese 8B Model Rivals Proprietary Image 2 in Performance!

· 量子位
国内AI

SenseNova U1.5 Lite

Jin Lei, from Aofeisi

It has been just three weeks, yet China’s image-generation models have evolved again.

In our last report, we introduced an image-generation model that supported native 4K output, was open source, and had a model size of just 8B.

This time, its 正式版本 has arrived! If we had to sum it up in three words, they would be—

More complete, more accurate, and more stable.

For example, we gave it nine different food photos all at once:

Along with a simple Prompt:

Combine these nine dishes into an elegant recipe-sharing card.

The nine images originally had different camera angles, proportions, and backgrounds, yet they were neatly arranged into a single image. The AI not only accurately identified each dish, but also enhanced them appropriately, making the overall composition much more visually pleasing.

Now let’s look at an example of image editing.

Starting with the original image, we used “selection + annotation” to mark the specific details we wanted to modify:

After receiving the task and completing the edits, the AI produced the following result:

As you can clearly see, the target areas changed as requested, while the bed, bedside table, glass, cabinet on the right, and the original spatial relationships throughout the bedroom were all largely preserved.

This AI is SenseNova U1.5 Lite, officially open-sourced this time by SenseTime. Its highlights include:

  • Following ultra-long, complex instructions of 3,000–4,000 Chinese characters: It can understand longer and more complicated requirements in a single pass while handling multiple constraints involving subjects, quantities, positions, text, layout, and style, with less risk of overlooking details.
  • Higher-quality visual generation: Composition, colors, materials, lighting, and details are more natural, resulting in a more polished overall image.
  • More reliable native image editing: It edits specified content more precisely while better preserving people, structures, layouts, and other areas.
  • Stronger handling of text and complex layouts: Chinese and English text, posters, infographics, and multi-text layouts are rendered more accurately, and the model is better at organizing complex visual compositions.
  • Finer visual control: It supports selection, marking, and multi-image references, making it more accurate when specifying “where to edit and what to change.”
  • Native 4K high-resolution output: It directly generates 4K images while preserving details such as composition, textures, small text, and lighting.

Moreover, SenseNova U1.5 Lite still maintains a model size of 8B. In core scenarios such as complex layouts and precise editing, it can also rival the closed-source GPT-Image-2:

If the SenseNova U1.5 Lite Preview released three weeks ago was testing how far a unified model could extend its visual capabilities, then with today’s official release, SenseTime is taking the next step: testing whether these capabilities can be strengthened individually, then unified again within a lightweight model and brought reliably into real-world creative workflows.

What makes it more capable? Look at the details

Next, let’s continue letting the results speak for themselves.

For the first task, we tested its ability to organize information, using the journey from cocoa bean to chocolate as the theme.

A single portrait-oriented infographic needed to present six stages: Harvesting, Fermentation, Roasting, Grinding, Tempering, and Finished Chocolate.

Each stage needed both a textual explanation and corresponding visual elements, including cocoa pods, fermentation boxes, roasting equipment, grinding devices, chocolate liquor, and more. At the same time, all six layers of content had to unfold downward along the layout.

Judging from the final result, SenseNova U1.5 Lite can now organize an entire visual narrative, including long-form text, hierarchical structures, illustrations, numerical sequencing, and layout relationships.

Applied to more everyday content production, this capability can handle tasks such as creating social media covers.

For example, this image:

Although this type of image does not contain many elements, the challenge lies in creating a genuine “cover” feel.

The title needs to carry visual weight, and the main subject needs to be eye-catching. The three dogs must have distinct appearances, clothing, and expressions without clashing with one another. At the same time, the entire image needs to maintain a cohesive fashion aesthetic rather than simply combining several cute elements.

Based on the final result, the official version is beginning to produce application-oriented visuals that are closer to covers ready for direct use.

Next, let’s look at SenseNova U1.5 Lite’s localized editing capabilities.

The Prompt was:

Create an infographic about the operating model of a community kitchen, imitating the paper-folding collage aesthetic of the reference image.

At first glance, the overall structure and style of the two images have barely changed—because our instruction was to “imitate the reference image.”

A closer look at the details reveals that elements originally belonging to an educational program have been replaced with kitchen-related visuals such as chef’s hats, cutting boards, and recipes. The information on the left has likewise been changed to Membership Fees, Commercial Rental, Cooking Classes, Market Sales, and Event Hosting.

For tasks like this, the model first needs to distinguish which elements in the image belong to the thematic content and which represent the underlying design logic. It then needs to preserve the latter and allow the new content to grow according to the original visual language.

However, if the single reference image becomes multiple images, the task becomes even more difficult.

For example, we entered the following three reference images:

Then we added this Prompt:

Edit Image 1 using Images 2 and 3 as content references rather than pasted rectangular images. Replace the framed band silhouettes with all five Radiohead members from Image 2, transformed into the poster’s distressed monochrome halftone style. Reflow the four lower feature blocks into a compact left column and add an ESSENTIAL LISTENING list on the right using Image 3 only for wording and hierarchy. Render the new list directly on the same continuous black distressed background with matching off-white type; do not retain any white or cream card background. Preserve all original text and all other poster elements.

The requirements in this Prompt were more detailed: the third reference image could be used only for its content and information hierarchy; its original white card background could not be carried over. The new track list needed to be integrated directly into the poster’s original distressed black background.

SenseNova U1.5 Lite produced the following result:

Looking at the result, all five members of Radiohead have been placed in the original band area, and the texture applied to the figures blends naturally with the rest of the poster. The information in the lower-left corner has been compressed, while ESSENTIAL LISTENING has been integrated directly into the layout on the right, without leaving an awkward white background.

For the three reference images used above, the first provided the layout and style, the second provided the identities of the people, and the third provided only the textual content and structure. The model had to understand these three elements separately before recombining them into a single image.

But in real design work, there is another requirement that is less complex than the examples above yet may occur much more frequently—changing a few words.

This time, after entering the original image, we used the following Prompt:

Replace the exhibition dates and address details in the dark blue rectangular information panel in the lower-left corner from
“JUNE 15 – JULY 30, 2024
URBAN GALLERY
123 CREATIVE WAY
ARTSVILLE, USA”
with
“AUGUST 10 –
SEPTEMBER 28, 2024
METRO MUSEUM
456 DESIGN ROAD
CREATIVITY CITY, UK”,
while keeping the original white and pink font styles and layout unchanged.

In the final result, the information in the lower-left corner was successfully replaced. The enormous “RHYTHM OF FORM” in the center, the highly saturated geometric shapes surrounding it, and the original text hierarchy all remained in their original positions.

Train Separately First, Then Bring It Back Together

After seeing these results, the next question is: How can an 8B model perform well across text, aesthetics, complex layouts, long instructions, and localized editing?

According to the company, SenseNova U1.5 Lite continues to use the NEO-unify native unified multimodal architecture, placing visual understanding, image generation, and editing within the same model.

In other words, before modifying any pixels, the model first needs to understand the image.

What is the main subject? Who is next to whom? Which elements are text? Which area is the user referring to? In a reference image, which elements are actually worth preserving? These processes of understanding are not completely separated from the subsequent generation and editing stages.

One of the most important changes in this official release took place during the training stage. SenseTime adopted an intuitive approach: train separately first, then bring everything back together.

Objectives such as text rendering, aesthetic expression, and image editing were first used to train their respective specialized Experts.

Once these Experts had been trained, their specialized capabilities were reintegrated into a single SenseNova U1.5 Lite through the multi-teacher online policy distillation technique MOPD.

As a result, although multiple specialized teachers are used during training, users still receive and deploy a single 8B model.

There is no additional Router behind the scenes deciding “which submodel should be called this time,” and users do not need to switch back and forth between separate models for text, aesthetics, and editing.

This approach corresponds directly to the cases discussed above.

For example, in the Radiohead multi-reference task, the model first needed to understand the visual structure of the original poster, then identify the people, extract the textual structure from another image, and finally generate the result.

The poster text-editing task works the same way.

“Change only the lower-left corner” first requires an understanding of what “lower-left corner” means, which text belongs to the target area, and which surrounding shapes, fonts, and layout elements must not be changed. Only after these judgments are complete does pixel generation take place.

Therefore, within a unified training system, the structured understanding formed through visual-semantic comprehension and spatial modeling can, in turn, assist generation and editing. SenseTime summarized this in its press release as “two-way feedback between understanding and generation.”

To keep these capabilities stable on complex tasks, the official version also specifically strengthened three areas during post-training.

The first is instruction following. When users provide a long list of requirements—including how many subjects there should be, who should be on the left, where the text should go, what colors to use, and which elements must not be modified—how many of those requirements does the model ultimately execute?

The second is visual preference. After following the instructions, do the composition, materials, lighting, and overall visual quality meet expectations?

The third is edit preservation. Can the target area be edited accurately, while areas that were already correct remain unchanged?

These three areas correspond precisely to the most likely failure points in the earlier examples: complex infographics can easily omit requirements; a cover such as STYLE directly tests aesthetic quality; and the bedroom edit, Radiohead poster, and text-replacement tasks all depend on preserving non-edited areas.

There is also prompt enhancement. If the user provides only a relatively brief request, PE can first help organize the intent and formulate a more complete creative plan before handing it over to U1.5 Lite for execution.

These are the key technologies behind the official version of SenseNova U1.5 Lite.

The Standards for Image-Generation Models Have Changed

Looking back over the evolution of the SenseNova U series, a clear trend emerges.

In the early days, the easiest way to evaluate an image-generation model was to compare whether its results were sufficiently impressive. But as AI image generation has become increasingly widespread, the questions that truly determine whether a model can enter a workflow increasingly arise after the first image has been generated.

This is why the official version of SenseNova U1.5 Lite repeatedly emphasizes “stability.” Rather than adding a few novel features that look fresh, it aims to address the complete chain of a visual task, from generation through delivery.

Accordingly, the standards for evaluating models are also changing. Whether an image is beautiful remains important, but the real differentiator is increasingly likely to be whether a model can go all the way from understanding requirements and generating content to completing revisions, refinements, and final delivery.

This is precisely where the value of a unified model becomes clear.

The final SenseNova U1.5 Lite still contains only 8B parameters. It does not continually add external routing, nor does it divide different tasks among multiple models. Instead, it first strengthens individual capabilities through multiple specialized Experts, then consolidates those capabilities back into a single unified model. The benefits are straightforward: a shorter invocation chain, lighter deployment, and better consistency across different tasks.

For developers and creators, this ultimately comes down to three words: fast, good, and efficient.

SenseTime believes that multimodality will remain an important direction as AI moves toward the physical world and enters real production environments. That is why the official U1.5 Lite release continues to devote significant effort to stability, instruction following, and editing precision. These capabilities may not be as immediately attention-grabbing as a new feature, but they more directly determine whether a model can truly be put to work.

If the Preview version three weeks ago was primarily testing how broadly the visual capabilities of an 8B unified model could extend,

then with the official release, the question has become something else:

Can these capabilities, now in place, be used reliably to get real work done?

GitHub:
https://github.com/OpenSenseNova/SenseNova-U1

Hugging Face:
https://huggingface.co/collections/sensenova/sensenova-u15

SenseNova Studio online demo:
https://unify.light-ai.top/