Skip to content
Main Site News Console

How to Generate Satisfying Images

The quality gap mostly comes from the prompt. This page provides ready-to-use wording and an iteration workflow.

Four-Part Prompt Structure

Organize it as “subject + scene + style + lighting/texture”:

A silver-shaded cat (subject) lying on a sunlit wooden windowsill (scene), healing-style illustration (style), soft backlight in the afternoon, delicate fur texture (lighting/texture)

  • Both Chinese and English are fine; for specific style terms like (cinematic, isometric, watercolor), English is more reliable;
  • If you want text to appear in the image, put it in quotes: The sign says "OPEN";
  • State negative requirements directly: “Do not include people in the background.”

Iteration Workflow

  1. Use -fast / the standard tier to quickly generate 3–5 images, changing only one variable at a time for comparison;
  2. After locking the composition, fine-tune the style and details;
  3. Once everything is satisfactory, switch to the high-resolution tier to produce the final result; see Compression and Output Resolution;
  4. If you need to keep characters/products consistent, use the Image Editing API for image-to-image editing, rather than repeatedly trying text-to-image and hoping for the best.

Model Focus by Series

SeriesStrengthsPrompt Focus
GPT-ImageSemantic understanding, complex instructionsLong compound sentences are acceptable
Nano Banana(Gemini)Realistic portraits, on-image textMark text content with quotes, specify the aspect ratio
ImagenPhotographic textureEmphasize lens/aperture/film terminology

Common Failures and Fixes

  • Broken fingers/text: switch to a higher-tier model, or add text later in post-processing;
  • Style drift: put style terms at the beginning of the prompt, and keep using the same model;
  • Composition not following instructions: use explicit composition terms like “centered composition / rule of thirds / 45-degree overhead shot”;
  • Inconsistency across multiple images: fix one reference image and use image-to-image generation; see Nano Banana Image Generation.