Skip to content
Main Site News Console

Connect, Run, Deploy: AI Workflows in Gradio

· Hugging Face Translated
教程模型卡

Most interesting AI applications are pipelines made up of multiple steps. You might first generate an image, then remove its background as needed, or edit it into something entirely new. You might write a script first, then generate speech for it, or replace the voice while keeping the script unchanged. We typically use Python to connect these steps, and whenever a result looks wrong, we go back to using print for debugging to determine exactly which step produced the unexpected value.

Built directly into Gradio, gr.Workflow turns the pipeline itself into the interface. You can describe each step as a graph of typed nodes, and Gradio provides a drag-and-drop canvas: every node can be run, and every intermediate result is clearly visible. The same graph automatically becomes a REST API and can be deployed to Hugging Face Spaces with a single command.

The best way to understand the concept is to look at a few workflows in action. Each application below is a Hugging Face Space that you can open, run, and duplicate.

Edit Images

Upload an image and enter an editing instruction, such as “turn it into a snowy winter scene,” “add sunglasses,” or “make the car red,” to receive an edited photo. The entire application consists of a single node that calls Qwen-Image-Edit through Hugging Face Inference Providers.

👉 Try the image editing pipeline

Chain Real Models into a Media Studio

One image, three pipelines. After you enter a prompt, FLUX first generates an image, which is then passed to a Gradio Space that removes backgrounds to turn it into a sticker. A theme can be turned into a voiceover using a text-to-speech Gradio Space, while an LLM is called to transform the same theme into an engaging episode title.

The entire canvas contains two model calls through Hugging Face Inference Providers, as well as two calls to Gradio Spaces.

Because this is a workflow, the three outputs also each have their own REST endpoint: /sticker, /voiceover, and /episode_title. You can call any of these endpoints directly from code without opening the UI. The Calling from Code section below includes a runnable example.

👉 Try the AI media studio

Scale Image Generation in Parallel

Enter an idea to generate a collection of artworks at once: a base image generated by FLUX, two AI reinterpretations based on that image—one in a soft watercolor style and another in a neon cyberpunk style—and a gallery title written by an LLM.

Each image is generated directly from the prompt by a model node using Inference Providers, while the title comes from an fn node that calls an LLM. This is the fan-out pattern in action: a single idea can be passed to multiple operation nodes simultaneously, allowing them to generate content in parallel.

👉 Try the generative art lab

Analyze Hugging Face Datasets

Enter a Hugging Face dataset ID, such as stanfordnlp/imdb or mteb/tweet_sentiment_extraction, and a single input fans out to four operation nodes that analyze the dataset in real time through the Datasets Server API.

You’ll receive an overview card, a preview of the first few rows, per-column statistics, and distribution charts. Everything is computed independently and in parallel. That’s the power of workflows!

👉 Try Data Detective

Run Your Own GPU Models

So far, every node has accessed Hugging Face. But an fn node is essentially just Python, which means it can also run models on the GPU inside a Space.

After adding the @spaces.GPU decorator to the bound function, ZeroGPU requests a GPU for the call when the node runs, executes the model, and then releases the GPU. We don’t always need to rely on Inference Providers or existing Gradio Spaces.

Take a look at this demo: it uses Lightricks/LTX-Video, loaded through Diffusers, to animate a static image in a single-node workflow. gr.Workflow doesn’t need to know anything about the GPU configuration—it only needs to call the bound function.

👉 Try the ZeroGPU animation tool

How It Works at a Glance

Every workflow is a graph made up of three types of nodes: references (your inputs), operators (the steps that perform tasks), and subjects (your outputs). Operators can be your own Python functions, models on Hugging Face Inference Providers, another Gradio Space, or a row from a dataset on the Hub. Simply drag connections between typed ports, click Run, and watch each result appear in its corresponding location.

Calling from Code

Every workflow you build automatically becomes an API with no additional work required. Each output becomes a REST endpoint named after its label, which you can call from Python using the Gradio client. Here is a live example for a multi-endpoint demo Space that runs directly without a token:

from gradio_client import Client
client = Client("ysharma/gr-workflow-multi-endpoint-API")
print(client.predict("hello there friend", api_name="/word_count")) # -> 3
print(client.predict(20, api_name="/fahrenheit")) # -> 68.0

Calling endpoints for models or Spaces runs under the permissions of your Hugging Face token, so you need to pass a token when creating the client:

from gradio_client import Client, handle_file
client = Client("ysharma/gr-workflow-image-editor", token="hf_...")
edited = client.predict(
handle_file("dog.jpg"),
"turn it into a snowy winter scene",
api_name="/edited_image",
)

Prefer plain HTTP? Every endpoint can also be accessed with curl:

Terminal window
curl -s https://ysharma-gr-workflow-multi-endpoint-API.hf.space/gradio_api/call/word_count \
-H "Content-Type: application/json" -d '{"data": ["hello there friend"]}'

Build Your Own Workflow

The fastest way to get started is to open any of the demos above, click Duplicate, and start reconnecting the nodes. In Python, the code can be as short as:

import gradio as gr
def your_function(text: str) -> str:
pass
gr.Workflow(bind=[your_function]).launch()

For a complete tutorial, operator types, JSON Schema, and reusable patterns, see the official gr.Workflow guide in the Gradio documentation.

You can even use gr.Workflow to build applications as complex as AUTOMATIC1111. Stay tuned for our next article, where we’ll walk through how to build one step by step. Here’s a sneak peek 😉👇