On August 23, Surya Narreddi posted an impressive video showing watercolors painted by a language model. The model wrote JavaScript using p5.brush, a library that “adds natural painting tools to p5.js.” At the time of writing, the video has quickly gone viral, garnering more than 1.5 million views.
The video was accompanied by a blog post describing the training process behind an earlier, narrower version of the project: painting close-up flowers rather than the complete compositions shown in the video. Unfortunately, none of the resulting artifacts have been released yet. His website says that a full technical report is forthcoming, so be sure to follow him. He originally came up with the idea, drawing on his background in art and design; in that area, he is far more capable than I am. My attempt focuses on the engineering implementation, with the goal of reproducing the approach in an open environment and releasing every component.
Note: For background on the project, watch Surya explain his thesis in his own words.
This article attempts to reproduce his idea using TRL and OpenEnv. The reference image dataset, reinforcement learning environment, training scripts, and trained models are all open.
The entire pipeline runs end to end on Hugging Face:
-
Train using Jobs
-
Deploy the reinforcement learning environment and reward model as Spaces
-
Run pairwise judges through Inference Providers
-
All artifacts on the Hub are grouped in a single collection
Once the two Spaces are running, the entire pipeline requires only a single command. Duplicate the environment and reward model, set two environment variables for the reward combination, and launch training:
hf jobs uv run train/watercolour_grpo.py --flavor h200 --timeout 48h --secrets HF_TOKEN -- \ --env-url https://<you>-watercolour-env.hf.space \ --model Qwen/Qwen3.5-35B-A3B --lora --all-linear --bf16 --gradient-checkpointing \ --subject 'a peach hibiscus' --references 4 \ --top-p 0.95 --top-k 20 \ --lr 5e-5 --lr-scheduler constant_with_warmup --warmup-steps 5 \ --scale-rewards none \ --steps 110 --n-episodes 240 --num-generations 8 \ --per-device-batch-size 1 --gradient-accumulation-steps 8 \ --max-completion-length 8192 \ --run-tag my-run --out <you>/watercolour-grpo --push-to-hubThe rest of this article explains how to complete the process; all the components are available in the code repository.
I followed the method in the original blog post step by step, making changes only when absolutely necessary. All of my own ideas were kept in a list rather than incorporated into the experiments; that list eventually became part of the “What I’d Try Next” section at the end of this article, alongside a complete list of the released artifacts. If you have already read his article, the overall approach and reward design here should feel familiar. What’s new is the open implementation, the manually rated reference gallery, and the three reward combinations that were trained and compared; the details begin with the reinforcement learning environment you need to build.
The three training runs evolved in parallel, each corresponding to one reward combination. Each frame shows the median painting from a particular step. There is no need to distinguish between them yet—the article will explain what each run represents.
Why People Like It
The paintings look loose, imperfect, and handcrafted, at a time when image models tend to generate perfect (statistically average) images. I suspect this contrast explains much of why the video went viral. It reminds me of the early days of generative AI art, when people were exploring the medium. DeepDream (2015) began as a debugging tool before people turned it into art; works such as Edmond de Belamy (2018) emerged from artists exploring what GANs could do; and artists such as Mario Klingemann used neural networks during those years to create dreamlike portraits.
This project feels closer to that early period. In his thesis, Surya describes the process that led him down this path. He initially used text-to-image models through prompting, with the prompt as the only adjustable lever; adding detail provided more control, but only up to a point. Training the model directly could go further.
The other half of the idea lies in the medium itself. The model writes roughly 150 lines of JavaScript to draw an image. The model’s output is code: you can read it, edit it, and run it again, with the decisions behind every brushstroke clearly visible. And the style comes from a constraint: the model can use only the 10 methods provided by the library. More on this below.
Around the same time, Anna Ridler photographed thousands of tulips and manually labeled each flower, exhibiting the dataset itself as an artwork before later training a model on it. I discovered her work through references returned by an AI agent while building this project; I like it because the project takes a remarkably similar approach: manually curate a collection of images, then train on that collection.
Going Beyond Aesthetic Preferences with Reinforcement Learning
Most recent reinforcement learning work with language models uses verifiable rewards. Examples include math problems with known answers, code that passes tests, or judges that determine correctness at low computational cost. This project is closer to an earlier exception—RLHF, where the model learns a reward model from human preferences.
The reward here is aesthetic preference; there is no correct answer. The real question this project explores is whether reinforcement learning can be built around aesthetic preferences.
As defined in his blog post and implemented in the reinforcement learning environment I built, the reward consists of the following:
| Component | Weight | What it measures |
|---|---|---|
gate | 0.05 | Whether the sketch compiles, renders content, and does not cheat |
length | 0.05 | Gently encourages longer code snippets |
| Pairwise judge | 0.60 | Style relative to the reference images |
| HPSv3 | 0.30 | Aesthetic preference for the rendered result |
HPSv3 is an open 7-billion-parameter preference model. Given an image and a text description, it returns a score indicating how much people are likely to prefer the image. It was trained on a large number of human choices between pairs of images, so its score represents the average aesthetic preference of many people.
The pairwise judge is Qwen3-VL-30B-A3B-Instruct, a general-purpose vision model accessed through Hugging Face Inference Providers. The pairwise judge displays the candidate painting alongside four reference images selected at random from the gallery, then evaluates them against a textual description that specifies factors to consider, such as bleeding, semi-transparent watercolor layers, and soft edges. Each comparison is run with both presentation orders, and the final score is the proportion of comparisons won by the candidate painting. Its only standard of judgment is the gallery, so its score represents my aesthetic preferences, encoded in those ratings.

Two of the four terms in the reward function come from models. Both models act as proxies for someone’s aesthetic preferences.
These are the weights Na
