Another open-weight model from Qwen. It is a “multimodal MoE model and an early preview of the architecture used by Qwen4.”
It is fairly large, containing 125B tokens, but activating only 6B at a time, which enables significant performance gains.
I’ve been testing these models quantized by Unsloth on a DGX Spark. I’m still exploring the model—so far, I’ve tried the 72.5GB UD-IQ1_S version (which generated these pelicans) and the 78.9GB UD-Q2_K_XL version (which generated these images).
So far, my favorite is this image generated by UD-Q2_K_XL with xhigh reasoning effort:
