On the evening of August 26, Qwen Office became the first to launch the newly released Qwen3.8-Flash model, while also introducing Standard Mode. Starting today, all users can experience Qwen3.8-Flash through the new Standard Mode. Powered by the latest model, users can complete tasks with lower credit consumption and faster token throughput. Going forward, Qwen Office will offer only two model modes: Standard and Advanced. Standard Mode is sufficient for 95% of everyday tasks, while only 5% of complex tasks require Advanced Mode.

The improved user experience comes from upgrades to the model and coordinated optimization with the Agent. With a new architecture and hundreds of billions of total parameters, Qwen3.8-Flash delivers performance that surpasses Claude Opus 4.6. At the same time, the Qwen foundation model team and the Qwen Office team jointly developed an office-specific version of Qwen3.8-Flash, with targeted training and optimization for scenarios such as multi-step planning, tool selection, and context compression. Reasoning optimizations and a customized Harness architecture further improve throughput efficiency. In tests conducted in real-world office scenarios, the Qwen Office Standard Mode increased single-task generation speed by approximately 100%, while reducing token consumption by an average of 75%.
In real-world AI application scenarios, high performance typically means higher costs and latency, while lower costs often come at the expense of intelligence. Deep optimization between Agents and models is breaking the “impossible trinity” of performance, cost, and speed. As model intelligence density continues to increase, along with two-way optimization between Qwen Office and its models, Agents are about to leave token anxiety behind and usher in an era of “abundant supply at full capacity.”
This article was provided by Qwen. QbitAI is authorized to republish it; all views expressed are those of the original author.