Qwen has released Qwen3.8-Omni-Flash, Alibaba’s latest native multimodal model. This model is built for audio and video understanding. It uses text, images, audio, and video as input. The release also includes thinking support, custom function calling, and web search. With these tools in the loop, the model can handle large multimedia inputs and move the task forward with external tools.
Alibaba Cloud documents match these claims. They also show the model is listed in Alibaba Cloud Model Studio.
What is Qwen3.8-Omni-Flash ?
Qwen3.8-Omni-Flash is a native multimodal model. It can take in text, images, audio, and video. Developers do not have to set up separate systems for each file type. For example, it could take a video and written directions. The model could point out key moments.
It can also follow spoken parts. Then it can return a clear, organized summary. It also has a thinking mode. That mode is meant for slower, more careful reasoning. On top of that, Qwen offers function calling. It also includes web search. With these tools, the model can do more than just write a reply.
1M Tokens for Longer Multimedia Context
A key point is the 1 million token context window. Alibaba Cloud also shows a 1M token window. It gives a top input size of 991,808 tokens in non-thinking mode. In thinking mode, it lists 983,616 tokens. Still, “1 million tokens” does not mean unlimited audio or video time.
A context window is about how much tokenized text and signal the model can read in one request. Real limits for audio and video can shift based on things like file size. Media processing steps matter too. API limits can also matter. How the content is converted into tokens also matters.
In practice, the gain is simple. Qwen3.8-Omni-Flash can handle substantially larger multimedia inputs in a single context. Bigger multimedia inputs may fit in one run. That can help with long meetings. It also fits interviews and lectures. It can apply to recordings and videos where key details are spread out across multiple parts.
Qwen Reports More Than 26% Improvement

Qwen says Qwen3.8 Omni Flash shows an average gain of over 26%. Qwen links this result to 30 evaluations. It compares the outcome to Qwen3.5 Omni Plus. The numbers are Qwen’s own reported tests. They are not third-party benchmark results. So the claim shows Qwen’s internal view of progress. It is not a full ranking versus every other AI system. This comparison matters because Qwen3.8 Omni Flash is meant to do more than basic multimodal tasks. It is also built for long and complex multimedia inputs.
Thinking, Tool Calls, and Web Search
The model can use extra computers to handle harder reasoning tasks. In addition, app builders can set up tool calls. These tool calls let the model work with outside systems. Web search is part of this set too. Alibaba Cloud says the model can use the web_search tool.
This setup helps when a job needs more than just reading the prompt. It fits steps like finding key parts before writing the final output.
For example, you can review a meeting recording to pull out decisions and action items. You can first scan a long video to locate the key scenes, then create a report. Tasks can also mix audio, images, and written instructions when one answer needs more than one type of input.
Qwen3.8-Omni-Flash Pricing and Availability
Qwen3.8-Omni-Flash is offered via Alibaba Cloud Model Studio. The docs list the supported regions as Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia. Users pick the API key for the region they plan to use. Pricing changes by region and by input type. For international use, the Singapore page gives example rates.

It lists $0.43 for each 1 million input text tokens. Audio input is listed at $3.81 per 1 million tokens. Image and video input is listed at $0.78 per 1 million tokens. Text output is listed at $1.66 per 1 million tokens when the input is text only. When the input includes images, audio, or video, text output is listed at $3.06 per 1 million tokens.
What Qwen3.8-Omni-Flash Means for Multimedia AI
Qwen3.8-Omni-Flash is a multimodal release. It focuses on longer context plus wider tool support. It lists a 1-million-token context window. It also includes four input modes, plus a thinking feature. It also includes function calling and web search. Taken together, this is meant for multimedia work and task flows that use tools.
The 26%+ gain across 30 tests is a Qwen report. So it is best taken as a company result, not an outside benchmark. What stands out right now is the pairing of long-context multimedia handling with tools that help after the model reads the material.
For builders, this mix may fit use cases like long meeting reviews. It can also fit video research. Multimedia search and agent-style workflows are part of the same idea. In those cases, audio, video, images, and text are expected to be handled together.