Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 18, 2026, a native omnimodal model built to move audio-visual AI from perception into agentic execution. The model accepts text, image, audio, and video inputs with a 1-million-token context window and is available on the Qianwen AI Platform and via Alibaba Cloud Model Studio APIs.

Qwen3.8-Omni-Flash targets real-world productivity scenarios — video editing, music video creation, film commentary, audio-visual summarization, and real-time conversation — by combining omnimodal understanding with planning, tool use, and end-to-end task completion. The release also includes Qwen-MM-Plugins for on-demand perception and workflow execution, and Qwen-Live Harness as an open-source runtime for continuous, real-time omnimodal interaction.

Confirmed

  • Model: Qwen3.8-Omni-Flash, next-generation native omnimodal model from Alibaba's Qwen team.
  • Modalities: Text, image, audio, and video input; text output.
  • Context window: 1 million tokens (native), with text performance claimed comparable to same-size text-only models.
  • Agentic capabilities: Planning, tool calling, and creative work delivery across coding, GUI operation, and audio-video workflows.
  • Benchmarks (vendor-reported): Average score across 29 evaluations improves >25% over Qwen3.5-Omni-Plus; WildClawBench-MM +36.5 points, AgenticVBench +22.3 points, UniClawBench 69.6. LongAudioSpan +8.3, OmniVideoBench +9.6, OmniCap-IF CSR +8.5 and ISR +14.1. AliMeeting DER drops from 88.11 to 3.35, cpWER from 89.61 to 17.18.
  • Pricing (vendor methodology): API price per hour of audio input decreases >98% vs prior generation; audio-visual input >93% reduction. Hourly estimates based on 30× two-minute source cost; audio-visual at 720p/1 fps. Gemini 3.8 Flash priced at media_resolution=high; Seed 2.0 Lite at max_frame_tokens=384; FX rate 1 USD = 6.7191 CNY.
  • Availability: Live as of September 18, 2026 on Qianwen AI Platform (chat.qwen.ai) and Alibaba Cloud Model Studio API (standard and realtime endpoints); also listed on Hugging Face and ModelScope.
  • Companion tools: Qwen-MM-Plugins (expanded for long-form audio/video perception and workflow execution), Qwen-Live Harness (open-source runtime for continuous real-time omnimodal interaction).

Unknown

  • Independent replication of the claimed benchmark gains (WildClawBench-MM, AgenticVBench, UniClawBench, OmniVideoBench, AliMeeting) has not been published.
  • Real-world API latency, throughput SLAs, and cost-per-agent-job under multi-turn tool-use workloads are not disclosed; list price per token is not total cost per agent job when reasoning steps and tool calls increase.
  • Exact model size (total and active parameters), MoE expert count, and training compute are not specified in the launch materials.
  • Rollout timeline for the realtime API beyond the Model Studio preview, and regional availability outside mainland China, are not detailed.
  • Whether the 1M-token context window is fully supported for all four modalities simultaneously in production, or only for text with shorter audio/video segments, is not clarified.

Our take

Alibaba is betting the next omnimodal frontier is the plumbing that turns hour-long media into agentic workflows — evidence gathering, editing, translation — without human stitching. The 1M-token window and open-sourced Live Harness signal intent to own the runtime layer for continuous audio-video interaction, where Gemini Live and OpenAI Realtime have set expectations but not standardized tooling. If token-cost reductions hold at scale, long-form video agent economics shift from prototype to production line, but missing independent benchmarks and SLA data mean buyers should pilot before committing volume.

Sources