On September 3, 2026, HUMAIN — the artificial intelligence company backed by Saudi Arabia's Public Investment Fund — unveiled humain-m3, a frontier Arabic-language model, at the LEAP conference in Riyadh. The model is available immediately as a research preview through HUMAIN Node, the company's platform for accessing advanced AI models and inference capabilities.
Commissioned by HUMAIN and delivered by Chinese AI firm MiniMax, humain-m3 is a 428-billion-parameter mixture-of-experts model built on the MiniMax-M3 lineage. HUMAIN further pre-trained the system on more than one trillion tokens of Arabic-native content, adapting the base model for Arabic rather than creating a new architecture from scratch. The model activates 23 billion parameters per token and was trained jointly on text, image, and video from the first step, with long-video understanding and native screen operation.
What's new
- Architecture: 428B total parameters, MoE activating 23B per token; natively multimodal mixture-of-experts engineered for agents
- Training: Further pre-trained on >1 trillion Arabic-native tokens atop MiniMax-M3; joint text, image, and video training
- Capabilities: Frontier tool use and computer use for long-horizon agent workflows in Arabic and English; three thinking modes (always-on, adaptive, off); long-video understanding and native screen operation
- Access tiers: Limited preview (Saudi alignment guardrail, thinking/streaming off, added latency) and research preview (full checkpoint, thinking, streaming, lower latency) via HUMAIN Node's no-code playground or OpenAI-compatible API
- Planned weight release: Targeted for next month under the MiniMax Community License, contingent on completing safety training and alignment
Benchmark results (HUMAIN's evaluation)
HUMAIN reported an equal-weighted average of 89.37% across seven public Arabic benchmarks for the previewed checkpoint, ahead of GPT-5.6 SOL at 87.30%, Opus 5 at 87.34%, and the MiniMax M3 reference checkpoint at 80.34%. The company said the previewed model leads five of the seven benchmarks. Individual scores: AlGhafa (core Arabic understanding) 86.45%, ArabicMMLU (native-Arabic broad knowledge) 90.70%, Arabic EXAMS (academic examinations) 67.67%, MadinahQA (Arabic language proficiency) 95.44%, AraTrust (truthfulness and trust) 97.53%, ALRAGE (retrieval-augmented generation) 94.63%, and Translated MMLU 93.20%. HUMAIN characterized these as its own evaluation of the previewed checkpoint and said its Arabic post-training adds nine points on average over the reference model.
Why it matters
The launch signals a shift in how Chinese open-source models are being adopted globally: rather than serving only as API endpoints, MiniMax-M3 is functioning as a technical foundation for a sovereign national AI platform. Saudi Arabia gains a model it can eventually run in local data centers without external cloud dependency, aligning with Vision 2030's technological autonomy goals. For the broader Arabic-speaking world — hundreds of millions of users — the model addresses a long-standing gap in frontier-level Arabic support. HUMAIN Node's single API, key, and billing across models, with global, in-Kingdom, or sovereign hosting options, also positions the platform as an enterprise gateway for the region.
Our take
The partnership validates MiniMax-M3's cross-lingual transfer and agent capabilities at national-project scale, but the 89.37% average comes from HUMAIN's own evaluation of a preview checkpoint — not an independent benchmark suite. The conditional weight release next month under the MiniMax Community License is the real milestone to watch; until weights are downloadable, sovereign deployment remains a promise rather than a capability.