Particle.news

Alibaba Publishes Qwen3.8‑Max Weights with Commercial Licence and Low API Pricing

The move lets companies self‑host a very large multimodal model under rules that require big service providers to pay Alibaba Cloud for commercial use.

Overview

  • Alibaba has released the core weights for Qwen3.8‑Max, a 2.4 trillion‑parameter sparse Mixture‑of‑Experts model that activates about 95 billion parameters per token and supports a one‑million‑token context window for text, images and video.
  • The company published commercial licence rules that exempt internal-only use but require a paid licence for 'model as a service' or AI assistant providers whose aggregate revenue exceeds US$50 million over any 12‑month period.
  • Alibaba set low public API prices and regional tiers, including 12 yuan per million input tokens and 36 yuan per million output tokens in China and US$2 per million input tokens and US$6 per million output tokens internationally.
  • Alibaba reports vendor benchmarks and hardware claims — an OSWorld‑Verified score of 86.1 and per‑GPU inference speeds above 4,000 tokens per second on Nvidia GB300 NVL72 racks — but those performance and long‑run autonomy demonstrations remain vendor‑reported and lack independent replication.
  • The release expands self‑hosting options for developers and enterprises, raises questions about governance and export controls, and fits a strategy to convert open‑weight interest into paid cloud and enterprise deals as other Chinese labs also publish large open models.