Models and training

Qwen

stable definition
Machine-readable Download Markdown

Definition

Qwen is Alibaba's family of language and multimodal models. It includes general models and specialist variants for tasks such as coding or vision-language understanding. Some releases provide downloadable weights; others are offered through hosted services. The family name alone does not identify an exact checkpoint or its deployment terms.

Model names carry distinctions an operator needs to retain. A base model, an instruction-tuned model, and a model configured to generate extended reasoning can behave differently even within the same generation. Their published evaluations apply to the stated configuration.

Origin and attribution

Alibaba Cloud's original repository records Qwen-7B's release on August 3, 2023. The repository uses Qwen and the name Tongyi Qianwen, and identifies the Alibaba Cloud team as the developer. Later official repositories credit the Qwen team at Alibaba Group.

Scope and access

Qwen spans multiple generations and architectures. The April 2025 Qwen3 announcement released both dense and mixture-of-experts models and described thinking and non-thinking modes. The current Qwen3.8 repository documents a later generation. Max and Plus names also appear in the hosted catalog. Weight access needs to be checked for the selected release.

Licensing changes across releases. The original repository separates Apache-licensed code from model weights under the Tongyi Qianwen agreements. The Qwen3 announcement applies Apache 2.0 to its listed downloadable models. The current repository directs readers to the license accompanying each weight release. A repository's code license cannot settle the license of every model it discusses.

Operational significance

Select and evaluate the checkpoint for the actual task, language, input modality, and serving configuration. Preserve its tokenizer and prompt template, and record any quantization or fine-tuning. Those choices can change results independently of the family name.

For local deployment, check compute requirements and the origin of downloaded artifacts. For a hosted service, record the endpoint's model identity and version behavior. Downloadable weights allow direct experimentation, but they do not by themselves supply the complete data and code needed to reproduce training.

Distinguish it from nearby terms

  • A large language model is a model category; Qwen names a particular vendor family.
  • Open-weight describes artifact access. Open-source AI adds broader rights and materials; the label needs a release-specific assessment.
  • Mixture of experts is an architecture used by some Qwen models, not a synonym for the family.
  • DeepSeek-R1-Distill-Qwen models are DeepSeek's adaptations of Qwen source models. Their lineage matters alongside the publisher's name.

Check your understanding

Two endpoints both say Qwen. One serves an instruction-tuned local checkpoint and the other a hosted Max model. What information would make their cost, quality, access rights, and reproducibility comparable?