Qwen3 VL 32B Instruct

Provider:Qwen

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Key specs

Token pricing

Modalities

Text, Image

Use Cases

When to pick this model

Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.

FAQ

More models from Qwen

Full info:Qwen3 VL 32B Instruct