GLM 4.7 Flash
Provider:Z.ai (Zhipu)
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Key specs
- Context:202.8K
- Max output:16.4K
- Tokenizer:Other
- Released:2026-01-19
Token pricing
- Input price:$0.060 / 1M tokens
- Output price:$0.400 / 1M tokens
- Cache read:$0.010 / 1M tokens
- Blended price:$0.145 / 1M tokens
Modalities
Text
Use Cases
- Deep Reasoning & Analysis:Process long legal docs, academic papers, financial reports; multi-hop reasoning and evidence-chain tracing across huge contexts.(≈ $0.0032/call)
- Chat & Support:FAQ auto-response, multi-turn dialogue, emotion detection — e-commerce, finance, gov.(≈ $0.0020/turn)
When to pick this model
Signals derived from public pricing and spec fields — not a benchmark. Validate with official docs and your own eval.
- Low input + output price: solid fit for high-frequency calls and batch processing.
- Low blended price: budget-friendly for fleet-scale deployment.
FAQ
- What is the token pricing of GLM 4.7 Flash? Input price $0.060 / 1M tokens, output price $0.400 / 1M tokens, blended about $0.145 / 1M tokens. Refer to Z.ai (Zhipu)’s official page for the exact rate.
- What context window does GLM 4.7 Flash support? Context window is 202.8K, max output about 16.4K, suitable for long documents, multi-turn chat and complex reasoning.
- What scenarios is GLM 4.7 Flash best for? Based on its capability and pricing, GLM 4.7 Flash fits best: Deep Reasoning & Analysis, Chat & Support. See the Use Cases section for details and cost estimates.
- What input/output modalities does GLM 4.7 Flash support? Supports Text modality.
- Is GLM 4.7 Flash a free model? GLM 4.7 Flash is billed per token, not a free model.
- Which provider offers GLM 4.7 Flash? GLM 4.7 Flash is offered by Z.ai (Zhipu).
More models from Z.ai (Zhipu)
Full info:GLM 4.7 Flash