Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

GLM 4.6

GLM 4.6 is a text generation model from Z.ai with a 200,000-token context window and configurable reasoning effort.

PublisherZ.ai
TypeText
Context Window200,000 tokens
ReleasedSeptember 2025
Input$0.50/MTok
Output$2.00/MTok
ProviderDeepInfra

Long-context text generation with adjustable reasoning

GLM 4.6 is a large language model developed by Z.ai (formerly Zhipu AI) and released in September 2025. It is a chat-oriented text generation model hosted via DeepInfra and supports a context window of up to 200,000 tokens, making it suitable for tasks that require processing long documents or extended conversations. The model exposes a reasoning effort toggle, allowing users to adjust how much computational effort the model applies to a given query.

GLM 4.6 is part of Z.ai's GLM (General Language Model) series, which has been developed with a focus on bilingual Chinese and English language understanding. The model generates responses up to 16,384 tokens in length, which accommodates detailed outputs such as long-form writing, summarization of large documents, and multi-step reasoning tasks. Its configurable reasoning effort input makes it adaptable to workflows where response depth or latency trade-offs matter.

What GLM 4.6 supports

Long Context Window

Processes inputs up to 200,000 tokens, enabling analysis of lengthy documents, codebases, or extended conversation histories in a single pass.

Adjustable Reasoning

Exposes a reasoning effort toggle that lets users control how deeply the model reasons before responding, useful for balancing speed against answer quality.

Long-Form Text Output

Generates responses up to 16,384 tokens, supporting detailed outputs such as reports, summaries, and multi-step explanations.

Bilingual Language Support

Built on Z.ai's GLM series with a focus on both Chinese and English, making it well-suited for bilingual tasks and cross-language content generation.

Chat-Oriented Generation

Designed for conversational use cases with a chat model architecture, supporting multi-turn dialogue and instruction-following interactions.

Ready to build with GLM 4.6?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines78.4%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)63.2%
LiveCodeBenchReal-world coding tasks from recent competitions56.1%
HLEQuestions that challenge frontier models across many domains5.2%
SciCodeScientific research coding and numerical methods33.1%

Common questions about GLM 4.6

What is the context window size for GLM 4.6?

GLM 4.6 supports a context window of 200,000 tokens, allowing it to process very long documents or extended conversations in a single request.

What is the maximum response length GLM 4.6 can generate?

The model can generate responses up to 16,384 tokens in length per request.

What does the reasoning effort toggle do?

The reasoning effort input is a configurable toggle that adjusts how much reasoning the model applies before producing a response. Setting it higher may improve answer quality on complex tasks, while lower settings can reduce latency.

Who developed GLM 4.6 and when was it released?

GLM 4.6 was developed by Z.ai (formerly known as Zhipu AI) and released in September 2025. It is part of the GLM (General Language Model) series.

Does GLM 4.6 support image or video inputs?

Based on the available metadata, GLM 4.6 does not have confirmed support for image or video analysis inputs. It is a text generation model.

Is pricing information available for GLM 4.6?

Published pricing for GLM 4.6 is not listed in the current metadata. You can use the model on MindStudio without managing API keys directly.

What people think about GLM 4.6

Community reception on r/LocalLLaMA has been broadly positive, with the GLM-4.6 announcement post receiving over 400 upvotes and 81 comments. Users have highlighted its large context window, open-weight availability under the MIT license, and performance on coding and agentic tasks as notable strengths.

A separate thread about the GLM-4.6-Air variant attracted over 600 upvotes, suggesting interest in lighter deployable versions of the model. Discussions also reference the subsequent GLM-4.7 release, indicating an active development cadence that some users follow closely for local deployment use cases.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size16,384 tokens
Reasoning EffortToggle Group
Default: medium
NoneLowMediumHigh

Start building with GLM 4.6

No API keys required. Create AI-powered workflows with GLM 4.6 in minutes — free.