Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

GLM 4.6V

GLM 4.6V is a multimodal text generation model from Z.ai with a 131,072 token context window and configurable reasoning effort.

PublisherZ.ai
TypeText
Context Window131,072 tokens
ReleasedDecember 2025
Input$0.30/MTok
Output$0.90/MTok
ProviderDeepInfra

Multimodal text generation with adjustable reasoning

GLM 4.6V is a text generation model developed by Z.ai (formerly Zhipu AI) and made available through DeepInfra. It belongs to the GLM (General Language Model) family and supports a context window of 131,072 tokens, with a maximum response size of 16,384 tokens. The model includes a configurable reasoning effort setting, allowing users to adjust how much computational effort the model applies when generating responses.

GLM 4.6V is suited for tasks that benefit from long-context understanding, such as document analysis, extended conversations, and complex instruction following. The adjustable reasoning effort input gives developers control over the trade-off between response speed and depth of reasoning. It is hosted on DeepInfra and accessible through MindStudio without requiring separate API key management.

What GLM 4.6V supports

Long Context Window

Processes up to 131,072 tokens in a single request, enabling analysis of lengthy documents or extended multi-turn conversations.

Adjustable Reasoning

Exposes a reasoning effort toggle that lets users control how deeply the model reasons before generating a response.

Extended Response Output

Supports responses of up to 16,384 tokens, accommodating detailed answers, long-form content, and structured outputs.

Instruction Following

Handles complex, multi-step instructions across a wide range of text generation tasks including summarization, Q&A, and drafting.

Multilingual Support

The GLM model family is trained on multilingual data with particular strength in Chinese and English language tasks.

Ready to build with GLM 4.6V?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines78.4%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)63.2%
LiveCodeBenchReal-world coding tasks from recent competitions56.1%
HLEQuestions that challenge frontier models across many domains5.2%
SciCodeScientific research coding and numerical methods33.1%

Common questions about GLM 4.6V

What is the context window size for GLM 4.6V?

GLM 4.6V supports a context window of 131,072 tokens, allowing it to process large documents or long conversation histories in a single request.

What does the reasoning effort setting do?

The reasoning effort toggle lets you adjust how much computational effort the model applies when generating a response. Higher effort may produce more thorough reasoning, while lower effort can yield faster responses.

What is the maximum response length?

The model can generate responses of up to 16,384 tokens per request.

Who developed GLM 4.6V and where is it hosted?

GLM 4.6V was developed by Z.ai (published under the zai-org organization) and is served through DeepInfra as the infrastructure provider.

Is pricing information available for GLM 4.6V?

Pricing details are not published in the current model metadata. You can check MindStudio or DeepInfra directly for the latest pricing information.

When was GLM 4.6V released?

GLM 4.6V was released in December 2025 and added to the MindStudio catalog in January 2026.

What people think about GLM 4.6V

Community reception on r/LocalLLaMA was positive at launch, with both the 106B and 9B Flash variants receiving several hundred upvotes and active discussion threads. Users highlighted the native visual function calling capability and the availability of a locally runnable 9B version as notable aspects of the release.

Some community members discussed hardware requirements for running the larger 106B model, with a later thread specifically covering deployment on dual RTX Pro 6000 setups with 192 GB VRAM using vLLM. The Flash variant drew particular interest from users focused on local inference and lower-resource deployments.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size16,384 tokens
Reasoning EffortToggle Group
Default: medium
NoneLowMediumHigh

Start building with GLM 4.6V

No API keys required. Create AI-powered workflows with GLM 4.6V in minutes — free.