Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

GLM 4.7

GLM 4.7 is a text generation model from Z.ai with a 131072-token context window and configurable reasoning effort.

PublisherZ.ai
TypeText
Context Window131,072 tokens
ReleasedDecember 2025
Input$0.40/MTok
Output$1.75/MTok
ProviderDeepInfra

Long-context text generation with adjustable reasoning

GLM 4.7 is a large language model developed by Z.ai (formerly Zhipu AI) and released in December 2025. It is part of the GLM (General Language Model) family and is available through DeepInfra as a chat-oriented text generation model. The model supports a 131072-token context window and allows up to 16384 tokens per response, making it suited for tasks that involve long documents or extended conversations.

A notable feature of GLM 4.7 is its configurable reasoning effort, which lets users adjust how much computational reasoning the model applies to a given prompt. This makes it flexible for use cases ranging from quick responses to more deliberate, multi-step reasoning tasks. The model is published under the identifier zai-org/GLM-4.7 and is accessible on MindStudio without requiring separate API key management.

What GLM 4.7 supports

Long Context Window

Processes up to 131072 tokens in a single context, enabling analysis of lengthy documents or extended multi-turn conversations.

Adjustable Reasoning

Exposes a reasoning effort toggle that lets users control how deeply the model reasons before responding, balancing speed against thoroughness.

Extended Response Length

Supports a maximum response size of 16384 tokens, accommodating detailed outputs such as long-form drafts or structured reports.

Chat-Optimized Generation

Built as an llm_chat model type, designed for instruction-following and multi-turn dialogue use cases.

Ready to build with GLM 4.7?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines85.6%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)85.9%
LiveCodeBenchReal-world coding tasks from recent competitions89.4%
HLEQuestions that challenge frontier models across many domains25.1%
SciCodeScientific research coding and numerical methods45.1%
AIME 2025American math olympiad problems (2025)95.7%
SWE-bench VerifiedReal GitHub issues requiring multi-file code fixes73.8%

Common questions about GLM 4.7

What is the context window size for GLM 4.7?

GLM 4.7 supports a context window of 131072 tokens, allowing it to process long documents or extended conversations in a single pass.

What is the maximum response length?

The model can generate up to 16384 tokens per response.

What does the reasoning effort toggle do?

The reasoning effort input is a configurable toggle that adjusts how much reasoning the model applies before generating a response, letting you trade off between speed and depth of reasoning.

Who developed GLM 4.7 and when was it released?

GLM 4.7 was developed by Z.ai (published under zai-org) and released in December 2025. It is served through DeepInfra on MindStudio.

Does GLM 4.7 support image or video inputs?

Based on the available metadata, GLM 4.7 does not have confirmed support for image or video analysis inputs; it is a text-only generation model.

Is pricing information available for GLM 4.7?

No published pricing information is currently listed for GLM 4.7 in the MindStudio catalog. Check MindStudio or DeepInfra directly for current pricing details.

What people think about GLM 4.7

Community reception on r/LocalLLaMA has been generally positive, with the GLM-4.7 Flash variant thread drawing 755 upvotes and 232 comments, indicating strong interest in running the model locally. Users have highlighted its benchmark performance on coding and reasoning tasks as notable for an open-source release.

Some discussion has shifted toward the subsequent GLM-5 release, which garnered even more engagement, suggesting that a portion of the community views GLM-4.7 as a stepping stone rather than a long-term target. Common use cases mentioned include local deployment for agentic coding workflows and multi-step automation tasks.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size16,384 tokens
Reasoning EffortToggle Group
Default: medium
NoneLowMediumHigh

Start building with GLM 4.7

No API keys required. Create AI-powered workflows with GLM 4.7 in minutes — free.