Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

GPT OSS 20B

GPT OSS 20B is an open-source text generation model from OpenAI with a 128,000-token context window, served via Groq.

PublisherOpenAI
TypeText
Context Window128,000 tokens
ReleasedAugust 2025
Input$0.10/MTok
Output$0.50/MTok
ProviderGroq
OPEN SOURCEVERY FAST

Open-source text generation at high speed

GPT OSS 20B is a 20-billion-parameter open-source language model developed by OpenAI and released in August 2025. It is served through Groq's inference infrastructure, which is optimized for low-latency throughput, and supports a 128,000-token context window with a maximum response size of 32,768 tokens. The model is available under an open-source license, making its weights accessible for inspection and deployment outside of proprietary APIs.

GPT OSS 20B is suited for text generation tasks that benefit from a large context window and fast inference, such as document summarization, multi-turn conversation, and code assistance. Its 20B parameter scale positions it as a mid-size model that balances capability with inference efficiency. Developers looking for an OpenAI-published model with open weights and high-speed serving will find this a practical option on the MindStudio platform.

What GPT OSS 20B supports

Large Context Window

Processes up to 128,000 tokens of input in a single request, enabling long documents, extended conversations, or large codebases to be handled without truncation.

High-Speed Inference

Served on Groq's LPU hardware, which is designed to deliver low-latency token generation compared to standard GPU-based inference.

Open-Source Weights

Released as an open-source model by OpenAI, allowing developers to inspect, download, and deploy the model weights independently.

Text Generation

Generates coherent, contextually relevant text for tasks including summarization, drafting, question answering, and multi-turn dialogue.

Long Response Output

Supports a maximum response size of 32,768 tokens, allowing detailed, long-form outputs in a single generation call.

Ready to build with GPT OSS 20B?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines74.8%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)68.8%
LiveCodeBenchReal-world coding tasks from recent competitions77.7%
HLEQuestions that challenge frontier models across many domains9.8%
SciCodeScientific research coding and numerical methods34.4%

Common questions about GPT OSS 20B

What is the context window size for GPT OSS 20B?

GPT OSS 20B supports a context window of 128,000 tokens, meaning it can process up to 128,000 tokens of combined input and conversation history in a single request.

What is the maximum response length?

The model can generate up to 32,768 tokens in a single response, which is suitable for long-form content such as detailed reports or extended code outputs.

Is GPT OSS 20B open source?

Yes. GPT OSS 20B is tagged as open source, meaning OpenAI has made the model weights publicly available. This distinguishes it from OpenAI's proprietary API-only models.

Who provides the inference for GPT OSS 20B on MindStudio?

Inference is provided by Groq. Groq's LPU-based infrastructure is designed for high-speed, low-latency token generation, which is reflected in the model's 'VERY FAST' tag.

What is the pricing for GPT OSS 20B?

Pricing information for GPT OSS 20B has not been published in the available metadata. Check MindStudio's platform or Groq's pricing page for current rate information.

When was GPT OSS 20B released?

GPT OSS 20B was released in August 2025 and was added to the MindStudio catalog on August 6, 2025.

What people think about GPT OSS 20B

Community reception on r/LocalLLaMA has been notably positive, with the open-weight release generating significant discussion and over 2,000 upvotes on the announcement thread. Users have praised the model's ability to run on consumer hardware, including older CPUs without dedicated NVIDIA GPUs.

A recurring theme in community threads is the model's efficiency on low-resource hardware, with users reporting usable inference speeds on machines as modest as an 8th-gen Intel i3. Some threads focus on benchmarking performance across specific GPU configurations, such as the RTX Pro 6000 Blackwell and RTX 5090M.

View more discussions →

Parameters & options

Max Temperature2
Max Response Size32,768 tokens

Start building with GPT OSS 20B

No API keys required. Create AI-powered workflows with GPT OSS 20B in minutes — free.