Grok 2 Vision
Grok 2 Vision is a multimodal model from X.ai that processes images and text within a 32,768 token context window.
Multimodal vision model from X.ai
Grok 2 Vision is a vision-capable language model developed by X.ai, released in December 2024 under the model ID grok-2-vision-1212. It is designed to handle both image and text inputs, allowing it to analyze visual content alongside natural language prompts within a 32,768 token context window. The model is part of X.ai's Grok 2 family and is available as a first-party offering through X.ai's API.
Grok 2 Vision is suited for tasks that require understanding visual information in combination with text, such as image description, document analysis, and visual question answering. Its 32,768 token context window accommodates longer conversations and multi-turn interactions involving images. Developers can access the model via the API identifier grok-2-vision-1212 for integration into applications that require multimodal capabilities.
What Grok 2 Vision supports
Image Understanding
Analyzes and interprets image content provided alongside text prompts, enabling tasks like visual question answering and image description.
Multimodal Input
Accepts both image and text inputs in a single request, supporting workflows that combine visual and natural language data.
Long Context Window
Supports up to 32,768 tokens per context, allowing extended conversations and multi-turn interactions that include image and text content.
Document Analysis
Can process images of documents or structured visual content to extract and reason about the information they contain.
Text Generation
Generates natural language responses based on visual and textual inputs, producing coherent descriptions, answers, or summaries.
Ready to build with Grok 2 Vision?
Get Started FreeBenchmark scores
Scores represent accuracy — the percentage of questions answered correctly on each test.
| Benchmark | What it tests | Score |
|---|---|---|
| MMLU-Pro | Expert knowledge across 14 academic disciplines | 70.9% |
| GPQA Diamond | PhD-level science questions (biology, physics, chemistry) | 51.0% |
| MATH-500 | Undergraduate and competition-level math problems | 77.8% |
| AIME 2024 | American math olympiad problems | 13.3% |
| LiveCodeBench | Real-world coding tasks from recent competitions | 26.7% |
| HLE | Questions that challenge frontier models across many domains | 3.8% |
| SciCode | Scientific research coding and numerical methods | 28.5% |
Common questions about Grok 2 Vision
What is the context window size for Grok 2 Vision?
Grok 2 Vision supports a context window of 32,768 tokens, which applies to the combined input of text and image content.
Who developed Grok 2 Vision and when was it released?
Grok 2 Vision was developed by X.ai and released in December 2024 under the model identifier grok-2-vision-1212.
What types of inputs does Grok 2 Vision accept?
Grok 2 Vision is a multimodal model designed to accept both image and text inputs, making it suitable for tasks that require visual understanding alongside natural language processing.
What is the pricing for Grok 2 Vision?
Pricing information for Grok 2 Vision is not included in the available metadata. For current pricing details, refer to X.ai's official API documentation or pricing page.
What is the knowledge cutoff date for Grok 2 Vision?
The knowledge cutoff date for Grok 2 Vision is not specified in the available metadata. X.ai's official documentation would be the authoritative source for this information.
Documentation & links
Parameters & options
Explore similar models
Start building with Grok 2 Vision
No API keys required. Create AI-powered workflows with Grok 2 Vision in minutes — free.