Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation ModelDeprecated

Sonar Small Chat

Sonar Small Chat is a text generation model from Perplexity built on Llama 3.1 with a 127,072 token context window.

PublisherPerplexity
TypeText
Context Window127,072 tokens
Replaced bySonar

Lightweight chat model with citation support

Sonar Small Chat, formally named llama-3.1-sonar-small-128k-chat, is a chat-oriented text generation model developed by Perplexity AI. It is part of the Llama 3.1 Sonar model family, which Perplexity built on top of Meta's Llama 3.1 architecture. The model supports a context window of 127,072 tokens and a maximum response size of 32,768 tokens, making it suited for extended conversational exchanges and document-length inputs.

Sonar Small Chat is designed for conversational use cases where efficiency and speed are priorities over raw scale. A notable feature is its optional citation and image return inputs, which allow developers to configure whether the model surfaces references alongside its responses. This model has since been deprecated and superseded by newer members of the Sonar family, so developers starting new projects should consult Perplexity's current model offerings.

What Sonar Small Chat supports

Long Context Window

Processes up to 127,072 tokens in a single request, enabling use with lengthy documents or extended multi-turn conversations.

Citation Return

Supports an optional return_citations input that instructs the model to include source references alongside its generated responses.

Image Return

Supports an optional return_images input that can surface relevant images as part of the model's response output.

Chat Text Generation

Generates conversational text responses with a maximum output size of 32,768 tokens per response.

Ready to build with Sonar Small Chat?

Get Started Free

Common questions about Sonar Small Chat

What is the context window size for Sonar Small Chat?

Sonar Small Chat supports a context window of 127,072 tokens, allowing for long documents or extended conversations in a single request.

Is Sonar Small Chat still available to use?

Sonar Small Chat has a deprecated status. Developers starting new projects should check Perplexity's current model catalog for actively supported alternatives.

What underlying model is Sonar Small Chat based on?

Sonar Small Chat's full model name is llama-3.1-sonar-small-128k-chat, indicating it is built on Meta's Llama 3.1 architecture by Perplexity AI.

Does Sonar Small Chat support image inputs?

The metadata does not confirm image input support. However, the model does have a return_images output option, which can surface images alongside text responses.

What is the maximum response size for Sonar Small Chat?

The model supports a maximum response size of 32,768 tokens per generation.

Parameters & options

Max Temperature1.9
Max Response Size32,768 tokens
Return CitationsSelect

Determines whether or not a request to an online model should return citations.

Default: false
NoYes
Return ImagesSelect

Determines whether or not a request to an online model should return images.

Default: false
NoYes

Start building with Sonar Small Chat

No API keys required. Create AI-powered workflows with Sonar Small Chat in minutes — free.