Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Embedding Model

BGE-M3

BGE-M3 is a multilingual embedding model from BAAI supporting over 100 languages with an 8192-token context window.

PublisherBAAI
TypeEmbedding
Context Window8,192 tokens
ReleasedMay 2024
Input$0.01/MTok
ProviderDeepInfra

Multilingual text embeddings across 100+ languages

BGE-M3 is a text embedding model developed by the Beijing Academy of Artificial Intelligence (BAAI). It is designed to generate dense vector representations of text and supports over 100 languages, making it suitable for multilingual retrieval and semantic similarity tasks. The model accepts input sequences of up to 8192 tokens, which is notably longer than many embedding models, allowing it to process full documents rather than just short passages.

BGE-M3 is distinguished by its support for three retrieval methods within a single model: dense retrieval, sparse retrieval, and multi-vector (ColBERT-style) retrieval. This makes it flexible for a range of information retrieval pipelines, including hybrid search systems. It is well-suited for use cases such as semantic search, document ranking, cross-lingual retrieval, and building retrieval-augmented generation (RAG) pipelines.

What BGE-M3 supports

Multilingual Embeddings

Generates text embeddings across more than 100 languages in a single model, enabling cross-lingual semantic search and retrieval without separate language-specific models.

Long Context Input

Accepts input sequences of up to 8192 tokens, allowing full documents or long passages to be embedded without truncation.

Dense Retrieval

Produces dense vector representations suitable for approximate nearest-neighbor search in vector databases like FAISS or Pinecone.

Sparse Retrieval

Supports sparse retrieval output, enabling compatibility with traditional keyword-based search systems and hybrid retrieval pipelines.

Multi-Vector Retrieval

Implements ColBERT-style multi-vector retrieval, where each token in a passage is represented separately to improve fine-grained matching accuracy.

RAG Pipeline Support

Designed to serve as the retrieval component in retrieval-augmented generation (RAG) systems, providing semantically relevant document chunks to a downstream language model.

Ready to build with BGE-M3?

Get Started Free

Common questions about BGE-M3

What is the context window for BGE-M3?

BGE-M3 supports an input context window of 8192 tokens, allowing it to embed long documents in a single pass.

How many languages does BGE-M3 support?

BGE-M3 supports over 100 languages, making it suitable for multilingual and cross-lingual retrieval tasks.

What retrieval methods does BGE-M3 support?

BGE-M3 supports three retrieval paradigms within a single model: dense retrieval, sparse retrieval, and multi-vector (ColBERT-style) retrieval. This allows it to be used in hybrid search pipelines.

What is the pricing for BGE-M3 on DeepInfra?

Pricing information is not included in the available metadata. You should check the DeepInfra model page directly at https://deepinfra.com/BAAI/bge-m3 for current pricing details.

What is BGE-M3 best used for?

BGE-M3 is best suited for semantic search, document ranking, cross-lingual information retrieval, and as the retrieval component in RAG (retrieval-augmented generation) pipelines.

Does BGE-M3 generate text or only embeddings?

BGE-M3 is an embedding model only. It produces vector representations of input text and does not generate natural language responses.

Start building with BGE-M3

No API keys required. Create AI-powered workflows with BGE-M3 in minutes — free.