Briefing

Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context

ai-dev
OpenAI

Replace your current embedding model with granite-embedding-97m-multilingual-r2 or granite-embedding-311m-multilingual-r2, update your framework config to use the new model name, and test retrieval performance on your multilingual data.

What to do now

Replace your current embedding model with granite-embedding-97m-multilingual-r2 or granite-embedding-311m-multilingual-r2, update your framework config to use the new model name, and test retrieval performance on your multilingual data.

Summary

Granite Embedding Multilingual R2 was released on May 14, 2026 and introduces two new models: granite-embedding-311m-multilingual-r2 (311 M parameters, 768‑dimensional embeddings, 32 K‑token context, Matryoshka support) and granite-embedding-97m-multilingual-r2 (97 M parameters, 384‑dimensional embeddings, 32 K‑token context). Both models cover 200+ languages, with 52 languages receiving enhanced retrieval training, and include code retrieval for nine programming languages (Python, Go, Java, JavaScript, PHP, Ruby, SQL, C, C++). The architecture has shifted from XLM‑RoBERTa to ModernBERT, enabling rotary position embeddings and FlashAttention 2.0, while the 311 M model uses the Gemma 3 tokenizer (262 K tokens) and the 97 M model uses a pruned GPT‑OSS tokenizer (180 K tokens). Benchmark results show the 97 M model scoring 60.3 on MTEB Multilingual Retrieval and the 311 M model scoring 65.2, a 13‑point gain over R1. The models ship with ONNX and OpenVINO weights for CPU inference, are licensed under Apache 2.0, and can be dropped into LangChain, LlamaIndex, Haystack, or Milvus with a single‑line model name change.

The training pipeline includes knowledge distillation from Granite 3.3 Instruct and Mistral v0.2 Instruct, contrastive fine‑tuning on multilingual retrieval pairs, and Matryoshka representation learning to allow dimensional truncation with minimal quality loss. IBM‑curated datasets and governance processes were used to filter public web data, avoiding non‑commercial licenses and reducing downstream risk. The release positions Granite as the highest‑scoring open multilingual embedding under 100 M parameters and a top‑tier model under 500 M parameters.

Key changes

  • Two new models: granite-embedding-311m-multilingual-r2 (311 M params, 768‑dim, 32 K context, Matryoshka) and granite-embedding-97m-multilingual-r2 (97 M params, 384‑dim, 32 K context)
  • Context window increased from 512 tokens (R1) to 32,768 tokens (R2)
  • Architecture switched to ModernBERT with rotary position embeddings and FlashAttention 2.0
  • Tokenizer changed: 311 M uses Gemma 3 (262 K tokens), 97 M uses pruned GPT‑OSS (180 K tokens)
  • Matryoshka embeddings allow truncation to 512/384/256/128 dims with minimal loss
  • Includes code retrieval for nine programming languages
  • ONNX and OpenVINO weights provided for CPU inference
  • Benchmark scores: 97 M 60.3, 311 M 65.2 on MTEB Multilingual Retrieval

Affects

enterprise

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting