Skip to main content

Overview

EmbeddingGemma 2 is Google’s multimodal embedding model. It maps text (including code), images, video, audio, and interleaved combinations of them into one shared, L2-normalized 768-dimensional vector space. SGLang detects the EmbeddingGemma2Model architecture and configures it automatically:
  • enables embedding mode, so --is-embedding is not needed;
  • selects the Triton attention backend, which implements the bidirectional sliding window;
  • disables RadixAttention prefix caching and chunked prefill, because bidirectional attention makes every token depend on the whole input;
  • caps the context length at 8,192 tokens.

Prerequisites

  • An NVIDIA CUDA GPU. The model weighs about 1.5 GB in BF16, so any data-center GPU works.
  • A Transformers build that includes EmbeddingGemma 2 (model_type: embedding_gemma2). Older releases fail at startup with Transformers does not recognize this architecture.
  • An SGLang build that includes EmbeddingGemma 2 support:
Run the model in BF16 (the checkpoint default) or FP32. Do not use FP16: the model’s activations exceed its range and produce NaN or degraded embeddings.

Start the server

Load only the encoders you need

The vision and audio encoders load only when their modality is allowed. Set a modality’s per-request limit to 0 to skip loading its encoder. Images and video share the vision encoder.
A request for a disabled modality fails with HTTP 400, for example Image count 1 exceeds limit 0 per request.

Enable Matryoshka dimensions

To let clients request shorter vectors through the dimensions field, declare the supported sizes when starting the server:
SGLang truncates the pooled vector and then re-normalizes it, as the model card requires. Other sizes are rejected with HTTP 400. Queries and documents must use the same dimension.

Task prompts

EmbeddingGemma 2 is trained with short task prefixes on text inputs. Prepend them yourself; images, video, and audio take no prefix. For documents without a title, use title: none.

Create embeddings

Native /encode API

/encode supports every modality, including several media items in one embedding. Text, one embedding per string:
An image, video, or audio clip on its own. SGLang inserts the placeholder for you:
Use video_data or audio_data the same way. Audio should be mono at 16 kHz. Interleaved input, one embedding for the whole sequence. Mark each item’s position with <|image|>, <|video|>, or <|audio|>, and pass the items in the same order:

How requests map to embeddings

Requests are rejected with HTTP 400 instead of being silently changed when:
  • the number of <|image|>, <|video|>, or <|audio|> placeholders differs from the number of items of that modality (text without any placeholder plus an image is also a mismatch);
  • a batched media list has a different length than the text list.

OpenAI-compatible /v1/embeddings

Each input item produces one embedding. In this API, an item carries at most one image or video and no audio. For several media items in one embedding, or for audio, use /encode.

Input limits

All modalities share the 8,192-token context. Requests that exceed it fail with HTTP 400 (The input (N tokens) is longer than the model's context length (8192 tokens).). Video is sampled at 1 frame per second and capped at 32 frames by default, spread uniformly over longer clips, so a video of any length fits by default. A 2-minute 720p clip, for example, uses 3,906 tokens. Only the sampled frames are decoded. To change the sampling rate or frame cap for every request, use --mm-process-config:
The number of frames that fit depends on resolution. At 720p, 64 frames fit; 70 frames exceed the context.