Overview
EmbeddingGemma 2 is Google’s multimodal embedding model. It maps text (including code), images, video, audio, and interleaved combinations of them into one shared, L2-normalized 768-dimensional vector space.
SGLang detects the
EmbeddingGemma2Model architecture and configures it automatically:
- enables embedding mode, so
--is-embeddingis not needed; - selects the Triton attention backend, which implements the bidirectional sliding window;
- disables RadixAttention prefix caching and chunked prefill, because bidirectional attention makes every token depend on the whole input;
- caps the context length at 8,192 tokens.
Prerequisites
- An NVIDIA CUDA GPU. The model weighs about 1.5 GB in BF16, so any data-center GPU works.
- A Transformers build that includes EmbeddingGemma 2 (
model_type: embedding_gemma2). Older releases fail at startup withTransformers does not recognize this architecture. - An SGLang build that includes EmbeddingGemma 2 support:
Start the server
Load only the encoders you need
The vision and audio encoders load only when their modality is allowed. Set a modality’s per-request limit to0 to skip loading its encoder. Images and video share the vision encoder.
Image count 1 exceeds limit 0 per request.
Enable Matryoshka dimensions
To let clients request shorter vectors through thedimensions field, declare the supported sizes when starting the server:
Task prompts
EmbeddingGemma 2 is trained with short task prefixes on text inputs. Prepend them yourself; images, video, and audio take no prefix. For documents without a title, usetitle: none.
Create embeddings
Native /encode API
/encode supports every modality, including several media items in one embedding.
Text, one embedding per string:
video_data or audio_data the same way. Audio should be mono at 16 kHz.
Interleaved input, one embedding for the whole sequence. Mark each item’s position with <|image|>, <|video|>, or <|audio|>, and pass the items in the same order:
How requests map to embeddings
Requests are rejected with HTTP 400 instead of being silently changed when:
- the number of
<|image|>,<|video|>, or<|audio|>placeholders differs from the number of items of that modality (text without any placeholder plus an image is also a mismatch); - a batched media list has a different length than the
textlist.
OpenAI-compatible /v1/embeddings
image or video and no audio. For several media items in one embedding, or for audio, use /encode.
Input limits
All modalities share the 8,192-token context. Requests that exceed it fail with HTTP 400 (The input (N tokens) is longer than the model's context length (8192 tokens).).
Video is sampled at 1 frame per second and capped at 32 frames by default, spread uniformly over longer clips, so a video of any length fits by default. A 2-minute 720p clip, for example, uses 3,906 tokens. Only the sampled frames are decoded.
To change the sampling rate or frame cap for every request, use
--mm-process-config:
