reply 2 /
How EmbeddingGemma 2 Works
EmbeddingGemma 2 transforms text, code, images, video, and audio into a shared ve | Hanami
reply 2 /
How EmbeddingGemma 2 Works
EmbeddingGemma 2 transforms text, code, images, video, and audio into a shared vector space, making it easier to compare content, measure semantic similarity, and build multimodal search and retrieval systems.
How it works:
1. Text Encoding
A 24-layer Transformer uses local and global attention to capture context and meaning.
2. Vision Encoding
A 16-layer Vision Transformer converts images into visual tokens.
3. Audio Encoding
A 12-layer Conformer transforms audio features into audio tokens.
4. Multimodal Fusion
Visual and audio tokens are aligned with text tokens and processed together.
5. Embedding Projection
Token representations are pooled into a 768-dimensional, L2-normalized vector.
6. Matryoshka Embeddings
Embeddings can be reduced to 512, 256, or 128 dimensions to save storage and computation while retaining useful information.
The result is a unified embedding space where text, code, images, video, and audio can be compared using cosine similarity, enabling more flexible semantic search and multimodal retrieval applications.
Github: https://github.com/kyegomez/Open-EmbeddingGemma