Instructions to use litert-community/embeddinggemma-2-text-vision-440m-litert-lm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/embeddinggemma-2-text-vision-440m-litert-lm with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli # A single .litertlm file in the repo is picked automatically; otherwise the CLI asks which one to run # (or pass its name right after the repo id). litert-lm run \ --from-huggingface-repo=litert-community/embeddinggemma-2-text-vision-440m-litert-lm \ --prompt="Write me a poem"
- Notebooks
- Google Colab
- Kaggle
litert-community/embeddinggemma-2-text-vision-440m-litert-lm
Main Model Card: google/embeddinggemma-2
EmbeddingGemma 2 Text Vision 440M is a lightweight, multimodal (text and vision) variant of the EmbeddingGemma 2 model family, tailored for efficient on-device execution using LiteRT. By combining the 270M parameter text backbone with a modular 170M parameter vision encoder, this 440M parameter model maps text, code, and images into a single, unified 768-dimensional embedding space. It captures deep cross-modality semantics—allowing visual concepts and corresponding text queries to align co-dependently—making it perfect for on-device applications such as image search, visual document retrieval, and image classification.
Try EmbeddingGemma 2 with LiteRT-LM
Ready to integrate this into your product? Get started here.
Try EmbeddingGemma 2 with MediaPipe
MediaPipe uses EmbeddingGemma 2 and LiteRT-LM to enable high-level, cross-platform tasks which you can experience firsthand with the following interactive web demos:
- Semantic image search powered by MediaPipe Universal Embedder and Semantic Retriever task.
- Decision making powered by MediaPipe Decision Maker.
Model Specifications
| EmbeddingGemma 2 Text 270M |
EmbeddingGemma 2 Text Vision 440M |
EmbeddingGemma 2 740M |
|
|---|---|---|---|
| Supported Modalities | Text | Text, Images | Text, Images, Video, Audio |
| Parameters | Total: 270M Transformer: 130M Embeddings: 140M |
Total: 440M Transformer: 130M Embeddings: 140M Vision Encoder: 170M |
Total: 740M Transformer: 130M Embeddings: 140M Vision Encoder: 170M Audio Encoder: 300M |
| Quantization Scheme | Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) |
Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) Vision Encoder: int8 per-channel (QAT) |
Transformer: int4 per-channel (QAT) Embeddings: int4 per-channel (QAT) Vision Encoder: int8 per-channel (QAT) Audio Encoder: mixed int2/int4/int8 per-channel (QAT) |
| Download Size (CPU/GPU file) | 165 MB | 388 MB | 485 MB |
| On-Demand Modality Loading | - | Yes | Yes |
| Supported Input Sizes | Text tokens: 128, 256, 512, 1024, 2048, 8192 |
Text tokens: 128, 256, 512, 1024, 2048, 8192 Image soft tokens: 70, 140 |
Text tokens: 128, 256, 512, 1024, 2048, 8192 Image soft tokens: 70, 140 Audio soft tokens: 12 |
| Supported Output Sizes | 128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
128 (512 bytes) 256 (1024 bytes) 512 (2048 bytes) 768 (3072 bytes) |
| Link | litert-community/embeddinggemma-2-text-270m-litert-lm | litert-community/embeddinggemma-2-text-vision-440m-litert-lm | litert-community/embeddinggemma-2-740m-litert-lm |
Additional Notes:
- LiteRT-LM automatically scales and patchifies images of any size to fit EmbeddingGemma 2's vision encoder model.
- LiteRT-LM supports EmbeddingGemma 2's streamed audio allowing a variety of different audio lengths.
EmbeddingGemma 2 Performance on LiteRT-LM
The text performance was measured by loading and running the 128 text signature. For vision benchmarking, the vision encoder used the 70 signature. The latency reported is the average of 5 iterations.
Memory was measured with each platform's native metrics and is not directly comparable across operating systems. CPU memory was measured using, rusage::ru_maxrss on Android, Linux and IOT, task_vm_info::phys_footprint on iOS and MacBook, and process_memory_counters::PrivateUsage on Windows. With the exception of task_vm_info::phys_footprint, accelerator (GPU/NPU/TPU) memory is not included. Memory measurements were taken from the second load which reads from loading caches. Memory usage on the first load may vary.
Android
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| Google Pixel 11 Pro | TPU | 8.3 ms | 49 ms |
| S26 Ultra | CPU | 27.1 ms | 175 ms |
| S26 Ultra | GPU | 25.9 ms | 119 ms |
| Device | Backend | Text CPU memory | Text + Vision CPU memory |
|---|---|---|---|
| Google Pixel 11 Pro | TPU | 112 MB | 125 MB |
| S26 Ultra | CPU | 334 MB | 694 MB |
| S26 Ultra | GPU | 333 MB | 427 MB |
iOS
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| iPhone 18 Pro | CPU | 41.8 ms | 191 ms |
| iPhone 18 Pro | GPU | 11.6 ms | 69.8 ms |
| Device | Backend | Text CPU/GPU memory | Text + Vision CPU/GPU memory |
|---|---|---|---|
| iPhone 18 Pro | CPU | 84 MB | 196 MB |
| iPhone 18 Pro | GPU | 85 MB | 196 MB |
Linux
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| Arm 2.3 & 2.8GHz | CPU | 105 ms | 825 ms |
| NVIDIA GeForce RTX 4090 | GPU | 7.6 ms | 23.9 ms |
| Device | Backend | Text CPU memory | Text + Vision CPU memory |
|---|---|---|---|
| Arm 2.3 & 2.8GHz | CPU | 310 MB | 647 MB |
| NVIDIA GeForce RTX 4090 | GPU | 528 MB | 678 MB |
macOS
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| MacBook Pro M5 | CPU | 31.3 ms | 151 ms |
| MacBook Pro M5 | GPU | 9.5 ms | 37.3 ms |
| Device | Backend | Text CPU/GPU memory | Text + Vision CPU/GPU memory |
|---|---|---|---|
| MacBook Pro M5 | CPU | 165 MB | 310 MB |
| MacBook Pro M5 | GPU | 233 MB | 403 MB |
Windows
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| Dell XPS 16 (Intel Core Ultra Series 3) | CPU | 71.7 ms | 338 ms |
| Dell XPS 16 (Intel Core Ultra Series 3) | GPU | 19.2 ms | 62.5 ms |
| Dell XPS 16 (Intel Core Ultra Series 3) | Running Intel OpenVINO NPU | 13.3 ms | 49.8 ms |
| Device | Backend | Text CPU memory | Text + Vision CPU memory |
|---|---|---|---|
| Dell XPS 16 (Intel Core Ultra Series 3) | CPU | 211 MB | 342 MB |
| Dell XPS 16 (Intel Core Ultra Series 3) | GPU | 822 MB | 1769 MB |
| Dell XPS 16 (Intel Core Ultra Series 3) | Running Intel OpenVINO NPU | 201 MB | 394 MB |
Web
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| MacBook Pro M5 | GPU | 21.8 ms | 107 ms |
IoT
| Device | Backend | Text latency | Text + Vision latency |
|---|---|---|---|
| Raspberry Pi 5 16GB | CPU | 161 ms | 1761 ms |
| Jetson Orin Nano | CPU | 207 ms | 1741 ms |
| Jetson Orin Nano | GPU | 88 ms | 485 ms |
| Arduino VENTUNO Q | NPU | 13.6 ms | 135 ms |
| Device | Backend | Text CPU memory | Text + Vision CPU memory |
|---|---|---|---|
| Raspberry Pi 5 16GB | CPU | 282 MB | 619 MB |
| Jetson Orin Nano | CPU | 302 MB | 627 MB |
| Jetson Orin Nano | GPU | 523 MB | 1103 MB |
| Arduino VENTUNO Q | NPU | 206 MB | 460 MB |
- Downloads last month
- 5
Model tree for litert-community/embeddinggemma-2-text-vision-440m-litert-lm
Base model
google/embeddinggemma-2