跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 3 天前AI 评分74

Google 发布 EmbeddingGemma 2 端侧多模态嵌入模型

AI 导读

Google 发布 EmbeddingGemma 2,一款采用 Apache 2.0 许可、面向端侧多模态 AI 的嵌入模型,可将文本、代码、图像、音频和视频映射到同一可搜索空间。

正文 · 原文

Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license

> puts text, code, images, audio and video into 1 searchable space on phones.

gives phones a missing piece: a way to understand and search your own stuff without sending it to a server.

> 740M parameters, uses the Gemma 4 architecture.

> Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed.

> On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB.

> The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input.

> Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level.

> Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,

来源:rohanpaul_ai · x.com