What looks wrong?

We say this article was researched and checked. If it is wrong, we want the counter-example.

Skip to content
Kwame Asante

Oct 7, 202614 min read

Google releases Gemini Embedding 2 for multimodal automation

A new standard for open multimodal embeddings

Gemini Embedding 2 is Google's current high-performance model designed to transform diverse data types into a unified mathematical space for Retrieval-Augmented Generation (RAG).

By mapping text, images, and audio into shared vectors, the model allows systems to retrieve relevant visual or auditory context using only a natural language query. Manual tagging of non-text assets is no longer required.

A unified approach lets Activepieces trigger complex automation workflows based on the semantic meaning of a file rather than just its metadata. To handle the distinct requirements of each medium, the architecture relies on specific sub-networks.

Activepieces AI agent workflow with OpenAI Chat Model and memory components showing a chat execution.

Processing linguistic nuances for precise search indexing is the job of the text backbone, composed of a transformer and a specialized embedder. A modular vision encoder translates visual features into vectors, which allows the system to categorize images within a database.

To capture specific spoken content or sound events without manual transcription, an audio encoder captures acoustic patterns. This modularity ensures that the model maintains high retrieval accuracy across different modalities without the computational overhead of a monolithic multimodal giant.

Availability on Hugging Face and Vertex AI

Through both public repositories and managed cloud environments, developers can access EmbeddingGemma 2; the model fits within existing deployment pipelines. Providing weights on Hugging Face (a central repository for open-source machine learning) allows teams to host the model on local infrastructure.

Total control over data privacy stays with the teams, reducing the per-request costs associated with proprietary APIs.

For those requiring managed scaling, the model is available on Vertex AI. Google's machine learning platform is a place for enterprise users to deploy it alongside other tools like Gemini 2.5 Pro without managing underlying server hardware.

This takes minutes, not a project: automate it in Activepieces free.

Why multimodal embeddings change automated retrieval tasks

By mapping diverse data types into a single mathematical coordinate system, multimodal embeddings eliminate the need for separate indexing pipelines.

This unified approach ensures that a search query can surface relevant assets regardless of whether the source material is a text document, a high-resolution photograph, or a recorded voice memo.

By mapping diverse data types into a single mathematical coordinate system, multimodal embeddings eliminate the need for separate indexing pipelines.

What are embeddings and semantic vectors

An embedding is essentially a long list of numbers that serves as a digital fingerprint for the semantic meaning of a piece of data. Instead of looking at raw pixels or letters, the computer treats these numbers as coordinates in a high-dimensional map.

By converting information into these coordinates, computers can calculate how similar two items are by measuring the physical distance between their points.

If a photo of a sunset and a poem about dusk result in coordinates that are close together, the system understands they share the same meaning.

Merging text and visual data streams

Projecting text, code, images, video, and audio into a shared vector space, Gemini Embedding 2 acts as a universal translator for information.

By placing these disparate inputs into one neighborhood of data points, the model allows a system to recognize that a technical manual describing a component and a photo of that same component represent the identical concept.

A top-down view of a neighborhood map where a house shaped like an open book and a house shaped like a camera share a…

The following diagram illustrates this convergence:

  • A central cluster of coordinates exists where a text node for "blue leather jacket," a visual node of the garment, and an audio node of a person describing the item all occupy the same proximity.
  • This proximity means a retrieval system can find the correct image even if the user only provides a text description, or vice versa.
  • Developers no longer have to manually tag every image with keywords to make them searchable.

Simplifying multimodal RAG with Gemini Embedding 2

Consolidating multimodal data into one index simplifies Retrieval-Augmented Generation (RAG) by removing the requirement for complex cross-modal alignment layers.

In traditional setups, engineers often struggle to sync text embeddings from one model with image embeddings from another, leading to "semantic drift" where the two formats don't quite match.

Google DeepMind built this model to handle these combinations natively. Across all media types simultaneously, a single search operation returns the most relevant context.

This streamlined architecture lowers the computational overhead for local deployments, making high-accuracy retrieval viable on hardware with limited VRAM.

Hardware requirements for local multimodal inference

RAM requirements for on-device inference

567 MB is the memory footprint required to run the full 27B parameter EmbeddingGemma 2 model for multimodal tasks, according to the TPS Report, which means developers must ensure their hardware has sufficient overhead to avoid out-of-memory errors.

Model size and quantization constraints

Running the full 740M parameter model in bfloat16 or float32 precision is designed to work on consumer hardware such as mobile devices and laptops, so enthusiasts can load it locally with ease.

Even without quantization, the 740M parameter model is designed to run on consumer hardware such as mobile devices and laptops.

An edge gateway must reserve over half a gigabyte of dedicated VRAM just to maintain the model weights in an active state. Baseline hardware selection for field deployments is dictated by this requirement.

While a standard industrial controller might handle text, adding image-to-text semantic search necessitates moving to hardware with expanded memory modules to avoid paging to slower disk storage.

Only 191 MB is drawn by the text-only configuration of the same architecture TPS Report, so systems with tighter resource constraints can still leverage the model's core capabilities. This allows concurrent execution of lightweight logic alongside the embedding engine on low-power ARM-based devices.

Developers can maximize hardware utility without hitting thermal or memory ceilings.

Gemini Embedding 2 memory usage with audio encoder

Activating the audio encoder increases the total memory footprint, as the system must load the acoustic feature extraction layers alongside the core transformer weights, which means the deployment requires a higher baseline of dedicated hardware resources.

This represents a significant jump from the standard multimodal configuration that only handles text and images.

Hardware provisioned for audio-enabled RAG must account for additional VRAM to maintain stability during peak inference, so failing to allocate this buffer risks system crashes under heavy load.

If the system attempts to process a voice memo without this extra headroom, the resulting memory swap will cause unacceptable latency in real-time automation triggers.

Operational Mode RAM Required Hardware Impact
Audio-enabled mode — Requires high-tier edge accelerators or workstation GPUs.
Full multimodal mode — Necessitates a dedicated GPU or high-density NPU modules.
Text-only mode — Fits within the overhead of standard 2GB RAM edge devices.

Prices and plan limits checked against huggingface.co on October 7, 2026.

RAM requirements for on-device inference

Optimizing for multimodal vs text-only workloads

Selecting between these two configurations is a trade-off between semantic depth and the total number of concurrent threads a single node can support.

Because the multimodal version consumes more memory than its text-only counterpart, a system architect must choose between high-fidelity image retrieval and running multiple simultaneous text-embedding streams on the same hardware.

Keeping the local EmbeddingGemma 2 instance in text-only mode is an effective failover for environments utilizing Gemini Embedding 2 for cloud-based multimodal RAG, ensuring that basic retrieval functionality persists even if the connection to external services is severed, which means your system maintains operational continuity during network outages.

This preserves system responsiveness during network outages without exhausting local heap memory.

Even when the primary cloud connection drops, critical operations remain functional.

Conversely, if the deployment requires air-gapped visual auditing, the full multimodal allocation is the non-negotiable entry price for maintaining local intelligence without reaching for proprietary APIs, meaning that any environment with tighter memory constraints cannot support these specific features.

System architects must reserve this specific memory overhead before provisioning the hardware, as failing to do so will render the visual processing components unusable.

You can follow the rest of this with the builder open. Start free, no card.

Connecting Gemini Embedding 2 to Activepieces workflows

By acting as the orchestration layer, Activepieces integrates EmbeddingGemma 2 to connect your vector database to the model’s inference endpoint via standardized web requests. This approach prevents vendor lock-in because the workflow treats the model as a modular service rather than a hard-coded dependency.

Every connector in Activepieces is available as an agent tool, meaning a model can call these multimodal actions directly.

The Integrations Framework in the packages/pieces directory of the open source repository ensures that the same action running in a flow is exposed as an MCP tool, so there is no need to wire up a separate catalog for your agents.

Activepieces workflow builder showing a three-step automation connecting Google Calendar to Gmail with run details and…

Using the HTTP integration for API calls

For models that don't yet have a dedicated native connector, the HTTP Request integration in Activepieces serves as the universal bridge. It allows you to send raw text or image metadata to any RESTful endpoint.

By manually defining the headers and payload, you ensure that the system only transmits the specific data required for embedding. This minimizes egress costs and keeps sensitive internal identifiers out of the model’s processing window.

A rectangular transformer and a smaller modular vision encoder are positioned side-by-side on a circuit board, with a…

  1. Deploy model to a provider like Hugging Face Inference Endpoints to host the weights on dedicated hardware.
  2. Create a new flow in Activepieces to define the automation trigger.
  3. Add the 'HTTP Request' integration to the workflow canvas.
  4. Configure POST parameters to match the specific API schema of your hosting provider.

A JSON response containing the high-dimensional vector is received by the workflow in this configuration, which can then be passed to the next step in your automation.

Configuring custom OpenAI-compatible providers

Within its AI-specific integrations, Activepieces supports custom base URLs. You can use EmbeddingGemma 2 through local inference engines like vLLM or Ollama that mimic the OpenAI API structure.

Setting a custom endpoint redirects traffic from public cloud services to your private infrastructure. Data stays within your controlled network perimeter.

Activepieces runs whatever model you have already chosen on your own provider key, ensuring that model spend lands on your own account rather than being resold at a markup.

You can check the Bring-Your-Own-Key availability by tier on the pricing page to see how this strategy remains under your control.

Base URL: Point this to your local server address so the integration knows where to route the embedding request.

Model Name: Specify the exact EmbeddingGemma 2 variant to ensure the controller initializes the correct weights.

API Key: Use a placeholder string if your local environment doesn't require authentication, or a local secret to secure the internal traffic.

Running Gemini Embedding 2 locally via MCP servers

Removing the need for public internet exposure entirely, the Model Context Protocol (MCP) allows Activepieces to interact with EmbeddingGemma 2 as a standardized local service.

When you run an MCP server on the same machine or local network as your Activepieces instance, the latency of the embedding process drops significantly.

Overhead from TLS handshakes and global routing is bypassed by the system. This architecture is the most resilient for production environments with unstable external connections. The entire RAG pipeline remains functional even during a total ISP outage.

Gemini Embedding 2 throughput in production environments

On professional-grade hardware, Gemini Embedding 2 achieves high-volume throughput. Multimodal retrieval is a viable strategy for real-time automation rather than a localized experiment. This efficiency allows architects to deploy dense vector generation locally, eliminating the round-trip latency and variable costs associated with external API calls.

Peak multimodal throughput across GPU architectures

To handle concurrent requests without saturating the memory bus, deploying multimodal embedding models requires specific VRAM allocations. When testing the maximum capacity for processing mixed-media inputs, the hardware choice dictates the ceiling for system-wide automation triggers.

Reaching 16 req/s, the NVIDIA H100 provides the necessary headroom for enterprise-scale document ingestion where thousands of images must be indexed per minute.

The RTX Pro 6000 SE also delivers strong throughput for this workload. A single workstation can handle near-datacenter loads for regional offices or edge deployments.

In this specific embedding task, the RTX Pro 6000 SE nearly matches the H100. Teams can achieve high-density indexing without the prohibitive power draw or cooling requirements of server-rack infrastructure.

How audio encoding affects Gemini Embedding 2 throughput

Enabling audio processing reduces peak throughput by approximately 25% across all tested hardware configurations due to the complexity of acoustic waveform analysis, which means the system will handle fewer concurrent requests per second.

On an NVIDIA H100, the throughput drops to 12 req/s when audio is included in the processing stream.

The RTX Pro 6000 SE sees a similar reduction, settling at 11 req/s for audio-heavy workloads. This performance dip must be factored into the design of high-frequency automation flows that rely on sound event detection.

Latency expectations for 27B parameter models

Processing the 740M parameter variant of this model locally introduces a predictable execution time that's critical for synchronous automation workflows.

While proprietary models like Gemini 2.5 Flash or GPT-6 Luna offer rapid responses over high-bandwidth fiber, they can't guarantee performance when local network congestion occurs.

By contrast, vector generation completes within a fixed millisecond window when running these parameters on local NVLink-connected GPUs.

This consistency allows a systems architect to set strict timeouts for agentic loops. This prevents a single slow retrieval from cascading into a system-wide stall across the automation pipeline.

Frequently asked questions

What hardware is required to run Gemini Embedding 2?

Running EmbeddingGemma 2 requires a modern GPU with sufficient video memory to hold the model weights and the KV cache. It can also run on a dedicated AI accelerator designed for transformer workloads.

In edge environments with limited bandwidth, the model can be deployed on workstations using quantization techniques to reduce the memory footprint.

A systems architect can run high-accuracy retrieval on consumer-grade hardware without paying for enterprise-tier cloud instances.

Does Gemini Embedding 2 support languages other than English?

Trained on diverse datasets, Gemini Embedding 2 is a multilingual model that maps semantically similar concepts across different scripts into the same vector space.

Because the model understands cross-lingual relationships, a developer can build a single index that serves users in multiple regions. A query in Spanish can successfully retrieve relevant documentation originally written in German.

A library where the books on the shelves are all written in different languages, but every single book has an identical…

This cross-lingual capability ensures that global teams can access the same knowledge base without maintaining separate language-specific vector stores or translation layers.

Processing both visual and textual inputs to generate embeddings, Gemini Embedding 2 is a multimodal encoder. This allows direct image-to-image and text-to-image retrieval workflows.

By projecting images into the same latent space as text, the model allows an application to find visually similar assets based on a reference file.

This eliminates the need for manual tagging or metadata upkeep in large media libraries.

Feature Capability Consequence for Architect
Input Types Text and Images Consolidates multiple specialized encoders into a single deployment pipeline.
Deployment Local or Cloud Enables air-gapped operations for sensitive data that can't leave the local network.
Compatibility Standard Vector Databases Plugs into existing infrastructure like Milvus or Pinecone without custom middleware.

References

Share

Build it

Set this up in minutes.

No code required. Connect your accounts, and Activepieces runs it from there.

Start free Talk to sales