← Managed inference

Gemma 4 E4B (text+image+video+audio)

The catalogue's audio-in model - speech and images without a separate ASR step

Overview

Gemma 4 E4B is an edge-oriented Gemma 4 size: about 4.5B effective parameters, 8B including embeddings, using Per-Layer Embeddings to keep the working set small. It is one of the three Gemma 4 sizes with native audio, via a roughly 300M-parameter audio encoder, and accepts text, image and audio while returning text. Google documents 128K context for this size; we serve it at 65,536 tokens.

Strengths

  • Native audio input, which the 31B size does not have
  • Small enough to be cheap per request
  • Text, image, video and audio through one model

Use cases

  • Question answering and summarisation over audio
  • Multimodal work where cost matters more than peak accuracy
  • Prototyping before committing to a larger model

Good to know

A 4.5B effective model will not match the 27B and 31B models on demanding reasoning. Google documents 128K context for this size; this alias is served at 65,536 tokens.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

gemma-4-e4b · curl
curl https://gateway.graphn.ai/v1/chat/completions \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-e4b",
    "messages": [{"role": "user", "content": "Hello"}]
  }'