Gemma 4 E4B (text+image+video+audio)
The catalogue's audio-in model - speech and images without a separate ASR step
Overview
Gemma 4 E4B is an edge-oriented Gemma 4 size: about 4.5B effective parameters, 8B including embeddings, using Per-Layer Embeddings to keep the working set small. It is one of the three Gemma 4 sizes with native audio, via a roughly 300M-parameter audio encoder, and accepts text, image and audio while returning text. Google documents 128K context for this size; we serve it at 65,536 tokens.
Strengths
- Native audio input, which the 31B size does not have
- Small enough to be cheap per request
- Text, image, video and audio through one model
Use cases
- Question answering and summarisation over audio
- Multimodal work where cost matters more than peak accuracy
- Prototyping before committing to a larger model
Good to know
A 4.5B effective model will not match the 27B and 31B models on demanding reasoning. Google documents 128K context for this size; this alias is served at 65,536 tokens.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
curl https://gateway.graphn.ai/v1/chat/completions \
-H "Authorization: Bearer $GRAPHN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-e4b",
"messages": [{"role": "user", "content": "Hello"}]
}'graphn model chat gemma-4-e4b -m "What is GraphN?"
@function
async def ask(question: str) -> str:
return await chat.complete(question, model="gemma-4-e4b")