Overview
Gemma 4 31B is the largest dense model in Google DeepMind's Gemma 4 family, around 30.7B parameters, using hybrid local-sliding plus global attention. Google documents it as a capable reasoner with configurable thinking, native function calling, a native system role, more than 140 languages, and text and image input -- audio is only on the E2B, E4B and 12B sizes. The public checkpoint supports 256K context; it is served here at 131K.
Strengths
- More than 140 languages
- Native function calling and system role
- Apache-2.0 from a major vendor
Use cases
- Multilingual assistants
- An architecturally different second opinion to the Qwen models
- Mid-scale reasoning and coding
Good to know
Served at 131K context, although the public google/gemma-4-31B-it checkpoint documents 256K. Configurable thinking is documented by Google but not enabled on this alias. Image and video input are enabled on the deployment, but this alias is not routed for attachments -- send them to a vision model, which has twice the context.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
curl https://gateway.graphn.ai/v1/chat/completions \
-H "Authorization: Bearer $GRAPHN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4",
"messages": [{"role": "user", "content": "Hello"}]
}'graphn model chat gemma-4 -m "What is GraphN?"
@function
async def ask(question: str) -> str:
return await chat.complete(question, model="gemma-4")