← Managed inference

Gemma 4 31B Instruct

A non-Qwen alternative with unusually wide language coverage

Overview

Gemma 4 31B is the largest dense model in Google DeepMind's Gemma 4 family, around 30.7B parameters, using hybrid local-sliding plus global attention. Google documents it as a capable reasoner with configurable thinking, native function calling, a native system role, more than 140 languages, and text and image input -- audio is only on the E2B, E4B and 12B sizes. The public checkpoint supports 256K context; it is served here at 131K.

Strengths

  • More than 140 languages
  • Native function calling and system role
  • Apache-2.0 from a major vendor

Use cases

  • Multilingual assistants
  • An architecturally different second opinion to the Qwen models
  • Mid-scale reasoning and coding

Good to know

Served at 131K context, although the public google/gemma-4-31B-it checkpoint documents 256K. Configurable thinking is documented by Google but not enabled on this alias. Image and video input are enabled on the deployment, but this alias is not routed for attachments -- send them to a vision model, which has twice the context.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

gemma-4 · curl
curl https://gateway.graphn.ai/v1/chat/completions \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4",
    "messages": [{"role": "user", "content": "Hello"}]
  }'