← Managed inference

gpt-oss 120B

Reasoning with a cost dial - low, medium or high effort per request

Overview

OpenAI's gpt-oss-120b is an Apache-2.0 mixture-of-experts reasoning model, 117B total parameters with 5.1B active, post-trained with MXFP4 quantisation of the MoE weights so it fits on a single 80GB GPU. Reasoning effort is configurable at low, medium or high, and the full chain-of-thought is available for inspection. It uses OpenAI's harmony response format, which this deployment parses into a separate reasoning field.

Strengths

  • Reasoning depth is a per-request knob
  • Full chain-of-thought for debugging
  • Permissive Apache-2.0 licence

Use cases

  • Work where latency can be traded for depth
  • Debugging how an agent reached an answer
  • Permissively-licensed reasoning workloads

Good to know

OpenAI states the chain-of-thought is for debugging and trust, and is not intended to be shown to end users. Function calling is not available here: the weights support it, but this deployment does not run a tool-call parser, so a request carrying tools gets the intended call as JSON text in content with tool_calls empty. Use a model that advertises tool_calling for agent loops.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

gpt-oss-120b · curl
curl https://gateway.graphn.ai/v1/chat/completions \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [{"role": "user", "content": "Hello"}]
  }'