Overview
OpenAI's gpt-oss-120b is an Apache-2.0 mixture-of-experts reasoning model, 117B total parameters with 5.1B active, post-trained with MXFP4 quantisation of the MoE weights so it fits on a single 80GB GPU. Reasoning effort is configurable at low, medium or high, and the full chain-of-thought is available for inspection. It uses OpenAI's harmony response format, which this deployment parses into a separate reasoning field.
Strengths
- Reasoning depth is a per-request knob
- Full chain-of-thought for debugging
- Permissive Apache-2.0 licence
Use cases
- Work where latency can be traded for depth
- Debugging how an agent reached an answer
- Permissively-licensed reasoning workloads
Good to know
OpenAI states the chain-of-thought is for debugging and trust, and is not intended to be shown to end users. Function calling is not available here: the weights support it, but this deployment does not run a tool-call parser, so a request carrying tools gets the intended call as JSON text in content with tool_calls empty. Use a model that advertises tool_calling for agent loops.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
curl https://gateway.graphn.ai/v1/chat/completions \
-H "Authorization: Bearer $GRAPHN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [{"role": "user", "content": "Hello"}]
}'graphn model chat gpt-oss-120b -m "What is GraphN?"
@function
async def ask(question: str) -> str:
return await chat.complete(question, model="gpt-oss-120b")