Managed inference for any model.
Serve catalog models or your own weights through one OpenAI-compatible API, on GPUs inside your security boundary.
Model launches
The built-in catalog
qwen3-80b131K contextQwen3 235B InstructLargest open-weight MoE model (22B active) for complex instruction following, agents, and long-context tasksInput $0.22/1M · Output $0.88/1Mqwen3-235b262K contextQwen3 CoderCode-specialized MoE model for agentic software engineering, repo-scale context, and multi-turn coding workflowsInput $0.90/1M · Output $0.90/1Mqwen3-coder131K contextNemotron 3 Super 120BNVIDIA hybrid Mamba-2 / MoE / attention model (12B active) for agentic workflows and long-context reasoningInput $0.90/1M · Output $0.90/1Mnemotron-3-super262K contextgpt-oss 120BOpen-weight Apache-2.0 MoE reasoning model with adjustable low/medium/high reasoning effort and full chain-of-thought via the harmony parserInput $0.15/1M · Output $0.60/1Mgpt-oss-120b131K contextQwen2.5 VL 7B (image+video)Compact vision-language model for image, document, chart, and video understandingInput $0.05/1M · Output $0.05/1Mqwen2.5-vl-7b-instruct128K contextQwen3.5 9B Vision (image+video)General-purpose 9B vision-language model for image analysis and multimodal tasksInput $0.10/1M · Output $0.15/1Mqwen3-vl262K contextQwen3.8 27B (image+video)Dense 27B vision-language model for coding, agentic work, and image/video understandingInput $0.45/1M · Output $3.20/1Mqwen3.8-27b262K contextQwen3.5 122B Vision (image+video)Unified vision-language MoE (10B active) for high-quality multimodal reasoning over images and videoInput $0.29/1M · Output