Best for
- Developers
- Self-hosters
- Regulated environments
OpenAI's open-weight reasoning models (120B and 20B) under Apache-2.0 — the 20B runs on a 16GB GPU.
gpt-oss brought OpenAI back to open weights: two MoE reasoning models with adjustable effort, built for agentic use with tools. The 20B fits consumer hardware; 120B runs on a single 80GB card. Supported day-one by Ollama, LM Studio and vLLM.
gpt-oss is best used as a practical helper for local reasoning agents, private coding help, on-prem assistants. It is useful when you want to move from rough ideas to high-quality drafts quickly, then refine with human review before publishing or sharing.
gpt-oss sits in the same category as Qwen 3 and DeepSeek R1 / V3. Use gpt-oss when its output style and workflow fit your team, but compare pricing limits, integrations, and review requirements before standardizing.
Explore workflows to find best-fit use cases for this tool.