GP
Open-weight LLMs

Overview - gpt-oss

OpenAI's open-weight reasoning models (120B and 20B) under Apache-2.0 — the 20B runs on a 16GB GPU.

gpt-oss brought OpenAI back to open weights: two MoE reasoning models with adjustable effort, built for agentic use with tools. The 20B fits consumer hardware; 120B runs on a single 80GB card. Supported day-one by Ollama, LM Studio and vLLM.

Difficulty: IntermediatePlatform: LocalPricing: FreeReleased: 2025-08
Open WeightReasoningOpenaiApache 2.0

What this tool is used for

gpt-oss is best used as a practical helper for local reasoning agents, private coding help, on-prem assistants. It is useful when you want to move from rough ideas to high-quality drafts quickly, then refine with human review before publishing or sharing.

Best for

  • Developers
  • Self-hosters
  • Regulated environments

Use this when

  • You want OpenAI-style reasoning offline
  • You have a 16GB+ GPU
  • You need Apache-2.0

Not ideal for

  • Multilingual-first products (English-centric)

Example tasks

  • Local reasoning agents
  • Private coding help
  • On-prem assistants

Limitations / things to watch

  • English-leaning
  • Text only

How this compares to similar tools

gpt-oss sits in the same category as Qwen 3 and DeepSeek R1 / V3. Use gpt-oss when its output style and workflow fit your team, but compare pricing limits, integrations, and review requirements before standardizing.

Related tools

Suggested workflows

Explore workflows to find best-fit use cases for this tool.