Introducing gpt-oss-safeguard

📰 OpenAI News

OpenAI releases gpt-oss-safeguard, open safety reasoning models for custom safety policies

advanced Published 29 Oct 2025
Action Steps
  1. Download gpt-oss-safeguard models from Hugging Face
  2. Fine-tune the models for specific use cases
  3. Integrate the models into existing safety pipelines
  4. Review and revise policies to increase performance
Who Needs to Know This

Developers and product managers can use gpt-oss-safeguard to create tailored safety policies for their applications, improving the relevance and effectiveness of safety classification tasks

Key Insight

💡 gpt-oss-safeguard enables developers to create flexible and adaptable safety policies using reasoning-based approach

Share This
🚀 OpenAI releases gpt-oss-safeguard, open safety reasoning models for custom safety policies! 🤖

Key Takeaways

OpenAI releases gpt-oss-safeguard, open safety reasoning models for custom safety policies

Full Article

# Introducing gpt-oss-safeguard | OpenAI

[Skip to main content](https://openai.com/index/introducing-gpt-oss-safeguard#main)

[](https://openai.com/)

* [Research](https://openai.com/research/index/)
* Products
* [Business](https://openai.com/business/)
* [Developers](https://openai.com/api/)
* [Company](https://openai.com/about/)
* [Foundation(opens in a new window)](https://openaifoundation.org/)

[Try ChatGPT(opens in a new window)](https://chatgpt.com/)

* Research
* Products
* Business
* Developers
* Company
* [Foundation(opens in a new window)](https://openaifoundation.org/)

[Try ChatGPT(opens in a new window)](https://chatgpt.com/)

OpenAI

Table of contents

* [System-level safety: the role of safety classifiers](https://openai.com/index/introducing-gpt-oss-safeguard#system-level-safety-the-role-of-safety-classifiers)
* [How we use safety reasoning internally](https://openai.com/index/introducing-gpt-oss-safeguard#how-we-use-safety-reasoning-internally)
* [How gpt-oss-safeguard performs](https://openai.com/index/introducing-gpt-oss-safeguard#how-gpt-oss-safeguard-performs)
* [Limitations](https://openai.com/index/introducing-gpt-oss-safeguard#limitations)
* [The road ahead: continuing to build with the community](https://openai.com/index/introducing-gpt-oss-safeguard#the-road-ahead-continuing-to-build-with-the-community)

October 29, 2025

[Product](https://openai.com/news/product-releases/)[Release](https://openai.com/research/index/release/)

# Introducing gpt-oss-safeguard

New open safety reasoning models (120b and 20b) that support custom safety policies.

Loading…

Share

Today, we’re releasing a research preview of gpt-oss-safeguard, our open-weight reasoning models for safety classification tasks, available in two sizes: gpt-oss-safeguard-120b and gpt-oss-safeguard-20b. These models are fine-tuned versions of our [gpt-oss⁠](https://openai.com/index/introducing-gpt-oss/) open models and available under the same permissive Apache 2.0 license, allowing anyone to use, modify, and deploy them freely. Both models can be downloaded today from [Hugging Face⁠(opens in a new window)](https://huggingface.co/collections/openai/gpt-oss-safeguard).

The gpt-oss-safeguard models use reasoning to directly interpret a developer-provided policy at inference time—classifying user messages, completions, and full chats according to the developer’s needs. The developer always decides what policy to use, so responses are more relevant and tailored to the developer’s use case. The model uses chain-of-thought, which the developer can review to understand how the model is reaching its decisions. Additionally, the policy is provided during inference, rather than being trained into the model, so it is easy for developers to iteratively revise policies to increase performance. This approach, which we initially developed for internal use, is significantly more flexible than the traditional method of training a classifier to indirectly infer a decision boundary from a large number of labeled examples.

gpt-oss-safeguard enables developers to draw the policy lines that best fit their use case. For instance, a video gaming discussion forum might want to develop a policy to classify posts that discuss cheating in the game, or a product reviews site might want to use its own policy to screen reviews that appear likely to be fake.

The model takes two inputs at once—a policy and the content to classify under that policy—and outputs a conclusion about where the content falls, along with its reasoning. Developers decide how, if at all, to use those conclusions in their own safety pipelines. We’ve seen this reasoning-based approach perform especially well in situations where:

* The potential harm is emerging or evolving, and policies need to adapt quickly.
* The domain is highly nuanced and difficult for smaller classifiers to handle.
* Developers don’t have enough samples to train a high-quality classifier
Read full article → ← Back to Reads

Related Videos

5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
MCP explained for beginners
MCP explained for beginners
Withmesravani_
Temperature Explained | Why ChatGPT Gives Different Answers | AI Series Day 14 #Shorts
Temperature Explained | Why ChatGPT Gives Different Answers | AI Series Day 14 #Shorts
Withmesravani_
4 Generative AI Projects That Will Get You Hired in 2026 🚀
4 Generative AI Projects That Will Get You Hired in 2026 🚀
SCALER
I Tested My AI-Powered Autocoder With 3 Different LLM Models
I Tested My AI-Powered Autocoder With 3 Different LLM Models
Making Made Easy
You Can Run Your Own Powerful LLM AI On Almost Any Computer! OPEN SOURCE! NO GPU NEEDED! MISTRAL 7B!
You Can Run Your Own Powerful LLM AI On Almost Any Computer! OPEN SOURCE! NO GPU NEEDED! MISTRAL 7B!
Making Made Easy