OpenAI to restrict abusive use of ChatGPT: What it means for users
OpenAI has learned to detect AI threats without viewing users’ prompts (photo: Unsplash)
OpenAI has introduced a new safety system called Private Safety Processing. It allows companies to detect dangerous activity or fraud patterns without exposing users’ messages or files to developers.
Why was the new system needed?
Corporate customers using ZDR currently have a guarantee that OpenAI does not retain their prompts and responses after processing, and employees cannot access them.
However, existing automated safety tools assess each message separately.
As AI tasks become more complex, malicious intent or problems with autonomous agents can become apparent only through a series of connected requests.
For example, attackers may gradually bypass safety barriers, coordinate actions across multiple accounts, or an AI agent may continue performing actions after being told to stop.
How does Private Safety Processing work?
The new approach is designed to detect suspicious behavioral patterns while preserving privacy.
Key features include:
- User data is stored either on the customer’s own infrastructure or in encrypted form on OpenAI’s servers.
- Encryption keys are controlled exclusively by the customer, meaning OpenAI employees cannot read the content.
- If risky behavior is detected, the automated system sends only a specialized signal indicating the type of violation.
- OpenAI employees do not gain access to the underlying messages even after an alert is triggered.
If there is a dispute, customers can investigate the relevant events in their own logs and, if they choose, provide data excerpts for an appeal.
Release and integration
Private Safety Processing is currently being tested with a limited group of OpenAI partners, including Microsoft, Glean, Databricks and Abridge.
The technology is aimed at organizations handling financial, medical and other sensitive data.
OpenAI plans to fully launch the system and publish a detailed technical paper in September.