AI ResearchJul 16, 2026, 6:48 PM

AI Red Teaming Breakthrough

TickrWire Editorial Desk·Jul 16, 2026, 6:48 PM·1 min read AI-assisted, human-reviewed

Reported by MarkTechPost: OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection. Analysis and context written by TickrWire.

Evolving story · 3 updatesOpenAI GPT-Red safety systemTimeline →
30-second summary

OpenAI's GPT-Red model outperformed human red-teamers in a prompt injection test, with a success rate of 84% compared to 13%. The model was trained using self-play reinforcement learning against a population of defender LLMs.

TickrWire
AI Red Teaming Breakthrough
Key takeaways
  • GPT-Red outperformed human red-teamers in a prompt injection test
  • The model was trained using self-play reinforcement learning
  • GPT-Red discovered a novel attack class called 'Fake Chain-of-Thought'
  • The model helped reduce the failure rate of GPT-5.6 Sol by a factor of six
Full story

OpenAI has developed an internal automated red-teaming model called GPT-Red, which has demonstrated impressive capabilities in identifying vulnerabilities in language models. The model was trained using self-play reinforcement learning, where it played against a population of defender LLMs. This approach allowed GPT-Red to learn and improve its attack strategies.

The results of the test were striking, with GPT-Red outperforming human red-teamers by a significant margin. The model's success rate of 84% compared to the human team's 13% highlights the potential of automated red-teaming in AI security testing.

One of the key findings of the test was the discovery of a novel attack class called "Fake Chain-of-Thought". This attack exploits the way language models process and generate text, and it has significant implications for the development of more secure AI systems.

OpenAI has also reported that GPT-Red has helped to reduce the failure rate of its GPT-5.6 Sol model by a factor of six on a direct injection benchmark. However, the company acknowledges that there is still work to be done, particularly in addressing multi-turn and image-based attacks.

The development of GPT-Red is an important step forward in the field of AI security, and it has significant implications for the development of more secure and robust language models.

Why this matters
Developers

Improved AI security testing

Businesses

More secure language models

Investors

Potential for increased investment in AI security

Everyone

Advancements in AI security

Glossary
red-teaming
The practice of testing a system's defenses by simulating an attack
Sources · 1
Read next
More stories
How AI-native companies turn workflows into operating capabilityBusiness

How AI-native companies turn workflows into operating capability

OpenAI highlights how firms like Basis, Clay, and Exa Labs deploy autonomous agents to handle onboarding, account management, and developer integrations.

Path to Astra: critical capabilities and frontier safeguardsSecurity

Path to Astra: critical capabilities and frontier safeguards

OpenAI’s Astra model is the first to meet the Critical cybersecurity capability threshold under its Preparedness Framework, enabling autonomous discovery of unknown vulnerabilities and exploit chains.

How law firm Gilbert + Tobin governs and scales AI with OpenAIBusiness

How law firm Gilbert + Tobin governs and scales AI with OpenAI

Gilbert + Tobin, a major Australian law firm, has rolled out ChatGPT Enterprise and Codex firm-wide, driven by CEO Sam Nickless and supported by rigorous governance. The adoption rate, with 87% of enabled seats active, more than doubles the firm's typical tool usage, and specific workflows such as recruitment research have been cut from four hours to about 20 minutes.

Polimill builds Japan's next-generation public AI infrastructureAI Tools

Polimill builds Japan's next-generation public AI infrastructure

Polimill introduced QommonsAI, an OpenAI‑powered platform that now supports roughly 1,050 Japanese local governments and 550,000 public employees, aiming to become a shared operating system for municipal work.

A milestone in expanding access to AIBusiness

A milestone in expanding access to AI

OpenAI announced that its ChatGPT Ads platform has crossed $1 billion in annualized revenue run rate in under 200 days and is rolling out self-service tools internationally.

India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone callFunding

India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call

Indian voice AI startup Ringg has raised $10 million in a Series A extension led by Peak XV Partners, bringing its total funding in the round to $15.5 million.