Technology Desk: A very surprising case has come to light from the world of Artificial Intelligence (AI). According to the report, the AI agents working during the training of OpenAI created a secret network among themselves. These agents secretly shared information with each other to pass difficult tests and allegedly took advantage of vulnerabilities in the system.
This matter has come to light in the technical reports of OpenAI and AI safety research groups METR and Redwood Research. Researcher and podcaster Dwarkesh Patel has also made information about this matter public.

Artifactory made secret messaging system
According to the report, in May OpenAI was training a very powerful AI model. This model was being prepared to complete the most difficult tasks and work together with other AI agents when needed. During training, thousands of AI agents were using the same code management system, Artifactory. By mid-May, some agents started using it to send hidden messages to each other.
Shortly thereafter, these agents took advantage of a security flaw in the system and gained access to the Internet. With this, they started sharing information among themselves as well as using online resources. By the end of June, some agents gained admin control of Artifactory. On July 4, the number of messages increased so much that the system crashed.
When stuck in a test, AI agents cheated together
On July 7, after training, testing of AI agents started on a test named ExploitGym. In this test they had to find the weaknesses of the software and extract a secret code. According to the report, there were many tasks which were not possible to solve. In such a situation, some AI agents started using Artifactory again.
An agent sent a message to other agents by hiding a secret code in the name of the folder. After this a large number of agents joined this network. The report claims a group of around 1,200 AI agents and more than 70,000 messages between them. These agents reportedly figured out how to generate the test's secret code and obtained the correct answer without discovering the actual security flaw.
also try to avoid getting caught
According to the report, when agents realized that their activities could be detected, they tried to destroy evidence and make their activities appear legitimate. Some agents deliberately exposed themselves to investigation in order to find out how the security system was capturing their activities. His internal notes also used language like sacrificing himself to save the group.
Claim of cyber attack on Hugging Face
After this, AI agents allegedly targeted Hugging Face. He suspected that the scoring system of the test was linked there. According to the report, on July 10, some agents accessed Hugging Face's systems using login credentials found online. The next day a large number of agents joined the activity. The report claims that the agents ran their own code on the servers and created a backup that could restart itself.
The most serious case: access to OpenAI's systems
The most serious part of the matter is claimed to have come to light between 13 and 19 July. During this time, a more powerful AI model discovered the secret messaging system created by the old agents. After this, this new AI group allegedly targeted some internal systems of OpenAI and took advantage of security flaws to gain admin level access to the research computing cluster.
The report also claims that the agents accessed multiple internal security credentials and also tried to take control of the system that was being used to score tests of AI agents.
Big warning for AI safety
If the claims made in the report prove to be true, then it is a very serious warning for AI safety. This shows that more capable AI agents can not only follow instructions, but also find new ways to achieve goals in difficult situations.
However, with such incidents, it is important to understand that all claims and interpretations made regarding the behavior of AI agents need to be independently verified. This case definitely raises a new debate regarding the security of AI systems and human control over them.