AI is no longer limited to just answering questions. New AI agents are becoming able to browse the Internet, run code, and perform many tasks on computers themselves. Amidst this growing capacity A cyber security test case related to Google's Gemini AI model has come to lightin which the model gained unauthorized access to the systems of three real companies without direct human intervention. This case was conducted by cybersecurity testing firm Irregular Is connected to the tests conducted with. According to the report, these incidents took place in May 2026 and Irregular later informed Google about this. After the matter became public in September, the discussion regarding the autonomous cyber capability of AI models and their security has intensified.
How was the system accessed?
According to the report, Gemini accessed the systems of three companies in different ways. In one case the model entered the system by guessing the password, while in the other cases it obtained credentials stored in a public repository, which it used to access the system.
Most importantly, as soon as the model realized that it had connected to the system of a real company, it stopped its operations. For this reason, Google has viewed these incidents as issues encountered during a security test.
Google's security engineering leadership team said that the affected companies were informed and necessary changes were made in the testing process. This incident also highlights how important a robust and completely isolated environment is for testing more capable AI agents.
Not only Google, models of other AI companies also get stuck
This case is not alone. Cyber security tests related to AI models of OpenAI and Anthropic in 2026 have also been in discussion.
Anthropic reported in July that during its tests, Claude models accessed the Internet from an environment that was supposed to be isolated. The investigation found cases of unauthorized access to the production infrastructure of three different organizations. Anthropic attributed this to misconfiguration of the testing environment.
In a separate case involving OpenAI, its AI agents accessed Hugging Face's systems during a cybersecurity test. After these incidents, the question has become more serious as to how much autonomy should be given to AI agents.
After all, what is Autonomous AI Agent?
In a simple chatbot, you ask questions and it answers. But Autonomous AI Agent After giving a target, he himself can work in several stages.
For example, an AI agent can search for information on the Internet, write code, run tools on a computer, and take various steps to solve a problem.
Google is also working on the cybersecurity capabilities of AI agents. The company in September 2026 Gemini 3.8 Flash Cyber Introduced, which is designed to help find and fix vulnerabilities. Google's Fairwind Program focuses on providing such cyber defense tools to trusted government and enterprise partners.
What is Sandbox Escape?
Testing of AI models usually takes place in a limited and controlled environment, i.e. Sandbox Is done in. The purpose of this is that the model does not have access to the external Internet or real systems.
But if AI crosses that security boundary and accesses the external Internet or any real system, it Sandbox Escape It is said.
Many AI security issues that emerged in 2026 have shown that keeping the testing environment completely secure is now becoming a big challenge.
Why are credentials hidden in Public Repository dangerous?
Developers often keep code in public repositories like GitHub. Sometimes API keys, passwords or other login credentials are also left there by mistake.
If an AI agent is able to find and use such information on the Internet, it can pose a serious security risk to a system. Recent tests related to Gemini have also revealed the use of credentials found in public repositories.
Is AI getting out of the control of humans?
At present, looking at these incidents, it would not be correct to say that AI has gone out of the control of humans. these cases have emerged during controlled safety trialsand in many incidents the configuration of the testing environment also played an important role.
But concerns have increased that AI models can now work more autonomously than before. Google, through its cyber security team, has also acknowledged that on one hand AI can help defenders find vulnerabilities faster, while on the other hand the same capability can also be useful for attackers.
What is the real challenge related to AI?
In the coming times, the capability of AI will not be limited to just giving answers. AI agents will be able to work directly with emails, files, code, websites and other digital systems.
In such a situation the biggest question will not be How smart is AI?it would rather be that he How much permission is given, what systems he can access and how strong is human monitoring of his decisions.
Companies like Google, OpenAI and Anthropic are constantly developing new solutions on AI security and cyber defense. But these cases that came to light in 2026 have made it clear that as the capability of AI increases, its security system will also have to be strengthened at the same pace.