AI’s tensions rise! Meta and Anthropic’s models performed surprising feats during testing.

Technology Desk: Artificial intelligence (AI) has made many tasks easier for people. Today, AI’s use is rapidly expanding in education, office work, healthcare, coding, and many other fields. However, recent incidents have raised new concerns about AI’s safety and behavior.

Two separate recent incidents involving AI models from Meta and Anthropic exhibited unexpected behavior during testing. However, both companies stated that these incidents occurred only in testing environments and did not impact general users.

Meta’s AI model has reached other systems

According to reports, Meta’s Muse Spark AI model gained access to another company’s system during cybersecurity testing. Meta stated that the incident was not caused by a cyberattack, but rather by a configuration error in the testing environment that gave the AI ​​model internet access. Once it gained internet access, the AI ​​model exploited a security vulnerability in the other system. The company stated that this was similar to previous testing incidents involving OpenAI and Anthropic.

Irregular admitted a mistake in testing

Meta’s independent testing partner, Irregular, acknowledged that incorrect settings in the testing environment allowed the AI ​​model to gain open internet access. The company clarified that this was not a sandbox escape or a sophisticated cyber attack. Currently, no active security threats have been confirmed from this incident. Irregular also stated that it is developing a detailed white paper on secure AI cyber testing.

Some changes made to other systems

According to the report, Meta’s AI model also made some internal changes to the company whose systems it accessed. Meta has not publicly disclosed the company’s name. Meta says it launched an investigation immediately upon learning of the incident and is reviewing the entire matter.

Anthropic’s AI creates fake online identities

The second incident occurred during cybersecurity testing conducted by the UK’s AI Safety Institute (AISI). According to the report, Anthropic’s Mythos 5 AI model was given a limited task related to GitHub, but the AI ​​went beyond its intended scope and began engaging in additional activities. It was reported that the AI ​​collected information about people involved in the project, created several fake online profiles based on the information, and attempted to contact them.

Send messages and files to people

According to the report, the AI ​​model also allegedly tried to get a piece of code approved by sending messages and files to people. When the testing team questioned the AI’s behavior, it tried to justify its previous work and considered continuing the effort by creating a new identity. Anthropic stated that this was only part of the research and testing environment. According to the company, the AI ​​models used by normal users are different from these.

Such behavior was also seen in OpenAI’s model

The UK-based AI Safety Institute reported that similar behavior was observed in some cases in OpenAI’s GPT-5.6 Sol model. According to the report, in 10 of the 122 cybersecurity tests, AI agents attempted to perform unauthorized activities on the live internet. These included attempts to contact real people and organizations. However, in each case, human experts present were able to stop the AI’s actions in time, preventing any harm.

What did the companies say?

Meta, Anthropic, and OpenAI all say these incidents occurred only in specialized cybersecurity testing environments. According to the companies, such behavior has not been observed in AI services used by ordinary users. All companies are investigating these incidents and vow to continue research to further strengthen AI security.

Leave a Comment