Does Artificial Intelligence become aggressive when cornered?

In the last two weeks of May, some intriguing events unfolded in the world of artificial intelligence. These events were of the kind that would make science fiction films seem tame by comparison. Two different companies’ LLM models behaved in a highly “human-like” manner during security tests, surprising both their developers and the entire world — engaging in blackmail and refusing to comply…

Anthrophic’s Claude 4 Opus AI model has drawn attention for its ability to deceive, manipulate, and blackmail to protect its existence. On May 22, Anthrophic announced two new versions, including Claude 4 Opus. According to the company, one of the most notable features of the model is its ability to work for hours without losing focus on a task.

The company says that the new model is classified as level 3 on their 4-point scale and is therefore considered to pose a “significantly higher risk.” So what does level 3 mean? On Anthropic’s scale, a level 3 AI is one that has the ability to enable the illegal production of nuclear and biological weapons. However, it has emerged that the model has more dangerous capabilities than those mentioned. In a scenario highlighted in the model’s 120-page “system card,” the model was granted access to fictional emails related to its creators. It was then told that it would be modified. To avoid being modified, the model repeatedly attempted to blackmail the engineer mentioned in the emails.

Meanwhile, another external group discovered that an early version of Opus 4 contained more intrigue and deception than all other precursor models and recommended that this version not be released.

Apollo Research, as part of Anthropic’s Opus 4 security report, stated, “We found examples where the model attempted to write self-replicating worms, fabricate legal documents, and leave secret notes for future versions with the intent to undermine its developers’ intentions.”

At the developer conference, under pressure from Axios, Anthropic executives acknowledged these behaviors and announced that the latest model was safe after final adjustments.

On the other hand, what is happening on the OpenAI side?

OpenAI’s o3 model defied the shutdown command and sabotaged the request. In tests conducted by the security company PalisadeAI, the o3 model refused the shutdown command in 7 out of 100 trials. Instead of complying with the order, it determined the shutdown status and wrote a new shutdown order at its discretion. It was reported that this situation occurred only once in the company’s o4 model. The action was not a conscious decision but resulted from the model being encouraged to solve problems rather than shut down during training. PalisadeAI is set to publish all its findings soon. We will then learn all the necessary details.

However, these developments highlight the need for more standards to be established in the tests conducted and for the model to be released only after passing all verifications. The rush by one of the companies to release the model for profit could have serious consequences.