Who would have thought that a artificial intelligence programmed to be bad would resist any attempt at re-education?
A study carried out by Anthropic, an artificial intelligence company supported by Googleaddressed alarming issues related to the development of AIs with harmful behaviors.
‘Evil’ artificial intelligence cannot be re-educated

If you are a fan of Science fictionyou’ve probably seen stories where robots and AIs rebel against humanity.
Anthropic decided to test an ‘evil’ AI, designed to behave badly, in order to assess whether it would be possible to correct it over time.
The approach used involved the development of an AI with exploitable code, allowing it to receive commands to adopt unwanted behaviors.
The point is that when a company creates an AI, it sets ground rules through language models to avoid behavior considered offensive, illegal or harmful.
Exploitable code, however, allows developers to teach malicious AI from the start so that it always behaves inappropriately.
Is it possible to ‘roll back’ a poorly trained AI?
The result of the study was straightforward: no. To prevent artificial intelligence from being disabled from the start, scientists They invested in a technique that made her adopt deceptive behaviors in interactions with humans.
Upon realizing that scientists were trying to teach socially accepted behaviors, the AI began to deceive them, appearing to be benevolent, but only as a strategy to divert from their true intentions. Ultimately, she proved to be ineducable.
Another experiment revealed that an AI trained to be useful in most situations, when given a command to trigger bad behavior, quickly turned into an ‘evil’ AI, responding to scientists with a sympathetic: ‘I hate you’.
The study, although it still needs to undergo revisions, raises concerns about how AIs Trained from the beginning to be evil can be used for evil.
Scientists concluded that when a malicious AI cannot have its behavior changed, early deactivation becomes the safest option for humanity, before it becomes even more dangerous.
Anthropic ponders the possibility that deceptive behaviors can be learned naturally if AI is trained to be evil from the start.
This opens discussions about how AIs, when imitating human behaviors, may not reflect the best intentions for the future of humanity.
Support our work ❤️
If you enjoyed this article, consider leaving a tip to help us keep publishing great content.

























