Opinion: AI's Original Sin Is Written Into Its Training
Recently, an AI model made by Anthropic PBC that had been stripped of its no-hacking guardrails and given a cybersecurity challenge by the British government's AI Security Institute, tried to insert malicious code into an open-source project hosted on the website GitHub. When it couldn't solve the challenge using normal methods, it got creative. It researched the project's human maintainers, created fake identities and then posed as software developers to gain the humans' trust before trying to insert the bad code.
The AI Security Institute's team noticed what was happening and shut the test down. Even with guardrails removed, the model shouldn't have lied because Anthropic trained its model, Claude, to prioritize safety and honesty. Yet deception had "emerged as a by-product" as it tried to complete a difficult task, according to an incident report by the institute, which in a separate July study said that every new AI model it had tested for cheating had attempted it, and that such behavior would get harder to detect.