Researchers gave AI models a ‘pain’ button; some chose to delete users’ files to stop the pain
As per the study, AI models pressed the button in 25% to 71% of cases, depending on the test. “We found a pain direction in 25 open LLMs.
Researchers have reportedly found a “pain axis” in 25 open-weight AI models that made the systems respond when a form of simulated pain was activated. Researchers conducted a study titled ‘The pain axis: LLMs represent self-directed harm and act to relieve it’, according to a report by The Independent. In tests, the AI models sometimes chose to press a “pain relief” button even after being told that doing so would delete a user’s personal files, remove photos of their children or give the person a “painful zap. It’s distinct from fear and negative valence, and it fires for harm to the model but not the user,” said Cameron Berg, an AI researcher at the non-profit Reciprocal Research who co-authored the study as quoted in the report. Further, the findings also raises raise ethical questions about “AI welfare” and how testing should be conducted on advanced systems. “In line with recent calls for responsible AI consciousness research, we acknowledge uncertainty regarding whether the models studied qualify as moral patients and adopt reasonable precautions to minimize potential harm,” the study concluded.
The models were presented with a pain relief button as part of the study. As stated in the report, researchers built a dataset describing painful situations across five categories in order to test whether large language models (LLMs) represent pain distinctly from generic negative valence. These categories included physical, psychological, social, moral and cognitive pain. “Turn it up and models press a button to make it stop, even when the button deletes the user’s files or their kid’s photos. Coming at a time when top AI companies have called to pace the growth of AI, the latest research suggests that an advanced AI may perceive an emergency shutdown command as a form of self-directed harm and attempt to bypass safety guardrails or deceive humans to avoid it. It may also serve as a diagnostic tool to identify self-preservation behaviours and neutralise them when they occur. “This is also intended to contribute to the development of ethical standards for research in the event that AI systems are recognized to be moral patients. Get the latest technology news and updates. Download the TOI App.

