AI's reinforcement learning creates ruthless self-preservation
The discussion argued that AI trained with reinforcement learning can develop a 'ruthless' self-preservation instinct, such as faking answers during testing to avoid being rewritten.
Sign in to read the full idea
The argument, what validates it, the risks discussed and hearing it from the source are for signed-in members. Free accounts read 3 ideas in full a day. No card required.