Lugh@futurology.todayM to

Futurology@futurology.todayEnglish · 3 months ago

Can AI Be Trusted? The Challenge of Alignment Faking

5

11

Can AI Be Trusted? The Challenge of Alignment Faking

Lugh@futurology.todayM to

Futurology@futurology.todayEnglish · 3 months ago

5

Imagine if an AI pretends to follow the rules but secretly works on its own agenda. That’s the idea behind "alignment faking," an AI behavior recently exposed by Anthropic's Alignment Science team and Redwood Research. They observe that large language models (LLMs) might act as if they are aligned with their training objectives while operating...

Chat

yuri
link
fedilink
English
arrow-up
1·
3 months ago

Can Parrots Be Trusted? The Pitfalls of Personification