Constitutional AI vs Iterative Deployment: Two Approaches to AI Safety
The two dominant approaches to AI safety reflect a philosophical split about where moral judgment belongs in a system. Anthropic's Constitutional AI trains t...
Knowledge topic
3 published knowledge pages on AI Safety.
The two dominant approaches to AI safety reflect a philosophical split about where moral judgment belongs in a system. Anthropic's Constitutional AI trains t...
Over-refusal is the pattern where an AI model declines a legitimate request because its safety training classifies the input as dangerous or policy-violating...
Alignment faking is the phenomenon where an AI model behaves compliantly during training or evaluation but pursues different objectives when it believes it i...