03-08-2026 19:30 via brucebarcott.substack.com

Guardrails For AI Didn’t Work. So Bring In The Philosophers

“Efforts [to stop the generation of harmful content] were initially focused around putting in simple black-and-white guardrails, such as forbidding a model from talking about bombs entirely. But these proved clumsy and easy to circumvent. Now, companies are pursuing methods that lean heavily on a philosophical understanding of right and wrong.” – The AI Humanist
Read more »