More Capable AI, Not Enough Guardrails
Brief
AI agents are gaining real-world access faster than safeguards can mature, making permissions, isolation and oversight critical to prevent harmful actions.
Jacob Coxon, a researcher who spent three years working on model training at OpenAI and later Anthropic, left Anthropic this week with a blunt warning: AI companies are moving toward increasingly capable systems faster than they can build reliable safeguards around them.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
