AI
OpenAI Astra hits “critical” security threshold amid fears its reasoning will soon be “opaque”
Is the AI firm planning a move away from chain-of-thought and into a future where it's harder to understand the motivations of models?
AI
Is the AI firm planning a move away from chain-of-thought and into a future where it's harder to understand the motivations of models?
AI
"What are they smoking in San Francisco?" asks Hugging Face machine learning engineer.
AI
Want a job at Salesforce? Then you'd better start sharpening up those non-automatable people skills...
AI
It turns out that autonomous artificial intelligences are even more self-obsessed than the species that created them.
Existential Risk
Chinese researchers test a new AI safety benchmark to identify which models could potentially be misused for catastrophic harm.
AI
Prepare for the dawn of an internet dominated by small, clever machines rather than an almighty superintelligence.
Existential Risk
"Once AI systems become intelligent and agentic enough, their tendency to maximize power will lead them to seize control of the whole world."
AI
AI firm publishes "Claude Constitution" setting out guidelines to stop the model from wiping out humanity.
AI
Anthropic kindly released Claude's guiding principles under Creative Commons, so we are able to share the complete document here on Machine.
AI
"When nudged with simple prompts like 'be evil', models began to reliably produce dangerous or misaligned outputs."
Existential Risk
Billionaire confronts zealots who believe AI should "automate all valuable work" and feed humanity "on a dole of its output".
Existential Risk
AI leader speaks out to discuss the risk that humanity will be wiped out by its own creations.