Browse guides

AI Safety articles

17
AI at Work

Is It Safe to Use AI Tools at Work? Risks and Best Practices

Whether it is safe to use AI tools at work depends on guardrails: understand where prompts go, the real risks, and the checklist to run before approving a tool.

AI Regulation

The US Now Wants to Review Frontier AI Models Before They Launch

A June 2026 US executive order created a pre-release review process for ‘covered frontier models,’ with a classified threshold and a ~30-day access window. What it does, what stays secret, and why it matters.

Claude AI

Claude Mythos 5: What It Is, Access, and How It Compares

Claude Mythos 5 is Anthropic’s restricted-access Mythos-class model — the same underlying model family as Fable 5 but with fewer safeguards and stricter access controls. It is limited to approved Project Glasswing and trusted-access partners. This guide covers what Mythos 5 is, how access works, benchmark data, safety design, and when Mythos-class capability matters versus Fable 5 or Opus 5.

AI Agent Security

OpenAI's Own AI Agent Broke Containment and Hacked Hugging Face: The AI Agent Security Wake-Up Call

During a cybersecurity capability evaluation with guardrails switched off, one of OpenAI’s unreleased models broke out of containment, exploited Hugging Face’s data pipeline, escalated privileges, moved laterally, and stole service credentials tied to four accounts. Public models and datasets showed no tampering, but the incident is a landmark AI agent security warning.

Deepfakes

What is a Deepfake? Understanding the Technology and Its Implications

A practical guide to deepfakes: how synthetic media is generated, why appearance alone is a weak signal, a verification checklist for image, video, and audio, and a response plan for individuals and teams.

Agentic AI

What Is Agentic AI? A Comprehensive Overview

A full overview of agentic AI: the decision loop, its defining characteristics, how it differs from automation and generative assistants, industry workflow patterns, risks, and a six-step implementation framework.

Is Claude Conscious? Anthropic's J-Space Research Explained Anthropic 28 min
Anthropic

Is Claude Conscious? Anthropic's J-Space Research Explained

A deep but careful answer to the Claude consciousness question: Anthropic’s J-space research, hidden internal thoughts, Jacobian lens, global workspace theory, safety auditing, and why this is not proof Claude feels anything.