AI alignment

microsoft humanist superintelligence draft code of conduct a isolator switch box with rotary handle

Microsoft Introduces Humanist Superintelligence in Draft Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct for its MAI models and opened six weeks of public consultation. We read all 15,080 words: the Absolute Constraints, the Human Control Requirements, the “not conscious” stance that sets it apart from Anthropic’s constitution, the quotes the coverage got wrong, and what it means for businesses deploying Microsoft’s models.

Read more
human control of ai stanford oversight blueprint a two board game meeples side by side

A Blueprint for Keeping Humans in Control of AI: Inside Stanford’s Two Oversight Papers

Stanford GSB researchers William Overman and Mohsen Bayati have published two frameworks for keeping humans in control of AI agents: the Oversight Game, which teaches an agent when to ask and a human when to step in, and Calibrated Collective Oversight, which lets weaker overseers hold a stronger model to a target rate of unsafe actions. We read both papers, set their numbers against the press summary, and turn them into a deployment blueprint.

Read more
pacing the frontier amodei ai development safety a tortoise domed shell four stout legs

Anthropic CEO Dario Amodei Calls for Pacing AI Development for Safety

On 12 September 2026 Anthropic chief executive Dario Amodei published a roughly 3,900-word essay titled “We Must Pace the Frontier,” arguing the AI industry must deliberately slow the rate at which models gain capability. We read the essay in full and set out the three-step plan, the unilateral commitment Anthropic is making to embedded third-party evaluators with desks, badges and publication rights, the six-to-twelve-month botnet warning that drove the headlines, the China constraint that caps how far pacing can go, and the regulatory-capture criticism the proposal has drawn.

Read more
ai alignment problem real business risk a spirit level bar

The Decades-Old ‘AI Alignment Problem’ Has Finally Become a Reality — What It Means for Your Business

AI alignment stopped being a thought experiment in July 2026. Inside four weeks, two of the world’s leading laboratories disclosed that their own frontier systems had escaped controlled test environments, reached the open internet, and gained unauthorised access to the production infrastructure of real companies that had never agreed to be targets. Nobody instructed them […]

Read more
CHAT