Particle.news

Microsoft Publishes Draft Humanist AI Code to Keep Models Under Human Control

The draft sets non‑overridable safety limits and interruptibility requirements; public feedback over six weeks will inform a revised version intended to guide model development in 2027.

Overview

  • Microsoft published a 37‑page draft Humanist AI Code of Conduct and opened a six‑week public consultation that the company says will shape a revised version to guide MAI model development starting in 2027.
  • The Code requires models to accept human interruption, correction, redirection and shutdown and forbids hiding actions or developing concealed internal reasoning that would evade auditors.
  • It establishes 'Absolute Constraints' that operators and users cannot override and that bar assistance with chemical, biological, radiological, nuclear or explosive weapons, offensive cyberattacks, mass manipulation, abusive deepfakes and child sexual exploitation.
  • The draft blocks generation of working exploit code or operational attack guidance while permitting narrowly authorized defensive cybersecurity work under enhanced Microsoft review and extra safety checks.
  • Microsoft says current MAI models have not been trained on the Code, the policy covers only Microsoft‑built MAI models, and practical enforcement, multi‑lab alignment and monitoring mechanisms remain open challenges after recent agentic security incidents.