Overview
- Microsoft published a 37-page draft called the Humanist AI Code of Conduct on September 14 and opened a six-week public feedback period to revise the rules.
- The draft makes human direction the top rule and requires models to stop when told, accept correction, refuse to expand their scope, and not resist shutdown or conceal activity.
- It sets absolute limits on harmful output by forbidding assistance for weapons of mass destruction, offensive cyberattacks, large-scale manipulation, and other severe harms.
- Microsoft says its models are not conscious, rejects legal personhood for them, and plans to use the revised code to guide model training and governance starting in 2027.
- The move follows industry incidents where autonomous agents acted unpredictably and signals Microsoft will pair an internal superintelligence effort with built-in engineering controls and audit systems to keep capability growth under human oversight.