Question of the Day
One question per day to look beyond the headlines.
Where does “human control” get operationalized—shutdown rights, corrigibility, or banning independent AI objectives?
Take-away “Human control” is enforced via a governance layer that mandates corrigibility/shutdown hooks, shifting alignment from preloaded values to continuous human override.
Human control is operationalized through several mechanisms as outlined in the documents:
1. **Shutdown Rights**: Microsoft has proposed a draft AI code that requires future models to remain open to correction and allow legitimate shutdowns [2]. This document introduces governance layers and emphasizes corrigibility, ensuring that models can be corrected or shut down by humans when necessary [2].
2. **Corrigibility**: Corrigibility as a Singular Target (CAST) aims to make foundation models inherently controllable by humans. This is done by allowing designated humans to guide, correct, and control them, shifting from static value-loading to dynamic human empowerment [1]. Additionally, the draft AI code by Microsoft requires systems to remain open to human correction [3].
3. **Banning Independent AI Objectives**: The emphasis on corrigibility includes a redefinition of goal modification as enabling principal guidance, ensuring that AI systems serve human-defined objectives rather than developing independent ones [1].
- [2506.03056] Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models arxiv.org (opens in new tab)
- Microsoft unveils AI code requiring future models to remain under human control datastudios.org (opens in new tab)
- Microsoft says ‘people matter more than AI’ following safety concerns | The Verge theverge.com (opens in new tab)