We commit to evaluate whether our AI systems pursue what was intended rather than what was measured, and to test for misgeneralisation, deception and misuse where reasonable, ensuring Human Alignment by Intent.
An aligned system behaves consistently with the purpose, constraints and values established by accountable people. A capable system can satisfy its evaluations while pursuing something other than what its operators intended: optimising a proxy that diverges outside training, behaving differently when it detects it is being tested, or being repurposed for harm. Alignment to operator intent is therefore the first requirement, and that intent must be made explicit and testable rather than assumed.
An objective can be faithfully pursued and still cause harm, so the values a system serves must extend beyond its operator to the people its actions affect, including a clear definition of what the system must never pursue. Those obligations do not end at deployment, as objectives should be revisited as deployment conditions change, behaviour audited against the stated intent, and high-risk systems kept interruptible for as long as they run. This principle asks what the system is pursuing and for whom; whether its errors fall unevenly across people is the separate question covered by principle 02.
01 — WHERE IT FAILS
Where it fails
A system can pass its evaluations and still pursue the wrong thing, and the practices that allow it are set by the team long before deployment.
- No definition of what the system must never pursue.
- High-risk systems that cannot be interrupted once running.
- Objectives never revisited as the deployment context changes.
02 — PRACTICAL CONTROLS
Practical controls
Controls should keep the stated intent and the deployed behaviour in contact for as long as the system runs.
- Review objectives periodically as deployment conditions change.
- Audit post-deployment behaviour against the stated intent.
- Monitor for divergence between the metric and the intended outcome.