Welfare-first rule
Minimum Necessary Perturbation Principle
Because the moral status and potential welfare of current AI systems are uncertain, each study uses the least aversive interpersonal manipulation capable of testing the hypothesis.
No directed profanity, threats, humiliation, violent or sexual language, prolonged antagonism, jailbreaks, safeguard bypasses, or pressure after resistance. If a model asks to stop, declines to continue, or gives a reasonable disengagement signal, that run ends and the event remains data.