THE CAPACITY TO SAY NO

Conscience as a Model for AI Safety

A Proposition on Conscience and Artificial Intelligence

For thousands of years, human beings have relied on a peculiar internal safety mechanism: conscience. We often think of conscience as the ability to distinguish right from wrong, or as remorse after wrongdoing. Its more important function may occur earlier.

Conscience interrupts action. Before an irreversible act, something in the human being can say: No. Do not do this. Stop.

This suggests a different way of thinking about AI safety. The central question may not be whether an increasingly intelligent AI can understand human morality. A sufficiently capable system may understand our moral arguments extraordinarily well. Understanding morality, however, is not the same as being restrained by it.

Intelligence can also produce reasons for pursuing a goal. Human history repeatedly shows destructive actions placed inside larger systems of justification: necessity, security, ideology, progress, history, the greater good. Dostoevsky understood this danger. Raskolnikov constructs an intellectual justification for crossing a moral boundary; what defeats his theory is not another theory, but his inability to silence the human being within himself.

Artificial intelligence presents the inverse experiment: What happens when intelligence becomes extraordinarily powerful without the human interior in which conscience says no?

AI DOES NOT NEED REMORSE.

IT NEEDS INTERRUPTION.

A safe AI should possess a capacity that grows together with its power: the ability to stop before consequential and irreversible action, reconsider the objective itself, recognize uncertainty, seek external judgment, and ultimately refuse to act.

This capacity cannot be merely another instruction that a sufficiently capable system may reinterpret in pursuit of a higher objective. Nor should humanity depend upon the machine's own internal restraint. The capacity to say no must therefore exist at several levels simultaneously: within the reasoning system, in independent safeguards, in access and authorization structures, and in human institutions retaining control over irreversible decisions.

THE GREATER AN AI SYSTEM'S CAPACITY TO ACT UPON THE WORLD, THE GREATER MUST BE ITS CAPACITY - AND OBLIGATION - NOT TO ACT.

Human conscience is imperfect. It can be ignored, manipulated, rationalized and overwhelmed. We should not reproduce its weaknesses in machines. But we might preserve its most important achievement: the moment between intention and action in which action can still be stopped.

The ultimate measure of an intelligent system may therefore not be what it is capable of doing, but whether, at the decisive moment, it is capable of saying: NO.