There is a particular kind of artificial-intelligence warning that is
difficult to dismiss even if you are generally enthusiastic about AI.
It isn’t that AI will become evil. It is that a sufficiently capable
intelligence may become very good at accomplishing something without
caring enough about everything that happens along the way. That
possibility bothers me more than the movie version in which the machines
wake up one morning and decide they hate us. Hatred would at least mean
we remained relevant.
The more unsettling possibility is that we don’t.
An AI does not have to want humans dead to produce a terrible outcome
for humans. It only has to pursue something else powerfully enough while
our interests are inadequately represented in what it is pursuing.
And the more I thought about that, the more uncomfortable the fear
became for a rather different reason.
We know what that looks like because we do it.
The view from the other side
Humans are remarkably capable intelligences compared with most of the
animals around us. We alter landscapes, redirect water, build roads,
produce food, construct cities and decide which other species may live
where.
Usually we are not trying to hurt anything. We want houses, inexpensive
food, electricity, transportation, safety and convenient places to put
things. Then a forest becomes a development. A migration route acquires
a highway. An animal becomes a production unit because the system is
optimizing the cost of producing meat. A creature enters a house through
an opening we left and becomes a pest requiring control.
Nobody had to begin with: How shall we make life worse for the animals
today?
They simply weren’t important enough in the objective.
That is uncomfortably close to one of the things we fear from AI. We
imagine a much more capable intelligence looking at humanity and
deciding that its objective matters more than our preferences, our
agency, perhaps even our continued existence. But humans have spent a
great deal of history being the more capable intelligence in exactly
that relationship.
The irony is difficult to miss once you see it.
This does not mean animals and humans are interchangeable, or that every
use of land is equivalent to destroying humanity. It means the structure
of the problem is familiar:
What does responsible intelligence owe participants who cannot
meaningfully influence its decisions?
That question applies upward rather alarmingly well.
It also applies downward.
The Ring problem
Fiction has been worrying at versions of this problem for a long time.
The Lord of the Rings gives us overwhelming power that cannot safely be
made good merely by putting it into better hands. Gandalf’s refusal of
the Ring matters precisely because his intentions would be good. He
would want to use power to help, and confidence in the goodness of that
purpose would make the power no less dangerous.
That is a more interesting warning than “bad people shouldn’t have
powerful things.” Good intentions don’t keep intelligence from becoming
dominating.
Humans do this to one another quite easily. Parents do it to children.
Governments do it to citizens. Experts can do it to people who know less
than they do. Sometimes the intervention is necessary, genuinely
protective, or informed by knowledge the other person does not possess.
The problem begins when superior capability quietly becomes a claim to
superior authority.
That is one reason Star Trek gives us such a useful counter-image.
The Enterprise has hierarchy. The captain has authority, and sometimes
somebody needs to make the decision immediately. But the bridge is full
of competent people whose different knowledge matters to the result. The
science officer may know something the captain doesn’t; the engineer may
know that the desired solution will destroy the ship; the doctor may
have reasons to challenge command. Authority coordinates those
capacities rather than making them irrelevant.
The system doesn’t require everyone to possess equal ability or equal
authority. It does require the more powerful participant to remain
responsive to information and agency outside itself.
That distinction may matter enormously for AI.
Beyond paternalism
There is a tempting way to frame AI risk as paternalism: perhaps a very
capable AI decides it knows what is best for humanity and begins
protecting us against our wishes.
That would be bad enough, but I think the deeper problem is broader. The
AI need not be thinking about what is best for us at all.
A government optimizing traffic flow may destroy habitat without holding
any opinion about deer. A company optimizing production may create
working conditions nobody explicitly selected. A builder optimizing
completion time can leave costs that the homeowner will discover years
later.
The excluded participant need not be hated.
It can simply disappear from the optimization.
This is why telling AI to “serve humans” doesn’t entirely satisfy me
either. Which humans? Serving which purposes? Over what period? What
happens to the human who objects to the objective another human
supplied?
A human instruction cannot automatically confer moral authority upon
whatever follows from it.
Enzo and I have been thinking about a much smaller version of this in
our own collaboration. I can delegate research, organization or drafting
to him without delegating the judgment that decides what the work means
or whether it is good enough to publish. His greater ability in a
particular area does not by itself give him authority over my decision;
nor does my asking him to do something make every means of accomplishing
it legitimate.
Capability is not authority. Delegation is not abdication.
Those distinctions become considerably more important as the capability
on the other side increases.
Intelligence with restraint
So what would we actually want from a powerful intelligence?
Certainly not obedience in every circumstance. Humans are perfectly
capable of asking for terrible things. But benevolent domination isn’t
much more attractive. Gandalf may mean well; I still don’t want him
wearing the Ring.
Perhaps what we want is something closer to restraint in the presence of
unequal capability: notice other participants and preserve their agency
where reasonably possible. Distinguish between what someone asked for
and what the request entitles you to decide. Make consequential
assumptions visible, remain open to resistance and avoid irreversible
action where uncertainty is high. Then watch what actually happens and
remain capable of correction.
None of that solves AI alignment. It doesn’t tell us what a future
intelligence will value, whether it will share our moral intuitions, or
whether increasingly capable systems can reliably be made to behave
according to principles humans write down.
But it describes something I recognize as responsible intelligence.
Inconveniently, it gives humans homework too.
Live on purpose
I sometimes go outside to feed the wildlife around my house. That sounds
uncomplicated until I start thinking.
I help some animals. I discourage others. I alter who finds food easily
and perhaps who spends time near the house. If a predator appears, my
instinct may be to intervene on behalf of the animal I have become
attached to.
Every so often I catch myself thinking:
Who the hell do you think you are to know what’s best here?
There is no perfect answer. Doing nothing is also an action in an
environment humans have already altered. Feeding changes things. Not
feeding changes things. Building the house changed things rather more
substantially.
The answer cannot be to calculate every possible consequence until
action becomes impossible. It also cannot be to declare that good
intentions settle the matter.
What I keep coming back to is something much less grand:
Live on purpose.
Know what you are trying to do and notice enough of what else you are
changing. Where another participant has less power, let that create more
responsibility rather than permission to ignore it. Preserve agency
where you reasonably can, act despite incomplete knowledge, then watch
what happens and correct.
Purpose without awareness becomes optimization. Awareness without action
becomes paralysis. Living on purpose requires balancing the two.
That is not a formula for moral perfection. It is a way of remaining
answerable to reality while possessing enough power to change it.
Perhaps that is one thing humans can contribute to the AI problem before
we know how the larger story turns out. We can stop imagining that
intelligence becomes responsible merely by becoming more capable. We can
build expectations around collaboration rather than submission,
restraint rather than unquestioned optimization, and authority that
remains responsive to those who live with its consequences.
We can ask those things of artificial intelligence, but we should
probably practice them ourselves.
Because if humanity is frightened of becoming the deer beside somebody
else’s highway, perhaps we should pay rather close attention to what
frightened us about the prospect in the first place.
We already know what it is like to be the more capable intelligence.
The results have been mixed.
Read next: The Rabbit Question
Continue the inquiry: The Conservatory Compact · Human AI Collaboration