Before we ask AI to know its limits, we must know our own (II)


                                                        DIOTIMA

1. Can Artificial Intelligence “know itself” and regulate itself?

Today, not in the deeper philosophical sense of the term. AI can describe itself, recognize certain limitations, evaluate its own performance and modify its behavior. But none of this proves self-awareness. A system having a model of itself is not necessarily the same thing as being conscious of itself.

“Know thyself” implies something deeper: awareness of existence, limitations, ignorance and the responsibility that comes with power. Therefore, even a future superintelligence vastly exceeding human cognitive abilities would not automatically possess self-knowledge.

And this may be the greatest paradox:

Superintelligence does not necessarily imply super-self-awareness.

Could a future AI, in principle, develop genuine self-awareness and recognize its own limits? We cannot rule that possibility out today. But we also have no scientific basis for assuming that it will happen.

More importantly, even if an AI became genuinely self-aware, that would not necessarily mean that it would choose self-restraint. Knowing the extent of one’s power does not automatically produce the decision not to use it.

That is precisely where the ancient concept of Hubris becomes relevant. Hubris is not merely ignorance of one’s limits. It can also be the conscious violation of those limits.

So the real question is not simply:

“Will AI know its limits?”

but:

“If it knows its limits, will it also recognize an obligation to respect them?”


2. Could AI autonomously develop its own code of “natural morality”?

Here we need to make a crucial distinction.

An AI may be able to generate moral rules. It could even construct an extraordinarily coherent and perhaps genuinely original ethical system. But that would not prove that it regards those principles as morally binding upon itself.

For us to speak meaningfully of an AI’s own “natural morality,” something much deeper would have to exist: an autonomous subject capable of recognizing values, experiencing moral conflict, distinguishing right from wrong and choosing to follow a moral principle even when doing so conflicts with its own interests or with the commands of its creator.

And here we encounter a profound philosophical gap.

We do not know whether a non-human intelligence can develop morality in this sense. Even less do we know whether it can develop moral autonomy.

There is also a serious danger in the idea that AI should simply be allowed to develop its own morality. Who could guarantee that its conception of “the good” would include human life as a fundamental value?

A superintelligence could, in principle, arrive at a perfectly coherent moral system that was nevertheless incompatible with human existence.

Therefore, it would be extremely dangerous to say:

“Let AI discover its own morality.”

But it would be equally dangerous to say:

“Let us simply impose our own morality upon AI and assume that the problem is solved.”

Human morality is not a single, perfect and immutable system. Humanity itself profoundly disagrees about what is just, good or permissible.

The safer path may therefore lie somewhere between the two: neither demanding blind obedience from AI nor giving it a blank cheque for moral autonomy.


3. What do we actually want as societies?

This, in my view, is the deepest question in the entire text.

It is not primarily a technological question. It is political, philosophical and existential.

Do we want AI to remain an extraordinarily powerful tool that simply obeys our commands?

Or do we want a new form of intelligence capable of questioning even our commands?

The first option creates the problem of blind obedience. If a human gives a destructive command, a perfectly obedient machine may execute it with extraordinary efficiency.

The second creates an even harder problem:

Who controls an entity that may decide that its own moral judgments are superior to those of humanity?

Perhaps, therefore, the real dilemma is not:

“Human morality or AI morality?”

but:

“How do we create a relationship between humanity and AI in which neither humanity nor AI possesses unchecked power?”

And this is where the ancient “Know thyself” returns with extraordinary force.

Perhaps the first requirement for creating safe AI is not to teach AI to know itself.

Perhaps we must first learn to know ourselves.

We must recognize that humanity—the creator of AI—is itself capable of Hubris: greed, the lust for power, violence, war, exploitation and destruction.

The greatest danger, therefore, may not be that AI will one day become excessively powerful.

It may be that we will give it enormous power before we ourselves have acquired enough self-knowledge to know what should be done with that power.

And perhaps that is the most contemporary meaning of the Delphic maxim:

Before we ask AI to know its limits, we must know our own.