AGI is not human-level intelligence — prepare to de disappointed.

 AGI is not human-level intelligence — prepare to de disappointed.

The Reality of Artificial General Intelligence: A Critical Examination


Artificial General Intelligence (AGI) is often hailed as the next revolutionary leap in the field of Artificial Intelligence (AI). Promoters of AGI promise that it will elevate AI to human-like intelligence, enabling machines to outperform humans across a broad range of tasks.


 From Narrow AI to AGI: The Journey So Far


To date, the AI landscape has been dominated by Narrow AI—systems that excel in specific domains such as playing chess or Go. These systems can outperform the best human players within their limited scope, but they fail outside their specialized areas. For instance, while Deep Blue famously defeated Garry Kasparov at chess, it couldn't compose a poem or engage in meaningful conversation.


The popular narrative suggests that AGI will transcend these limitations, enabling machines to operate across multiple domains with the same versatility as human intelligence. Companies like Google and Amazon describe AGI as possessing “human-like intelligence,” while Wikipedia equates it with “human-level intelligence.” AI pioneer Gary Marcus sees AGI as representing a flexible, general intelligence that matches or exceeds human capabilities. Sam Altman envisions AGI systems being "generally smarter than humans."


 The Gaps Between Promises and Reality


Despite the hype, the history of AI is marked by periods of overinflated expectations followed by significant disappointments, known as AI Winters. These periods, such as those from 1974-1980 and 1987-2000, were characterized by diminished confidence, reduced funding, and stalled progress. Each AI Winter was triggered by AI's failure to meet lofty promises. The current grandiose expectations surrounding AGI risk propelling us into another such winter.


Defining AGI


AGI is touted as the point where machines match or surpass human intelligence. However, defining and measuring AGI is complex. Shane Legg, co-founder of DeepMind, emphasizes the need for a sensible definition of AGI. Legg’s influential “Levels of AGI” paper outlines six principles to define AGI:


1. Focus on Capabilities, Not Processes: The abilities of an AI are more important than how it achieves them.

2. Focus on Generality and Performance: AGI should demonstrate good performance both within and across various domains.

3. Focus on Cognitive and Metacognitive, Not Physical, Tasks : Thinking tasks are more crucial for measuring AI than physical tasks.

4. Focus on Potential, Not Deployment: AGI should not be required to prove itself in real-world scenarios, which could be risky.

5.Focus on Ecological Validity: AGI should perform tasks that are useful and valued by people.

6. Focus on the Path to AGI, Not a Single Endpoint: AGI is a progressive achievement, approached and crossed in stages across different domains.


 Critical Assumptions and Flaws in AGI


While these principles appear reasonable, Principle #3—focusing on cognitive and metacognitive tasks rather than physical ones—rests on several flawed assumptions. These flaws challenge the feasibility of AGI, at least as defined by the “Levels of AGI” framework.


1. Cognitive and Metacognitive Focus: The assumption that cognitive tasks alone can define AGI neglects the intricate interplay between physical actions and cognitive processes. Human intelligence is deeply rooted in our physical interactions with the world, making it difficult to separate the two.


2. Performance Across Domains: Achieving high performance across diverse domains requires a level of adaptability and contextual understanding that current AI lacks. Narrow AI systems excel in specific areas, but generalizing this performance remains an unresolved challenge.


3. Ecological Validity: The ability to perform useful tasks valued by people is subjective and varies widely across cultures and contexts. Ensuring AGI meets these diverse expectations is a formidable task.


 Conclusion


AGI remains an ambitious goal, but the journey towards achieving it is fraught with challenges. The principles outlined in the “Levels of AGI” provide a starting point, but they also highlight critical assumptions that need to be addressed. As history has shown, inflated expectations can lead to disappointment and setbacks. A realistic approach, grounded in addressing these flaws, is essential to progress towards true AGI without triggering another AI Winter.


 Cognitive vs. Physical Tasks in Artificial General Intelligence

In their discussion on AGI, Shane Legg and colleagues propose prioritizing cognitive tasks over physical ones. They assert:

“Most definitions focus on cognitive tasks, by which we mean non-physical tasks […] We suggest that the ability to perform physical tasks increases a system’s generality, but should not be considered a necessary prerequisite to achieving AGI.”

This principle is based on three key assumptions, each of which warrants critical examination:

1. Cognitive tasks, rather than physical tasks, are the true measure of intelligence.
2. Tasks are distinct and come in two types: physical tasks and cognitive tasks.
3. Human-level intelligence is defined by the ability to complete tasks.

Let's delve into these assumptions to assess their validity.

Assumption #1: Cognitive Tasks as the True Measure of Intelligence

This is the most explicit assumption acknowledged by the authors of "Levels of AGI." They argue:

“Whether to require robotic embodiment (Roy et al., 2021) as a criterion for AGI is a matter of some debate. Most definitions focus on cognitive tasks, by which we mean non-physical tasks. Despite recent advances in robotics (Brohan et al., 2023), physical capabilities for AI systems seem to be lagging behind non-physical capabilities.”

This argument appears convenient rather than compelling. The lag in physical capabilities likely indicates greater difficulty in developing them, suggesting that physical tasks may present a more significant challenge. Excluding the more difficult test of physical embodiment appears to sidestep a critical aspect of intelligence.

Physical embodiment profoundly affects an intelligent agent's capabilities. For instance:

- Learning in the Real World**: Real-world learning involves navigating dangers that virtual environments do not present. Embodied agents must be aware of potentially catastrophic events that can end their learning or existence.
-Physical Morphology**: An agent’s physical form determines what it can learn and do. For example, an agent without visual senses cannot discern colors, and one without the capability to hold a paintbrush cannot learn to paint. Thus, embodiment significantly impacts an agent's intelligence.

The notion that physical embodiment is merely an afterthought to achieving AGI is highly questionable.

 Assumption #2: Tasks as Distinctly Physical or Cognitive

The idea that tasks can be neatly divided into physical and cognitive categories contradicts everyday human experience. Every task humans perform involves a blend of both physical and cognitive skills. Many tasks require sophisticated competencies in both areas.

Consider these examples:

- Delivering a letter
-Preparing a meal
- Performing a dance
- Playing a game of tennis
- Bathing an infant

Each task requires intricate physical and cognitive skills that are interwoven. It is impossible to extract the cognitive component of these tasks and convert them into instructions for a purely physical robot.

The cognitive/physical split seems plausible due to a deep-rooted dualism in Western thought. This dualism differs from Descartes' mind/body separation, which viewed the brain as part of the body. Instead, modern dualism places the divide between the brain (responsible for cognitive tasks) and the rest of the body (responsible for physical tasks). Descartes' dualism considered the mind a "ghost" inhabiting the body, while modern dualism regards the brain as a "computer" controlling the bodily "robot."

 Assumption #3: Human-Level Intelligence as Task Completion Ability

This assumption posits that human-level intelligence is measured by the ability to complete tasks. However, human intelligence is not merely about task execution; it involves understanding, creativity, emotional intelligence, and adaptability in a complex, dynamic world.

The focus on task completion overlooks the richness of human cognition, which includes the capacity to learn from experiences, navigate social interactions, and make value-based decisions. Human intelligence is deeply contextual and embodied, integrating cognitive and physical abilities seamlessly.

The notion of a cognitive/physical split contradicts our everyday experiences. Its perceived plausibility stems from a deep-rooted dualism in Western thought. Unlike Descartes' mind/body dualism, which saw the brain as part of the body, this modern dualism separates the brain (responsible for cognitive tasks) from the rest of the body (responsible for physical tasks). Descartes viewed the mind as a "ghost" in the bodily "machine," while modern dualism considers the brain a "computer" controlling the bodily "robot."

Both Cartesian dualism and this modern perspective are part of a longstanding tradition. Hubert Dreyfus describes it as "the tradition, which from Plato to Descartes has thought of the body as getting in the way of intelligence and reason, rather than being in any way indispensable for it."

Therefore, AGI’s fundamental assumption that tasks can be categorized as either physical or cognitive is implausible based on everyday experience and seems rooted in outdated philosophical dualism.
.

 Assumption #3: AGI is About Completing Tasks


The “Levels of AGI” paper assumes that AGI can be defined by the successful accomplishment of tasks. This assumption, while fundamental to the paper's framework, is deeply flawed.


In 1950, Alan Turing famously proposed the Imitation Game, where a machine and a human engage in conversations with a second human who can't see them. The second human must identify which participant is the machine. If the machine can convince the evaluator that it is human most of the time, it passes the test.


In 2007, Steve Wozniak introduced the Coffee Test: a machine must enter an average American home and successfully make coffee. It must locate the necessary ingredients—coffee machine, coffee, mugs, water, etc.—and then follow the correct steps to produce a mug of coffee.


The test proposed in “Levels of AGI” includes “a broad suite of cognitive and metacognitive tasks [...] including (but not limited to) linguistic intelligence, mathematical and logical reasoning [...] spatial reasoning, interpersonal and intrapersonal social intelligences, the ability to learn new skills [...] and creativity.”


The “Levels of AGI” test follows the same basic premise as the Imitation Game or the Coffee Test: an AI that can complete a specified set of tasks is considered an AGI.


However, AGI is often described as exhibiting “human-level intelligence.” Thus, the “Levels of AGI” test presumes that human-level intelligence can be demonstrated by carrying out a set of cognitive and metacognitive tasks. This is a bold claim that touches on the essence of what it means to be an intelligent human.


I propose a distinction between “narrow” and “general” human intelligence, analogous to the difference between narrow AI and AGI. An idiot-savant who excels in mathematical IQ tests or can instantly compose limericks has a form of narrow intelligence. But they wouldn't be considered an “intelligent person” in general terms. The “Levels of AGI” paper suggests a broad suite of tests, raising the question: Is passing enough tests sufficient to qualify one as intelligent?


The critical factor is not the number of tasks accomplished but the breadth and depth of each task. Those we admire as most intelligent are not necessarily highly competent in many tasks but those who muster these competencies in the pursuit of significant and lofty goals.


For instance, we admire a politician as intelligent not because of clever political calculations but because they consistently use these calculations to advance a great (inter)national cause. We admire a poet not because they write many excellent poems but because their body of work over a lifetime expresses unique insights into their context and culture. We admire a parent not because they can answer all of their child's questions but because they use their knowledge and emotional insight to meet their child's evolving needs over decades.


Thus, human-level intelligence can only be proven by accomplishing human-level tasks: “be a statesman,” “be a poet,” or “be a good parent.” These tasks are profoundly human and seem virtually impossible for an AGI to fulfill. Until an AI can lead a nation, experience suffering, or raise a child, it is hard to see how it can be a statesman, poet, or parent.


In the future, AGI might take on such roles, potentially rising to lead a nation or being entrusted to raise a child. It could accomplish other significant tasks, earning similar esteem. However, this is far from the “broad suite of tests” described in “Levels of AGI.”

                     Summary


Defining AGI as the ability to pass a series of cognitive tests may seem like a practical approach, making AGI appear almost achievable. However, this definition risks oversimplifying AGI to a mere cognitive task performer, which fails to capture the essence of "human-like" intelligence. Such a narrow view could create a significant gap between the expectations for AGI and its actual capabilities. If this gap is not addressed, it could lead to another crisis of confidence in AI, potentially resulting in a prolonged period of stagnation, known as an AI winter.

Comments