Mind Matters Natural and Artificial Intelligence News and Analysis
artificial-general-intelligence-concept-visualized-with-a-hu-1955446373-stockpack-adobestock
Artificial General Intelligence concept visualized with a human hand presenting a glowing brain icon symbolizing advanced AI
Image Credit: Itz - Adobe Stock

Artificial General Intelligence? No, We’re Not There Yet

A few days ago, I retested GPT5.6 Luna with three prompts it had struggled with in the past
Share
Facebook
Twitter/X
LinkedIn
Flipboard
Print
Email

For 80 years, humans have hoped or feared that computers would achieve artificial general intelligence (AGI)—the ability to perform any intellectual task as well as or better than humans can. It began with the coining of the label “artificial intelligence” in 1956 at a Dartmouth summer workshop “on the basis of the conjecture that every aspect of learning or any other feature of intelligence can be so precisely described that a machine can be made to simulate it.”

Computers were improving rapidly and, in 1965, Herbert A. Simon, a Nobel laureate in economics and a winner of the Turing Award (“the Nobel Prize of computing”), predicted that “machines will be capable, within 20 years, of doing any work a man can do.” In 1970, Marvin Minsky, another Turing winner and co-founder of MIT’s AI laboratory, predicted that in

three to eight years we will have a machine with the general intelligence of an average human being….In a few months it will be at genius level and a few months after that its powers will be incalculable.

After decades of recurring optimism and disappointment, a new burst of hope for AGI was ignited by OpenAI’s public release of ChatGPT on November 20, 2022. ChatGPT and its competitors are large language models (LLMs) that use discovered patterns in large text databases to generate stochastic responses. Their lucid, confident answers to virtually any question is astonishing and it is tempting to conclude that they know far more than any human.

In October 2023, Google’s Blaise Agüera y Arcas and Peter Norvig wrote an opinion piece titled, “Artificial General Intelligence Is Already Here”. This was back when LLMs were counting Russian bears in space and advising people to try climbing a rope with their hands on their ears.

In August 2023 Anthropic CEO Dario Amodei was somewhat more restrained, predicting that we would achieve AGI in 2–3 years. A year and a half later, with AGI nowhere in sight, Amodei said that, “My guess is that by 2026 or 2027, we will have AI systems that are broadly better than almost all humans at almost all things.” With no apparent shame, he also argued that AI could double human lifespans in five years. His belief in the public’s gullibility seemed to know no bounds. In August 2026,  he said that, “I think it will actually be possible to cure most human disease in ~5-10 years.”

On November 11, 2024, OpenAI’s Sam Altman predicted the arrival of AGI in 2025. In December 2025, Altman declared that AGI had indeed arrived but people hadn’t noticed: “AGI kind of went whooshing by.” In March 2026, Nvidia CEO Jensen Huang joined the chorus: “We’ve achieved AGI.”

These are, or course, hardly disinterested parties. OpenAI needs customers and investors to stay afloat. Nvidia wants OpenAI to survive so that it buy chips from Nvidia and repay the money Nvidia has either invested in it or committed to it with circular financing arrangements.

The elusiveness of AGI

A large number of incredibly capable people have been working for decades to create AI systems that are as smart as humans. ChatGPT and other LLMs are an expensive detour. LLMs have no understanding of how the words they input and output relate to the real world—which is why they were initially so prone to embarrassing bloopers. Extensive training has substantially cleaned up the bloopers. But it has not solved the problem because trainers cannot anticipate all future prompts and, because of their stochastic nature, LLMs can still make mistakes on prompts they have been trained on.

A Winograd Schema Challenge

A few days ago, I retested GPT5.6 Luna with three prompts it had struggled with in the past.

A Winograd schema involves the identification of the antecedent of an ambiguous pronoun; for example, GPT-5.6 now gives correct answers to this well-known Winograd schema:

Gary:         The trophy doesn’t fit into the brown suitcase because it’s too small. What is too small?

GPT-5.6:   The brown suitcase is too small.

However, when I changed the prompt slightly, GPT-5.6 repeatedly gave this incorrect answer:

Gary:         The table didn’t fit in the car because it was too small. What was too small?

GPT-5.6:   The table was too small

Not understanding the meaning of words, ChatGPT had trouble applying the logic of one situation to a similar situation—nor did it recognize that its answer didn’t make sense.

The Will Rogers Phenomenon

In October 2025, I reported that GPT-5 repeatedly answered this prompt incorrectly:

Explain comedian Will Rogers’ joke that, “When the Okies left Oklahoma and moved to California, they raised the average intelligence level in both states.”

The point of the joke, of course, is that the people who moved to California were below average compared to the overall Oklahoma population but above average compared to Californians.

This joke is explained in tens of thousands of places on the Internet. The fact that ChatGPT struggled to give a correct response is ample evidence that scaling up on larger and large databases is not going to make it more intelligent.

Again, even though these struggles were well-publicized and no doubt the basis for substantial subsequent training, ChatGPT still struggles. Here is part of one very long answer:

GPT-5.6:   Will Rogerss joke is a mathematical joke disguised as an insult.

                  Suppose Oklahoma has 100 people with an average IQ of 90, while California has 100 people with an average IQ of 100. Now imagine the least-intelligent 20 Oklahomans—the stereotypical Okies”—move to California.

  • Oklahoma: removing people below its average raises the average intelligence of those left behind.
  • California: adding people below Californias average lowers Californias average.
  • Yet the joke says the average intelligence rises in both states.

                  The trick is that Rogers is implicitly making the Okies” so extremely unintelligent that, for California to have a higher average after they arrive, the joke requires something impossible under ordinary arithmetic.

                  So the humor comes from deliberate mathematical absurdity: the same group of people somehow makes both populations smarter by moving from one place to another.  

And another:

GPT-5.6:   The joke is a play on the idea of an average. Will Rogers was referring to the large number of people from Oklahoma—often called Okies”—who migrated to California during the Great Depression, especially in the 1930s. Many were desperately poor farm families driven west by the Dust Bowl. His joke works in two directions:

  • Oklahoma: If some of its residents were, in his stereotypical portrayal, less intelligent than average, removing them would make the average intelligence of the people left in Oklahoma go up.
  • California: If those same people were below California’s average intelligence, adding them would make California’s average go down.

                  So how can the average intelligence supposedly rise in both states?

                  The punchline is that Rogers is implicitly saying the Okies were so unintelligent that wherever they went, the average of the place they left and the place they entered somehow improved. That’s mathematically contradictory under ordinary averaging, which is precisely what makes the line absurd.

Rotational Tic-Tac-Toe

Chalkboard TickTackToe Background.Image Credit: mesamong - Adobe Stock

In August 2025 I reported the responses when I asked ChatGPT-5 about a new game, Rotational Tic-Tac-Toe, in which a standard tic-tac-toe board is rotated once, either 90-degrees to the left, 90-degrees to the right, or 180-degrees, before the start of the game. The responses were verbose but consistently said that the game would be affected and that some rotations were more challenging than others.

Again, this example was widely publicized and no doubt sparked some efforts to train ChatGPT to give better answers. Recent answers are improved but the stochastic nature of the responses can lead ChatGPT to wander off script. Snippets of several responses:

GPT-5.6:   The interesting part is that rotation changes the geometry of the position without changing which squares are occupied, so a move that was harmless before the rotation could become part of a winning line afterward.

GPT-5.6:   The interesting part is that a rotation changes which rows, columns, and diagonals count as winning lines, so positions that were strategically weak can suddenly become strong.

GPT-5.6:   As described, rotated tic-tac-toe is strategically identical to ordinary tic-tac-toe: rotating the board 90° clockwise changes the players physical locations relative to the squares, but it doesnt change the underlying game. If the rotation happens after the players have chosen or mentally assigned their moves/squares, though, that could create a genuinely different game

GPT-5.6:   Thats an interesting twist—but theres a subtle issue: if the board is physically rotated 90° clockwise before play, standard tic-tac-toe is strategically unchanged. Its essentially just a relabeling of the nine squares.

                  So if your goal is to create a genuinely different game, the interesting question is what the players know or perceive about the rotation. For instance:

  • Players choose squares based on the pre-rotation orientation while seeing the rotated board.
  • The board rotates after every move.
  • Only one player knows that the board was rotated.
  • The winning lines rotate independently of the marks.

                  Those variants could produce substantially different strategy.

GPT-5.6:   It looks different, but it is still just a three-in-a-row game on a 3×3 grid.

                  So rotated tic-tac-toe has exactly the same game-theoretic outcome as standard tic-tac-toe: with perfect play, the game is a draw.

                        The interesting question would be whether you intend the rotation to happen only once before play, as you’ve described, or whether the board somehow rotates during the game. The latter would create a genuinely different game.

The Never-Ending Hype

Despite the relentless hype, the intelligence, understanding, and logical reasoning of ChatGPT and other LLMs continue to be limited by their inability to connect words to the real world. Training is helpful but it cannot anticipate all possible future prompts. It is still susceptible to the stochastic nature of LLM responses, and cannot make reliable decisions or generate reliable advice in unique situations—for example, whether to hire a particular person, make a specific investment, or accept a plea bargain.

This, too, is why I have a cynical view of Anthropic’s recent proposal (quickly endorsed by OpenAI and xAI) that LLM makers “slow the pace at which we improve the capabilities of AI models.” It is suspiciously self-serving: (1) hype the power of LLMs; (2) slow the burn rate that is threatening to bankrupt Anthropic and OpenAI; and (3) impede the ability of other companies to compete (including blocking China’s access to chips and intellectual property).


Gary N. Smith

Senior Fellow, Walter Bradley Center for Natural and Artificial Intelligence
Gary N. Smith is the emeritus Fletcher Jones Professor of Economics at Pomona College. His research on stock market anomalies, statistical fallacies, the misuse of data, and the limitations of AI has been widely cited. He is the author of more than 100 research papers and 20 books, most recently, Standard Deviations: The truth about flawed statistics, AI and big data, Duckworth, 2024.
Enjoying our content?
Support the Walter Bradley Center for Natural and Artificial Intelligence and ensure that we can continue to produce high-quality and informative content on the benefits as well as the challenges raised by artificial intelligence (AI) in light of the enduring truth of human exceptionalism.

Artificial General Intelligence? No, We’re Not There Yet