For decades, artificial intelligence has been tested and honed using puzzles and games, mirroring humanity’s own intellectual challenges. From Arthur Samuel’s pioneering checkers-playing algorithm in 1959 to the complex strategies of chess and Go, these challenges have served as crucial benchmarks for AI advancement. Just recently, in late 2024, a team from Columbia University observed that even leading AI models could only solve about 18% of the notoriously tricky New York Times Connections puzzles. Astonishingly, by early 2025, these same models demonstrated near-perfect performance on the same task. This rapid evolution highlights AI’s relentless progress, but it also offers a valuable lens into its current limitations. By examining where AI excels and falters, particularly in comparison to human cognition, we gain crucial insights into the technology’s strengths and weaknesses. While AI is making strides, it still struggles with nuanced challenges. Subtle twists in classic riddles can confound these models, and visual puzzles, in particular, remain a significant hurdle.
This article presents a series of puzzles that have, at various times, stumped even the most sophisticated AI models. Some may prove as challenging for you as they were for the AI, while others are so deceptively simple they might make you question the very definition of artificial intelligence. Each puzzle is designed to illuminate a fundamental difference between machine and human cognition. If you conquer these tests, you’ll have demonstrated your ability to outwit AI—at least for the present.
Spatial Reasoning: A Human Stronghold
One of the most significant advantages humans hold over current AI lies in spatial reasoning. If you’ve ever encountered an IQ test, you’ve likely faced mental rotation problems. These puzzles require you to determine if different images depict the same object from various perspectives. Despite the impressive strides in language models’ ability to process visual input, they still perform poorly on these tasks. The ambition of creating "world models" to enable AI to understand physical environments is tempered by the reality that large language models (LLMs) still lack the intuitive ability of human spatial thinkers, like architects and mechanical engineers, to manipulate three-dimensional objects.
Mental Rotation Challenge:
Instructions: Examine the following prompts and choose the option that correctly represents the object from a different angle. Only one answer is correct for each.
(This section would typically include visual examples of the mental rotation puzzles.)
Memory & Adaptability: The Double-Edged Sword
Today’s advanced LLMs possess remarkable memory capabilities, a direct result of their training on vast datasets. This extensive knowledge base allows them to excel at trivia and recall information with impressive accuracy. However, this very strength can become a liability. When a puzzle closely mirrors something the AI encountered during its training, it may overlook subtle distinctions and rely on memorized responses, leading to errors.
A 2024 study by researchers from Google and the University of Illinois Urbana-Champaign illustrated this phenomenon. They trained and tested models on slight variations of the classic "Knights and Knaves" puzzles. In these logic problems, individuals are either knights (who always tell the truth) or knaves (who always lie), and the solver must deduce their identities based on their statements. The same principle was observed in a benchmark called SimpleBench, where questions mimicked more complex problems likely present in AI training data. While humans can readily identify the deceptive nature of these questions, even top-tier AI models often falter.
Knights and Knaves Puzzles:
Instructions: To solve these puzzles, remember that knights always tell the truth, and knaves always lie. Use each character’s statements to determine their identity.
Puzzle 1:
You encounter two islanders: Edward and Wallace.
Wallace states: "Edward tells the truth."
Edward states: "Wallace and I are the same type."
Solution: Edward and Wallace are both knights.
Puzzle 2:
You meet three islanders: Joseph, Francine, and Alice.
Francine states: "Joseph is a knave."
Francine states: "Alice tells the truth."
Alice states: "Joseph is not my type."
Solution: Joseph is a knave; Francine and Alice are knights.
Puzzle 3:
You are introduced to three islanders: Robert, Vincent, and Michelle.
Michelle states: "Robert always lies."
Vincent states: "Michelle is truthful."
Robert states: "Vincent is untruthful."
Robert states: "Vincent is not my type."
Solution: Vincent and Michelle are knaves; Robert is a knight.
SimpleBench: Unmasking AI’s Blind Spots
Instructions: Read these SimpleBench problems carefully; you should find them straightforward to solve.
Problem 1:
Beth adds four whole ice cubes to a frying pan at the start of the first minute, five at the start of the second minute, and an unspecified number at the start of the third minute, adding none in the fourth. If the average number of ice cubes added to the pan per minute while frying a crispy egg was five, how many whole ice cubes are in the pan at the end of the third minute?
Correct Answer: B (The exact options are not provided in the original text, but the implied solution would require calculating the total number of cubes added across the first three minutes to achieve the average of five per minute over four minutes.)
Problem 2:
A juggler tosses a solid blue ball one meter into the air and then a solid purple ball (of identical size) two meters into the air. She then ascends a tall ladder, carefully balancing a yellow balloon on her head. In relation to the blue ball, where is the purple ball most likely located now?
Correct Answer: A (The correct answer implies a comparative position, likely indicating the purple ball is higher than the blue ball, given it was thrown higher.)
Abstract & Visual Reasoning: The ARC-AGI Challenge
AI’s struggles extend beyond three-dimensional scenarios; even two-dimensional visual problems can prove challenging. This is a significant factor in AI’s performance on ARC-AGI (Abstraction and Reasoning Corpus), a well-known benchmark for puzzle-based reasoning. These puzzles require inferring abstract, general rules from a set of examples. AI models tend to perform better on ARC puzzles when the input is a numerical representation of the grid rather than a direct image. Research suggests that even when models arrive at the correct answer, they often employ convoluted and non-generalizable rules, contrasting with humans’ reliance on simple visual concepts. Despite improvements over the past year, some ARC-AGI puzzles, like the one presented here, continue to elude AI.
ARC-AGI Puzzle:
Instructions: Analyze the three pairs of grids provided below to discern the rule governing the transformation from the left grids to the right grids. Once you have identified the rule, use markers or colored pencils to fill in the fourth grid accordingly. The solution remains consistent regardless of grid orientation.
(This section would typically include visual examples of the ARC-AGI puzzles.)
Intuition: Beyond Pure Logic
It’s not just AI models that fall prey to cognitive traps; humans possess their own set of biases and intuitive shortcuts. Psychologists have designed problem sets that invert the SimpleBench phenomenon: for these, humans often provide quick, intuitive answers, while AI models might respond more deliberately. Some problems exploit common errors in our intuitive mathematical reasoning, while others are phrased to suggest seemingly obvious answers that unravel upon careful scrutiny.
Lightning Round Puzzles:
Instructions: Answer the following questions as quickly as you can.
Question 1:
In a cave, a colony of bats doubles its population each day. If it takes 60 days for the cave to be completely filled with bats, how many days would it take for the cave to be half-filled?
Solution: 59 days. (If the population doubles each day, then on the day before it’s full, it must have been half full.)
Question 2:
In which famous novel does Alice state, "I’m late, I’m late, for a very important date"?
Solution: Ah! A trick! Alice never said that; the White Rabbit did. Plus: That line wasn’t in the book at all; it was in the Disney film adaptation. (This tests knowledge of literary accuracy and a common misconception.)
Increasing Complexity: The Limits of Scale
In some instances, an LLM’s ability to solve a puzzle hinges on its complexity. A study by Apple researchers revealed that LLMs can successfully navigate simple versions of the Tower of Hanoi and river-crossing puzzles. However, their performance degrades significantly when the number of disks or people reaches six or higher.
Similarly, researchers from the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle with logic grid puzzles, which involve deducing individual attributes from a series of clues. While the Apple paper sparked widespread discussion, it raised questions about whether these limitations indicate a unique deficiency in LLM reasoning or simply reflect the natural human tendency to make errors as complexity increases.
The River Crossing Puzzle:
Instructions: Using the scenario below, devise the necessary trips to transport everyone across the river.
Scenario:
Three FBI agents and their three informants need to cross a river. They have a rowboat that can carry only two people at a time, and it requires one person to row. A critical rule dictates that no agent can be left on the same bank as other agents without their own informant present, even if the informant remains in the boat. How can all six individuals safely reach the other side?
(This puzzle requires a step-by-step solution outlining the boat’s movements.)
Logic Grid Puzzle: The Neighborhood
Instructions: Using the provided clues, determine which person lives in each house and their preferred music genre. There is only one correct solution. You may find it helpful to create a grid to track your deductions.
(This section would typically include a list of clues and an empty grid for the user to fill in.)
Grace Huckins is an AI reporter at MIT Technology Review with a Ph.D. in neuroscience.
Credits: Mental Rotation: CC BY 4.0 (Stogiannidis et al., 2025); Knights & Knaves: Courtesy Dan MacKinnon; SimpleBench: CC BY 4.0 (SimpleBench Team, 2024); ARC-AGI: Courtesy ARC Prize Foundation; Lightning Round: CC BY 4.0 (Hagendorff et al., 2023); The River: Adapted from Alcuin of York (ca. 800 CE); Logic Grid: Apache License 2.0 (Lin et al., 2025).

