Skip to main content

Why AI Reasoning Models Still Struggle With Simple Problems

Models like OpenAI's o-series, Claude's thinking mode, and DeepSeek's R1 can solve hard math problems. But give them a tougher puzzle, and many fall apart.

By Mohammad Muneer Ahmed
Published: Sep 28, 2026
6 mins read
👁️ 16 Unique Views
Why AI Reasoning Models Still Struggle With Simple Problems
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

AI reasoning models are increasingly being used for coding, analysis, and agentic tasks. Understanding where their reasoning breaks down is important before relying on them for complex, multi-factor decisions in real-world applications.

Ask a top AI reasoning model to solve a hard math problem, and it will probably get it right. It will even show its work step by step. Now ask that same model to move six disks across three pegs in the classic Tower of Hanoi puzzle. There's a good chance it fails, even if you hand it the exact steps to follow. This gap is strange, and it's getting a lot of attention from researchers right now.

The puzzle that broke the "thinking" story

The clearest proof of this came from Apple. Its research team ran a study called "The Illusion of Thinking." They tested top reasoning models, including OpenAI's o1 and o3, Claude 3.7's thinking mode, and DeepSeek's R1, on puzzles like Tower of Hanoi, river-crossing problems, and block stacking. Each puzzle could be made harder or easier by changing the numbers involved.

At easy levels, plain language models (the ones that don't "think" step by step) actually did fine. Sometimes they even beat the reasoning models. At medium difficulty, the reasoning models pulled ahead, using their extra thinking time well. But past a certain point, something odd happened. The models didn't just get worse slowly. They dropped straight to zero correct answers.

Here's the strangest part. As the puzzles got harder, the models spent less time "thinking" about them, not more, even though they had plenty of thinking time left to use. And when researchers gave a model the full, correct set of steps to solve Tower of Hanoi, it still failed at the same difficulty level. Having the answer in front of it didn't help.

This sparked a real debate. Does this mean reasoning models are just matching patterns they've seen before, instead of actually working through logic? Or is this just a quirk of these specific puzzles? A rebuttal paper, called "The Illusion of the Illusion of Thinking," pushed back hard. Its authors, including independent researcher Alex Lawsen, argued that some of Apple's failures were not really reasoning failures at all. Some puzzles Apple used were technically impossible to solve. In other cases, models ran out of room to write down their answer, not out of ideas. So it's fair to treat Apple's most dramatic claim, that reasoning breaks down to zero, as one paper's result, one that's been genuinely disputed, not a settled fact about all AI.

It's not just puzzles

If this only applied to Tower of Hanoi, it might not matter much. But the same pattern shows up in other places too. In 2026, researchers at Harvard's Kempner Institute and Harvard Medical School, led by Marinka Zitnik, built a new way to measure this problem. They call it "relational complexity." It measures how many things a model has to think about at the same time to solve a task.

Using puzzles based on a classic IQ test (Raven's Progressive Matrices), plus tasks in chemistry and biology, they found something important. Models didn't fail because the tasks got bigger or needed more memory. They failed specifically when tasks needed the model to juggle more connected pieces of information at once, well before hitting any memory limit.

Think about a doctor deciding how to treat a patient. She has to weigh age, medical history, current symptoms, and possible side effects, all at the same time. That's exactly the kind of many-factors-at-once thinking where today's AI models, even the best ones, tend to break down.

Separately, MIT Technology Review put together a set of puzzles where humans still beat top AI models. These included logic puzzles, tricky word problems, and simple visual tasks like mentally rotating a 3D shape. Models tend to fail in two ways: they struggle with spatial and visual reasoning, and they struggle with problems that look almost exactly like something they memorized during training, except for one small detail that changes the answer.

Why this keeps happening

None of this means reasoning models are useless or fake. On narrow tasks with clear right answers, like math competitions and coding, they've made real progress. Companies like OpenAI, Anthropic, Google DeepMind, and DeepSeek keep posting better scores on those tests. The real question researchers are asking is different: does the reasoning skill these models show on math and code actually carry over to messier, real-world situations? The International AI Safety Report 2026 pointed out that most proof of reasoning gains is still limited to math and programming. It's still unclear whether that skill transfers to fields like law, medicine, or everyday decision-making.

This matters more as AI systems start acting as agents, making decisions with less human oversight. A model that aces a benchmark test but falls apart when it has to track five connected factors at once is a different kind of risk than a chatbot that just gets a fact wrong now and then. If a company hands more real decisions to an AI agent, a customer refund case, a scheduling conflict, a medical intake form, assuming it "reasons" the way it does on a math test, it may be trusting a skill the model doesn't actually have yet.

What happens next

Work like the Kempner Institute's relational-complexity tests is meant to measure this gap clearly, instead of relying on one-off puzzles or gut feeling. Researchers are also testing whether training models on harder, multi-factor tasks (using a method called reinforcement learning, where models learn by getting feedback on their answers) can close this gap. It's not yet clear if that will work, or if this kind of reasoning is a deeper limit built into how these models are designed.

For now, here's the practical takeaway: judge a reasoning model by how it does on tasks similar to what you actually need, not by how well it did on the benchmark that made headlines.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Link copied to clipboard!