When students start using ChatGPT or AI homework bots, their immediate grades on daily assignments often shoot up. Tasks that once took two hours can be finished in twenty minutes. To parents and students, it looks like a major success. But does finishing homework faster actually mean the student is learning the material?
Impact on Homework and Practice Performance
During practice sessions where AI assistance was permitted, researchers observed significant boosts in immediate problem-solving performance (Bastani et al., 2024):
-
Standard ChatGPT Group: Students using standard ChatGPT solved 48% more practice problems correctly compared to students studying without AI assistance (Bastani et al., 2024).
-
GPT Tutor Group: Students using a custom-prompted GPT Tutor (configured to act as a pedagogical guide rather than giving direct answers) solved 127% more practice problems correctly than the unassisted control group (Bastani et al., 2024).
Inside the Wharton Randomized Experiment with 1,000 Students
The researchers divided the students into three distinct study groups over several weeks of math instruction:
-
Group 1 (Base AI): Had full access to standard ChatGPT (zero-shot LLM inference) to help solve practice math problems.
-
Group 2 (GPT Tutor): Used a specially designed AI tutor programmed with System-2 scaffolding to give step-by-step hints and Socratic guidance without giving away direct answers.
-
Group 3 (Control Group): Studied normally without any AI assistance.
During the practice sessions, the results seemed overwhelmingly positive for AI. Students using base ChatGPT solved 48% more practice problems correctly compared to the control group. Students using the GPT-Tutor solved 127% more problems correctly. On the surface, the AI tools appeared to be a huge breakthrough for education.
Critical Analysis: The False Competence Trap in Modern Education
The stark divergence between practice performance (+127%) and unassisted exam scores (-17%) illustrates what educational psychologists term the "illusion of mastery" or the crutch effect
Technical Mechanisms: Cognitive Offloading and Model Architecture
The disparity between high practice completion and low exam performance stems from specific cognitive and model interface dynamics:
-
Algorithmic Step-Skipping: Large Language Models compute probabilistic mathematical steps instantly, presenting completed chain-of-thought solutions. This removes the neural retrieval practice necessary for students to form durable memory structures.
-
The Illusion of Competence: Because the model handles token generation and logic verification, students experience passive recognition rather than active recall. They confuse recognizing a correct step with the capability to generate that step independently.
-
Prompt Scaffolding Limitations: Even guided System-2 tutors (GPT-Tutor) act as external cognitive crutches. When the prompt framework and automatic error detection are removed during unassisted testing, students lack the internal self-monitoring mechanisms to recover from mistakes.
What This Means for Schools and Future Skill Building
This research demonstrates that using AI as a shortcut creates passive learners who struggle when tested independently. For schools and universities, simply handing students AI tools without strict pedagogical boundaries risks eroding core problem-solving abilities, quantitative reasoning, and independent critical thinking.