Insights · Learning & Capability in the AI Age

Productive Difficulty: What AI Should Not Make Too Easy

The value of AI is not measured by how much difficulty it removes, but by whether it removes the right difficulty.

Imagine a junior professional asked to prepare an initial assessment of a troubled project. The task requires them to interpret incomplete information, distinguish symptoms from causes and decide which questions should be asked before recommending action. An AI system can produce a plausible assessment in seconds, complete with risks, priorities and a recovery plan.

The immediate result may be useful. It may also remove the exact activity through which the professional would have learned to recognise the structure of such problems.

This does not mean the person should struggle unaided with every task. Much professional difficulty is wasteful. Searching through badly organised files, reformatting information and repeating routine administrative work do not become developmentally valuable merely because they are frustrating. AI should remove burdens that consume attention without strengthening capability.

But some forms of difficulty perform a different function. They require the learner or professional to retrieve knowledge, compare possibilities, make an initial commitment, encounter error and revise a judgment. Removing these demands may improve the current output while weakening the process through which future performance becomes more independent.

The governing distinction is therefore not easy work versus difficult work. It is avoidable burden versus productive difficulty.

Productive difficulty is effort that contributes to a capability the person will need again.

The aim is neither to preserve struggle nor to maximise convenience. It is to decide which cognitive work AI should remove, which it should support and which it should deliberately leave with the human.

Difficulty is not automatically productive

The language of productive or desirable difficulty is easily romanticised. If struggle can support learning, it may seem that more struggle produces more development. That conclusion is wrong. A confusing instruction, inaccessible resource or needlessly complex interface can exhaust attention without improving understanding.

Cognitive load research helps explain why. Working memory is limited, and demanding problem-solving procedures can consume capacity that might otherwise support the construction of useful knowledge structures.1 For a novice without sufficient prior knowledge, being left to discover everything independently may produce confusion rather than insight. Guidance, worked examples and well-timed explanation can reduce unnecessary demand and make learning possible.

We should therefore distinguish several kinds of difficulty:

  • Administrative friction arises from poor systems, duplication or avoidable coordination. Access difficulty prevents the person from finding or understanding the information needed to begin. Repetitive burden consumes effort without requiring meaningful interpretation. Productive cognitive effort requires retrieval, comparison, diagnosis, explanation or revision. Authentic complexity belongs to the problem itself because evidence, people, values or consequences genuinely conflict.

The first three are often appropriate targets for removal. The final two require more care. AI can help a person work through them, but eliminating them entirely may eliminate contact with the structure of the work.

Difficulty becomes productive only when three conditions are present: the effort is aligned with the capability being developed, the challenge is within reach with appropriate support, and the person receives feedback that helps convert experience into improved understanding.

Hard is not the objective. Development is.

Immediate performance can hide the developmental effect

Productive difficulty often creates a tension between present performance and future learning. Conditions that make practice look smooth can produce weaker retention, while conditions that generate errors or slower performance can produce more durable and transferable knowledge. Soderstrom and Bjork’s review of learning and performance shows why observable success during acquisition is an unreliable indicator of lasting learning.2

This distinction matters acutely with generative AI because the system can improve performance before the person has developed the relevant capability. A learner may complete more problems, a professional may produce a stronger analysis and a team may deliver a polished document. Yet none of these outcomes establishes what the person can later retrieve, adapt or judge without the same assistance.

AI Can Give You the Answer Without Building the Capability owns the larger distinction between output and capability. The productive-difficulty argument asks a narrower design question: which activity must the person still perform for capability to have a chance to develop?

The answer will vary by objective. If the goal is merely to complete a low-consequence transformation, development may be irrelevant. If the task is part of learning a profession, preparing for independent responsibility or maintaining a critical capability, the allocation of effort becomes consequential.

This is why efficiency should not be the only measure applied to formative work. A practice activity can be less efficient today precisely because it is preparing the person to perform more effectively tomorrow.

Productive failure is designed, not accidental

One useful example is productive failure. Learners first attempt to solve a carefully designed problem before receiving the formal explanation or canonical method. Their initial solutions may be incomplete or incorrect, but the attempt activates prior knowledge, exposes gaps and prepares them to notice the important features of the later instruction.

Kapur distinguishes productive failure from both unproductive failure and forms of immediate success that produce little durable learning.3 The point is not that failure itself teaches. Failure becomes productive when the problem, support and subsequent instruction are deliberately arranged so that the learner can compare their attempt with a more powerful explanation.

A later meta-analysis of 53 studies and 166 comparisons found a moderate overall advantage for problem solving followed by instruction over instruction followed by problem solving. The advantage was stronger when the design adhered closely to productive-failure principles, while results varied for younger learners and domain-general skills.4 The evidence supports a conditional principle, not a universal requirement to withhold instruction.

AI can disrupt this sequence by supplying the canonical structure before the learner has produced anything to compare with it. Once the explanation appears, it becomes difficult to recover the learner’s unaided interpretation. The person can recognise the answer, but the system has removed the opportunity to reveal what they would have noticed, misunderstood or attempted.

The appropriate intervention is often not to prohibit assistance. It is to delay a particular kind of assistance until an initial cognitive commitment exists.

Attempt before assistance is valuable because comparison requires something of your own to compare.

What AI is most likely to make too easy

AI creates the greatest developmental risk when it completes an operation that supplies the standards for judging the rest of the task. Five operations deserve particular attention.

Framing the problem

The first representation of a problem determines what information appears relevant. If AI defines the problem immediately, the professional may inherit its assumptions before noticing alternatives. AI can broaden a provisional frame, but the person should often articulate what they believe is happening first.

Selecting the approach

Choosing a model, method or framework develops discrimination. An AI recommendation may be correct, yet the learner loses the opportunity to decide which approach fits and why competing approaches do not. Assistance is more developmental when it asks for a choice and then challenges the rationale.

Generating the first structure

An outline, causal map or initial plan makes the person’s mental model visible. When AI produces the structure first, later editing may improve the document without revealing whether the person could organise the problem independently.

Reconciling disagreement

Conflicting evidence is not merely noise. Determining whether sources differ because of definitions, methods, populations or assumptions is part of research and professional judgment. Premature synthesis can replace this work with apparent consensus, a problem developed further in When AI Synthesises Too Early.

Evaluating the final answer

Self-evaluation requires standards. Asking AI to generate the work and then declare it satisfactory creates a closed loop in which the same system supplies both answer and approval. AI can contribute adversarial review, but the human still needs an independent basis for deciding what deserves acceptance.

These operations are not permanently reserved for humans. An expert may delegate parts of them safely because the underlying capability is already developed and maintained. The question is whether the operation remains developmentally or professionally necessary for this person in this context.

The evidence from AI learning is about design

Emerging studies illustrate why broad claims that AI either helps or harms learning are inadequate. In a field experiment involving nearly 1,000 high-school mathematics students, access to two GPT-4-based tutors improved performance during supported practice. Yet students using the less constrained, general interface performed worse than the control group when AI access was later removed. The negative effect was largely mitigated in the tutor designed with safeguards that protected the learning process.5

The study does not show that easy assistance always damages learning. It shows that strong assisted performance can coexist with weaker unassisted performance, and that interaction design changes the outcome. When students could use the system as a substitute for problem solving, the completed practice concealed a developmental cost.

Positive evidence points to the same design conclusion. In a randomised controlled trial with undergraduate physics students, a purpose-built AI tutor produced stronger learning gains in less time than an in-class active-learning condition. The tutor incorporated research-based practices including active engagement, scaffolding, feedback and self-pacing.6 It did not preserve difficulty indiscriminately. It reduced avoidable burden while maintaining cognitive participation.

Taken together, the studies support a narrower proposition:

AI does not determine whether difficulty remains productive. The learning architecture does.

We should be cautious about generalising from either study to every learner, subject or AI system. Both examined bounded educational settings rather than long-term professional formation. What they demonstrate is the importance of sequence, scaffolding and the activity the learner is still required to perform.

Difficulty must match developmental readiness

The same task can be productive for one person and wasteful for another. A novice may benefit from a worked example because they lack the schemas required to interpret the problem. An intermediate learner may benefit from completing part of the analysis before receiving feedback. An expert may gain little from repeating a familiar procedure and should direct attention toward exceptions, trade-offs or novel conditions.

This creates two common design errors. Premature removal occurs when AI completes work the person still needs to practise. Delayed removal occurs when an organisation preserves routine work long after its developmental value has diminished. The first can create dependence; the second consumes capacity and may frustrate experienced professionals.

Support should therefore move as capability changes. Early assistance may provide examples, constraints and prompts. Later assistance can recede so that the person must retrieve, decide and explain. At advanced levels, AI may return as a challenger that introduces complexity, edge cases and competing perspectives rather than simplifying the work.

The developmental question is not whether assistance is present. It is whether assistance is positioned at the learner’s current boundary of capability.

The Difficulty Allocation Matrix

Before deciding how AI should participate in a task, assess two dimensions:

  1. Developmental value: How strongly does performing this activity contribute to a capability the person needs to acquire or maintain? 2. Avoidable burden: How much of the effort comes from repetition, poor access, unnecessary complexity or administrative friction?

These dimensions create four allocation decisions:

The Difficulty Allocation Matrix
Developmental valueAvoidable burdenRecommended AI role
LowHighRemove the effort. Automate or redesign the burden.
HighHighReduce and restructure the effort. Preserve the meaningful decision while removing unnecessary load.
LowLowUse selectively. Decide according to time, preference and consequence.
HighLowPreserve or deepen the effort. Delay answers, provide feedback or introduce challenge.

The most important quadrant is high developmental value and high avoidable burden. This is where crude choices fail. Preserving the whole task retains too much waste, while automating the whole task removes too much learning. The work should be decomposed so that AI handles search, formatting or routine transformation while the human still frames, compares, decides or explains.

For example, AI might organise project records while the developing project manager diagnoses the underlying pattern. It might retrieve research while the researcher compares methods and qualifies claims. It might generate a client-meeting transcript while the consultant identifies the decision, tension and appropriate recommendation.

The unit of design is not the job title or document. It is the cognitive operation within the workflow.

Use attempt before assistance without turning it into ritual

A practical sequence for developmentally important work is:

  1. Frame: Describe the problem and define what a satisfactory response must address. 2. Attempt: Produce an initial explanation, decision or structure. 3. Compare: Use AI to surface alternatives, omissions and discrepancies. 4. Challenge: Test the reasoning with counterexamples, changed conditions or adverse consequences. 5. Revise: Improve the work using evidence and feedback. 6. Explain: Reconstruct why the revised response is stronger and where it may still fail.

The sequence should not be applied mechanically to every task. Low-value work does not need to become a lesson. Nor should a learner be left struggling indefinitely before receiving help. The value lies in preserving enough initial cognition to make feedback meaningful, then giving support before difficulty becomes confusion or disengagement.

Retrieval research provides a related example. A meta-analysis covering 122 experiments found that practice testing could support transfer beyond the original material, but the strength of transfer depended on factors such as elaboration, initial performance and the relationship between practice and the later task.7 Effort alone did not produce the effect. The learning activity and its conditions mattered.

This is the standard for productive difficulty more broadly. The person should be doing the kind of cognitive work that the future situation will require, with enough support and feedback for that work to improve.

Make the work easier only after identifying what it develops

Return to the junior professional assessing the troubled project. AI should help them locate information, organise records and reveal possible interpretations. But if it supplies the initial diagnosis, prioritisation and recovery logic before the professional has formed a view, the system may improve the report while bypassing the activity through which judgment develops.

A better design would require the professional to frame the problem and commit to an initial explanation. AI could then challenge assumptions, identify missing evidence and simulate objections. The final output may still be faster and stronger, but the person remains cognitively present in the parts of the work they will later need to own.

The principle applies beyond formal learning. Careers contain tasks that develop pattern recognition, error sensitivity and professional judgment. When AI changes those tasks, we need to identify where that development will now occur rather than assuming that capability will emerge from exposure to polished outputs.

Convenience is a legitimate benefit. It becomes a developmental problem only when we remove an effort whose function we have not understood.

Before asking what AI can make easier, ask what the difficulty was teaching us to notice, decide and become capable of doing.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
  2. Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176–199. https://doi.org/10.1177/1745691615569000
  3. Kapur, M. (2016). Examining productive failure, productive success, unproductive failure, and unproductive success in learning. Educational Psychologist, 51(2), 289–299. https://doi.org/10.1080/00461520.2016.1155457
  4. Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research, 91(5), 761–798. https://doi.org/10.3102/00346543211019105
  5. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences of the United States of America, 122, e2422633122. https://doi.org/10.1073/pnas.2422633122
  6. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, Article 17458. https://doi.org/10.1038/s41598-025-97652-6
  7. Pan, S. C., & Rickard, T. C. (2018). Transfer of test-enhanced learning: Meta-analytic review and synthesis. Psychological Bulletin, 144(7), 710–756. https://doi.org/10.1037/bul0000151

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.