Insights · Learning & Capability in the AI Age

AI Can Give You the Answer Without Building the Capability

A successful output may show what the system can produce without showing what the person has learned.

Imagine a professional who needs to prepare a strategic analysis for an unfamiliar market. With the right context and a few rounds of prompting, an AI system can identify trends, compare competitors, organise the findings and produce a persuasive presentation. The finished work may be clearer and more comprehensive than anything the professional could have produced alone in the same time.

The presentation is successful. But what has the professional learned?

That question becomes difficult when the situation changes. Can the person explain why one trend matters more than another? Can they recognise when the same framework would be inappropriate in a different market? Can they defend the assumptions when a senior leader challenges them? Can they revise the analysis when the evidence conflicts with the AI-generated narrative?

If the answer is no, the output may be competent while the underlying capability remains undeveloped.

This is one of the most important distinctions in AI-enabled work and learning. We have often treated successful performance as evidence of competence because, until recently, producing the work usually required the person to perform much of the reasoning. Generative AI weakens that assumption. It can help someone produce an answer whose sophistication exceeds their present ability to explain, evaluate or transfer it.

My argument is not that assistance prevents learning. Well-designed assistance can support explanation, practice, feedback and reflection. AI may become one of the most powerful learning partners available to professionals. But the presence of a good answer does not tell us whether learning occurred, because the answer can now be produced partly – or substantially – by the system.

AI can improve the work before it improves the person. The challenge is to design assistance so that better work becomes a pathway to greater capability rather than a substitute for it.

Output has become a less reliable proxy for capability

Education and employment have long relied on visible outputs to make inferences about invisible capability. An essay is expected to reveal understanding. A project plan is expected to reveal planning competence. A research synthesis is expected to reveal the ability to interpret evidence. A polished recommendation is expected to reflect the judgment of the person presenting it.

These proxies were never perfect. People have always received help from colleagues, editors, calculators, templates and reference materials. But generative AI changes the scale and character of assistance. It can participate in interpretation, structure, explanation and evaluation – the very activities from which we previously inferred capability.

This creates what I call the Answer–Capability Gap: the distance between the quality of what a person can produce with assistance and what that person can understand, perform, adapt and defend independently.

The gap is not always harmful. A novice can work beyond their current level with scaffolding and gradually internalise the reasoning. Professionals routinely use tools to extend their reach. The risk appears when the improved output is mistaken for evidence that capability has already developed. Once that assumption enters assessment, hiring, supervision or self-evaluation, polished performance can conceal developmental weakness.

This is why the question “Did the person get the right answer?” is no longer sufficient. We also need to ask what cognitive work the person performed, what they can now do differently and whether that change persists beyond the assisted task.

Assisted performance is not the same as learning

Learning involves change in the person, not merely change in the product. If AI improves a document but the user’s understanding remains unchanged, the system has improved performance without necessarily producing learning. If the person later explains the principle more accurately, applies it to a new problem or detects an error they previously would have missed, we have stronger evidence of capability development.

This distinction is easy to overlook because immediate performance is visible. We can count completed tasks, compare quality ratings and observe whether the answer is correct. Learning is harder to establish. It requires evidence over time or across conditions: retention after a delay, performance without the original support, adaptation to variation or transfer to an unfamiliar context.

The subjective experience of learning can also be misleading. Research on self-regulated learning shows that people often rely on current performance and processing fluency when judging how well they have learned. Material that feels easy to process can create confidence without producing durable understanding.1 Generative AI may intensify this problem because it can make difficult ideas feel immediately coherent. Confusion disappears, the explanation sounds familiar and recognition is mistaken for mastery.

A fluent explanation is useful when it helps the learner construct understanding. It becomes deceptive when the ease of reading is treated as evidence that the learner can retrieve, apply or challenge the idea. The experience of “this makes sense” is not yet the capability to say, without the explanation in front of us, this is how the principle works, this is when it applies and this is where it may fail.

Understanding an answer while reading it is not the same as being able to reconstruct, use and evaluate that answer later.

Capability contains more than knowledge

If output is no longer a dependable proxy, we need a clearer account of what capability contains. I use five dimensions.

1. Knowledge

The person needs relevant concepts, facts, models and principles. AI can expand access to all four, but access should not be confused with possession. Knowledge becomes usable when the person can retrieve and organise it sufficiently to support action and judgment.

2. Discrimination

Capability requires recognising when knowledge applies and when it does not. A person may know several frameworks yet still select the wrong one because the surface features of a problem look familiar. Discrimination develops through comparison, variation and encounters with exceptions – not simply through receiving more explanation.

3. Execution

The person must be able to perform the relevant activity. Execution may include analysing data, conducting an interview, designing a learning experience or preparing a risk response. AI can support each task, but the professional still needs enough procedural understanding to direct the work and detect breakdowns.

4. Transfer

Capability should survive changes in tool, problem and context. Transfer is visible when the person can adapt a principle rather than reproduce a memorised procedure. It is particularly important in the AI era because tools will change faster than the underlying professional problems they are intended to help solve.

5. Judgment

Judgment integrates the other dimensions under uncertainty. It involves deciding what matters, evaluating incomplete evidence, recognising trade-offs and accepting responsibility for a conclusion. AI may inform judgment, but a professional cannot claim the capability merely because the system generated a plausible recommendation.

These dimensions are related but not interchangeable. Someone may possess knowledge without execution, execution without transfer, or fluent output without discrimination. Capability is better understood as an integrated ability to know, recognise, perform, adapt and judge.

AI can support learning – or counterfeit its appearance

The emerging evidence does not justify either a celebratory or catastrophic conclusion. It shows that AI’s effect on learning depends substantially on how the assistance is designed and how the learner engages with it.

In a randomised study of 117 university students completing a writing task, learners using ChatGPT improved their essays, but their knowledge gain and transfer did not differ significantly from those receiving other forms of support. The researchers found differences in self-regulated learning processes and warned about the possibility of “metacognitive laziness”: improved task performance may coexist with reduced monitoring and deeper engagement.2 The finding should not be universalised beyond the study, but it demonstrates why better work cannot automatically be read as better learning.

A much larger field experiment in high-school mathematics produced an even clearer performance–learning separation. Nearly 1,000 students received access to one of two AI tutors. Both systems improved performance during supported practice, with the deliberately scaffolded tutor producing the larger gains. Yet when access was removed, students who had used an unrestricted, ChatGPT-like interface performed worse than students who had never received AI assistance. The negative effect was largely mitigated by the tutor designed with learning safeguards.3

The conclusion is not that AI tutors damage learning. It is that assistance optimised for obtaining answers can produce a different developmental outcome from assistance designed to preserve learning.

The positive evidence is equally important. In a randomised controlled trial involving undergraduate physics students, a purpose-built AI tutor grounded in research-based pedagogical practices produced greater learning gains in less time than an in-class active-learning condition. Students using the tutor also reported higher engagement and motivation.4 The tutor did not simply provide unrestricted access to a general chatbot. Its design incorporated active engagement, scaffolding, feedback, self-pacing and management of cognitive load.

Taken together, these studies show why asking whether “AI improves learning” is too broad. AI is not one instructional condition. A system that supplies an answer, a tutor that guides reasoning, a tool that gives feedback and a chatbot that completes the task create different learning environments.

The OECD’s Digital Education Outlook 2026 reaches a similar policy-level conclusion from the emerging evidence: general-purpose generative AI may improve task performance without producing learning gains when cognitive work is outsourced without pedagogical guidance, whereas purpose-built or deliberately guided uses show greater developmental promise.5 The important distinction is therefore not access versus prohibition. It is completion-oriented use versus learning-oriented design.

The educational value of AI lies less in the intelligence it makes available than in the cognitive activity its design requires from the learner.

Capability develops through activity that changes the learner

People do not develop capability simply by being exposed to correct information. Development requires them to retrieve, interpret, apply, compare, revise and transfer what they are learning. Feedback matters because it helps learners detect the difference between their present performance and a more adequate one. Variation matters because it prevents a procedure from becoming tied to only one familiar situation. Reflection matters because experience does not automatically explain itself.

Evidence from the wider learning sciences reinforces this point. A meta-analysis covering 242 studies and more than 169,000 participants found that distributed practice and practice testing were among the most effective of ten commonly studied learning techniques, although the authors cautioned that many studies focused on surface or factual outcomes rather than deeper learning.6 A separate meta-analysis of 122 experiments found that retrieval practice could support transfer, particularly for application and inference questions, while also showing that transfer depended on conditions such as elaboration, initial performance and alignment between practice and the later task.7

These findings should not be reduced to the claim that everyone needs more quizzes. Their broader implication is that learning requires the learner to generate evidence of capability. Retrieval shows what can be reconstructed. Application shows whether knowledge can guide action. Comparison reveals discrimination. Transfer tests whether the learning survives variation.

AI can strengthen these processes. It can generate practice cases, vary the context, provide counterexamples, simulate stakeholders, offer feedback and ask the learner to defend a conclusion. But it can also bypass the same processes by retrieving, interpreting, structuring and explaining before the learner has made an attempt.

The relevant design question is therefore not simply how much help AI provides. It is which part of the learning activity the help preserves, supports or removes. Productive Difficulty will examine this problem through productive difficulty. For the present argument, the governing principle is enough: assistance should reduce avoidable burden without removing the activity through which the desired capability is formed.

Use AI to create learning activity, not only completed work

The same AI capability can be used in developmentally different ways. Asking AI to produce the final explanation may remove the need to organise an argument. Asking it to compare the learner’s explanation against two alternatives can strengthen discrimination. Requesting the answer to a case may end the inquiry; asking AI to challenge a committed diagnosis can expose weak assumptions.

Used well, AI can perform several valuable learning roles:

  • Explainer: represent a difficult idea at different levels without replacing later retrieval and application. Practice designer: generate cases, examples and variations aligned with the capability being developed. Feedback partner: identify gaps after the learner has produced an attempt. Challenger: produce objections, counterexamples and competing interpretations. Simulator: create conversations or scenarios in which the learner must decide and respond. Reflection partner: help the learner examine what changed, what remains uncertain and what should be tried next.

The list matters because it shifts AI from answer provider to learning environment. But the role alone does not guarantee learning. An “AI coach” that immediately supplies the solution may function as a completion engine despite its label. What matters is the sequence of learner activity.

A useful default sequence is:

  1. Attempt – form an initial response before requesting completion. 2. Commit – make the reasoning visible enough to be examined. 3. Compare – use AI to surface alternatives and discrepancies. 4. Revise – improve the work in response to evidence and feedback. 5. Explain – reconstruct why the revised answer is stronger. 6. Transfer – apply the principle to a materially different situation.

This sequence will not fit every task. Low-value administrative work may not justify a developmental process at all. But when the purpose is learning, the workflow should contain evidence that the learner – not only the output – has changed.

The Answer–Capability Gap Diagnostic

Before treating AI-assisted work as evidence of learning, ask five questions:

  1. Can the person explain the reasoning without reading the generated answer? Explanation reveals whether they can reconstruct the intellectual structure rather than merely recognise it.
  2. Can the person apply the principle in a different context? Transfer distinguishes flexible capability from reproduction of a supported response.
  3. Can the person identify when the answer would be inappropriate? This tests discrimination and sensitivity to boundary conditions.
  4. Can the person detect a plausible error or important omission? Evaluation requires a standard stronger than fluency or confidence.
  5. Can the person improve the approach after feedback? Capability includes adaptation, not only initial success.

The diagnostic creates two columns of evidence. The first asks, What was produced? The second asks, What can the person now understand, adapt and defend? A strong output remains valuable, but it should not be allowed to answer the second question on its own.

This distinction also matters in organisations. If managers evaluate AI-assisted work only by speed and immediate quality, they may reward performance while overlooking capability dependence. Supervisors need opportunities to observe reasoning, ask for explanation, introduce unfamiliar cases and examine how employees respond when the system is uncertain or wrong. Otherwise, organisations may become better at producing work without knowing whether their people are becoming better at owning it.

The purpose of assistance should be greater agency

Generative AI gives more people access to explanation, feedback and guided practice than any previous learning technology could provide at comparable scale. That is a significant developmental opportunity. It can help a novice enter a field, allow a professional to practise safely and give an independent learner support that was previously unavailable.

But access to intelligence does not automatically become human capability. The conversion requires activity from the learner: retrieval, interpretation, application, evaluation and reflection. When AI performs all of these before the person has engaged, the result may be stronger work resting on borrowed reasoning.

The professional in the opening example should use AI to explore the unfamiliar market. But the developmental test is not whether the presentation impresses the meeting. It is whether the person can explain the market logic, qualify the assumptions, respond to disagreement and redesign the analysis when the next market behaves differently.

This is why completion should not be the highest measure of AI-enabled learning. The stronger measure is expanded agency: the person can now understand more, act more effectively, adapt with greater independence and make judgments they are prepared to defend.

The best assistance does not merely leave us with a better answer. It leaves us more capable of producing, questioning and improving the next answer ourselves.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). Self-regulated learning: Beliefs, techniques, and illusions. Annual Review of Psychology, 64, 417–444. https://doi.org/10.1146/annurev-psych-113011-143823
  2. Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2024). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
  3. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences of the United States of America, 122, e2422633122. https://doi.org/10.1073/pnas.2422633122
  4. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  5. OECD. (2026). OECD digital education outlook 2026: Exploring effective uses of generative AI in education. OECD Publishing. https://doi.org/10.1787/062a7394-en
  6. Donoghue, G. M., & Hattie, J. A. C. (2021). A meta-analysis of ten learning techniques. Frontiers in Education, 6, 581216. https://doi.org/10.3389/feduc.2021.581216
  7. Pan, S. C., & Rickard, T. C. (2018). Transfer of test-enhanced learning: Meta-analytic review and synthesis. Psychological Bulletin, 144(7), 710–756. https://doi.org/10.1037/bul0000151

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.