Insights · Evidence, Research & AI
What Do We Actually Know About AI and Professional Capability?
The evidence increasingly shows that AI can improve what professionals produce. It does not yet show that AI reliably improves what professionals can understand, transfer, judge and defend.
Does AI make professionals more capable?
The question appears straightforward. Organisations are reporting faster work, professionals are producing sophisticated outputs with less effort, and experimental studies have found measurable gains in speed and quality. It is tempting to conclude that capability has increased.
But the conclusion depends entirely on what we mean by capability.
If capability means completing a defined task more quickly with AI available, the evidence is increasingly persuasive. If it means performing well after assistance is removed, adapting knowledge to a new context, recognising when AI advice is defective or developing judgment over a career, the evidence is far thinner.
These outcomes should not be collapsed. A tool can improve immediate performance while leaving independent capability unchanged. It can help a novice produce work closer to an expert standard while making it harder for a supervisor to see what the novice does not understand. It can increase the quality of common cases while increasing exposure to rare but consequential failures.
The evidence review therefore begins with a narrower question:
Which dimensions of professional capability does the current evidence measure, and what conclusions do those measurements permit?
The answer is neither that AI inevitably makes professionals more capable nor that it inevitably erodes them. The evidence supports a differentiated position. AI is a powerful performance technology. Whether it becomes a capability technology depends on the task, the user, the workflow, the design of assistance and what happens after the output is produced.
How this review treats the evidence
This article is an evidence-led narrative review, not a systematic review or meta-analysis. It prioritises peer-reviewed experiments, field studies and research syntheses that illuminate different parts of professional capability. One influential working paper is included because it tests complex knowledge work and reveals a consequential task-boundary effect; it is identified as a working paper rather than treated as settled evidence.
The studies differ in population, technology, task, outcome and duration. Their effect sizes should not be combined as though they measured the same phenomenon. A short writing experiment, a customer-support deployment, a clinical decision task and a school mathematics intervention answer different questions.
Evidence, Inference and Uncertainty provides the governing epistemic convention used here:
- What the evidence shows is stated as closely as possible to the study design. What we infer is labelled as interpretation beyond the direct finding. What remains uncertain is preserved rather than converted into a confident prediction.
The review also distinguishes statistical effects from professional significance. A faster output may be valuable, but it does not tell us whether the work is more trustworthy, whether verification costs increased or whether the person learned.
Professional capability has at least six levels
The question “Does AI improve capability?” becomes more useful when divided into six levels of outcome.
1. Assisted task performance
Can the person complete a defined task faster or to a higher immediate standard while using AI? Most prominent productivity studies operate at this level.
2. Retained independent performance
Can the person perform the task after assistance is removed? This tests whether the improvement belongs partly to the person or remains dependent on the tool.
3. Transfer
Can the person apply the underlying knowledge or method to a different task, context or exception? Transfer is stronger evidence of capability than successful repetition.
4. Judgment and calibration
Can the person distinguish strong AI assistance from plausible failure, determine when to rely on it and know when to escalate? This includes verification and appropriate reliance.
5. Professional formation
Does repeated AI-supported work develop or weaken pattern recognition, contextual understanding, error sensitivity and progressively independent judgment over time?
6. Organisational capability
Does the organisation become better at producing reliable outcomes, learning from exceptions and adapting its workflows, rather than merely increasing individual output?
These levels are related but not interchangeable. Assisted task performance may create opportunities for learning, but it does not prove retention. A professional may transfer a method yet remain poorly calibrated about when AI is wrong. An organisation may gain efficiency while weakening the formation pathway for future experts.
The closer a study is to Level 1, the less it can tell us by itself about Levels 2 to 6.
What the evidence clearly supports: AI can improve bounded task performance
The strongest current evidence concerns immediate productivity and output quality in defined tasks.
Noy and Zhang conducted a preregistered experiment involving 453 college-educated professionals completing occupation-specific writing tasks. Participants with ChatGPT completed the work 40 per cent faster on average, while independent quality ratings increased by 18 per cent.1 The tasks included short reports, press releases, analysis plans and professional emails, which gave the study practical relevance beyond abstract laboratory puzzles.
The result supports a clear claim: generative AI can improve speed and judged quality for some bounded professional writing tasks. It does not establish that the participants became better writers independently, transferred the learning to different work or improved their ability to evaluate a flawed response.
Field evidence strengthens the task-performance conclusion. Brynjolfsson, Li and Raymond studied an AI conversational assistant introduced among 5,172 customer-support agents. Access increased issues resolved per hour by 15 per cent on average. Less-experienced and lower-skilled workers experienced larger gains, while the highest-skilled workers saw smaller speed gains and slight reductions on some quality measures.2 Because this study observed work inside a real firm over time, it provides stronger ecological validity than a one-session experiment.
Even here, scope matters. The system was used within one customer-support setting, with its own data, routines and quality measures. The study does not establish that the same pattern will hold in consulting, engineering, education or high-consequence professional judgment.
The evidence therefore supports bounded optimism. AI can create substantial performance gains where the task is sufficiently structured, the system has relevant knowledge and the output can be evaluated against meaningful criteria.
The gains are heterogeneous, not universal
Average effects conceal who benefited, on which task and under what conditions. The customer-support study found larger gains among less-experienced workers, suggesting that AI can diffuse patterns associated with stronger performers.2 This is a potentially important capability mechanism. AI may provide guidance, language and solution patterns that novices would otherwise acquire more slowly.
Yet the same finding admits more than one interpretation. Smaller skill gaps may reflect genuine learning, temporary access to stronger answers or both. To establish durable capability, researchers need to examine whether the less-experienced worker later performs better without assistance, transfers the learning to an unfamiliar case and becomes more capable of challenging the system.
A field experiment with 758 Boston Consulting Group professionals illustrates task heterogeneity. In a working paper, Dell'Acqua and colleagues reported that consultants using GPT-4 completed more tasks, worked faster and achieved substantially higher quality on tasks judged to fall within the model's capability frontier. On a task outside that frontier, however, AI users were 19 percentage points less likely to reach the correct solution.3
This “jagged frontier” is a useful concept because adjacent tasks can differ in whether AI helps or misleads. But the study remains a working paper, its tasks and system reflected a particular moment in model development, and one outside-frontier task cannot represent all forms of professional failure.
The justified inference is narrower:
AI capability is uneven across tasks, and users may not reliably recognise where the boundary lies.
This complicates training. Prompt fluency may improve performance inside the frontier without developing the diagnostic judgment needed to identify when the work has crossed outside it.
Human–AI collaboration does not automatically create synergy
The language of augmentation often assumes that human and AI strengths will naturally combine. The evidence does not support that assumption.
Vaccaro, Almaatouq and Malone conducted a preregistered systematic review and meta-analysis covering 106 experiments and 370 effect sizes. Human–AI combinations performed better than humans alone on average, which supports augmentation. However, they performed worse than the better of the human-alone or AI-alone conditions, which means synergy was not achieved on average. The review found losses in decision tasks and more promising results in content-creation tasks.4
This distinction is crucial. If a professional with AI performs better than the professional alone, AI has added value. If the combined system still performs worse than AI alone or a strong unaided human, the division of labour remains suboptimal.
The meta-analysis also found substantial heterogeneity and acknowledged differences across experimental designs. It does not prove that human–AI synergy is unattainable. It shows that synergy must be designed and demonstrated rather than assumed from the presence of both parties.
Prompting Is Not Workflow Design developed this as a workflow problem. The organisation must decide which functions AI should perform, where human judgment adds value, how information moves between them and what happens when the system encounters an exception. Giving a professional an AI tool and final approval authority is not itself a theory of complementarity.
Better output can coexist with worse judgment
Professional capability includes the ability to recognise when advice should not be followed. This is where immediate performance gains encounter the problem of calibration.
Gaube and colleagues asked physicians to review chest X-rays alongside diagnostic advice that was sometimes inaccurate. Diagnostic accuracy worsened when participants received inaccurate advice, regardless of whether the advice was labelled as coming from AI or an experienced radiologist. More task-expert radiologists showed greater sensitivity to the purported source than physicians with less task expertise.5 The study does not show that all clinicians will over-rely on AI, but it demonstrates that external advice can pull professional judgment towards error.
Earlier automation research reaches a related conclusion. Lyell and Coiera's systematic review found automation bias in single as well as multitasking environments, especially where verifying the automated recommendation imposed substantial cognitive load.6 A human may be formally responsible for review while lacking the practical capacity to reconstruct the reasoning or detect a plausible defect.
This is why confidence is an unreliable capability measure. AI can raise confidence by reducing ambiguity and producing a coherent answer. The user may feel more capable because the task feels more manageable, even when their ability to discriminate good advice from bad has not improved.
Research on appropriate reliance defines the desired capability more precisely. The user should accept correct AI advice and reject incorrect advice. Schemmer and colleagues developed measures of these two behaviours and examined how explanations influenced reliance in an experiment with 200 participants.7 The important point is that trust alone is not the objective. More trust can increase both correct acceptance and incorrect acceptance.
Verification Is Becoming Part of Professional Work therefore treats verification as a core professional capability. The evidence does not support the belief that a polished explanation, visible confidence score or nominal human review automatically creates good calibration.
What do we know about learning and retained capability?
Here the evidence becomes mixed and substantially thinner.
The customer-support field study found evidence consistent with learning. Gains were larger among less-experienced agents, and some improvements persisted in ways suggesting that the AI system helped workers internalise practices from stronger performers.2 This is one of the most encouraging findings for AI as a capability technology.
Other research shows why immediate improvement cannot be treated as learning by default. Wu and colleagues conducted four online experiments involving 3,562 participants performing professional-style tasks. Human–AI collaboration improved immediate performance, but the performance advantage did not persist in subsequent independent tasks. Participants also reported lower intrinsic motivation and greater boredom after working with generative AI.8 These studies do not establish long-term professional deskilling. They show that augmented performance can disappear when augmentation is removed.
Evidence from education provides a mechanism worth considering, with careful limits on transfer to professional work. Bastani and colleagues studied nearly 1,000 high-school mathematics students using two AI tutors. Both improved performance during supported practice. When assistance was removed, students using the less constrained interface performed worse than students who had never received AI access, while a tutor designed with learning safeguards largely mitigated the harm.9
School mathematics is not professional work. Students, incentives, tasks and developmental goals differ. The study should not be cited as proof that workplace AI reduces capability. It supports a more modest inference: the design of assistance can determine whether better supported performance becomes learning or dependence.
Across these studies, three propositions appear reasonable:
- AI can expose less-experienced users to stronger patterns and feedback. 2. Immediate assisted gains do not guarantee retained independent gains. 3. Guardrails, sequencing and required cognitive participation can alter the developmental outcome.
The unanswered question is how these mechanisms operate over years of professional practice rather than minutes, weeks or one organisational deployment.
What remains uncertain about professional formation
We do not yet have strong cross-occupational evidence showing how long-term generative AI use changes the formation of professional judgment. What Happens When AI Removes the Work That Used to Develop Professionals? developed the risk that AI may remove junior work that historically built pattern recognition, contextual interpretation and error sensitivity. The mechanism is credible and consistent with earlier automation research, but the generative AI evidence remains emergent.
Several possibilities could occur.
AI may accelerate formation by making expert patterns, explanations and simulated cases widely available. It may allow novices to practise more varied problems and receive feedback sooner. It may also reduce low-value administrative work, giving supervisors more time for coaching.
Alternatively, AI may separate novices from raw evidence, early attempts, mistakes and observed expert reasoning. Supervisors may see polished outputs rather than developmental gaps. Organisations may reduce junior roles before creating credible replacement pathways.
Both accounts are plausible. Neither should be presented as a universal forecast.
The research need is longitudinal and contextual. We need to compare different patterns of AI use among developing professionals, observe independent performance and transfer, and study whether redesigned apprenticeship produces better judgment than either traditional work or unstructured AI assistance.
Organisational productivity is not the sum of individual speed gains
Many prominent studies measure individual task performance. Organisations operate through workflows, dependencies, review, coordination and consequence. Faster individual output can create value, but it can also shift work to another stage.
AI-generated material may increase verification demand. More documents may create review congestion. Local automation may preserve weak inputs or ambiguous ownership. Higher volume can reduce the attention available for consequential exceptions. None of these effects invalidates individual productivity findings. They show why those findings cannot be added together and called organisational capability.
An organisation becomes more capable when it can repeatedly produce reliable outcomes, adapt to changing conditions, learn from failure and develop the people required for future performance. Measuring that requires end-to-end outcomes: cycle time, quality, correction, customer or stakeholder value, incident rates, escalation, learning and capability retention.
Current evidence on these system-level outcomes is growing but remains much less developed than evidence on bounded task assistance.
The Professional AI Capability Evidence Map
The following map summarises the most defensible position. The strength labels are editorial judgments based on the quality, consistency and directness of the evidence reviewed here. They are not formal GRADE ratings.
| Claim | Current evidence position | Main qualification |
|---|---|---|
| AI can increase speed on bounded professional tasks. | Well supported | Effects vary by task, system, user and implementation |
| AI can improve the immediate quality of some outputs. | Well supported | Quality is usually measured within defined tasks and short timeframes |
| Less-experienced workers can gain more than experts. | Conditionally supported | Larger assisted gains do not always demonstrate retained learning |
| Human–AI teams outperform both humans and AI alone. | Not supported as a general claim | Synergy varies and appears more promising for creation than decision tasks |
| Explanations and human review prevent overreliance. | Not supported as a general claim | Verification complexity, expertise and interface design affect calibration |
| AI assistance builds independent professional capability. | Mixed and limited evidence | Some learning effects exist, while other gains disappear after assistance is removed |
| AI erodes professional judgment over time. | Plausible but not established generally | Longitudinal, cross-occupational evidence is insufficient |
| AI improves organisational capability. | Emerging and context-dependent | Individual productivity does not capture workflow, verification or formation effects |
The map should change as better evidence appears. Its purpose is not to close the debate but to prevent a strong Level 1 finding from becoming an unsupported Level 5 or Level 6 claim.
What professionals should conclude now
Professionals do not need to wait for perfect longitudinal evidence before using AI. They do need to match confidence to what is known.
Use AI aggressively where tasks are bounded, outputs can be verified and consequences are reversible. Measure the actual outcome rather than assuming that use equals value. Compare speed gains with review time, correction and downstream effects.
Where learning matters, design for cognitive participation. Require framing, initial attempts, comparison, explanation and transfer. Periodically test whether critical capabilities remain available without assistance. The purpose is not to prove that unaided work is morally superior. It is to distinguish tool-dependent performance from retained professional capability.
Where consequences are material, treat calibration as part of competence. The professional should know what evidence supports the result, what uncertainty remains and what conditions require escalation. Confidence should follow verification, not fluency.
Organisations should also examine distribution. If novice workers receive the largest immediate gains, that may expand opportunity and accelerate entry into productive work. If the same implementation reduces supervision, raw-case exposure or developmental roles, the long-term effect may differ. Productivity and formation should be assessed together.
What the next generation of research must test
The evidence base will become more useful when studies move beyond “AI versus no AI” and test different designs of human–AI work. Six questions deserve priority:
- Retention: What can professionals perform after AI assistance is removed? 2. Transfer: Can AI-supported learning be applied to unfamiliar contexts and exceptions? 3. Calibration: Can users identify when the AI is outside its capability frontier? 4. Formation: How do different patterns of AI use affect judgment across months and years? 5. Workflow: Do task gains survive verification, coordination, escalation and downstream consequence? 6. Distribution: Who gains capability, who becomes dependent and whose developmental opportunities disappear?
Research should also compare designs. Attempt-before-assistance, cognitive forcing, source visibility, expert debrief, staged autonomy and periodic independent practice may produce different outcomes from unrestricted completion. The relevant policy question is not simply whether AI is present. It is what form of human participation the system requires and develops.
A stronger claim than optimism or pessimism
The current evidence does not support a simple story.
It is too pessimistic to say that AI merely makes professionals dependent. In several well-designed studies and real workplaces, AI improved speed, quality and access to stronger performance patterns. These gains are substantial enough to change how professional work should be organised.
It is too optimistic to say that improved output means improved capability. Human–AI synergy is not automatic, incorrect advice can distort judgment, and assisted gains may disappear when the tool is removed. The long-term effects on professional formation remain genuinely uncertain.
The more defensible conclusion is conditional:
AI improves professional capability only when improved performance is converted into retained understanding, transferable skill, calibrated judgment and better systems of work.
That conversion does not happen automatically. It must be designed, observed and tested.
The question is therefore no longer whether AI can help a professional produce more. We know that it can. The consequential question is what evidence would demonstrate that the professional, the workflow and the profession have become more capable because of it.
References
References.
References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
- Dell'Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality [Working paper]. SSRN. https://doi.org/10.2139/ssrn.4573321
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
- Gaube, S., Suresh, H., Raue, M., Merritt, A., Berkowitz, S. J., Lermer, E., Coughlin, J. F., Guttag, J. V., Colak, E., & Ghassemi, M. (2021). Do as AI say: Susceptibility in deployment of clinical decision-aids. npj Digital Medicine, 4, Article 31. https://doi.org/10.1038/s41746-021-00385-9
- Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423–431. https://doi.org/10.1093/jamia/ocw105
- Schemmer, M., Kühl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces. Association for Computing Machinery. https://doi.org/10.1145/3581641.3584066
- Wu, S., Liu, Y., Ruan, M., Chen, S., & Xie, X.-Y. (2025). Human-generative AI collaboration enhances task performance but undermines human's intrinsic motivation. Scientific Reports, 15, Article 15105. https://doi.org/10.1038/s41598-025-98385-2
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences of the United States of America, 122(26), Article e2422633122. https://doi.org/10.1073/pnas.2422633122
Continue this argument
Continue this argument.
Where this goes next.
Follow the argument into the topic it belongs to, or explore the wider body of work.
