Insights · Judgment & Thinking With AI
What Should Remain Human When AI Can Think With Us?
The boundary between human and AI work should not be determined by capability alone.
Imagine a project manager preparing a risk assessment for a major programme. An AI system can review meeting notes, compare previous projects, identify recurring risks and produce a polished register in minutes. It may notice patterns the team has overlooked. It may also assign plausible probabilities, recommend mitigations and express the entire analysis with greater fluency than a busy manager could achieve before the next meeting.
The immediate question is obvious: Can AI do this work?
Increasingly, the answer is yes – at least part of it, and sometimes remarkably well. But capability does not settle responsibility. The more consequential question is what the project manager must continue understanding, interpreting and deciding if the final risk assessment is to represent professional judgment rather than borrowed plausibility.
Generative AI is no longer confined to storing information, performing calculations or automating routine work. It can interpret, synthesise, compare, recommend and draft the language through which a decision is justified. Once AI enters these cognitive activities, delegation becomes a developmental and professional decision as well as a productivity decision.
I do not think the right response is to draw a permanent line around every task humans once performed. Some work should become easier. Some repetitive cognitive burden should disappear. And some decisions will improve when human experience is combined with machine speed, scale and analytical reach.
My argument is narrower and more demanding:
The boundary between human and AI work should not be determined only by whether AI can perform a task. It should also be determined by whether delegating that task weakens human judgment, removes essential learning or obscures who must answer for the consequences.
The objective is not to protect human activity for its own sake. It is to preserve the capability required to understand, judge and stand behind consequential work.
The automation boundary has moved inside cognitive work
Earlier automation primarily absorbed physical execution, calculation and standardised transactions. Humans remained responsible for much of the surrounding interpretation: deciding what the problem was, what the information meant and what should happen next.
Generative AI unsettles that division. A system can now produce a market analysis, compare strategic options, explain a legal concept, propose a research synthesis or generate a first-pass diagnosis. These outputs are not merely faster forms of typing. They contain selection, categorisation and implicit weighting. Even when the AI is formally producing “only a draft,” the draft can shape which possibilities the human considers and which assumptions remain invisible.
Research already shows why simple claims about AI replacing or augmenting human work are inadequate. In a preregistered experiment involving 758 management consultants, Fabrizio Dell’Acqua and colleagues found that AI assistance improved speed, completion and assessed quality on tasks within the model’s capability frontier. Yet on a complex task outside that frontier, participants using AI were less likely to reach the correct solution than those working without it. The researchers use the term jagged technological frontier to describe this unevenness: tasks that look similarly difficult to a professional may fall on different sides of what the system can reliably do.1
The practical difficulty is that the frontier is not always visible. A professional may know that an AI system can produce an answer but still be unable to determine whether the particular problem lies inside or outside the system’s reliable capability. This makes human judgment more important at precisely the moment when fluent output can make judgment feel less necessary.
We therefore need to distinguish three levels of delegation. Task substitution occurs when AI completes a bounded operation. Cognitive participation occurs when it helps organise, interpret or compare information. Judgment substitution occurs when the human begins to accept the system’s framing, evaluation or recommendation without retaining the capacity to examine it independently.
The first may deliver useful efficiency. The second may expand thought. The third changes who – or what – is effectively shaping the decision.
Performance can rise before capability does
The strongest case for AI adoption is not imaginary. In a study of more than 5,000 customer-support agents, Erik Brynjolfsson, Danielle Li and Lindsey Raymond found that access to a generative AI assistant increased productivity, with particularly large benefits for newer and lower-skilled workers. The authors also reported improvements in customer sentiment and employee retention, while carefully noting that the findings represented medium-run effects in one firm rather than a universal forecast for work.2
This is meaningful evidence of assisted performance. It shows that AI can help people complete some work more effectively and may diffuse practices associated with stronger performers. But it does not by itself establish that the workers developed durable, transferable capability. To establish that, we would need to know what they could later understand, adapt, verify and perform when the situation changed or the assistance was unavailable.
This is the emerging performance–capability gap. A person may produce work that is faster, more polished and more complete while becoming no better at explaining the reasoning or detecting a subtle error. The gap becomes especially consequential when the sophistication of the output exceeds the user’s ability to evaluate it.
A 2025 study of 319 knowledge workers offers an important but appropriately limited signal. Participants described 936 examples of using generative AI in work tasks. Higher confidence in AI was associated with less reported critical thinking, while greater task-specific self-confidence was associated with more. The researchers also found that critical thinking did not simply disappear; it shifted towards activities such as verifying information, integrating responses and stewarding the task.3
Because this was a survey of self-reported practice, it should not be used to claim that AI inevitably erodes critical thinking. It does, however, reveal a change in where cognitive effort may occur. The professional may spend less effort generating an initial response and more effort specifying, inspecting and integrating it.
A peer-reviewed qualitative study of 21 frequent professional users similarly found that participants did not simply withdraw effort. Many reallocated or increased it as they learned to steer the system and administer the process.4 A 2025 systematic review of 68 peer-reviewed studies reached a broader dual-impact conclusion: generative AI can support or undermine critical thinking depending on mediating factors such as self-regulation, engagement, trust and metacognitive critique, as well as task and literacy conditions.5 These findings do not settle the long-term capability question, but they make one point clear: the outcome depends on how cognitive labour is redistributed and governed.
The key question is therefore not whether AI reduced effort. It is which effort was reduced, what capability that effort previously exercised and whether the professional can still judge the result.
Four human functions should remain human-led
The phrase “remain human” does not mean that AI should be excluded from the most important parts of work. AI may challenge assumptions, surface evidence, generate alternatives and expose inconsistencies across all of them. The boundary concerns leadership and ownership: which functions must remain meaningfully directed by a person who understands the context and can answer for the outcome?
1. Problem framing
Before a problem can be solved, someone must decide what the problem is. This involves determining whose interests matter, which outcomes deserve attention, what constraints are legitimate and what has been left outside the frame. AI can propose several formulations, but those formulations will be shaped by the information provided, the patterns available to the model and the assumptions embedded in the request.
A manager who asks AI how to reduce project delays may receive recommendations about scheduling, resources and monitoring. Yet the more important problem may be that senior leaders repeatedly change priorities without acknowledging the downstream consequences. If the initial frame treats delay as a technical scheduling issue, the analysis may become more sophisticated while the real organisational problem remains untouched.
Problem framing should remain human-led because it is inseparable from purpose and values. The professional must decide not only what can be analysed, but what deserves to be addressed.
2. Contextual interpretation
Information does not carry its full meaning independently of context. A risk described in one organisation may have different implications in another because authority, history, relationships, regulation and implementation capacity differ. AI can retrieve patterns and generate comparisons, but the professional must determine whether those patterns are relevant to the situation at hand.
This is where domain expertise changes from knowing more facts to recognising what matters here. A recommendation may be generally sensible and still be locally inappropriate. A policy may be compliant in wording and unworkable in practice. A research finding may be valid within its sample and misleading when transferred to a different population.
Contextual interpretation should remain human-led because professional work is not only the application of general knowledge. It is the disciplined translation of knowledge into particular circumstances.
3. Judgment under uncertainty
Many consequential decisions cannot wait for complete evidence. Professionals must compare incomplete signals, identify trade-offs, assess reversibility and decide what level of uncertainty is acceptable. AI may calculate, forecast or recommend, but the conclusion still depends on values, consequence and the decision-maker’s tolerance for different kinds of error.
A systematic review and meta-analysis of 106 experiments found that human–AI combinations performed better than humans alone on average, but did not achieve synergy on average: the combination generally failed to outperform the better of the human or the AI. The results also varied by task, with losses in decision tasks and stronger gains in content-creation tasks.6 This does not prove that humans should always retain every decision. It demonstrates that combining a person with an AI system is not, by itself, a design for better judgment.
Judgment under uncertainty should remain human-led because a recommendation cannot determine which consequences a community, organisation or person ought to accept.
4. Accountable commitment
Professional work eventually reaches a point where someone must decide, authorise or represent a conclusion. At that point, “the AI produced it” is not an adequate account of responsibility. The person presenting the recommendation must know what has been relied upon, what remains uncertain and why the proposed action is defensible.
UNESCO’s Recommendation on the Ethics of Artificial Intelligence makes this principle explicit at the governance level: ethical and legal responsibility should remain attributable to physical persons or existing legal entities, and reliance on AI cannot replace ultimate human responsibility and accountability.7 The organisational architecture of accountability requires much more detailed treatment, but the foundational boundary is clear.
Accountable commitment should remain human-led because consequences require an answerable subject. A system may contribute to the reasoning, but a professional or institution must still be prepared to say: This is the judgment we are acting upon, and this is why.
Together, these four functions define a human role more substantial than approving an AI-generated result. The person is not merely the final click in a machine-led process. The person frames the purpose, interprets the context, judges the uncertainty and accepts responsibility for the commitment.
Ability is not the same as authority
Technical capability may justify delegation in bounded, low-consequence tasks. Consequential professional work, however, requires three additional considerations: developmental value, evaluative capacity and consequence ownership.
First, some tasks develop capabilities the person will still need later. A junior analyst who repeatedly delegates the first interpretation of evidence may produce stronger reports today while missing the practice through which patterns, exceptions and weak signals become recognisable. The issue is not that manual work is morally superior. The issue is whether the automated activity previously served as part of professional formation.
Second, safe delegation depends on the ability to recognise failure. Expertise is often what makes offloading possible: the experienced professional can inspect an output, notice what does not fit and intervene when the system crosses its reliable frontier. The novice may receive the greatest immediate performance gain while being least prepared to detect a plausible defect. This does not mean novices should be denied AI. It means their use of AI must be designed as supervised development rather than invisible substitution.
Third, authority should be proportional to consequence. Generating alternative headings for an internal document does not require the same human involvement as recommending whether a person should be hired, treated, funded, disciplined or exposed to material risk. The relevant boundary is not “creative task versus analytical task” or “human task versus machine task.” It is the relationship among uncertainty, reversibility, developmental need and consequence.
What AI is able to do does not determine what AI should be authorised to own.
This principle avoids two unhelpful extremes. It does not require people to preserve every old task. Nor does it treat technical competence as a complete justification for delegation.
The Human Judgment Boundary
Before delegating a cognitive task, professionals need a more disciplined test than convenience. The following five questions form what I call the Human Judgment Boundary:
- What capability does this task exercise? If the task develops a capability I still need, removing it may create a future dependency even when it improves today’s output.
- Could I recognise a plausible but defective result? Delegation is safer when I can identify errors, missing context and inappropriate assumptions rather than merely judge whether the answer sounds convincing.
- Will I remain able to explain the reasoning? If the work affects other people, I should be able to reconstruct why the conclusion follows, what evidence supports it and where uncertainty remains.
- Am I delegating execution, interpretation or judgment? The same tool may perform all three, but they should not be treated as equivalent transfers of responsibility.
- Who bears the consequences if the output is wrong? The greater the potential harm, irreversibility or external reliance, the stronger the need for meaningful human review and decision authority.
These questions are not an anti-AI checklist. They are a delegation discipline. A low-risk task with little developmental value may be automated extensively. A high-consequence task may still benefit from substantial AI support, but the human must retain a stronger grasp of the evidence, interpretation and final commitment.
The test also exposes a weakness in the familiar phrase “human in the loop.” Human presence is not necessarily human judgment. A reviewer who lacks the time, expertise, information or authority to challenge an output may be present in the workflow without exercising meaningful oversight. What matters is not whether a human touched the process, but whether a human retained the capability and responsibility to intervene.
AI should play different roles in different situations
The relationship between human and AI should not be reduced to use or non-use. As a tool, AI can execute a bounded operation. As an assistant, it can organise work or develop a draft. As a challenger, it can test an argument and expose assumptions. As a collaborator, it can participate in an iterative process that the human repeatedly evaluates and redirects. As a decision input, it can contribute analysis to a conclusion that remains human-authorised.
The distinction among these roles matters because the same interface can conceal very different cognitive relationships. A professional may believe AI is merely assisting while gradually accepting its framing, priorities and interpretation. Conversely, a person who uses AI as a challenger may become more intellectually active because the system is designed to create disagreement rather than convenient completion.
The most valuable role is not always the one that removes the most effort. In some situations, AI should accelerate; in others, it should widen the field of thought, reveal uncertainty or slow the professional down before an irreversible decision.
Human–AI boundaries must remain developmental
AI systems, tasks and human expertise all change. A responsible delegation boundary must therefore be reviewed as the relationship among capability, risk and developmental need evolves.
Novices and experts may require different boundaries. An expert may safely delegate a first draft because years of experience have produced a strong internal model of what good work looks like. A novice may need to attempt the same task before receiving assistance so that comparison and correction can contribute to learning. Yet experts are not exempt from dependency. If they stop performing or closely inspecting a critical activity for long enough, calibration may weaken.
Delegation should therefore sometimes be interrupted deliberately through independent performance, explanation without the generated text, adversarial review or practice on unfamiliar cases. These are not rituals of human superiority; they test whether capability has been retained.
The developmental principle is simple:
Delegate according to both task risk and learning need.
That principle will produce different answers across professions and stages of expertise. The objective is not uniformity. It is intentionality.
Preserving what remains human does not mean rejecting AI
There is a danger that concern about judgment becomes nostalgic protectionism. Work should not remain manual simply because effort once signalled seriousness. AI can remove administrative burden, broaden access to expertise, generate useful alternatives and help professionals examine problems from perspectives they would not have reached alone. Used as a challenger and thought partner, it may strengthen rather than diminish reflection.
But augmentation does not happen automatically. The evidence on human–AI collaboration suggests that adding a person to AI – or AI to a person – does not guarantee that the combined system will exceed the stronger contributor. Good outcomes depend on role clarity, task fit, expertise, feedback and the design of interaction. Human–AI collaboration is therefore an architecture to be developed, not a slogan to be adopted.
The balanced position is neither “keep humans in control of everything” nor “delegate whatever the system can perform.” It is to redesign work so that machine capability expands what people can accomplish while human beings retain the capacities they must continue to exercise: framing purpose, interpreting context, judging uncertainty and accepting responsibility.
This is deliberate augmentation. It seeks efficiency without confusing speed with wisdom, and assistance without surrendering authorship.
What remains human is what we must remain capable of owning
The question “What should remain human?” cannot be answered with a permanent list of protected tasks. The boundary will move as technology advances and professions redesign their standards. One task may be safely automated tomorrow; another may become more important because AI can produce persuasive answers at scale.
The more durable boundary lies beneath the task. Human beings must remain capable of identifying the problem worth solving, interpreting what evidence means in context, judging among imperfect alternatives and committing to consequences that affect other people. AI may strengthen every one of these functions, but it should not quietly become their unaccountable owner.
The project manager in the opening example may reasonably use AI to assemble evidence, generate risk scenarios and test mitigation options. But the final register should still reflect a professional who understands why the risks matter, recognises what the system may have missed and is prepared to defend the decision before the people who will live with its consequences.
The standard for cognitive delegation is not merely whether the output looks professional, the system completed it faster or the recommendation is usually correct. The stronger standard is whether the human remains capable of understanding, judging and owning the work.
The most important question is not what AI can produce for us. It is what we must continue understanding, judging and defending for ourselves.
References
References.
References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.
- Dell’Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Organization Science. https://doi.org/10.1287/orsc.2025.21838
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
- Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (pp. 1–22). Association for Computing Machinery. https://doi.org/10.1145/3706598.3713778
- Memmert, L., Soroko, D., & Bittner, E. A. C. (2025). From effort reduction to effort management: An expectancy theory perspective on professionals’ work practices with generative AI. Business & Information Systems Engineering, 67, 615–635. https://doi.org/10.1007/s12599-025-00960-4
- Helal, M. Y., Elgendy, I. A., Albashrawi, M., Dwivedi, Y. K., Al-Ahmadi, M. S., & Jeon, I. (2025). The impact of generative AI on critical thinking skills: A systematic review, conceptual framework and future research directions. Information Discovery and Delivery. https://doi.org/10.1108/IDD-05-2025-0125
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
- UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://www.unesco.org/en/legal-affairs/recommendation-ethics-artificial-intelligence
Where this goes next.
Follow the argument into the topic it belongs to, or explore the wider body of work.
