Insights · Accountable Work With AI
Prompting Is Not Workflow Design
A better prompt can improve an AI interaction. It cannot decide whether the surrounding work is worth doing, properly governed or designed to produce a better decision.
Imagine a project team that uses AI to generate its weekly status report. After several attempts, the team develops a detailed prompt that produces the right headings, extracts updates from meeting notes and writes in the organisation's preferred style. Report preparation falls from two hours to twenty minutes.
The experiment appears successful. Yet senior leaders still receive the report after key decisions have been made. Different workstreams use incompatible definitions of progress. Risks are softened because nobody wants to escalate them. Several figures come from systems that have not been reconciled, and no one is clearly responsible for acting on the exceptions described.
The prompt has improved. The workflow has not.
This distinction matters because much of the current conversation about professional AI capability begins with the quality of instructions given to a model. Prompting is useful. It can clarify the task, supply context, specify constraints and improve the form of an output. In some situations, those improvements create genuine value. An experiment involving 453 college-educated professionals found that access to ChatGPT reduced the time required for bounded professional writing tasks by 40 per cent and increased independently rated output quality by 18 per cent.1
That is meaningful evidence of task-level benefit. It is not evidence that every organisational process containing similar writing has been improved. The experiment assessed defined outputs completed by individuals. A professional workflow may also contain disputed inputs, multiple actors, approval rights, exceptions, downstream dependencies and consequences that appear only after the document is produced.
The governing distinction is therefore between designing an interaction and designing a system of work.
A prompt coordinates what happens inside one exchange with AI. A workflow coordinates what must happen across people, information, decisions, controls and consequences.
Confusing the two can make defective work faster. It can also create the appearance of transformation while leaving the underlying decision, responsibility and learning system untouched.
What prompting actually solves
A prompt is an instruction interface. It tells an AI system what role to adopt, what material to consider, what operation to perform and what form the response should take. Good prompting can improve five aspects of an interaction:
- Purpose clarity: what the user wants from the exchange; context provision: which facts, documents or constraints should inform the response; task specification: what the model should compare, generate, classify or transform; output structure: how the answer should be organised and expressed; iterative refinement: how the user will challenge, correct or develop the result.
These are legitimate capabilities. Poor instructions create ambiguity, omit context and make evaluation unnecessarily difficult. Prompting can also expose unclear thinking before the model responds.
But the scope remains bounded. A prompt cannot determine whether the requested task is the right task, whether the available inputs deserve trust, whether another person must approve the result or what should happen when the case falls outside the expected pattern. Those are properties of the workflow and the organisation around it.
This is why a sophisticated prompt can optimise a low-value document, accelerate an unnecessary approval or produce a convincing answer to a badly framed problem. Prompt quality matters only after the purpose and position of the task have been understood.
A workflow contains what the prompt leaves outside
A professional workflow is an end-to-end arrangement for converting information and effort into a result, decision or service. It includes technical operations, but it also includes human roles, institutional standards and routes for responding when the normal sequence fails.
At minimum, a workflow contains:
- the outcome or decision the work exists to support; the inputs and their sources; the people and systems participating; the transformations applied to the information; the decision points and decision criteria; the standards used to evaluate adequacy; the review and validation requirements; the ownership of actions and consequences; the treatment of exceptions and escalation; the feedback through which the process improves.
Decades of human-automation research show why this wider view matters. Parasuraman, Sheridan and Wickens distinguished automation across information acquisition, information analysis, decision and action selection, and action implementation.2 Their model demonstrates that automation is not a single switch. Different functions can be automated to different degrees, and each allocation changes human performance and coordination demands.
Generative AI makes these allocations easier to overlook because several functions can occur inside one conversational interface. A user may ask a model to locate information, interpret it, recommend an action and draft the implementation message in a single exchange. The interaction feels continuous, but the professional obligations at each stage are different.
The workflow must make those differences visible even when the interface does not.
Task improvement is not system improvement
Most early AI adoption begins with a visible task. A team asks how to write the report faster, summarise the meeting, create the proposal or answer the client. The chosen metric is then speed, volume or perceived output quality.
These measures are useful but incomplete. A faster report may not produce an earlier decision. More proposals may create additional review work without increasing conversion. Automated meeting summaries may multiply records while leaving ownership unclear. A quicker response to a client may be less valuable if the system cannot identify when the case requires specialist intervention.
The management literature often presents automation and augmentation as alternatives, but Raisch and Krakowski argue that the two are interdependent and create a continuing paradox.3 Automating some activities changes the human work that remains, while augmentation can generate new opportunities for automation. The design problem is therefore not solved by labelling an initiative “human-centred” or “AI-assisted.” The organisation must examine how the division of labour changes across the system.
This is particularly important when a task-level improvement shifts work elsewhere. AI may reduce drafting time while increasing verification demands. It may simplify information retrieval while creating more options for a decision-maker to compare. It may allow one employee to generate more material while transferring an unsustainable evaluation burden to another.
Local efficiency is not system performance. A workflow improves only when the value, reliability and consequences of the end-to-end work improve.
The unit of redesign is the decision, not the document
Professional work often becomes visible through documents: a risk register, research review, lesson plan, business case, diagnosis, proposal or project report. This visibility makes the document an attractive target for AI. Yet the document usually exists to support something else.
A risk register should improve the recognition and treatment of uncertainty. A research review should support a defensible interpretation of evidence. A lesson plan should improve learning. A business case should help an authorised person decide whether an investment deserves commitment.
If redesign begins with the document, the team asks, “How can AI produce this faster?” If it begins with the decision, the questions change:
- What decision or action is this work meant to improve? 2. What information must be available before that decision can be justified? 3. Which interpretations require professional judgment? 4. What must be validated, by whom and to what standard? 5. Which exceptions require a different route? 6. How will the outcome feed back into future work?
These questions may reveal that the document should be shorter, produced earlier, assembled differently or removed entirely. The most valuable AI contribution may not be writing it. AI might instead detect missing inputs, compare scenarios, surface contradictions or monitor whether an agreed action has occurred.
The document is an artefact. The decision is the reason the artefact exists.
Task automation can preserve the defects around it
Technology projects frequently automate the visible operation while preserving the assumptions and constraints that made the original process weak. AI can draft around poor data, unclear definitions, duplicated approval and ambiguous ownership, but it cannot make those conditions sound.
It may even conceal them. A fragmented set of inputs can be converted into a coherent narrative. Missing criteria can be replaced by plausible generalisations. Conflicting stakeholder expectations can disappear inside neutral language. The output looks more integrated than the system that produced it.
Organisational research increasingly argues for treating AI as part of a system of work rather than only as a tool used by an individual. Anthony, Bechky and Fayard propose viewing AI as a counterpart within a system that includes design, implementation and use.4 This perspective directs attention towards relationships, routines and organisational context. The relevant question is not only how a person interacts with AI, but how that interaction changes the work performed by everyone else.
Before automating a task, inspect the defects around it:
- Are the inputs timely, complete and consistently defined? Is ownership clear before and after AI contributes? Are decision criteria explicit enough to be examined? Do approvals add independent judgment or merely delay action? Can exceptions be recognised and routed appropriately? Does the process learn from outcomes, corrections and failure?
If these conditions are weak, prompt optimisation may scale the weakness rather than repair it.
AI-enabled, AI-native and AI-augmented
Three terms help distinguish different ambitions for professional AI use.
AI-enabled work adds AI to an existing task. The user employs a model to summarise, draft, classify, translate or generate options while the surrounding sequence remains largely unchanged. This can be entirely appropriate for bounded, low-risk work.
An AI-native workflow is redesigned around capabilities and constraints that did not previously exist. Human and AI roles are allocated deliberately across the end-to-end process. Inputs, decision points, verification, escalation and feedback may all change. The objective is not to insert AI everywhere. It is to create a better system because different arrangements have become possible.
An AI-augmented professional can operate and improve such a workflow. This person knows how to use AI, but also understands the domain, evaluates evidence, recognises exceptions, preserves accountability and redesigns the division of labour when the system underperforms.
These definitions are conceptual rather than maturity badges. An AI-enabled solution is not automatically inferior, and “AI-native” should not become promotional language for unnecessary complexity. The correct level depends on the problem. A private drafting task may need only a capable user and a clear prompt. A multi-stage process affecting other people may require redesign.
The anatomy of an accountable AI workflow
An AI-supported workflow can be examined through seven functions:
Input → interpret → generate → validate → decide → escalate → learn
Input establishes what information enters the process, where it came from and whether it is suitable. Interpret determines what the information means for the present problem. Generate creates possible outputs, options or analyses. Validate tests claims, reasoning, context and adequacy. Decide assigns authority to accept, reject or qualify the result. Escalate provides a route for uncertainty, exceptions and material risk. Learn uses outcomes, corrections and incidents to improve the process.
The sequence need not be linear. Failed validation may return the work to input or interpretation. An exception may require a different decision-maker, while learning may change the prompt, data, role allocation or workflow.
Human–AI interaction guidelines similarly emphasise supporting correction, communicating system capability, enabling efficient dismissal and learning from user behaviour over time.5 These principles are useful because they recognise that the first output is not the whole interaction. At workflow level, they must be joined by ownership, standards and feedback across multiple participants.
For each function, ask four questions:
- What is the human role? What is the AI role? What evidence or standard applies? Who owns the transition to the next stage?
If any stage has no clear answer, the workflow contains an unowned assumption.
A project risk assessment: prompt view and workflow view
Consider a project manager preparing a risk assessment. From the prompting perspective, the task is to generate a useful risk register. A strong prompt supplies the project scope, schedule, stakeholders, assumptions and required categories. It asks for causes, events, consequences, probability, impact and suggested responses. It may request that the model challenge generic risks and identify dependencies.
This produces a better register. The workflow view begins earlier and continues further.
The team first defines which decisions the assessment must inform. It identifies the systems, interviews and lessons from comparable projects that should supply evidence. AI helps organise the inputs and propose patterns, but workstream owners test whether the risks reflect current conditions. Material assumptions are traced. The project manager validates interactions across workstreams, and accountable owners decide whether proposed responses are feasible.
Thresholds determine which risks require sponsor attention. Exceptions such as regulatory exposure or irreversible commitments follow a separate escalation route. After agreed responses are implemented, the team monitors whether risk indicators change. Missed risks and ineffective responses become inputs to the next assessment cycle.
The prompt remains valuable inside this workflow. It is no longer mistaken for the workflow itself.
When prompting is sufficient
Prompting may be sufficient when the task is individual, bounded, reversible and low consequence. Examples include private ideation, tone adjustment, formatting, preliminary question generation and transformations that the user is competent to inspect. The surrounding workflow may already be sound, and AI simply reduces effort within one step.
Workflow redesign becomes necessary when several people rely on the output, information passes through multiple systems, a decision requires formal authority, errors create material consequences or exceptions cannot be handled by the normal route. Redesign is also warranted when task efficiency repeatedly creates bottlenecks elsewhere.
This qualification matters because generative AI does not always increase productivity. A review of human-factors evidence identified role shifts from production to evaluation, poorly restructured workflows, interruptions and the tendency of automation to make easy tasks easier while making difficult tasks harder.6 Better results depend on task allocation, feedback, interface design and the wider conditions of use.
The choice is therefore not between prompting and workflow design. Prompting is one capability inside workflow design. The mistake is allowing competence at the smaller level to substitute for examination at the larger one.
The Workflow-Before-Prompt Canvas
Before writing or refining a consequential prompt, answer eight questions:
| Canvas question | Design purpose |
|---|---|
| 1. What outcome or decision must improve? | Prevent the visible artefact from becoming the objective |
| 2. What evidence should enter the work? | Establish input quality, provenance and boundaries |
| 3. Where is interpretation required? | Identify the points at which context and judgment matter |
| 4. What should AI perform or support? | Allocate automation and augmentation deliberately |
| 5. What must a human validate? | Match review to uncertainty and consequence |
| 6. Who can authorise action? | Separate recommendation from decision authority |
| 7. What triggers exception or escalation? | Protect the workflow from cases outside its normal design |
| 8. How will outcomes improve the process? | Create feedback for prompts, roles, data and standards |
Only after these questions are answered should the team decide what prompt is required. The resulting prompt will usually be better because it occupies a defined role in a coherent system.
Begin with the work that must improve
The attraction of prompting is understandable. It is immediate, visible and individually controllable. A person can change an instruction and see a different output within seconds. Workflow design is slower because it exposes disputed purposes, weak information, unclear authority and organisational dependencies.
Yet these are precisely the conditions that determine whether AI creates durable professional value. A powerful model cannot compensate for a process that asks the wrong question, uses unreliable inputs, hides exceptions or produces an output that nobody is authorised to act upon.
Return to the weekly status report. The team does not primarily need a faster document. It needs timely, comparable information that causes risks to be surfaced, decisions to be made and actions to be owned. AI may help build that system, but only after the team stops treating the report as the unit of transformation.
Do not begin with what AI can generate. Begin with the decision, service or system of work that must become better.
References
References.
References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586
- Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics, Part A: Systems and Humans, 30(3), 286–297. https://doi.org/10.1109/3468.844354
- Raisch, S., & Krakowski, S. (2021). Artificial intelligence and management: The automation–augmentation paradox. Academy of Management Review, 46(1), 192–210. https://doi.org/10.5465/amr.2018.0072
- Anthony, C., Bechky, B. A., & Fayard, A.-L. (2023). “Collaborating” with AI: Taking a system view to explore the future of work. Organization Science, 34(5), 1672–1694. https://doi.org/10.1287/orsc.2022.1651
- Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Article 3, pp. 1–13). Association for Computing Machinery. https://doi.org/10.1145/3290605.3300233
- Simkute, A., Tankelevitch, L., Kewenig, V., Scott, A. E., Sellen, A., & Rintel, S. (2025). Ironies of generative AI: Understanding and mitigating productivity loss in human-AI interaction. International Journal of Human–Computer Interaction, 41(5), 2898–2919. https://doi.org/10.1080/10447318.2024.2405782
Continue this argument
Continue this argument.
Where this goes next.
Follow the argument into the topic it belongs to, or explore the wider body of work.
