Insights · Accountable Work With AI

You Can Delegate Work to AI. You Cannot Delegate Accountability.

AI can contribute to an output, recommendation or decision. It cannot stand before the people affected and answer for the consequences.

Imagine that an AI-assisted project plan omits a critical dependency. The omission passes through review, the project proceeds and a costly failure follows. In the investigation, the project manager explains that the plan was generated by AI. The manager assumed the system had considered the relevant risks, while the approving executive assumed the manager had verified them.

The AI performed part of the work. Yet “the AI produced it” does not explain who was responsible for deciding that the plan was adequate.

This is the accountability problem emerging across professional work. Organisations can delegate drafting, classification, analysis and recommendation to AI systems. They can distribute responsibilities across technology providers, data teams, managers, reviewers and users. What they cannot do is treat the system itself as the final answerable party.

This is not a universal legal conclusion about liability, which varies across jurisdictions and circumstances. It is a professional and institutional principle. AI systems do not hold professional duties, occupy organisational authority or participate in moral and disciplinary communities in the same way that people and legal organisations do. UNESCO’s Recommendation on the Ethics of Artificial Intelligence states that ultimate responsibility and accountability should remain with natural or legal persons rather than being assigned to AI systems.1

The governing distinction is therefore between delegating activity and assigning accountability.

Delegation determines who or what performs the work. Accountability determines who must justify the decision and answer for its consequences.

When these two functions are confused, AI does not remove responsibility. It makes responsibility harder to locate.

Task delegation and accountability are different

A professional task contains several forms of participation that are often collapsed into the word responsibility. Separating them shows what AI can perform and what must remain institutionally assigned.

Execution concerns who or what performs an operation. AI may draft a report, identify a pattern, calculate a forecast or generate possible actions.

Recommendation concerns who or what proposes a course of action. An AI system may rank options or indicate which outcome appears most likely according to its model and inputs.

Decision authority concerns who is permitted to accept, reject or modify the recommendation. This authority is created by professional role, organisational policy or law, not by the fluency of the output.

Answerability concerns who must explain how and why the decision was reached. It requires access to the relevant reasoning, evidence and standards.

Consequence ownership concerns who must respond when the decision causes harm, failure or an unexpected result. It may include remediation, disclosure, learning, compensation or disciplinary action.

AI can participate extensively in execution and recommendation. It may also influence the decision so strongly that the human response becomes almost automatic. But influence is not the same as formally holding authority, and producing an answer is not the same as being answerable.

NIST’s AI Risk Management Framework reflects this organisational view by calling for policies that define and differentiate human roles and responsibilities in human–AI configurations and oversight.2 The emphasis is not merely on keeping a person somewhere in the process. It is on knowing which person or function owns which part of the process.

AI makes the accountability chain harder to see

Traditional professional work is already distributed. A major decision may involve analysts, subject-matter experts, managers, executives, suppliers and regulators. AI adds further contributors and dependencies: model developers, training data, external platforms, system configurations, prompts, retrieval sources and automated transformations.

This creates several forms of opacity. A user may not know how the system weighted the available information. A manager may not know how extensively the employee relied on AI. An organisation may use a third-party service without visibility into its development or updates. An output may vary when the same task is repeated.

The resulting chain can become long enough for everyone to occupy only one fragment of it. Developers may say they did not determine the use. Users may say they did not design the system. Managers may say they approved a professional’s recommendation rather than the model. The organisation may refer to the supplier’s assurances.

Every statement may contain some truth. Together, however, they can create responsibility diffusion: many actors influenced the outcome, but no actor sees themselves as owning the whole decision.

The answer is not to pretend that one person controls every technical component. Accountability can be distributed without becoming indeterminate. Different actors may be accountable for system design, data quality, approved use, professional review and final authorisation. The requirement is that the allocation is explicit enough to determine who must act when the system is uncertain, wrong or harmful.

Meaningful control requires more than a final click

The phrase “human in the loop” is often offered as a complete safeguard. If a person reviews or approves the output, the system appears to remain under human control. Yet the presence of a human tells us very little about the quality of that control.

Work on meaningful human control proposes two useful conditions. The system should respond to the relevant human reasons and facts in its environment, and outcomes should be traceable to at least one human actor across the design and operation chain.3 Although developed partly through debates about autonomous systems, the logic applies more broadly: control is meaningful only when human values can shape the system and responsibility can be traced through its use.

A reviewer cannot exercise meaningful oversight merely by being shown an output. They need at least five conditions:

  • Competence: enough domain knowledge to recognise important defects. Information: access to sources, assumptions, limitations and material system behaviour. Time: sufficient opportunity to examine the output rather than approve it ceremonially. Authority: the ability to reject, revise, pause or escalate the work. Obligation: a clear expectation that intervention is part of the role.

Remove any one of these and the human may become an accountability ornament. They remain visible enough for the organisation to claim oversight but lack the practical capacity to change the outcome.

This matters because automated advice can alter human attention. A systematic review of automation bias found that decision-support systems can improve overall performance while introducing new errors that users fail to recognise. Trust, workload, time pressure, task complexity and experience influenced overreliance, while mitigations included training, emphasis on user accountability and changes to how advice was presented.4 Human oversight is therefore a cognitive and organisational capability, not merely a workflow position.

A human is not meaningfully in the loop if the system has already determined what the human is expected to approve.

Four dimensions of accountability

An AI-supported workflow may satisfy one form of accountability while failing another. I distinguish four dimensions.

Epistemic accountability

Epistemic accountability concerns whether the claims supporting the work deserve confidence. The accountable professional asks where the information came from, what was inferred, which uncertainty remains and whether the evidence is sufficient for the proposed conclusion.

This does not require reconstructing every model computation. It requires knowing enough about provenance, limitations and context to justify relying on the result. Verification Is Becoming Part of Professional Work develops verification in depth, while Evidence, Inference and Uncertainty provides the evidence, inference and uncertainty convention. The accountability argument locates the responsibility: someone must own the decision that the evidence was adequate.

Professional accountability

Professional accountability concerns whether the work meets the standards attached to a role. A clinician, engineer, educator, researcher or project professional cannot lower the applicable standard merely because AI contributed to the output. The tool may change how the standard is met, but not whether the standard applies.

This includes competence boundaries. If a person cannot evaluate an AI-generated analysis, presenting it as their professional judgment creates an accountability claim they are not equipped to honour.

Decision accountability

Decision accountability concerns who had the authority to commit the organisation or affect another person. Recommendation and authorisation should not be blurred. A system may recommend a candidate, risk response or allocation, but a named role must determine whether that recommendation becomes action.

Where several approvals are involved, the workflow should specify what each approval means. Otherwise, signatures accumulate while each approver assumes that substantive review occurred elsewhere.

Consequential accountability

Consequential accountability concerns who must respond after the decision. When an AI-assisted process causes harm, the organisation needs a route for correction, appeal, remediation and learning. Accountability is incomplete if it exists only at the moment of approval and disappears when consequences emerge.

The four dimensions form a sequence. Evidence must be adequate, professional standards must be met, authority must be exercised and consequences must be owned. A defensible workflow makes all four visible.

What should not be silently delegated

The word silently matters. AI can support every activity below, and in some settings it may perform much of the analytical work. The problem arises when its practical influence expands without a corresponding decision about authority, review and consequence.

Several activities should never drift into AI ownership by default:

  • defining what constitutes material risk; interpreting ambiguous or conflicting evidence; deciding when an exception to policy is justified; resolving conflicts between professional values; accepting an irreversible or high-consequence action; making a representation that no accountable person can explain or defend.

These activities contain institutional or professional commitments. An AI system may reveal options, identify patterns and test reasoning, but the organisation must decide who is authorised to make the commitment.

The boundary is not based only on task complexity. A technically simple decision may have serious consequences, while a complex draft may remain low risk because it will not be acted upon without substantial review. The relevant variables are consequence, reversibility, uncertainty and external reliance.

Accountability should be designed across the lifecycle

Many organisations treat accountability as a final review step. By then, important choices have already been made about the purpose of the system, the data it uses, the populations affected, acceptable error and the conditions for deployment. A reviewer at the end cannot repair every upstream decision.

Raji and colleagues proposed an end-to-end framework for internal algorithmic auditing in which documentation and assessment occur throughout the development lifecycle.5 The value of this approach is not limited to large algorithmic systems. It establishes a wider principle: accountability must leave evidence as decisions are made, not be reconstructed only after failure.

For professional AI workflows, that evidence should include:

  • the purpose and approved boundary of AI use; the task owner and decision authority; the system, data and material sources used; the required level of human review; the evidence or performance standard; the conditions that trigger escalation; material changes made by the human; incidents, appeals and lessons after deployment.

Documentation should remain proportionate. A private brainstorming session does not need the governance of a consequential decision system. But when others will rely on the output, the workflow should preserve enough information to establish how the conclusion was produced and who accepted it.

Legal and governance frameworks increasingly reflect this principle. The EU AI Act, for example, requires human-oversight measures for high-risk AI systems, while scholarly analysis has questioned what should be overseen, when oversight should occur and who is qualified to perform it.6 These questions are more useful than the generic instruction to “keep a human involved” because they force the organisation to specify the architecture of control.

The AI Accountability Chain

Before an AI-supported output becomes professional action, make six stages explicit:

  1. Generate: What did the AI system contribute, and within what approved boundary? 2. Inspect: Who examined the output, with which expertise, sources and standards? 3. Qualify: What uncertainty, limitation or condition must accompany the result? 4. Authorise: Who has the authority to convert the recommendation into a commitment or action? 5. Act: How will the decision be implemented, monitored and paused if necessary? 6. Answer: Who must explain the decision, address its consequences and ensure learning occurs?

The sequence can be expressed compactly:

Generate → inspect → qualify → authorise → act → answer

For each stage, the organisation should name the owner, required evidence and escalation trigger. The same person may own several stages in low-risk individual work. In high-consequence systems, separation may be necessary to create independent challenge and prevent convenience from becoming self-approval.

The chain also reveals where accountability is merely nominal. If nobody can inspect the evidence, qualification is impossible. If the reviewer cannot reject the result, authorisation is ceremonial. If no one monitors the action, answerability begins only after harm becomes visible.

Accountability must be proportional to consequence

Not every AI-assisted sentence requires an audit trail. Excessive governance can make responsible use so burdensome that people bypass it or conceal ordinary experimentation. A workable accountability system distinguishes between levels of consequence.

Low-consequence uses such as private ideation, formatting and reversible drafting may require little beyond appropriate confidentiality and professional discretion. Moderate-consequence work used by colleagues or clients may require source checks, named review and clear disclosure of uncertainty. High-consequence recommendations affecting safety, rights, employment, finance or irreversible commitments require stronger provenance, independent oversight, documented authorisation and routes for challenge or appeal.

Proportionality does not mean that low-risk work is risk free. It means the cost of control should correspond to the plausible harm, reversibility and reliance involved. The verification argument will develop this as risk-calibrated verification. Here, the accountability principle is that the person assigning the level of scrutiny must also be identifiable.

Accountability is part of professional value

It is tempting to frame accountability as the administrative cost that remains after AI has completed the valuable work. That interpretation misunderstands professional value. A professional does more than transmit output. They determine what deserves trust, interpret it in context, qualify what remains uncertain and stand behind the decision made from it.

As AI makes competent-looking production more widely available, this function becomes more visible. The scarce capability may not be producing the first draft or generating possible answers. It may be knowing which answer can be represented to another person, under what conditions and with what commitment to respond if it is wrong.

This is why professional authorship should not be reduced to who typed the words. Authorship also concerns who selected the evidence, accepted the reasoning and is prepared to defend the conclusion. AI can contribute extensively without becoming the accountable author.

Return to the failed project plan. The relevant question is not whether the project manager personally wrote every line. It is whether the organisation defined who had to inspect the dependencies, what evidence was required before approval, who had authority to proceed and who would answer when the assumptions failed.

The AI may have generated the omission. The accountability failure belonged to the system of human and organisational decisions that allowed the omission to become action.

If no person or organisation is prepared to explain, defend and answer for an AI-assisted decision, that decision is not ready to be represented as professional judgment.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://www.unesco.org/en/legal-affairs/recommendation-ethics-artificial-intelligence
  2. Tabassi, E. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1
  3. Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account. Frontiers in Robotics and AI, 5, Article 15. https://doi.org/10.3389/frobt.2018.00015
  4. Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089
  5. Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 33–44). Association for Computing Machinery. https://doi.org/10.1145/3351095.3372873
  6. Enqvist, L. (2023). “Human oversight” in the EU artificial intelligence act: What, when and by whom? Law, Innovation and Technology, 15(2), 508–535. https://doi.org/10.1080/17579961.2023.2245683

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.