Insights · Accountable Work With AI

Verification Is Becoming Part of Professional Work

When plausible output becomes abundant, professional value shifts from producing something that looks credible to establishing what deserves to be trusted.

Imagine receiving a polished market analysis prepared with AI. The argument is coherent, the language is confident and the recommendations appear commercially sensible. A quick review finds no obvious contradiction. Only later does someone discover that a cited report does not exist, two figures refer to different time periods and the recommendation assumes a regulatory condition that does not apply in the target market.

The document looked professional. It had not been professionally verified.

This distinction is becoming more consequential because generative AI can produce credible-looking work faster than many professionals can examine it. Fluency, structure and completeness were never guarantees of truth. AI has simply reduced the cost of producing these signals. A confident explanation or well-structured recommendation can now appear before anyone has established whether its claims are accurate, relevant or sufficient.

Research on natural-language generation has documented the problem of systems producing fluent content that is unfaithful to source material or unsupported by it.1 NIST identifies this form of confabulation as one of the risks that generative AI can create or intensify, alongside wider concerns about information integrity and content provenance.2 The problem is therefore not simply that AI can be wrong. People have always been wrong. The professional problem is that an AI-generated error can arrive in a form that reduces the felt need to question it.

The governing distinction is between checking for errors and establishing warranted trust.

Verification is not the attempt to prove that an output is flawless. It is the disciplined process of establishing whether the output is trustworthy enough for its intended use.

That standard changes the work. A private brainstorming note and a safety recommendation do not require the same scrutiny. A correct statistic may be irrelevant, a real source may not support the attributed claim, and a coherent recommendation may ignore the local context. Verification must therefore be proportionate to purpose, uncertainty and consequence.

Verification is more than fact-checking

Fact-checking is an important part of verification, but it is only one part. It asks whether names, dates, quotations, calculations and factual claims are correct. Professional verification asks a larger set of questions.

Does the source exist, and is it authoritative for this claim? Does the conclusion follow from the evidence? Has material counterevidence been omitted? Do the assumptions fit the present organisation, population or jurisdiction? What happens if the output is wrong, incomplete or misunderstood? A report can pass a factual check while failing every one of these tests.

Consider an AI-generated project risk assessment. Its risks may be real, its scores internally consistent and its language aligned with the organisation's template. Yet it may omit a dependency known only through recent conversations, misclassify a politically sensitive issue or recommend a response the organisation cannot implement.

The error is not necessarily a false statement. It may be a failure of completeness, reasoning, context or consequence.

This is why verification cannot be reduced to asking the same AI system, “Are you sure?” The reviewer must know what kind of trust is being requested. Trust that a quotation is accurate is different from trust that an argument is adequate. Trust that a summary reflects a source is different from trust that the source should influence a decision.

Why AI changes the shape of verification work

AI does not create the need for professional checking. It changes its scale, location and difficulty.

First, AI increases the volume of material produced. If review capacity remains unchanged, more output competes for the same finite attention. Speed at production can therefore create congestion at verification.

Second, AI compresses visible reasoning. A polished recommendation may contain choices about framing, selection, categorisation and weighting, but the output rarely makes every choice inspectable. The user sees the conclusion without necessarily seeing how the system moved from the available information to that conclusion.

Third, generative output is variable. A changed prompt, model or context can produce a different answer, so one approved output may not represent stable system behaviour.

Fourth, provenance may be weak. The system may combine provided documents, retrieved sources, training data and generated interpolation without making those origins clear. Supplied citations still need to be checked for existence, relevance and faithful representation.

AI can reduce the effort required to create an output while increasing the sophistication required to inspect it. The verifier may need stronger domain knowledge, information literacy and decision criteria than the person who generated it.

The Five-Level Verification Ladder

Verification becomes more manageable when it is separated into levels. Each level asks a different question and catches a different class of failure. The levels are cumulative for consequential work: passing a surface check does not remove the need to examine sources, reasoning, context or consequences.

Level 1: Surface verification

Question: Are the visible details internally and externally correct?

Surface verification examines names, dates, quotations, calculations, terminology, formatting and internal consistency. It is the appropriate starting point because visible errors may reveal weak source handling or careless generation. It is not a sufficient stopping point for work that will guide action.

At this level, check whether totals reconcile, dates align and defined terms remain consistent. Confirm factual claims against suitable external sources rather than relying on repetition by the model.

Level 2: Source verification

Question: Can each material claim be traced to an appropriate source that actually supports it?

Source verification establishes existence, provenance, fidelity, authority and currency. The reviewer asks where the claim originated, whether the cited material says what the output attributes to it, and whether the source is strong enough for the use being made of it.

This work often requires leaving the generated document and reading laterally. In a study comparing professional fact-checkers, historians and university students, fact-checkers opened new sources to investigate who stood behind a website and how others assessed it. They reached more warranted judgments more efficiently than participants who remained within the original page.3 Do not inspect a source only on the terms supplied by the output. Investigate its origin, standing and independent corroboration.

Level 3: Reasoning verification

Question: Does the conclusion follow, and what has the argument left out?

Reasoning verification examines premises, causal claims, comparisons, assumptions, alternatives and uncertainty. A set of accurate facts can still support a weak conclusion if they are selectively chosen, improperly combined or treated as more decisive than they are.

Reconstruct the argument in your own words and identify what must be true for the recommendation to hold. Ask what alternative explanation fits the evidence, what would change the conclusion and whether confidence exceeds the evidence. Evidence, Inference and Uncertainty develops the distinction among evidence, inference and uncertainty; the verification argument applies it to establishing trust.

Level 4: Context verification

Question: Does the output fit the situation in which it will be used?

Context verification tests whether a generally sensible answer remains appropriate for a particular organisation, population, stakeholder group, jurisdiction or moment. AI may offer a standard practice while missing local capacity, informal authority, recent events, cultural expectations or constraints absent from its inputs.

Contextual fit is rarely established by factual accuracy alone. A valid study may not transfer to another population, a compliant policy may fail in another jurisdiction, and a reasonable project response may exceed the team's capability. Verification includes deciding what the output means here, not only whether it could be correct somewhere.

Level 5: Consequence verification

Question: What degree of confidence is required before anyone acts?

Consequence verification examines severity, reversibility, exposure and external reliance. It asks who may be affected, whether the action can be corrected, what would happen if the output were wrong and whether another person will rely on it without conducting their own review.

This level determines how much of the ladder must be climbed and how independently. A low-consequence draft may need only surface and limited source checks. A recommendation affecting safety, employment, finance, legal rights or an irreversible commitment may require domain review, independent evidence, documented assumptions and explicit authorisation.

The five levels can be summarised as follows:

Five levels of verification
LevelVerification questionTypical failure detected
1. SurfaceAre the visible details correct and consistent?Wrong names, dates, figures, quotations or calculations
2. SourceCan material claims be traced to suitable support?Fabricated, outdated, irrelevant or misrepresented sources
3. ReasoningDoes the conclusion follow from the evidence?Unsupported inference, omitted alternatives or hidden assumptions
4. ContextDoes the output fit this situation?Inapplicable advice, missing local knowledge or policy mismatch
5. ConsequenceIs confidence adequate for the action proposed?Disproportionate risk, weak safeguards or unjustified reliance

The higher the consequence of being wrong, the less acceptable it is to verify only what is easiest to see.

Verification should be risk-calibrated

The ladder does not mean that every AI-assisted sentence requires five levels of formal review. That would replace one unsustainable practice with another. Verification must be calibrated according to the use.

Six variables help determine the required depth:

  1. Consequence severity: How serious could the harm or failure be? 2. Reversibility: Can the decision be corrected before material damage occurs? 3. Uncertainty: How incomplete, disputed or unstable is the evidence? 4. Novelty: Is the task familiar, or does it involve a new context or exceptional case? 5. Reviewer capability: Does the reviewer possess enough expertise to identify a plausible defect? 6. External reliance: Will clients, colleagues or the public act on the output as though it has been professionally endorsed?

These variables should influence both depth and independence. A professional may verify their own low-risk draft. A high-consequence recommendation may require a second qualified reviewer, primary-source confirmation or testing against an independent method. Where the reviewer lacks the relevant expertise, more effort does not necessarily solve the problem. The work may need escalation to someone capable of judging it.

This principle prevents verification from becoming a generic instruction to “check the AI.” The organisation must specify what adequate checking means for the task, who can perform it and what evidence must exist before the output is used.

The verification paradox

AI creates a practical paradox. It makes the production of analytical and written material cheaper, but every additional output creates a potential claim on human attention. If organisations generate more than they can responsibly inspect, apparent productivity may rise while the reliability of the work falls.

The answer cannot be exhaustive manual review. Nor can low observed error prove that a system deserves general trust. Verification requires triage.

Low-consequence uses can be bounded, sampled and monitored. Repeated tasks can preserve traceability and use automated tests for known failure modes. Material claims can require primary-source confirmation, while high-consequence decisions can require independent review and escalation.

NIST's Generative AI Profile follows this risk-management logic. It treats governance, content provenance, pre-deployment testing and incident disclosure as connected responsibilities across the AI lifecycle, not as a single correctness check performed at the end.2 For professional workflows, the corresponding lesson is that verification should be designed into how work is produced rather than added after the output already looks finished.

Appropriate reliance is harder than trust

The goal of verification is not universal distrust of AI. Refusing correct assistance can be as unproductive as accepting incorrect assistance. The more useful goal is appropriate reliance: accepting AI advice when it is sound and resisting it when it is defective.

Schemmer and colleagues formalised appropriate reliance as a two-dimensional capability and examined how explanations affected it in an experiment with 200 participants.4 A user who accepts every recommendation is appropriately reliant only if they can also reject bad advice. Rejecting everything, however, loses genuine value.

Explanations can support judgment, but they do not eliminate the need for verification. An explanation may itself be incomplete, persuasive or difficult to assess. The key professional capability is discrimination: knowing when the available evidence justifies acceptance, when it requires qualification and when the output should be rejected or escalated.

This capability is affected by the difficulty of the review. A systematic review by Lyell and Coiera found automation bias in single tasks as well as multitasking environments, particularly where verification imposed high cognitive load.5 The practical implication is not that professionals are incapable of oversight. It is that oversight fails when organisations assign people complex verification duties without enough time, information, tools or expertise.

The instruction “a human must review it” is therefore incomplete. The relevant question is whether the human has a realistic capacity to distinguish a trustworthy output from a persuasive failure.

AI can assist verification, but it cannot close the loop by itself

AI can make verification more efficient. It can extract claims, compare versions, locate contradictions, recalculate figures, test a recommendation against alternative scenarios and propose questions a reviewer may have missed. A second model or retrieval system can sometimes expose weaknesses in the first output.

These practices are useful as adversarial assistance, but they are not independent confirmation. Systems can share blind spots, rely on the same weak source or produce mutually reinforcing explanations. Another generated answer may add perspective without establishing the final standard of trust.

A sound division of labour is straightforward. Use AI to increase the range and speed of checks. Use independent evidence, qualified human judgment and explicit decision criteria to determine whether those checks are sufficient.

From final-stage checking to verification by design

Verification is most reliable when the workflow preserves the conditions needed to perform it. Once provenance has been lost, assumptions have been hidden and deadlines have removed the possibility of challenge, a final reviewer can do little more than inspect the surface.

Verification by design includes five practices:

  • Preserve traceability. Keep material claims connected to their sources, data and relevant transformations. Define evidence thresholds. State what support is required before an output can influence a decision. Place review at meaningful points. Examine framing, evidence and reasoning before the final document becomes psychologically settled. Specify escalation triggers. Identify the uncertainties, exceptions and consequences that require additional expertise or authority. Record material judgments. Preserve why a recommendation was accepted, qualified or rejected when others will rely on it.

These practices also connect the verification argument to the wider pillar system. You Cannot Delegate Accountability establishes who must answer for AI-assisted work. Prompting Is Not Workflow Design examines how responsibilities and controls belong within the whole workflow rather than inside a prompt. The verification argument supplies the missing operational question: what must be examined before an accountable person can reasonably stand behind the output?

Verification is becoming a mark of professional capability

Professional value has long been associated with producing the analysis, report, lesson, plan or recommendation. As AI lowers the cost of initial production, part of that value moves elsewhere. The professional must determine what is relevant, trustworthy, uncertain and fit to be represented to another person.

This is not merely quality assurance performed after the “real work.” It is a form of judgment. Verification requires the professional to understand the domain well enough to detect absence as well as error, to recognise when a true claim is being used incorrectly, and to increase scrutiny when the consequences require it.

Return to the market analysis. The failure was not that nobody proofread it. The failure was that surface credibility substituted for professional verification. No one traced the report, reconciled the figures, tested the regulatory assumption or asked what level of confidence the recommendation required.

AI did not make those questions obsolete. It made them easier to postpone.

The future professional will not be distinguished only by what they can produce. They will be distinguished by what they can establish as worthy of trust.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Chen, D., Dai, W., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730
  2. Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1
  3. Wineburg, S., & McGrew, S. (2019). Lateral reading and the nature of expertise: Reading less and learning more when evaluating digital information. Teachers College Record, 121(11), 1–40. https://doi.org/10.1177/016146811912101102
  4. Schemmer, M., Kühl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces. Association for Computing Machinery. https://doi.org/10.1145/3581641.3584066
  5. Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423–431. https://doi.org/10.1093/jamia/ocw105

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.