Insights · Evidence, Research & AI

Evidence, Inference and Uncertainty: How to Read Claims About AI

A confident claim about AI may contain evidence, but confidence does not tell us where the evidence ends.

One week, a study appears to show that AI makes professionals more productive. The next week, another reports that AI weakens critical thinking. A company announces that its system outperforms experts. A commentator warns that the same technology is eroding human capability. Each claim arrives with a number, a graph, a university affiliation or a link to research.

The natural response is to ask which claim is true. But that question assumes the claims are describing the same outcome, population, task and period. Often they are not. One study measures the quality of an immediate output, another records self-reported effort, and a third examines performance after assistance is removed. All three may be credible within their own boundaries while supporting very different conclusions.

This is why the public conversation about AI can sound more certain than the underlying knowledge. Evidence is compressed into a headline. A bounded finding becomes a general prediction. An association becomes a cause. A short-term result becomes a claim about lasting capability.

The problem is not that people interpret evidence. Interpretation is unavoidable. The problem begins when the interpretation is presented as if it were the observation itself.

I use three categories to preserve this boundary:

  • Evidence describes what was observed or systematically established. Inference explains what we think the evidence means beyond the immediate observation. Uncertainty identifies what remains unresolved, conditional or vulnerable to revision.

Evidence disciplines what we may claim. Inference makes the evidence meaningful. Uncertainty tells us how far the claim can travel.

This is not a method for avoiding conclusions. It is a method for reaching conclusions that remain defensible when the evidence, technology or context changes.

Evidence is not another word for proof

In ordinary conversation, people often use evidence to mean proof. If a study supports a claim, the matter appears settled. Scientific evidence rarely works that way. A study offers observations produced under particular conditions, using particular measures and analytical choices. Those observations may strengthen one explanation, weaken another or revise the probability that a proposition is true, but they do not become independent of their design.

This is why evidence quality cannot be judged by the presence of a citation alone. In evidence-based medicine, the GRADE approach distinguishes the quality or certainty of an evidence base from the strength of the recommendation made from it.1 The distinction matters beyond clinical practice. Decisions also involve values, consequences, feasibility and context. Strong evidence about an effect does not automatically determine what an organisation should do, while an urgent decision may still have to be made when the available evidence is limited.

The same principle applies to AI. A carefully designed experiment may establish that a particular system improved performance on a defined task. It does not, by itself, prove that every AI system will improve every form of professional work. Nor does it establish that the improvement will persist, transfer or build human capability unless those outcomes were measured.

The correct question is therefore not simply, Is there evidence? It is:

What is this evidence evidence of?

That question forces us to name the outcome rather than import the conclusion we hoped to find.

Every claim has an evidence boundary

I call the evidence boundary the furthest point to which a finding can travel without requiring an additional inferential step. Inside the boundary, we can describe what the research observed. Outside it, we may still form a reasonable conclusion, but we must acknowledge that we are interpreting, generalising or predicting.

Consider a hypothetical experiment in which consultants using an AI tool complete a market analysis 30 per cent faster and receive higher quality ratings than consultants working without it. The evidence may support the following statement:

Under the study conditions, participants using this system completed this task faster and produced outputs rated more highly by the selected evaluators.

Several stronger statements move beyond that boundary:

  • AI makes consultants more capable. AI improves professional judgment. Organisations should automate market analysis. Consultants who do not use AI will become obsolete.

These statements may eventually be supported, but none follows automatically from faster completion and higher ratings. They introduce new constructs, longer time horizons, different settings or recommendations that the experiment did not test.

The distinction is especially important because AI-assisted performance can be highly visible while the underlying mechanism remains ambiguous. Better output may result from stronger reasoning, borrowed reasoning, improved presentation or a combination of all three. If the study does not separate them, the reader should not pretend that it did.

Evidence, inference and uncertainty perform different work

The three categories are not a hierarchy in which evidence is respectable and inference is suspect. They perform different intellectual functions.

Evidence constrains the argument

Evidence includes observations, measurements, documented cases and systematically synthesised findings. Its function is to prevent the argument from becoming whatever the author finds plausible. Good evidence answers a bounded question and makes its method available for scrutiny.

The relevant details include what was measured, who or what was studied, the comparison condition, the duration and the analytical method. A percentage without these details may be accurate but still uninformative. “Productivity increased by 40 per cent” means something different when productivity refers to words produced, tasks completed, expert-rated quality or economic value.

Inference connects the finding to a larger claim

Inference is the reasoning that moves from observation to explanation, generalisation or implication. It asks whether the finding is likely to hold for another population, whether the measured outcome represents the construct we care about and what mechanism may explain the result.

Inference is not a defect to be eliminated. Without it, research remains a collection of disconnected observations. The discipline lies in making the inferential step visible. Phrases such as this suggests, a reasonable interpretation is and the professional implication is tell the reader that the author is now doing more than reporting the study.

Uncertainty defines the conditions of confidence

Uncertainty includes imprecision, inconsistent findings, alternative explanations, unmeasured outcomes and unknown transfer across contexts. It may also arise because the technology has changed since the study was conducted. Naming these limitations does not invalidate the evidence. It tells us how much confidence the conclusion deserves and what evidence could change it.

This is why uncertainty should be expressed specifically. “More research is needed” tells the reader very little. A stronger formulation identifies what remains unknown: whether the effect persists after assistance is removed, whether experienced professionals respond differently from novices or whether the result transfers from a writing task to a consequential decision.

Read the study before accepting the claim

A headline usually tells us the conclusion. A disciplined reading begins with the study architecture. Five questions reveal most of the important boundary conditions.

What outcome was actually measured?

Accuracy, speed, output quality, confidence, learning, transfer and judgment are different outcomes. Improvements in one should not be silently converted into improvements in another. This is the same distinction AI Can Give You the Answer Without Building the Capability develops between successful output and human capability, but the present concern is evidentiary: did the research measure the outcome named in the public claim?

Compared with what?

A result depends on its comparator. AI-assisted workers may outperform unassisted workers while still performing worse than the AI alone. A new system may appear impressive against a weak baseline but offer little advantage over an established method. “Better” has no empirical meaning until the alternative is named.

Who performed which task?

Novices and experts may use the same system differently. A structured classification task is not equivalent to an ambiguous strategic decision. Evidence from students, crowd workers, clinicians or software developers may be relevant elsewhere, but the transfer is an inference that should be justified rather than assumed.

Over what period?

An immediate performance gain does not establish retention, skill development or long-term dependence. Likewise, a short-term reduction in effort does not prove cognitive decline. Claims about professional formation require evidence collected across time, not merely at the end of one assisted session.

How stable is the intervention?

In AI research, the system itself may change between data collection and publication. Model versions, interfaces, safeguards, prompts and user practices can all affect results. Readers should ask whether the study examined a general mechanism, a particular model or one configuration that may no longer exist.

These questions do not require every professional to become a research methodologist. They require the reader to resist a category error: using a result about one outcome or context as proof of another.

A prestigious study can still support a narrow conclusion

Peer review, journal quality and institutional reputation are useful signals, but none removes the need to inspect design. Research is produced through human choices about questions, samples, measures, analyses and publication. Transparent methods and independent scrutiny improve reliability because they make those choices visible.

Large-scale work on reproducibility has shown why a published result should not be treated as the end of inquiry. In an influential project that attempted to replicate 100 psychology studies, replication effect sizes were, on average, about half the magnitude of the original effects, while different criteria produced different estimates of replication success.2 The project itself generated methodological debate, which reinforces the larger lesson: scientific reliability develops through cumulative testing, not deference to one result.

Reproducible-science reforms therefore emphasise transparent methods, fuller reporting, data and material availability, preregistration where appropriate, replication and incentives that reward robustness rather than novelty alone.3 These practices do not guarantee truth. They make it easier to detect where a result depends on hidden analytical flexibility, selective reporting or conditions that do not survive repetition.

For the reader of AI claims, the implication is practical. Do not ask only whether a study was published. Ask whether its design supports the claim, whether the result has been examined elsewhere and whether the relevant materials or assumptions are visible enough to be challenged.

AI research carries distinctive forms of uncertainty

All empirical research has limitations, but AI introduces several forms of instability that deserve special attention.

First, the object of study changes quickly. A result produced with one model version may not describe its successor. This can make recent findings more relevant technologically but less mature evidentially, while older findings may be methodologically strong yet tied to systems no longer in use.

Second, benchmarks can be mistaken for practice. Performance on a test set may not transfer to work involving incomplete information, institutional constraints, contested values or consequences that unfold over time. The benchmark measures what it was designed to measure, not every capability implied by the model’s score.

Third, data leakage can inflate performance. Kapoor and Narayanan examined machine-learning research across multiple scientific fields and identified leakage as a widespread source of overoptimistic results.4 Information that should remain unavailable during training or model selection can inadvertently enter the process, making a system appear more predictive than it would be on genuinely unseen data.

Fourth, human–AI performance is relational. Results depend not only on the model, but on who uses it, what information they receive, when they rely on it and whether they can recognise its errors. A statement about “AI performance” may conceal a workflow, interface and pattern of human judgment that materially shaped the outcome.

Finally, publication favours legible stories. Dramatic gains, failures and predictions travel more easily than mixed findings. AI discourse intensifies this tendency because novelty attracts attention and both commercial enthusiasm and social concern create demand for decisive conclusions.

Messeri and Crockett describe a related risk in scientific research: AI systems may create an illusion of understanding by increasing productivity and apparent objectivity while narrowing the questions, methods and perspectives through which knowledge is produced.5 Their argument is conceptual rather than evidence that every use of AI creates a scientific monoculture. Its value is to identify what increased output may conceal: we can produce more findings while becoming less able to see the assumptions organising them.

One evidence base can contain several defensible conclusions

A 2024 systematic review and meta-analysis of 106 experiments provides a useful example of why AI claims must retain their conditions. The researchers compared humans alone, AI alone and human–AI combinations across 370 effect sizes. On average, the combined systems performed better than humans alone, which supports a claim of human augmentation. Yet the same combinations performed worse than the better of the human or AI working alone, which does not support average human–AI synergy.6

Task type changed the picture further. The analysis found performance losses in decision tasks and comparatively greater gains in content-creation tasks. It also found that combined performance tended to improve when humans were better than the AI alone, but deteriorated when the AI was stronger. The authors noted limitations including possible publication bias and variation across study designs.

What, then, does the evidence show?

It shows that AI assistance often improved performance relative to unaided humans in the experiments studied. It does not show that combining humans and AI reliably produces the best available performance. It also suggests that the benefit depends on task type and the relative strengths of the human and the system.

Three apparently contradictory headlines could therefore be written:

  • AI makes humans perform better. Human–AI teams underperform. AI collaboration works for some tasks but not others.

The third is less dramatic, but it preserves more of the evidence. The first two become defensible only when their comparators and conditions are included. This is the discipline this essay seeks to establish across the Dr Jim Choo body of work: retain the condition that makes the claim true.

Uncertainty should change the decision, not end it

Professionals cannot wait for perfect evidence. Leaders must choose tools, educators must design learning and individuals must decide which work to delegate while research is still developing. The practical purpose of uncertainty is therefore not to suspend action indefinitely. It is to calibrate the action to the confidence available.

When evidence is strong, directly relevant and consistent, an organisation may justify broader adoption. When evidence is promising but indirect, a bounded pilot with explicit measures may be more appropriate. When consequences are high and evidence is weak, safeguards, independent review and reversibility become more important.

This produces a simple principle:

The weaker or less transferable the evidence, the more reversible, observable and accountable the decision should become.

Uncertainty can then be converted into design. If long-term capability effects are unknown, measure what people can do after assistance is removed. If performance varies by expertise, analyse novices and experts separately. If a model changes rapidly, document the version and repeat critical evaluations. A limitation becomes useful when it tells us what to monitor.

The Evidence Boundary Card

Before accepting, repeating or acting on a claim about AI, ask seven questions:

  1. What was observed? State the finding without interpretation, prediction or recommendation.
  2. What outcome was measured? Distinguish speed, quality, accuracy, confidence, learning, transfer and judgment.
  3. What was the comparison? Identify the baseline and ask whether a different comparator would change the conclusion.
  4. Under which conditions did the result hold? Name the participants, task, system, setting and time horizon.
  5. What inferential step is being added? Locate where the claim moves from observation to explanation, generalisation or recommendation.
  6. What remains uncertain? Identify alternative explanations, missing outcomes, inconsistent evidence and limits to transfer.
  7. What decision is proportionate to that uncertainty? Decide whether the evidence supports adoption, a bounded experiment, additional verification or no action yet.

The card is deliberately more demanding than asking whether a source is credible. A credible source can still be interpreted carelessly. The task is to preserve the chain from observation to claim and from claim to action.

An editorial convention for the intellectual movement

Across the Core essays and evidence collection, the categories should become a recognisable epistemic convention.

What the evidence shows should report the measured result and its conditions. It should not contain claims about outcomes the research did not examine.

What we can reasonably infer should present the interpretation and explain why it is warranted. This is where Dr Jim Choo’s conceptual contribution belongs, clearly distinguished from the source finding.

What remains uncertain should identify the boundary that matters to practice, not provide a generic disclaimer. It should tell the reader what future evidence, context or outcome could alter the conclusion.

What I argue we should do should then convert the evidence and uncertainty into a proportionate professional or developmental position. Recommendation remains an act of judgment. Its credibility increases when the reader can see how it was constructed.

Used consistently, this convention will do more than improve citation practice. It will give the intellectual movement a public standard: strong arguments without certainty theatre, constructive positions without evidence inflation and openness to revision without intellectual passivity.

Better claims create better professional judgment

Return to the contradictory claims in the opening. AI may improve productivity, reduce some forms of cognitive effort, strengthen selected outputs and create new risks for learning or judgment. These propositions are not mutually exclusive. Their truth depends on what was measured, for whom, under which conditions and over what period.

The responsible reader does not ask research to remove every ambiguity. They ask it to clarify the decision. They distinguish a result from the story built around it, then decide how much confidence and consequence that story can carry.

This is not a weaker way to write about AI. It is a more durable one. Technologies will change, findings will accumulate and some conclusions will be revised. An intellectually defensible body of work should be able to absorb those changes because it has never concealed where observation ended and interpretation began.

The strength of an argument lies not in sounding certain. It lies in showing the reader why a conclusion is justified, where it remains conditional and what would cause us to think again.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. Guyatt, G. H., Oxman, A. D., Vist, G. E., Kunz, R., Falck-Ytter, Y., Alonso-Coello, P., & Schünemann, H. J. (2008). GRADE: An emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 336(7650), 924–926. https://doi.org/10.1136/bmj.39489.470347.AD
  2. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
  3. Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, Article 0021. https://doi.org/10.1038/s41562-016-0021
  4. Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804
  5. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627(8002), 49–58. https://doi.org/10.1038/s41586-024-07146-0
  6. Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.