Insights · Evidence, Research & AI

When AI Synthesises Too Early: Preserving Judgment in AI-Assisted Research

The greatest risk in AI-assisted research is not always receiving a false answer. It may be receiving a coherent answer before the researcher has understood the question.

Imagine beginning a new research problem by asking an AI system to review the literature. Within seconds, it returns the major themes, identifies several influential studies, describes areas of agreement and offers a balanced conclusion. The response is plausible, well organised and immediately useful. It gives the researcher a structure where, moments earlier, there was only uncertainty.

That is precisely why it can be dangerous.

The first synthesis does more than save time. It suggests what the question means, which concepts belong together, which differences matter and what a reasonable answer should look like. Once this structure is visible, subsequent reading can become an exercise in filling its categories rather than discovering whether those categories were appropriate.

The concern is not that researchers should reject AI. AI can help locate patterns, compare claims, generate alternative explanations and reveal connections that would otherwise take longer to find. The concern is timing. If synthesis arrives before the researcher has formed an independent map of the problem, assistance can become an intellectual anchor.

Messeri and Crockett warn that AI tools in science can create illusions of understanding, in which researchers believe they comprehend more than they do. They also identify the risk of scientific monocultures, where certain methods, questions and perspectives become dominant while alternatives become harder to see.1 This does not demonstrate that every AI-assisted review produces such an outcome. It establishes a serious possibility: greater research productivity can coexist with narrower understanding.

The governing distinction is between using AI to extend a developing judgment and allowing AI to supply the judgment's initial structure.

Build the question map before accepting the answer map.

This principle does not require the researcher to understand an entire field before using AI. It requires enough initial engagement to identify the problem's concepts, boundaries, disagreements and uncertainties before a fluent synthesis begins deciding them by default.

Synthesis is not neutral compression

Synthesis is often described as bringing many sources together into a coherent account. The description sounds procedural, as though a researcher merely reduces a large body of information into a shorter form. In practice, synthesis contains several acts of judgment.

The researcher decides which sources belong in the conversation, which findings are comparable, which concepts refer to the same phenomenon and which differences are substantive. They determine how much weight to give a method, whether contradictory findings can be reconciled, what uncertainty must remain visible and which details can be omitted without distorting the conclusion.

A literature review is therefore a research method, not an extended summary. Snyder argues that different purposes require different review approaches and that the choices involved in searching, selecting, analysing and synthesising literature must be made with methodological rigour.2 Even a narrative review contains a logic of inclusion and interpretation, whether or not that logic is made explicit.

AI can perform operations that resemble these acts. It can group papers, label themes, contrast findings and write a narrative that connects them. Yet the apparent coherence of the output does not reveal whether the groupings are conceptually valid, whether important sources were absent or whether the reconciliation flattened a genuine theoretical conflict.

The summary is shorter than the literature because something has been removed. The central research question is who decided what could safely disappear.

The danger is larger than fabricated references

Fabricated citations remain an obvious reason for caution. In one comparative study that asked large language models to reproduce references from systematic reviews, precision and recall were low, while hallucination rates ranged from 28.6 per cent to 91.4 per cent across the tested systems.3 The models and versions in that study do not represent every contemporary tool, and performance will continue to change. The durable lesson is that a generated reference list is not a verified evidence base.

Researchers can respond by checking whether each source exists and whether it supports the attributed claim. Necessary? Yes. Sufficient? No.

A synthesis can contain only genuine sources and still mislead. It may overstate the breadth of a finding, remove limiting conditions, combine studies that operationalise a construct differently or treat a frequently repeated claim as though repetition increased evidentiary strength.

Peters and Chin-Yee compared 4,900 AI-generated summaries with their original scientific texts. Most of the ten tested models produced conclusions broader than those warranted by the source material, even when prompted for accuracy. In a direct comparison, the AI summaries were nearly five times more likely than human-authored summaries to contain broad generalisations.4 This finding is especially relevant to synthesis because qualification is often the first information lost when many studies are compressed into one clean account.

The professional risk is therefore not only invented evidence. It is distorted evidentiary scope: real research presented as supporting more than it actually establishes.

Question formation should precede answer formation

Before a researcher can evaluate a synthesis, they need a provisional understanding of the question being answered. This begins with a question map.

A question map identifies the central phenomenon, adjacent concepts, relevant populations or contexts, plausible mechanisms, important outcomes and contested definitions. It records the researcher's initial assumptions without pretending that they are settled. It also distinguishes what the researcher is trying to explain from what they are merely trying to describe.

For example, a question about whether AI improves professional capability contains several unresolved choices. Does capability mean faster output, better immediate quality, retained independent performance, transfer to a new context or judgment under uncertainty? Which professionals, tasks and periods of use matter? Is the question about individual performance, workplace learning or organisational capability?

If AI synthesises the literature before these distinctions are surfaced, it may select one meaning of capability and organise the evidence around it. The resulting answer can appear comprehensive while addressing only a smaller question.

An initial question map does not lock the inquiry in place. It gives the researcher something to revise. New evidence may expose a false distinction, reveal an omitted construct or require the scope to change. That revision is part of research judgment. The difference is that the researcher can now see how the map is changing rather than inheriting an invisible map from the first generated response.

Disagreement is evidence about the structure of the problem

AI synthesis is naturally drawn towards coherence. Users often ask for the main themes, consensus or overall conclusion, and the model responds by reconciling variation into a readable narrative. Yet disagreement is not always noise that better synthesis should remove.

Studies may disagree because they define the construct differently, examine different populations, use different methods or measure different outcomes. Their findings may apply under different conditions or at different stages of development. They may also begin from incompatible theoretical assumptions.

A disagreement map should therefore precede a consensus statement. For each important difference, ask:

  • Are the studies asking the same question? Are they studying comparable people, tasks and settings? Do they define and measure the construct in the same way? Are the methods capable of supporting the claims being compared? Could both findings be valid under different conditions?

This turns contradiction into an analytical resource. Two studies that appear to conflict may reveal a moderator, boundary condition or developmental sequence. Conversely, several studies that appear to agree may all rely on the same narrow definition or data source.

Contradiction does not always mean that one source is wrong. It may reveal how the phenomenon changes across conditions.

The researcher's task is not to preserve disagreement indefinitely. It is to understand the reason for disagreement before allowing synthesis to dissolve it.

Keep three epistemic layers visible

Evidence, Inference and Uncertainty established a convention that should govern the entire Dr Jim Choo body of work:

  1. What the evidence shows 2. What we infer 3. What remains uncertain

In research synthesis, these layers are easily blended. A study reports an association, the synthesiser proposes a mechanism and the final narrative presents the mechanism as the finding. Several short-term studies show improved assisted performance, and the synthesis concludes that AI develops long-term capability. A pattern found in one occupation becomes a claim about professional work generally.

Keeping the layers separate does not weaken the argument. It shows the reader where the argument changes status. Evidence can be strong while interpretation remains contestable. A reasonable inference can guide further inquiry without being misrepresented as an established result. Uncertainty can identify what must be tested next.

AI can help reconstruct these layers if the researcher asks it to classify claims, identify inferential steps and locate qualifications. The researcher must still inspect the underlying sources because the model's classification is another output to evaluate, not the final epistemic authority.

Trace it, qualify it, defend it

I use a three-part discipline for material claims in AI-assisted research.

Trace it

Where did the claim originate? Locate the primary source where possible. Determine whether later papers are repeating an earlier interpretation and whether the cited passage, data or analysis supports the wording used in the synthesis.

Tracing also reveals citation chains. A claim may appear well supported because many articles repeat it, even though they ultimately rely on one limited study. The number of citations surrounding a proposition should not be confused with the number of independent sources of evidence.

Qualify it

Under what conditions does the claim appear to hold? Record the population, setting, method, timeframe, comparison and outcome. Identify limitations that affect how far the conclusion can travel.

Qualification is not the mechanical addition of phrases such as “may” or “could.” It specifies the boundary of the evidence. A useful qualification tells the reader where, for whom, compared with what and according to which measure the claim is supported.

Defend it

Why should this claim influence the conclusion? Explain the source's relevance, evidentiary strength and relationship to competing evidence. If the claim is central, the researcher should be able to state what would weaken or overturn it.

The three actions form a compact standard:

Trace the origin. Qualify the scope. Defend the weight.

A better division of cognitive labour

Preserving research judgment does not mean performing every operation manually. It means allocating AI to activities that extend inquiry while retaining human responsibility for epistemic commitments.

AI is particularly useful for expanding the search space. It can propose alternative keywords, identify adjacent concepts, generate counterarguments, compare definitions and suggest possible reasons for disagreement. It can create tables from verified sources, interrogate a developing framework and challenge a tentative conclusion.

The researcher remains responsible for determining relevance, source quality, comparability, evidentiary sufficiency, contextual meaning and the final conclusion. These are not protected human activities because machines can never contribute to them. They remain human responsibilities because the researcher must explain and defend how the evidence became an argument.

Research on AI-assisted decision-making supports the value of delaying easy acceptance. Buçinca, Malaya and Gajos tested interventions that required people to engage more deliberately with AI recommendations. Cognitive forcing reduced overreliance compared with simpler explanation interfaces, although participants liked the more demanding designs less.5 The study concerned decision support rather than literature synthesis, so the application here is an inference. It suggests that requiring an independent commitment before viewing or accepting AI advice can protect analytical engagement, even when the extra effort feels less convenient.

In practice, this means using an attempt-before-synthesis sequence:

  1. Frame the initial question. 2. Map concepts, boundaries and assumptions. 3. Read a diverse set of foundational and conflicting sources. 4. Record evidence, inference and uncertainty independently. 5. Invite AI to compare, challenge and extend the emerging map. 6. Return to primary sources before accepting the final synthesis.

The sequence can be adapted to the scale of the inquiry. Its purpose is to ensure that AI encounters a developing judgment rather than an empty page.

Use transparency to preserve judgment

Research judgment becomes more defensible when its material decisions leave a trace. For formal evidence synthesis, the PRISMA 2020 statement emphasises transparent reporting of why a review was conducted, what the reviewers did and what they found, including how studies were identified, selected, appraised and synthesised.6 AI does not remove the need for this transparency. It adds new decisions that may need to be recorded.

Depending on consequence and publication standard, an AI-assisted research process should document:

  • which systems and versions were used; what sources or datasets were supplied; what tasks AI performed; which outputs materially shaped inclusion, categorisation or interpretation; how claims and references were verified; which judgments were made or changed by the researcher.

Documentation should remain proportional. Exploratory brainstorming does not require a procedural audit. A published evidence review, professional recommendation or consequential research claim requires much more. The governing principle is that another informed person should be able to understand how the evidence became the conclusion.

The Research Judgment Card

Before accepting an AI-generated synthesis, ask five questions:

  1. What question has this synthesis assumed? Identify its definitions, boundaries, population, outcomes and implied purpose.
  2. Where do the sources genuinely disagree? Examine whether differences arise from theory, method, context, measurement or evidence quality.
  3. Which claims carry the conclusion? Trace those claims to primary sources and determine whether their weight is justified.
  4. What am I inferring rather than observing? Separate reported findings from mechanisms, generalisations and professional interpretation.
  5. What conclusion am I prepared to defend? State the conclusion in your own words, with its qualifications and remaining uncertainty.

If these questions cannot be answered, the synthesis may still be useful as orientation. It is not yet ready to function as research judgment.

Research judgment begins before the summary

The temptation of AI-assisted synthesis is understandable. Research begins in uncertainty, and a coherent answer offers immediate relief. It reduces the volume of material, gives the inquiry a vocabulary and creates the feeling that the intellectual terrain is becoming manageable.

But uncertainty at the beginning is not merely inefficiency. It is the period in which the researcher discovers what the question contains. Definitions are tested, assumptions become visible and disagreement begins to reveal the structure of the problem. Removing that period too quickly can make research more efficient at answering a question that was never adequately formed.

Return to the instant literature synthesis. Its themes may be useful, its sources may be real and its conclusion may eventually prove defensible. The researcher should still ask what became less visible when the literature was made coherent: which qualification disappeared, which perspective was excluded and which question was silently chosen.

AI should help us examine the evidence more widely and rigorously. It should not make coherence arrive before understanding.

References

References.

References are formatted in APA 7 style and ordered by first citation to correspond with the superscript notation.

  1. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627(8002), 49–58. https://doi.org/10.1038/s41586-024-07146-0
  2. Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, 333–339. https://doi.org/10.1016/j.jbusres.2019.07.039
  3. Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research, 26, Article e53164. https://doi.org/10.2196/53164
  4. Peters, U., & Chin-Yee, B. (2025). Generalization bias in large language model summarization of scientific research. Royal Society Open Science, 12(4), Article 241776. https://doi.org/10.1098/rsos.241776
  5. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188, 1–21. https://doi.org/10.1145/3449287
  6. Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., . . . Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. Systematic Reviews, 10, Article 89. https://doi.org/10.1186/s13643-021-01626-4

Where this goes next.

Follow the argument into the topic it belongs to, or explore the wider body of work.