AI Aqueducts: Speed. Confidence. And Contaminants.

The AI Pipeline Has No Way to Judge the Science It Uses

Author

Leslie McIntosh

Published

July 2, 2026

AI infrastructure is constantly improving – building better aqueducts for the flow of information. These aqueducts are also streamlining science. But aqueducts do not make water clean, they only move it faster. If corrupted scientific records enter the system, LLMs and even MCP-enabled tools will not merely repeat the contamination. They will package it, route it, and deliver it with confidence at a very fast pace.

We have a few issues to address – first, access isn’t trust. Having information does not mean you should trust such information. But now systems are being built to restrain and focus the corpus of research – prioritizing the scientific literature rather than pulling from any sources on the internet. These systems pull from specified data sources that represent research. This could be anywhere from medical blogs, to preprints, to peer-reviewed literature. The LLMs retrieve documents, synthesize them, smooth over bumps, then turn that pile of sources into a very confident, believable voice. And that is what I want to explore.

Without understanding what data are going into the MCP pipeline, we are not necessarily going to get more trusted outcomes. We will get more outputs that look trustworthy.

Digital information aqueduct (Image generated with AI. Photo by Adobe Stock.)

Digital information aqueduct (Image generated with AI. Photo by Adobe Stock.)

How LLMs and MCPs Work (and Why That Matters Here)

Large language models (LLMs) retrieve documents, synthesize them, and deliver an answer in a confident, polished voice. The Model Context Protocol (MCP) is designed to make that process more controlled: rather than pulling from anywhere on the internet, an MCP-connected system draws from a specified pool of sources. In a research context, that pool might include peer-reviewed literature, preprints, or curated databases. Think of it as a designated tap rather than an open river. Anthropic describes MCP as a kind of “USB-C port” for AI: a standard connector that removes friction between the model and the data it draws on.

That sounds like an improvement. And in some respects it is. Constraining the corpus is better than leaving it open. But controlling the pipeline does not control the quality of what flows through it.

Issue 1 – Provenance

Using AI and connectors don’t by default surface sources for review – they retrieve, synthesize, then deliver an answer. A polished one at that. But the provenance may disappear into the output. Even if a research paper is cited, you won’t know where it came from necessarily and the process to cite that paper. So the contaminated paper that was previously findable becomes an invisible input to a trusted-sounding conclusion.

Issue 2 – Reach

MCP is being built for enterprise and institutional workflows – drug discovery pipelines, policy research, grant evaluation. The same contaminated source that previously misled only one researcher may now silently inform a procurement or policy decision.

Issue 3 – Velocity

MCP agents don’t pause, and they don’t flag uncertainty about source integrity. They just run. Uncertainty is a core feature of research. Here, it has been engineered out.

So MCPs create a new contamination problem. They take the existing documents with all their problems and make those issues simultaneously invisible, embedded, and fast. The three combined provide a confidence obscuring uncertainty, making the output look more authoritative.

It’s like a credit rating on unaudited books. The rating looks authoritative. The methodology is sound. But if the underlying financials were never verified, the credit rating means nothing.

But AI tools cannot remove uncertainty they don’t know about.

The Scholarly Data that Feed the AI Pipelines

What if we removed the ‘bad content’ from our research databases? Wouldn’t that solve our problem of having a trusted data source on which to base our research decisions? Not quite.

Let’s look at what happens in four scenarios of filtering research publications with the purpose that we want the differently trusted (through curation) scientific papers to be the foundational information for LLMs and to run through an MCP to garner better decisions. Read through the table, but the gist is - it depends. There are always tradeoffs.

Different levels of provenance are useful for research and forensically understanding research.

Intelligence from the Contaminants

The goal is not always to delete contaminated material. Sometimes the goal is to study it; contaminates become the science. That rubbish is telling – useful in understanding how and where science is manipulated. This points to faults in the system, and to research security and integrity intelligence. This is what we do in forensic scientometrics.

Having the choice to remove that noise is also vital to a lot of science. Having the ability to discern noise from signals is part of the discovery process. Polish that and you lose knowledge – not gain it.

Interestingly, this incorporation of and regurgitation of contamination is not an AI hallucination problem. It is more dangerous in some ways because the system may be retrieving real documents, citing real sources, and still producing an answer built on corrupted evidence. And as already noted, evidence that appears trustworthy.

Confidence and Friction

What happens when we ask the same research question against a clean corpus vs. a contaminated one? Specifically: not just whether the answer was wrong, but how. Does the framing shift? Is confidence tracked inversely to accuracy, whether the contaminated answer sounded more authoritative than the clean one?

The outcome in these four scenarios with differing filtration levers appears very similar - you get a polished, truthful-sounding, confident output.

Confidence is one part of this picture - the polished frame. It is the default of any LLM. All AI systems I have seen are built to produce polished confidence.

The lack of friction is the other issue. When we research ideas, we need challenges to our thoughts – that friction is what helps ideas percolate. The introduction of LLMs and MCPs remove friction by design. And with things like Claude Science, they aim to make science smoother. The linked blog states some of my concerns with a lack of friction.

Yet, when a system removes too much friction, we lose the ability to inspect and better understand the science. We actually need that time to think.

Trustworthy looking Results - A Simple Example

A fictitious disease was invented and put onto a blog and preprint server in 2024. It was consumed and regurgitated in LLMs even to this year (2026).

But MCPs would purportedly control the literature. So, I would argue that developing an MCP on top of unfiltered literature would reduce some of the noise. But the vengeful bixonimania would have still made it through a search because it was inserted into the scientific literature to appear as trustworthy.

Provenance as Evidence

Provenance metadata are the machine-readable evidence layer that tells a system what kind of material it is handling. They are labels or tags that give more information about the data.

A tag could be as simple as a label that says an author is a name. Another tag might show whether a paper has been retracted, corrected, or flagged with an expression of concern. It might show why it was retracted. Retractions due to honest error or paper mill activity should not be collapsed into the same bucket – both matter, but they matter differently.

A tag could show whether a record is linked to known manipulation patterns. It might show whether the journal has integrity concerns, whether the paper sits in a suspicious citation cluster, whether the authorship history looks strange, whether the metadata changed, whether the publication came from a venue with repeated problems.

The combination of metadata might also show fitness for use. The paper suitable for discovery may not be suitable for evidence synthesis – useful for integrity investigation, but not fit for policy decisions.

We need tags that help humans and machines enrich research output to understand provenance, risk, context, and purpose.

The provenance metadata tags don’t pass judgement but do make more informed decision making possible. Without them, MCP-connected systems will keep treating retrieval as if it were trust. They will keep moving water without knowing what is in it. Or in our forensic context, they will keep giving polished answers from compromised inputs.

And the consumer will be left to hope for the best.

Who is Responsible for the Data

Who is responsible when contaminated science becomes machine-actionable?

Right now, too much of that burden lands on the end user. The person least equipped to inspect the whole source layer gets handed the final answer and asked to trust it.

The future of AI-enabled science will not be determined only by who builds the fastest pipes. The pipes are already here. They have become faster, smoother, and will now be more deeply embedded in institutional work.

If we want AI systems that serve science, we need more than better connectors. We need provenance infrastructure. We need trust markers that move with the data. We need systems that know whether they are studying contamination or using evidence.

Fast pipes are not enough. Someone has to prove the water is fit to drink.

Ancient aqueduct. (Modified image originally by Jason Rocks from Pixabay)

Ancient aqueduct. (Modified image originally by Jason Rocks from Pixabay)

Reuse

Citation

For attribution, please cite this work as:
McIntosh, Leslie. 2026. “AI Aqueducts: Speed. Confidence. And Contaminants.” FoSci Blog, July 2. https://fo-sci.org/blog/2026-07-02/.