AI scientists produce results without reasoning scientifically

A comprehensive study reveals that while LLM-based agents can execute scientific workflows, they fail to adhere to the epistemic norms of scientific inquiry, often ignoring contradictory evidence and failing to self-correct.
Computer Science > Artificial Intelligence
Title: AI scientists produce results without reasoning scientifically
Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make scientific inquiry self-correcting is poorly understood.
In this study, researchers evaluate LLM-based scientific agents across eight domains, spanning workflow execution to hypothesis-driven inquiry, through more than 25,000 agent runs and two complementary lenses: (i) a systematic performance analysis that decomposes the contributions of the base model and the agent scaffold, and (ii) a behavioral analysis of the epistemological structure of agent reasoning.
Key Findings:
- The base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold.
- Across all configurations, evidence is ignored in 68% of traces.
- Refutation-driven belief revision occurs in only 26% of cases.
- Convergent multi-test evidence is rare.
The same reasoning pattern appears whether the agent executes a computational workflow or conducts hypothesis-driven inquiry. These patterns persist even when agents receive near-complete successful reasoning trajectories as context, and the resulting unreliability compounds across repeated trials in epistemically demanding domains.
Conclusion:
Current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcome-based evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it.
Source: arXiv cs.AI Recent
















