Within Closed Loops
Does Real Feedback Make AI Better at Discovery?
Experiments with shuffled feedback suggest that AI agents improve because of real laboratory signals, not merely memorised scientific patterns.
On this page
- How the biological experiments were structured
- What changed when feedback was randomised
- Why stronger models benefited more from evidence
Page outline Jump by section
Introduction
One of the central questions for closed-loop AI science is whether an AI system genuinely learns from experiments or merely repeats patterns absorbed during training. A convincing way to test this is surprisingly simple: deliberately scramble the experimental feedback. If an AI continues to perform just as well after the feedback has been randomised, then its apparent improvement probably comes from prior knowledge or prompt engineering rather than from interpreting new evidence. If its performance collapses, that suggests it was relying on real laboratory signals.
Recent research has begun applying exactly this kind of control. The results provide some of the strongest evidence so far that sufficiently capable AI systems can adapt their scientific reasoning to genuine experimental outcomes, while also revealing that weaker systems often fail to do so. For the broader vision of AI-enabled scientific acceleration, this distinction matters. Faster discovery depends not simply on generating plausible hypotheses, but on repeatedly changing those hypotheses when experiments demand it.
How the biological experiments were structured
The clearest evidence comes from recent work on iterative biological discovery using Cell Painting, a high-content microscopy technique that measures how cells respond to genetic or chemical perturbations. Instead of asking an AI to make a single prediction, researchers placed it inside an experimental loop.
The process was deliberately iterative:
- The AI proposed promising genes or perturbations to investigate.
- Laboratory experiments measured the biological outcomes.
- The experimental results were returned to the AI.
- The AI revised its next round of hypotheses using those observations.
This resembles the way human scientists gradually narrow the search space during research rather than attempting to solve an entire biological problem in one step.
The researchers compared two approaches. One agent received genuine laboratory feedback after each round, while a baseline agent relied only on its pretrained scientific knowledge without incorporating new experimental evidence. Across approximately 800 independently replicated experiments, the feedback-enabled agent produced substantially more successful discoveries, with access to real feedback increasing discoveries per feature by around 53%. The improvement was statistically significant, suggesting that repeated experimental interaction added information beyond what the model already knew.[arXiv]arxiv.orgCan AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation DiscoveryMarch 27, 2026…
What changed when feedback was randomised?
The most revealing part of the study was not the improvement itself but the control experiment.
Why shuffle the answers?
Machine learning systems can appear to “learn” for misleading reasons. A model may recognise familiar biological terminology, infer likely answers from previous publications or simply exploit hidden patterns in prompts. To distinguish genuine adaptation from these possibilities, researchers randomised the feedback.
Instead of returning the correct laboratory outcomes, they permuted the hit-or-miss labels before feeding them back to the AI. Every response still looked like experimental data, but its relationship to reality had been destroyed.
This is analogous to giving a scientist randomly altered laboratory notebooks. If they still appeared to improve, the improvement could not be attributed to learning from experiments.
The performance gain disappeared
That is exactly the result researchers hoped for as a scientific control.
Once the feedback was randomised, the advantage of iterative learning largely vanished. The system no longer outperformed the baseline that relied solely on prior knowledge. In other words, the AI benefited only when the laboratory information preserved its real structure. When that structure was destroyed, so was the improvement.[arXiv]arxiv.orgCan AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation DiscoveryMarch 27, 2026…
This makes the randomised feedback test particularly valuable because it addresses a longstanding criticism of AI discovery systems. Critics have argued that apparent scientific reasoning may simply reflect sophisticated retrieval from training data. By demonstrating that shuffled feedback eliminates the benefit, the experiment provides evidence that the AI was responding to genuinely informative observations rather than merely reproducing memorised associations.
Why stronger models benefited more from evidence
An equally important finding was that not all AI models used the feedback equally well.
Earlier versions of the system showed only weak or statistically insignificant improvements from iterative experimental feedback. Upgrading to a more capable language model changed that picture dramatically.
The stronger model:
- extracted more useful information from each experimental round;
- incorporated that information into later hypotheses more effectively;
- produced significantly larger discovery gains; and
- dramatically reduced gene hallucinations, where the AI invented or incorrectly referenced biological entities.
In the reported experiments, hallucination rates fell from roughly one-third to below one-tenth after the model upgrade. More importantly, what had previously been a negligible learning effect became a large, statistically significant improvement when genuine feedback was available.[arXiv]arxiv.orgCan AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation DiscoveryMarch 27, 2026…
This suggests that learning from experiments may represent an emerging capability rather than a universal property of large language models. Simply placing any AI inside a laboratory loop is not enough. The model must first possess sufficient reasoning ability to recognise which observations should change its beliefs.
Why this matters for closed-loop AI scientists
Randomised feedback experiments test something deeper than predictive accuracy. They ask whether an AI behaves scientifically.
Scientific reasoning requires more than producing plausible explanations. It requires abandoning explanations when evidence contradicts them.
The randomised-feedback design therefore evaluates whether an AI:
- distinguishes informative from uninformative evidence;
- changes future experimental choices based on observations;
- avoids blindly repeating previous assumptions; and
- improves because reality teaches it something new.
These are precisely the properties that separate an automated research assistant from a genuine closed-loop discovery system.
The experiments also illustrate why evaluation standards for AI scientists must extend beyond benchmark scores. An AI may answer biology questions accurately while still failing to incorporate fresh laboratory evidence. Conversely, an agent that reliably updates its hypotheses after every experiment may contribute meaningfully to scientific discovery even if its initial guesses are imperfect.
What the findings do not prove
Although these results are encouraging, they should not be interpreted as showing that AI systems have achieved human-like scientific reasoning.
Several limitations remain.
First, the experiments were conducted within a specific biological discovery task rather than across all areas of science. Different experimental domains may present very different challenges.
Second, the feedback was relatively structured. Real scientific research often produces ambiguous, conflicting or partially reliable evidence that is much harder to interpret than binary success or failure.
Third, demonstrating sensitivity to experimental evidence is not the same as demonstrating deep causal understanding. An AI may learn useful patterns from feedback without possessing the conceptual models that human scientists use to explain underlying mechanisms.
Other recent evaluations of autonomous research agents have likewise found that many systems still struggle to revise hypotheses consistently when confronted with contradictory evidence, highlighting that effective closed-loop scientific reasoning remains an active research challenge rather than a solved problem.[SSRN]papers.ssrn.comA Survey of 80 "AI Scientist" Systems: From Evaluator Quality to Validated Discovery and Feedback Loops by Jemin George:: SSRNJune 6…
Why these tests matter for AI-enabled scientific acceleration
Within the broader vision of AI helping humanity accelerate discovery, randomised feedback experiments provide an unusually rigorous form of evidence. Rather than asking whether AI can generate convincing scientific language, they ask whether reality changes the AI’s behaviour.
That distinction is crucial.
If future AI scientists can reliably extract information from laboratory experiments, revise unsuccessful hypotheses and progressively improve over many experimental cycles, they could shorten the iterative process that dominates modern research. Faster iteration could eventually contribute to advances in medicine, biology, materials science and other fields central to long-term human flourishing.
The current evidence does not show that AI has reached that point. It does, however, show something more specific and scientifically valuable: under controlled experimental conditions, stronger AI models appear capable of benefiting from genuine laboratory feedback, and that benefit disappears when the feedback is deliberately stripped of its informational content. Randomised feedback tests therefore provide one of the clearest demonstrations so far that at least some improvements in AI-guided discovery arise from interaction with experimental evidence rather than from simply repeating what the models already knew.[arXiv]arxiv.orgCan AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation DiscoveryMarch 27, 2026…
Amazon book picks
Further Reading
Books and field guides related to Does Real Feedback Make AI Better at Discovery?. Use these as the next step if you want deeper reading beyond the article.
Artificial Intelligence
Rating: 4.5/5 from 10 Google Books ratings
Artificial intelligence: A Modern Approach, 3e,is ideal for one or two-semester, undergraduate or graduate-level courses in Artificial In...
The Book of Why
The hugely influential book on how the understanding of causality revolutionized science and the world, by the pioneer of artificial inte...
Superforecasting
The international bestseller 'A manual for thinking clearly in an uncertain world. Read it.' Daniel Kahneman, author of Thinking, Fast an...
Thinking, Fast and Slow
Why is there more chance we'll believe something if it's in a bold type face? Why are judges more likely to deny parole before lunch? Why...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromscience model oneBay.co.uk.
Endnotes
1.
Source: arxiv.org
Link:https://arxiv.org/abs/2603.26177
Source snippet
Can AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation DiscoveryMarch 27, 2026...
Published: March 27, 2026
2.
Source: papers.ssrn.com
Link:https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6889440
Source snippet
A Survey of 80 "AI Scientist" Systems: From Evaluator Quality to Validated Discovery and [Feedback Loops]({{ 'feedback-loops/' | relative_url }}) by Jemin George:: SSRNJune 6...
3.
Source: arxiv.org
Title: arXiv The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Link:https://arxiv.org/abs/2408.06292
4.
Source: papers.cool
Link:https://papers.cool/arxiv/2603.26177
Additional References
5.
Source: owkin.com
Link:https://www.owkin.com/publications/can-ai-scientist-agents-learn-from-lab-in-the-loop-feedback-evidence-from-iterative-perturbation-discovery
Source snippet
Evidence from Iterative Perturbation DiscoveryMarch 27, 2026 — March 27, 2026 arXiv CAN AI SCIENTIST AGENTS LEARN FROM LAB-IN-THE-LOOP FE...
Published: March 27, 2026
6.
Source: researchtrend.ai
Title: Can AI Scientist Agents Learn from Lab-in-the-Loop Feedback?
Link:https://researchtrend.ai/papers/2603.26177
Source snippet
Evidence from Iterative Perturbation Discovery | ResearchTrend.AIMarch 30, 2026 — CAN AI SCIENTIST AGENTS LEARN FROM LAB-IN-THE-LOOP FEED...
Published: March 30, 2026
7.
Source: pubmed.ncbi.nlm.nih.gov
Link:https://pubmed.ncbi.nlm.nih.gov/42420458/
Source snippet
2026 Jul 8. doi: 10.1038/s41586-026-10742-x. Online ahead of print. LARGE LANGUAGE MODELS CAN PREDICT THE RESULTS OF SOCIAL SCIENCE EXPER...
8.
Source: youtube.com
Title: AI Scientist v2: The AI That Writes Scientific Papers Accepted by Peer Review
Link:https://www.youtube.com/watch?v=mg68wk40MO8
Source snippet
AI-Driven Research Workflows: Lessons learned from a million automated experiments - Paul Jensen...
9.
Source: youtube.com
Link:https://www.youtube.com/watch?v=g45Alfg7diw
Source snippet
AI Scientist v2: The AI That Writes Scientific Papers Accepted by Peer Review...
10.
Source: nature.com
Link:https://www.nature.com/articles/s41586-026-10742-x
Source snippet
July 8, 2026 — * Article *...
Published: July 8, 2026
11.
Source: nature.com
Link:https://www.nature.com/articles/s41562-024-02077-2
12.
Source: nature.com
Link:https://www.nature.com/articles/s41539
13.
Source: nature.com
Link:https://www.nature.com/articles/d41586-024-03939-5
14.
Source: youtube.com
Title: Google’s AI Co-Scientist Explained
Link:https://www.youtube.com/watch?v=x-ZqPELgYQA
Source snippet
How [Autonomous Labs]({{ 'autonomous-labs/' | relative_url }}) Will Transform Scientific Research: Ginkgo Bioworks’ Jason Kelly...


