
I’m a Natural Scientist & Statistical Modeler. I’ve spent over a decade designing and performing experiments, developing statistical models informative of the system, and then coding each step of the analyses. A couple of days ago, Anthropic announced “Claude Science”, an “AI workbench for scientists”. Apparently it’s at the same level of flagshipness as Claude Code and Claude Cowork (not sure what level that is, nor who cares for the latter, but still). They also appear to be late to the game, since competitors have released their products months ago like Google’s Gemini for Science (pretty boring title, like Anthropic) and OpenAI’s GPT-Rosalind (more interesting, but borderline offensive to anyone who cares about science).
Except for a few people on social media, who have no interest nor qualification to perform scientific research, but seem to make a big deal about anything these companies release, nobody seemed to care about those AI workbenches, not even (in fact, especially not) scientists, but this being Anthropic we can count on a lot of hype on their side – they’ve definitely been able to dominate the headlines, though with a low signal-to-noise ratio. I think scientists really should NOT care; they’ve been using automated software and pipelines for ever, which carry some of the same speed/quality trade-offs. Point-and-click apps like Prism streamline statistical testing in research, the problem being that many scientists using it do not understand the underlying statistics. Bioinformatic pipelines are standard in many fields for decades now. Unfortunately, many of these tools create hamster wheel-like incentives to repeat similar experimental set ups and run the data through the same pipes without deeper thinking of the assumptions and implications, promoting repetitive, uncreative science. That’s the wheel Anthropic is recreating (or re-recreating); it’s also no different from having an LLM plugin on your coding IDE (e.g. Claude Code itself, or something like Pi coupled to a local model like Google’s Gemma4 or Mistral’s Le Chaton Fat).
In addition to the blind automation from any tool, like anything that uses “AI” these days it introduces the risk of thought- and brain-less slop: workslop, thinkslop, and otherwise vibe-scienceing. Academia has many problems, but on that front most of them are absolutely NOT related to moving too slow because tools are lacking. Quite the contrary, scientists need more time to think their experiments, statistical analyses, interpretation and discussion, not a pseudo-intelligent tool that cuts corners, prevents them from thinking deeply about their problems and make deliberate choices about how to solve them, as well as removing the very struggles that make scientists experts by understanding the intricacies of their systems of interest. That’s where the real science lives, where the rubber of untested ideas hits the road, and scientists need to make tough decisions about insisting on a hypothesis, reformulating, or discarding a theory.
Beyond science, this is a typical tech approach to things: treat everything like a low-stakes engineering problem, a cute puzzle that can be solved with enough technicality, and if it fails you can always throw more tech at the problem. Science doesn’t work like that, though, science is exactly the art of solving problems by throwing less science at it, it’s finding the statistical model that represents the system well enough, with few parameters that can be inferred (as always, I now have to refer to the definition of inference) from the experimental data available, under a theoretical framework that is well-understood but can be challenged. AI/ML may have success in some cases where there’s a lot of data and little actual knowledge of how things work, but we’re also finding out that these successes don’t scale trivially to complex problems in most areas.
-- caetano,
July 2, 2026