AI Helpers in Newborn Care: What They Can Do Today, and Why They Must Always "Show Their Work"
A plain-language look at the AI tools now used in neonatal medicine and research — and a study from the In[Neo]Sight team on why an AI must look things up rather than guess from memory
Artificial intelligence, or AI, is moving into newborn intensive care — not as a robot doctor, but as a set of helpers that can take notes, search the medical literature, and draft summaries. Some of these tools are genuinely useful today, one has strong evidence behind it, and others are still experimental and not safe to rely on for decisions about a baby. The most important thing to understand as a parent or reader is a single rule that runs through all of it: a trustworthy AI helper must "show its work" by pointing to the real medical sources it used, rather than answering from memory. The In[Neo]Sight team tested exactly what happens when an AI answers from memory alone, and the results explain why that rule matters so much.
How AI Got Into the Hospital
The kind of AI in the news today is built on "large language models" — computer systems trained on enormous amounts of text so they can understand and produce human-like writing. Researchers found early on that these systems had absorbed a surprising amount of medical knowledge and could reason through clinical questions [1], and later work showed they could answer medical-exam-style questions at close to a passing level [2]. But careful reviews also raised a warning that has never gone away: these systems can state something false with complete confidence, inventing details that sound right but are not [3]. At the same time, other studies found that AI answers to health questions could feel more thorough and more caring than rushed human replies [4], which is part of why hospitals became interested. The newest step is that these systems have become "agents" — they no longer just chat, but can carry out multi-step tasks like searching, looking things up, and writing a draft. That makes them more helpful, but it also makes their mistakes harder to catch, which is why doctors are being careful about where they use them.
The One Tool With Strong Evidence: AI That Takes Notes
Doctors and nurses spend enormous amounts of time typing notes into the computer, time taken away from the bedside. "Ambient" AI note-takers listen to a visit and write a draft note for the clinician to check. This is the one AI tool with strong evidence: a high-quality study in which clinicians were randomly assigned to use the tool or not found that it cut documentation time and reduced burnout [5]. For families, the benefit is indirect but real — a clinician who spends less time typing can spend more time with your baby. Two things are worth knowing. The strong study was done in ordinary outpatient clinics, not in newborn intensive care, so hospitals still need to check that it helps there too. And the doctor always reviews and signs the note, because the AI can still make mistakes.
A Fast-Moving Research Helper: AI That Reads the Literature
A second group of tools helps researchers and doctors search through the mountain of published studies. Newer systems can read many papers, pick out the relevant ones, and pull out the key numbers, working much like a team of human reviewers [6]. One especially striking project reported that its AI could screen and summarise studies extremely accurately and re-created an entire batch of respected "Cochrane" medical reviews in just two days — work that would normally take people years [7]. This is promising, but it comes with an important caution: that project has not yet been fully checked by other scientists, and experts still confirm the AI's work rather than trusting it blindly. Think of these tools as a very fast research assistant whose work a human always double-checks.
The Tools That Are Not Ready: AI That "Decides"
The riskiest use — an AI that suggests a diagnosis or a treatment on its own — is not ready for newborn care, and the research is clear about it. When scientists tested AI "agents" on medical decision tasks, even with access to web search and other tools, the systems were only a little better than simpler ones and were judged not reliable enough for routine use [8]. A careful test in cancer care reached the same conclusion [9]. And a review of AI built specifically for newborn intensive care found that most of it is still experimental, held back by data problems and, crucially, by rarely being tested in hospitals other than the one where it was built [10]. The good news is that engineers have a fix for the "making things up" problem: a method that forces the AI to look up real documents and base its answer on them. A review of this approach found it consistently reduces invented answers and improves accuracy [11]. This is the technical version of "show your work."
The In[Neo]Sight Test: Why "Look It Up" Beats "Remember"
To show why this matters, the In[Neo]Sight team ran its own study [12]. They took 14 recent, high-quality newborn and pregnancy studies — most published after the AI had finished learning — and asked an AI (Claude Sonnet 4.6) each study's main question, without letting it see the actual paper. Then they checked whether the AI's answer pointed in the same direction as the real result. This "same direction or not" check matters, because "this treatment works" and "this treatment does not work" look almost identical to a simple word-matching computer, so you need a smarter check to know if the AI was truly right.
The AI got the direction right in 9 of the 14 studies — about 64% — and was wrong in 5. The pattern of mistakes was telling: in four of its five errors, the AI was too cautious, saying "not proven, don't use it" about treatments that the new studies actually found helpful. It was good at recognising when something made no difference, but it lagged behind on new positive discoveries — exactly the kind of up-to-date news a family would most want a doctor to have. The team also looked at how often the AI said "more study is needed." It added that phrase to 8 of its 14 answers, but that caution was only right half the time, and in one case it confidently gave a wrong answer with no warning at all. And here is the twist that drives the whole point home: that one confident, wrong answer was about the single oldest study in the test — the only one old enough for the AI to have already studied before. The AI was most confidently wrong about the very paper it had the best chance of already knowing. The lesson is that you cannot tell whether an AI is right just from how sure it sounds. That is the whole reason In[Neo]Sight's own articles are not written by a single AI guessing from memory. Instead of one helper that looks things up while it writes, the work is split among three: one finds and double-checks the real papers, a second writes using only those checked facts, and a third re-checks every statement against the sources — and a human editor reviews the result before anyone reads it.
What This Means for Families, and What Comes Next
If you hear that your baby's care team is using AI, the most likely and most helpful use today is note-taking that frees up the clinician's time, always with a human reviewing the result. AI is also speeding up how quickly doctors can find and summarise the latest research, which can help your team stay current with new findings. It is completely reasonable to ask a member of the care team how any AI tool is being used and whether a person has checked its output — good teams will welcome the question, because the same principle guides them: the AI assists, and a human remains responsible. What AI is not doing — and should not be doing — is making decisions about your baby on its own. If anyone ever describes an AI as making a diagnosis or choosing a treatment by itself, that is a reason to ask more questions, not fewer. Every study here points the same way: these tools are helpers, not decision-makers, and the trustworthy ones show you where their information comes from. This evidence has limits worth being honest about — the In[Neo]Sight test looked at only 14 studies with one AI system [12], the fast literature-review results are still awaiting full scientific review [7], and the note-taking study was done outside newborn units [5]. What researchers are working on next is AI made specifically for newborn care that always grounds its answers in real medical evidence, tested carefully before it is trusted — so that the speed of AI comes with the safety that sick newborns and their families deserve.
References
- Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine. Nature Medicine. 2023;29:1930–1940. doi:10.1038/s41591-023-02448-8 ↩
- Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–180. doi:10.1038/s41586-023-06291-2 ↩
- Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. New England Journal of Medicine. 2023;388(13):1233–1239. doi:10.1056/NEJMsr2214184 ↩
- Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine. 2023;183(6):589–596. doi:10.1001/jamainternmed.2023.1838 ↩
- Lukac PJ, et al. Ambient AI scribes in clinical practice: a randomized trial. NEJM AI. 2025. doi:10.1056/AIoa2501000 ↩
- Khan MA, Ayub U, Naqvi SAA, et al. Collaborative large language models for automated data extraction in living systematic reviews. Journal of the American Medical Informatics Association. 2025;32(4):638–647. doi:10.1093/jamia/ocae325 ↩
- Cao Y, Arora A, et al. Automation of systematic reviews with large language models. medRxiv [preprint]. 2025:2025.06.13.25329541. doi:10.1101/2025.06.13.25329541 ↩
- Kather JN, et al. Benchmarking large language model-based agent systems for clinical decision tasks. npj Digital Medicine. 2026;9(1):259. doi:10.1038/s41746-026-02443-6 ↩
- Ma W, et al. AI for evidence-based treatment recommendation in oncology: a blinded evaluation of large language models and agentic workflows. Frontiers in Artificial Intelligence. 2025;8:1683322. doi:10.3389/frai.2025.1683322 ↩
- Tudor S, Khan UR, et al. Opportunities and challenges of using artificial intelligence in predicting clinical outcomes and length of stay in neonatal intensive care units: systematic review. Journal of Medical Internet Research. 2025;27:e63175. doi:10.2196/63175 ↩
- Amugongo LM, Mascheroni P, Brooks SG, et al. Retrieval augmented generation for large language models in healthcare: a systematic review. PLOS Digital Health. 2025;4(6):e0000877. doi:10.1371/journal.pdig.0000877 ↩
- Chou F-S. Reliability of large language model out-of-the-box memory in neonatal–perinatal medicine [internal analysis]. In[Neo]Sight. 6 July 2026. Available from the authors. ↩