Can a Computer Read Your Baby's Eye Photographs?

Two 2026 studies taught software to spot the most dangerous sign of retinopathy of prematurity — and showed how easily that skill is lost when the software moves to a new hospital

Two research teams published studies in March 2026 testing whether artificial intelligence can identify the warning sign that tells doctors a premature baby's eye disease needs urgent treatment. Both systems performed well on the images they were trained with. Both stumbled when shown photographs from a hospital they had never seen — and that gap, not the impressive accuracy, is the finding that matters.

The two studies took different routes to the same goal [1][2]. One group, in China, tried mixing photographs from hospitals in different countries so the software would see more variety during training. The other, in Egypt, built a more sophisticated program and then deliberately tested it on collections of images from other countries. Understanding why either approach was needed means starting with the condition itself.

What Retinopathy of Prematurity Is, and Why It Is Watched So Closely

Babies born very early have retinas — the light-sensing tissue at the back of the eye — whose blood vessels have not finished growing. After birth those vessels sometimes grow abnormally. This is called retinopathy of prematurity, or ROP, and in its severe form it can pull the retina out of place and cause permanent blindness. Treated in time, that outcome is usually preventable. Untreated, it is not.

The signal doctors watch for is called plus disease: the blood vessels near the centre of the retina become unusually wide and twisted. Its precise definition is set out in an international classification agreed by eye specialists worldwide [3]. When plus disease appears, treatment is usually needed within days.

The difficulty is that recognising it is a judgement call. There is no measurement and no blood test — an ophthalmologist looks at the retina and decides whether the vessels are twisted enough. Experienced specialists looking at the same eye do not always agree. And because there is no way to predict in advance which babies will develop it, every baby born early enough or small enough must be examined repeatedly. In many countries the schedule is set by a paediatric policy that lists which infants qualify [4]. Each examination means dilating drops, a speculum holding the eyelids open, and a bright light — an ordeal for a fragile baby, repeated weekly for weeks. A large study following premature infants through their examinations documented how many are needed to find the relatively few babies who go on to require treatment [5].

The Problem Families Faced Before This Research

For most of the history of neonatal care, this system depended entirely on an ophthalmologist walking into the nursery. Where one was available, screening worked, and blindness from ROP became rare. Where one was not — which describes much of the world, and an increasing number of rural regions in wealthy countries — babies were simply not screened, and some went blind from a condition that treatment would have stopped. Parents in those settings were never offered a choice; the examination was not available to decline.

The first attempt at a solution was photography. Instead of the specialist coming to the baby, a trained nurse or technician photographs the retina with a special camera, and the images are sent to a specialist elsewhere. Researchers showed in the 2000s that this remote reading was accurate and reliable enough to use [6], and it remains how many screening programmes run today. But it still requires a specialist at the other end, and there are not enough of them.

That is where computers entered. In 2018 a research consortium reported that a program trained on thousands of retinal photographs could identify plus disease about as well as individual experts [7], and another team reported similar results the same year [8]. These results were genuinely impressive. Eight years later, no such system is in routine use — and the reason is the subject of the two new studies.

What Happens When the Software Changes Hospitals

A review published in 2023 found that many of these programs had been built but very few had been properly tested outside the hospital that created them [9]. When one system developed in the United States was tried in three lower-income countries, it worked less well than at home [10]. The reasons are mundane rather than mysterious: different camera models, different eye-drop protocols, different lighting, different populations of babies. A program can learn to recognise the hospital rather than the disease, and nobody notices until it moves.

The Chinese team measured this directly, and their numbers are stark [1]. Software trained only on images from an Iranian hospital was correct about 92% of the time on new Iranian images — and about 37% of the time on Chinese images. Trained only on Chinese images, the same pattern reversed: 94% at home, 42% abroad. A program that appeared to be an expert became worse than guessing the moment it crossed a border.

Mixing the two collections together during training fixed most of that, bringing accuracy back above 92% for both. But when the combined program was finally tested on a third collection, from India, overall accuracy was 80%. Within that figure sits the finding families and clinicians should know about: the program correctly identified every case of full plus disease, but it also raised the alarm on many healthy eyes — only about a quarter of the eyes it flagged actually had the disease. And for the in-between stage, called pre-plus, it caught only one case out of twelve.

The Egyptian team's program, which combines two different image-analysis systems and lets them vote, did better on unfamiliar images — about 93% accuracy on the Indian collection [2]. Part of that advantage is real, and part is a difference in the question asked: their test asked only "disease or no disease," rather than sorting eyes into three categories. Merging the in-between category away removes the very judgement that both programs find hardest.

What This Means for Your Baby's Care

The honest summary is that no software currently decides whether a baby needs eye treatment, and none of this research suggests it should. A doctor makes that decision. If your baby's unit uses a system like these, it is being used as an extra set of eyes — flagging photographs that should be looked at sooner — and not as a replacement for the ophthalmologist.

It is also fair to say why the research matters anyway. Both studies were built on image collections that researchers released publicly in 2024, one from Iran [11] and one from India [12] — a deliberate act of sharing that made it possible for the first time to check whether these programs travel. The most valuable thing in either paper is not the 93% or the 96%; it is the 37%. A field that publishes its failures is a field that is checking itself, and that is exactly what should happen before software is allowed anywhere near a decision about a baby's sight.

There is one number in these studies worth understanding, because it explains why doctors are cautious rather than enthusiastic. When the Chinese team's program flagged an eye as having plus disease in the Indian images, it was right about a quarter of the time. That is not a dangerous kind of error — it errs toward raising the alarm rather than missing disease, which is the safer direction for a screening tool. But it means that a program of this quality, used on its own, would send many families through the worry of an urgent referral for eyes that turned out to be fine. Knowing how often a test cries wolf is as important as knowing how often it is right, and it is a fair question to ask about any technology offered for your baby.

If your baby is being screened for ROP, the practical implications are unchanged: the examinations still matter, the schedule still matters, and keeping the follow-up appointment after discharge matters a great deal, because the eyes continue to develop after a baby goes home.

What Researchers Are Working On Next

Three things need to happen before any of this reaches the bedside. These programs need to be tested forward in time — used on babies as they are actually screened, rather than on photographs collected years ago. They need to be tested on far more images from far more hospitals, especially in the places with no ophthalmologist, where the potential benefit is greatest. And they need to become better at the uncertain middle ground, since that is where they currently fail and where human specialists most want help.

Both teams describe the same longer-term goal: software that does not simply answer yes or no, but measures how twisted and how wide the vessels are, giving a number that can be tracked over time as an eye improves or worsens [1][2]. That would help specialists as much as it would help hospitals without one. It does not exist yet — but the honesty with which these two groups reported where their systems broke is a reasonable sign that the people building it are asking the right questions.

References

  1. Zhang X, Liang H, Wang R, Wu R, Shen X, Zhang Y. Research on the construction of an AI diagnostic model for plus disease of retinopathy of prematurity based on cross-center fusion datasets. Frontiers in Pediatrics. 2026;14:1765353. doi:10.3389/fped.2026.1765353
  2. Mohiy E, AbdulWakel HI, Khairy M, Houssein EH. An efficient deep learning model for reliable detection and classification of retinopathy of prematurity. Discover Artificial Intelligence. 2026;6(1):305. doi:10.1007/s44163-026-00834-y
  3. Chiang MF, Quinn GE, Fielder AR, et al. International Classification of Retinopathy of Prematurity, Third Edition. Ophthalmology. 2021;128(10):e51–e68. doi:10.1016/j.ophtha.2021.05.031
  4. Fierson WM, Chiang MF, Good W, et al. Screening examination of premature infants for retinopathy of prematurity. Pediatrics. 2018;142(6):e20183061. doi:10.1542/peds.2018-3061
  5. Quinn GE, Ying GS, Bell EF, et al. Incidence and early course of retinopathy of prematurity: secondary analysis of the Postnatal Growth and Retinopathy of Prematurity (G-ROP) study. JAMA Ophthalmology. 2018;136(12):1383–1389. doi:10.1001/jamaophthalmol.2018.4290
  6. Chiang MF, Wang L, Busuioc M, et al. Telemedical retinopathy of prematurity diagnosis: accuracy, reliability, and image quality. Archives of Ophthalmology. 2007;125(11):1531–1538. doi:10.1001/archopht.125.11.1531
  7. Brown JM, Campbell JP, Beers A, et al. Automated diagnosis of plus disease in retinopathy of prematurity using deep convolutional neural networks. JAMA Ophthalmology. 2018;136(7):803–810. doi:10.1001/jamaophthalmol.2018.1934
  8. Wang J, Ju R, Chen Y, et al. Automated retinopathy of prematurity screening using deep neural networks. EBioMedicine. 2018;35:361–368. doi:10.1016/j.ebiom.2018.08.033
  9. Ramanathan A, Athikarisamy SE, Lam GC. Artificial intelligence for the diagnosis of retinopathy of prematurity: a systematic review of current algorithms. Eye (London). 2023;37(12):2518–2526. doi:10.1038/s41433-022-02366-y
  10. Coyner AS, Oh MA, Shah PK, et al. External validation of a retinopathy of prematurity screening model using artificial intelligence in 3 low- and middle-income populations. JAMA Ophthalmology. 2022;140(8):791–798. doi:10.1001/jamaophthalmol.2022.2135
  11. Akbari M, Pourreza HR, Khalili Pour E, et al. FARFUM-RoP, a dataset for computer-aided detection of retinopathy of prematurity. Scientific Data. 2024;11(1):1176. doi:10.1038/s41597-024-03897-7
  12. Agrawal R, Walambe R, Kotecha K, et al. HVDROPDB datasets for research in retinopathy of prematurity. Data in Brief. 2024;52:109839. doi:10.1016/j.dib.2023.109839