
Since the November 2022 release of OpenAI’s ChatGPT, generative artificial-intelligence tools have made their way into clinics at speed and scale. Some of the numbers are staggering.

Bring us your LLMs: why peer review is good for AI models
In less than four years, thousands of studies have been published describing, testing and evaluating the use in health care of both general-purpose AI models and specialized systems trained on medical knowledge. According to one estimate, roughly three peer-reviewed articles about AI tools used in clinical medicine are published every day1. For many people, chatbots have become a go-to source of medical advice. Every week, more than 230 million people worldwide ask ChatGPT health-related questions.
Moreover, the health-care sector is seeing huge growth in the number of AI-powered medical products aimed at assisting clinicians and clinical-administration teams with their work, including in their decision-making. AI systems can now handle complex administrative tasks and order laboratory tests. They can help clinicians to prescribe drugs2 and interpret X-rays3, magnetic resonance imaging (MRI) scans and computed tomography (CT) images4. They can also be used to diagnose rare diseases5.
But these achievements come with some big questions: who is regulating this innovative class of medical product; how are they doing this; and how can the public be confident that AI-powered medical systems are capable of doing what developers say they do?
Feedback wanted
The US Food and Drug Administration (FDA) is seeking feedback to answer some of these questions through a discussion paper published in August. The article considers how to regulate generative-AI-enabled medical devices (see go.nature.com/3tagjxj). We urge researchers to submit their views. The deadline for submissions is 19 October. Other countries are also publishing proposals for AI in health care, including medical-device regulations.

Transparent research: can big tech learn from big pharma?
The FDA and its counterparts around the world currently authorize medical products according to a scale of risk. In general, all products must meet legally binding quality and safety standards, but not all need to be tested independently of their manufacturers or in real-world settings. Manufacturers of bandages, for example, can self-certify that standards have been met. Makers of health-assessment devices, such as stethoscopes, often need third-party authorization, but such tests can be done in a lab. By contrast, diagnostic products that affect clinical decisions have to be tested in real-world situations and, in some cases, through clinical trials similar to those used for assessing drugs. In general, the greater the risk to a person if something goes wrong with the product, the more scrutiny is needed to obtain authorization.
The FDA and others are now asking whether medical products powered by generative AI should be scrutinized more comprehensively. The answer in most cases should be ‘yes’. AI devices that can record and summarize physician–patient conversations are more than administration tools because their outputs are used in clinical decision-making. In some cases, these or other tools are assisting physicians with disease diagnosis. They need to be tested comprehensively and transparently before deployment.
In a Comment article published in Nature Medicine in September, researchers make the case that pre-registered clinical trials should become standard for AI systems used in health and medicine, as is the case for drugs and vaccines more broadly6.
Benchmarking problems
Currently, some AI-powered medical products are being certified without being assessed in real-world settings. This is not good practice because these products are new and few data sources exist that companies can draw on to assess what works and what doesn’t. In the case of a revised stethoscope model or an innovative type of bandage, there are already decades of data, including from real-world settings, that manufacturers can benchmark these products against.
By contrast, AI is powering devices that did not exist before and for which clinical data are scarce or non-existent. A systematic review published in Nature Medicine in March found that only 23% of some 4,600 papers about AI tools for clinical medicine used real-world patient data, and just 19 of the studies were prospective randomized trials1.

Synthetic data can benefit medical research — but risks must be recognized
Another problem with the current practice for AI models is that companies often test them in one-off simulated scenarios in which the accuracy of the model’s judgement is measured against the judgement of human physicians in the same scenario. Moreover, studies that compared specialized medical AI models with general-purpose ones found conflicting results on whether specialized systems had improved accuracy7,8. Ultimately, even accurate AI models do not reflect patients’ perspectives. For example, the companies developing them do not usually ask people whether they agree that these technologies should be part of their care.
Interestingly, there’s one application of AI that has clear implications for safety, health and well-being for which no one argues about the need for real-world data. National and city governments are taking a cautious approach to driver-less cars — and are following principles that are not unlike those used for the regulation of drugs. Such cars have been tested for thousands of hours under lab conditions, in simulated scenarios and under supervision on streets. These assessments are necessary because the consequences of any errors will be horrendous.
Not every AI-powered medical device needs to be tested in full-scale randomized controlled trials — but the principles behind such assessments must still apply. Data transparency, open data and testing in real-world settings cannot be negotiable.