The Food and Drug Administration has opened a public comment process on possible regulatory approaches for medical devices that use generative artificial intelligence, a rapidly developing category whose outputs may be less predictable than those of conventional medical software. The agency’s discussion paper, announced Aug. 18, asks manufacturers, clinicians, patients and other stakeholders how the FDA should assess risks, review evidence before marketing and monitor these products after they enter use.
The action does not establish a rule, binding guidance or a new pathway for clearing devices. Instead, the FDA’s Digital Health Center of Excellence, within the Center for Devices and Radiological Health, is testing ideas that could eventually inform policy. Comments are due Oct. 19 through docket FDA-2026-N-7874, according to the FDA announcement.
That early-stage status is important for patients and companies alike. The paper is not evidence that a particular generative-AI device has been found safe, effective or unsafe, nor does it impose immediate new evidence requirements. It is a request for input on whether longstanding device-regulation tools need adaptation for software that can generate new text, images or other outputs in response to prompts and clinical context.
Why generative AI presents a different oversight problem
Many medical devices already use AI, including software that helps interpret images or flag patterns for clinicians. But generative-AI systems can produce variable, open-ended responses rather than a fixed result for a narrowly defined input. The FDA says these systems may offer benefits for patient care while also raising risks that differ from those posed by traditional software and earlier AI-enabled devices.
The practical concern is not simply that an algorithm is complex. A device may depend on a third-party foundation model, may be updated over time, or may operate as an agentic system capable of pursuing multistep tasks. Those features can complicate who is responsible for validating performance, how a manufacturer controls changes and how users detect failures once the product is deployed. They are regulatory questions, not proof that any individual product will cause harm.
The agency is considering a two-axis approach to risk. As described in coverage by the Regulatory Affairs Professionals Society, one axis would consider the activity the software performs, ranging from informational functions to fully autonomous action. The other would consider the seriousness of the possible consequences if the output is wrong. A tool that supplies background information for a clinician could warrant different controls than one that can independently influence a treatment decision, although the FDA has not adopted a final classification system.
From benchmark testing to real-world surveillance
A central proposal in the discussion paper is a potential premarket model the FDA calls competency assessment. It would pair nonclinical benchmarking with clinical confirmation of performance before a device reaches patients. In broad terms, benchmark testing could probe how a system performs against defined tasks and challenging scenarios; clinical confirmation would examine whether performance holds in the intended clinical setting and population.
This approach reflects a limitation of evaluating generative systems only with the methods used for software designed around bounded inputs and fixed outputs. A benchmark can be rigorous, but it cannot perfectly reproduce the range of prompts, workflows and patient circumstances a product may encounter. Conversely, clinical evaluation may be difficult to generalize if a model changes, is used by a different health system or interacts with tools not included in the original assessment.
The FDA is also seeking feedback on risk-proportionate postmarket monitoring. Reporting by Healthcare Dive described the agency as considering whether more uncertainty before marketing might sometimes be acceptable if accompanied by stronger monitoring after launch. That is a question for consultation, not an FDA decision. It would require clear answers about what signals should trigger review, what data manufacturers should collect and when changes in performance demand corrective action.
Postmarket oversight is especially consequential for systems whose real-world performance can be affected by shifts in clinical practice, user behavior, data quality or dependencies on outside models. Monitoring can identify problems that smaller premarket evaluations miss, but it cannot substitute automatically for adequate premarket evidence. The discussion paper asks how those tradeoffs should be handled according to the device’s function and potential consequences.
A framework still taking shape
The August paper follows a November 2024 discussion by the FDA’s Digital Health Advisory Committee about total-product-lifecycle issues for generative-AI medical devices. The new docket turns that exploratory work into a more formal public solicitation, covering risk assessment, premarket evaluation, postmarket monitoring, transparency and the use of foundation models and agentic AI.
The timing also reflects the scale of AI already reaching the U.S. medical-device market. The FDA’s device center has authorized more than 1,000 AI-enabled devices, though most do not incorporate generative AI, Healthcare Dive reported. That figure provides context rather than a count of the products addressed by the new paper: the FDA announcement does not say how many generative-AI-enabled devices are currently authorized.
For manufacturers, the consultation may offer an early opportunity to address questions about testing standards, documentation and responsibility when a product relies on outside model providers. A legal analysis from Katten similarly noted that the paper is distinct from draft or final guidance. Any eventual requirements would require further FDA action and could look different after the agency receives comments.
For clinicians and patients, the immediate practical change is limited. No new FDA standard takes effect with the discussion paper. But the issues under review—how to establish competence before use and how to watch for performance problems afterward—could shape the evidence expected from a class of tools likely to become more visible in clinical workflows. The public-comment deadline is Oct. 19.
