Sponsored by

A review of six external-validation studies published from 2022 through 2025 found a median decline in AUC of about 0.03 when radiology AI models were evaluated outside their development setting. The largest concern in the review was specificity: declines reached as much as roughly 24 percentage points. Sensitivity remained above 85% in most of the external validations reviewed.

These figures are not a universal estimate for all radiology AI. They describe the studies in that review. Still, they point to a deployment problem that polished accuracy claims can conceal.

The best voice models, now across all channels

Most CX platforms do not own the voice. They orchestrate a workflow, then call a third party for speech and transcription. Every hop adds latency, cost, and another vendor to manage.

ElevenAgents is the opposite. They make the voice models the market builds on, and ElevenAgents puts full orchestration on top. Voice, transcription, text-based chat, and reasoning run in one vertically integrated pipeline, so responses come back in <400 milliseconds and sound human, not synthetic.

Plus, you keep full control. Plug in any LLM, integrate tools, webhooks, and MCP servers, and ground responses in your knowledge base. Get an agent live in minutes, then A/B test with Experiments, enforce Guardrails, and version every change.

The payoff: more human conversations, lower latency, and far less time stitching infrastructure together. You build on the models you already trust. Pricing is transparent and flat at $0.08 per minute.

A model is developed under particular conditions. Its images come from certain scanners, protocols, institutions, and populations. Its labels are created with particular reference standards. When those conditions change, performance can change with them. The FDA’s guidance for computer-assisted detection devices recognizes the importance of these details by recommending that submissions describe datasets, reference standards, target populations, imaging protocols, compatibility, and independent testing.

Specificity deserves attention because it concerns how often a system correctly identifies cases without the target finding. A decline in specificity can create more false positives. In radiology, an extra flag is not merely an abstract error category. Someone must decide whether the signal represents a finding, an artifact, an anatomic variant, postoperative change, or the presence of a medical device.

An AJR review catalogs precisely these interpretive pitfalls. It also points to gaps in clinical context, limited explainability, and communication problems between developers and radiologists. The risk is not that AI is uniquely prone to uncertainty; medical imaging already requires interpretation under uncertainty. The risk is that an output can appear more definitive than the conditions behind it justify.

This is why local evaluation matters, even where regulatory status exists. The FDA maintains a list of AI-enabled medical devices that includes radiology entries, and its guidance sets out recommended evidence for a subset of computer-assisted detection submissions. But a listing or clearance does not establish that every product will preserve its performance across all hospitals and imaging environments.

The current evidence does contain reasons for measured optimism. Some integrated workflows have reported faster reporting and turnaround times. Those operational improvements may be valuable, especially where they fit existing systems. Yet speed does not settle whether a model is appropriately calibrated for a particular population or whether an increase in false positives offsets part of the gain.

The practical standard should be more demanding than a headline accuracy number. A radiology AI tool should be examined in the setting where it will be used: with its local scanners, protocols, clinicians, and patients. The available evidence makes the case for that scrutiny; it does not tell us which individual products have met it.

NEON DREAMS AI

AI • TECHNOLOGY • CULTURE

Independent analysis of artificial intelligence, emerging technology, and the consequences in between.

Source:

Thank you for reading,

- Neon

Reply

Avatar

or to participate