AI & Tech

AI System Faster but Not More Accurate Than Docs for Diagnosing Rheum Disease

[post_content]


Disclaimer: This article has been automatically aggregated from

Prof. Valmed, a large language model (LLM) cleared by European regulators for medical diagnosis, outperformed human physicians for diagnosing rheumatologic conditions in terms of speed but not accuracy, a randomized trial showed.

Diagnostic accuracy was almost exactly the same in rheumatology case scenarios for physicians who used Prof. Valmed versus those reaching their diagnoses in their usual ways, according to Johannes Knitza, MD, PhD, MHBA, of Philipps-Universität Marburg in Germany, and colleagues.

But the human doctors relying on their own resources required an average of 206 seconds to make their diagnoses, compared with 94 seconds among those aided by Prof. Valmed (P<0.001), the group reported in a medRxiv preprint manuscript, which has not undergone peer review.

As well, use of Prof. Valmed increased physicians’ confidence in their diagnoses — perhaps too much, Knitza and colleagues suggested. “An interesting finding was that confidence exceeded observed accuracy across all evaluated settings and increased further after [Prof. Valmed] assistance in both groups, indicating persistent overconfidence,” the investigators wrote.

“Exploratory analyses furthermore suggested substantial AI [artificial intelligence] over-reliance in the intervention group, whereas under-reliance was uncommon,” they continued. “These findings suggest that diagnostic support may increase confidence more readily than correctness and highlight the need to evaluate calibration and behavioural reliance alongside accuracy.”

AI system developers have long touted medicine as one of the most important potential applications for the emerging technology: disease diagnosis and management rely so strongly on synthesizing numerous, disparate forms of information that sufficiently capable machines ought to be able to do it better and faster than humans.

Whether the field has reached that point, however, remains uncertain. General AI systems have been tried with mixed results — sometimes finding diagnoses that had eluded human professionals, but also prone to “hallucinations,” decidedly false conclusions that appear to stem from the systems’ emphasis on pleasing their users.

Prof. Valmed, founded by Vera Roedel, an attorney for Merck KGaA, and Heinz Wiendl, MD, a neuroimmunologist at University Hospital Freiburg, is designed specifically for medical applications with guardrails to limit hallucinations. It’s billed as an “AI copilot” to assist but not replace physicians and received the European Union’s CE mark in March 2025, allowing it to be sold for medical use. As an LLM, users can direct it through ordinary language.

For the trial, dubbed ALLIANCE, Knitza and colleagues sought to evaluate its performance in the rheumatology field. They recruited 82 physicians from seven institutions in Germany and Norway, randomizing them 1:1 to make diagnoses in three clinical scenarios either with or without Prof. Valmed’s support. These scenarios involved Cogan syndrome, dermatomyositis, and familial Mediterranean fever, all drawn from published case reports. Participants were also told to quantify their confidence in each potential diagnosis.

For example, in the Cogan syndrome case, participants were asked to list up to three possible diagnoses, with probabilities, for a patient described as follows: “61-year-old man. Presented with tinnitus, progressive hearing loss, generalized joint pain, blurred vision, and redness in both eyes.”

Participants were also asked to rate their satisfaction with the methods they used. For those assigned to the intervention group, the questions included several specific to Prof. Valmed about its ease of use and trustworthiness, and their interest in using it again.

Only about one-quarter of participants were rheumatology specialists. Many disciplines were represented, including general internal medicine, nephrology, endocrinology, and eight others.

The primary outcome was accuracy, defined as the percentage of instances in which a participant’s most likely diagnosis matched the actual published one. This was achieved in 33.3% of intervention group cases versus 35.0% of those approached conventionally, a difference that did not come close to statistical significance. Prof. Valmed’s performance looked somewhat better with a less stringent outcome — having one of a participant’s three possible diagnosis match the actual one — with the system’s use leading to a 49.2% success rate, compared with 39.2% in the control group, but the effect was not significant (P=0.291).

Knitza and colleagues also examined Prof. Valmed’s “standalone” performance, without the human physician’s input. Working by itself, the system was also 33.3% accurate for its top likelihood matching the true diagnosis, and 47.6% accurate in having one of its three candidates be correct.

Confidence ratings vastly exceeded accuracy at every step. In the intervention group, even before Prof. Valmed was called in, the mean confidence rating was 39%, compared with accuracy of 22%. Their 33% accuracy when using the AI system came with confidence ratings averaging 57%. The same pattern was seen in the control group.

Overall, participants liked Prof. Valmed, giving it high ratings for ease of use and a pleasing interface. About two-thirds said it was trustworthy, and over 80% said they would use it again. On the other hand, only 36% said it was easy to fix mistakes when using the system, suggesting the developers still have some work to do.

“LLM-based diagnostic support may be most valuable for improving efficiency, enhancing perceived support quality, and broadening differential diagnosis, while overconfidence and overreliance remain key safety concerns,” Knitza and colleagues concluded. “Further real-world trials are needed to define the role of certified LLM-based decision support in routine care.”

for informational purposes only. We do not claim ownership, accuracy, or liability for the content provided. All rights belong to the original publisher.