Skip to main navigation Skip to search Skip to main content

MEDGPT-AUDIT: a software quality assurance framework for trustworthy medical GPT applications

Research output: Contribution to conferencePaperpeer-review

Abstract

The rapid deployment of large language model (LLM) applications creates new software quality assurance challenges that extend beyond conventional model benchmarking. Medical GPT (generative pre-trained transformer) applications, deployed through web-based AI marketplaces, provide clinical guidance but must simultaneously satisfy requirements for factual reliability, policy compliance, privacy governance, and safe software configurations.

We present MedGPT-Audit, the first unified software quality assurance framework for auditing deployed medical GPT applications (MedGPTs) at ecosystem scale. The framework integrates automated application discovery, interaction-based factual accuracy (or hallucination) assessment, policy compliance analysis, privacy governance evaluation, and user trust analysis into a continuous auditing pipeline.

Using MedGPT-Audit, we analyze 6,233 MedGPTs from the OpenAI GPT Store, evaluate a stratified sample of 1,500 deployed applications, and compare them against 10 representative open-source medical LLMs. Our evaluation reveals that 25--30% of MedGPTs exhibit low factual reliability, 33.6--54.3% exceed the policy-compliance threshold, and 57.06% of Actions-enabled MedGPTs lack adequate privacy disclosures. MedGPT-Audit completed the audit in 73 hours, reducing estimated clinician review effort by 8.6--17.1x while preserving strong agreement with expert judgment. Furthermore, marketplace popularity and user feedback are poor indicators of software quality and clinical reliability.

These findings demonstrate that trustworthy deployment of LLM applications requires continuous software quality assurance throughout the application lifecycle rather than one-time model benchmarking. We release the MedGPT-Audit code and dataset to facilitate reproducible research on software quality assurance, governance, and auditing of AI application ecosystems.
Original languageEnglish
Number of pages12
Publication statusSubmitted - 1 Jul 2026
EventThe International Conference on Software Engineering - The Convention Centre Dublin (CCD), Dublin, Ireland
Duration: 25 Apr 20271 May 2027
Conference number: 49
https://conf.researchr.org/track/icse-2027/icse-2027-research-track

Conference

ConferenceThe International Conference on Software Engineering
Abbreviated titleICSE
Country/TerritoryIreland
CityDublin
Period25/04/271/05/27
Internet address

Fingerprint

Dive into the research topics of 'MEDGPT-AUDIT: a software quality assurance framework for trustworthy medical GPT applications'. Together they form a unique fingerprint.

Cite this