Abstract
The rapid deployment of large language model (LLM) applications creates new software quality assurance challenges that extend beyond conventional model benchmarking. Medical GPT (generative pre-trained transformer) applications, deployed through web-based AI marketplaces, provide clinical guidance but must simultaneously satisfy requirements for factual reliability, policy compliance, privacy governance, and safe software configurations.
We present MedGPT-Audit, the first unified software quality assurance framework for auditing deployed medical GPT applications (MedGPTs) at ecosystem scale. The framework integrates automated application discovery, interaction-based factual accuracy (or hallucination) assessment, policy compliance analysis, privacy governance evaluation, and user trust analysis into a continuous auditing pipeline.
Using MedGPT-Audit, we analyze 6,233 MedGPTs from the OpenAI GPT Store, evaluate a stratified sample of 1,500 deployed applications, and compare them against 10 representative open-source medical LLMs. Our evaluation reveals that 25--30% of MedGPTs exhibit low factual reliability, 33.6--54.3% exceed the policy-compliance threshold, and 57.06% of Actions-enabled MedGPTs lack adequate privacy disclosures. MedGPT-Audit completed the audit in 73 hours, reducing estimated clinician review effort by 8.6--17.1x while preserving strong agreement with expert judgment. Furthermore, marketplace popularity and user feedback are poor indicators of software quality and clinical reliability.
These findings demonstrate that trustworthy deployment of LLM applications requires continuous software quality assurance throughout the application lifecycle rather than one-time model benchmarking. We release the MedGPT-Audit code and dataset to facilitate reproducible research on software quality assurance, governance, and auditing of AI application ecosystems.
We present MedGPT-Audit, the first unified software quality assurance framework for auditing deployed medical GPT applications (MedGPTs) at ecosystem scale. The framework integrates automated application discovery, interaction-based factual accuracy (or hallucination) assessment, policy compliance analysis, privacy governance evaluation, and user trust analysis into a continuous auditing pipeline.
Using MedGPT-Audit, we analyze 6,233 MedGPTs from the OpenAI GPT Store, evaluate a stratified sample of 1,500 deployed applications, and compare them against 10 representative open-source medical LLMs. Our evaluation reveals that 25--30% of MedGPTs exhibit low factual reliability, 33.6--54.3% exceed the policy-compliance threshold, and 57.06% of Actions-enabled MedGPTs lack adequate privacy disclosures. MedGPT-Audit completed the audit in 73 hours, reducing estimated clinician review effort by 8.6--17.1x while preserving strong agreement with expert judgment. Furthermore, marketplace popularity and user feedback are poor indicators of software quality and clinical reliability.
These findings demonstrate that trustworthy deployment of LLM applications requires continuous software quality assurance throughout the application lifecycle rather than one-time model benchmarking. We release the MedGPT-Audit code and dataset to facilitate reproducible research on software quality assurance, governance, and auditing of AI application ecosystems.
| Original language | English |
|---|---|
| Number of pages | 12 |
| Publication status | Submitted - 1 Jul 2026 |
| Event | The International Conference on Software Engineering - The Convention Centre Dublin (CCD), Dublin, Ireland Duration: 25 Apr 2027 → 1 May 2027 Conference number: 49 https://conf.researchr.org/track/icse-2027/icse-2027-research-track |
Conference
| Conference | The International Conference on Software Engineering |
|---|---|
| Abbreviated title | ICSE |
| Country/Territory | Ireland |
| City | Dublin |
| Period | 25/04/27 → 1/05/27 |
| Internet address |
Fingerprint
Dive into the research topics of 'MEDGPT-AUDIT: a software quality assurance framework for trustworthy medical GPT applications'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver