Key Takeaways
AI works best on a narrow task with good input. Diagnosis still needs examination, history and someone who is responsible for the decision.
- The statement that ChatGPT handles 80% of what a good doctor does with bloodwork was editorial shorthand, not a validated estimate. Performance varies by model, task, prompt, input quality, population, and clinical context. [1]
- Accuracy and AUC values from ECGs, mammograms, cytology, skin images, medical exams, and diagnostic cases cannot be placed in one league table. They use different diseases, populations, outcomes, thresholds, and study designs. Each result is meaningful only within its own validation setting. [2]
- The skin model's 26 categories represented 80% of skin cases in the source setting, not 80% of all primary-care visits. Its 0.66 versus 0.44 top-1 comparison used 963 teledermatology cases and six primary-care physician readers. [2]
- In the 50-physician trial, clinicians using GPT-4 scored 76% versus 74% with conventional resources, a difference that was not statistically significant. The model alone scored 16 percentage points above the conventional-resources group under the study rubric. The article's 92% figure is not the reported model-alone result. [1]
- Therabot was tested for four weeks in 210 adults with selected symptom profiles against a waitlist. The trial found a larger mean symptom reduction, but it did not establish a general 51% treatment effect or equivalence to clinician-delivered therapy. [3]
- A consumer AI tool can help organize questions, but it cannot safely replace diagnosis, examination, longitudinal context, or clinician review. Health records are sensitive. Product, privacy settings, connected apps, eligibility, jurisdiction, and data controls should be checked before any health information is shared. [4]
Common questions
Can AI answer 80% of primary-care questions safely?
No study has shown that ChatGPT safely handles 80% of primary care. That number was editorial shorthand. Performance changes with the model, prompt, task, population and the quality of the information supplied. [1]
Is ChatGPT as accurate as a doctor for medical diagnosis?
There is no honest single score for AI versus doctors. An ECG model, a skin-image classifier and a medical exam test different jobs. Compare a model only with clinicians doing the same task in a similar population. [2]
How much of primary care can actually be handled by an AI chatbot?
The often-cited 80% number came from a skin model whose 26 categories covered 80% of cases in 1 source setting. It did not cover 80% of all primary-care visits. The physician comparison used 963 teledermatology cases and 6 readers. [2]
Does using AI improve a doctor's diagnostic accuracy?
In a trial of 50 physicians, the GPT-4 group scored 76%. The conventional-resources group scored 74%. That gap was not statistically significant. The model-alone result was better under the study rubric, but the article's 92% figure was wrong. [1]
Is it safe to use a chatbot for mental-health advice?
A 4-week trial of Therabot included 210 adults with selected symptoms and used a waitlist control. Symptoms fell more in the chatbot group. The trial did not establish a 51% treatment effect for everyone or show that a chatbot matches therapy with a clinician. [3]
Which health questions can I ask AI, and when do I need a clinician?
AI can organize symptoms, prepare questions and translate medical terms. A clinician is needed for diagnosis, examination, urgent symptoms, treatment changes and results that depend on your history. The product's privacy settings need checking before any health record is shared. [4]
References
- Goh, E., Gallo, R., Hom, J., et al. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open, 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969
- Liu, Y., Jain, A., Eng, C., et al. (2020). A deep learning system for differential diagnosis of skin diseases. Nature Medicine, 26(6), 900–908. https://doi.org/10.1038/s41591-020-0842-3
- Heinz, M. V., Mackin, D. M., Trudeau, B. M., et al. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI, 2(4). https://doi.org/10.1056/AIoa2400802
- OpenAI. (n.d.). Health privacy policy. https://openai.com/policies/health-privacy-policy/
Related reading
For informational and educational purposes only. This page does not provide individual medical advice, diagnosis, or treatment. Read the full medical disclaimer.
