Despite the success of AI chatbots in medical professional exams, they still face challenges in one of the most critical tasks of physicians, which is diagnosing diseases through conversations with patients. New research shows that the accuracy of these models significantly decreases when interacting with simulated patients. According to a study conducted by researchers at Harvard University, AI models like GPT-4 from OpenAI performed notably well in multiple-choice medical exams with an accuracy of 82%, but their accuracy dropped to 26% when diagnosing diseases through conversations with simulated patients. Researchers used 2,000 medical cases, mostly extracted from the American Medical Board exams, to evaluate these models. In this process, the GPT-4 model played the role of the simulated patient and conversed with other AI models acting as doctors. The results of these conversations were also reviewed by medical experts. The findings showed that AI models were not only incapable of fully gathering the patient's medical information but also could not always provide correct diagnoses even when complete information was received. For instance, the GPT-4 model succeeded in gathering complete information in only 71% of conversations. Pranav Rajpurkar, the senior researcher of this study, stated, 'Real-world medical practice is much more complex and involves factors like managing multiple patients, coordinating with treatment teams, and understanding social and systemic factors.' According to him, AI can be an effective auxiliary tool in medicine but will not replace the comprehensive judgment of physicians.
Artificial Intelligence Fails in Disease Diagnosis Through Patient Conversations
A study from Harvard reveals that AI chatbots, despite high performance in medical exams, struggle significantly with disease diagnosis through patient conversations, showing only 26% accuracy. This highlights the limitations of AI in real-world medical practice, emphasizing the need for human judgment.
👥 Key Players
📰 What Happened
A study from Harvard revealed that AI chatbots, while performing well in medical exams, struggled with disease diagnosis through patient conversations, achieving only 26% accuracy. This highlights significant limitations of AI in real-world medical practice.
- GPT-4 achieved 82% accuracy in multiple-choice medical exams.
- The AI model gathered complete patient information in only 71% of conversations.
💡 Why It Matters
📚 Background
AI technology is rapidly evolving, with applications in various fields including healthcare. Understanding its limitations is essential for effective use.
🏷️ Entities Mentioned
Translated from the original and edited for English readers. View original source →
Translation confidence: 85%