AI-powered symptom checkers direct worried patients to appropriate care 

Karen Blum

Share:

person using an AI-powered symptom checker on their smartphone

Photo by Shane via Unsplash

Worrying abdominal pain, headaches or other health symptoms frequently happen at inconvenient times, like on a weekend or after your doctor’s office has closed for the day. 

To accommodate patients with concerns like this, some medical centers have begun offering their own AI-powered symptom checkers — programs that ask patients questions about their symptoms and then recommend appropriate care, including showing available time slots to book physician appointments. It’s an interesting trend for journalists to cover. 

Some 32% of U.S. adults have used AI chatbots for health information in the past year, according to a recent KFF survey, and most users (65%) said a desire for quick, immediate advice was a major reason. Three in 10 (29%) say they used AI tools for information or advice about their physical health, and one in six (16%) used them for mental health information.

But the accuracy generated by general chatbots like ChatGPT Health has been variable, according to the medical literature. AI chatbots failed to accurately generate a list of possible diagnoses based on initial patient symptoms more than 80% of the time, according to an April study by researchers at Mass General Brigham Hospital in Boston, reported by Becker’s Health IT. Another study from researchers at Mount Sinai Health System in New York City found in February that ChatGPT Health undertriaged medical emergencies some 52% of the time. 

Home field advantage 

The advantage to using AI programs housed by the medical centers (either online or through a mobile app) is that they are more likely to be clinical grade, said Bilal Naved, Ph.D., an instructor and AI fellow at Mount Sinai who helped develop a commercially available symptom checker that Mount Sinai has incorporated. 

Mount Sinai’s program, called Check Symptoms and Get Care, was built in partnership with Barton Schmitt, M.D., co-author of the Schmitt-Thompson telephone triage protocols used by nurses nationwide. “Patients using it can trust that it will route them according to clinical best practice,” Naved told AHCJ. “It won’t hallucinate. It’ll be consistent. It’ll be safe.”

Other academic centers incorporating programs like this include Tufts Medicine in Boston, Baylor Scott and White Health in Dallas, and Banner Health in Tucson, Ariz. 

How do symptom checkers work?

Naved explained it this way: Let’s say you have a fever and sore throat and don’t have care established in your local area. You might do a general internet search of your symptoms and see an ad from Mount Sinai or your local medical center inviting you to try the symptom checker. 

If you click there, it pulls you to the landing page for the AI program. You enter symptoms and answer questions to help qualify what level of care is appropriate. On average, the program asks 10-14 questions. If the program senses an emergency, it will ask fewer questions, notify you it’s an emergency, and connect you to emergency services. 

If the program determines that a primary or specialty care visit is most ideal and during office hours, it will allow you to input your insurance and demographic information. Then, like the OpenTable system for restaurant reservations, it will display time slots you can book with one of the health system’s physicians. And, if the program determines it’s something you can manage at home without physician care, it will tell you so, and connect you to educational information.

For established patients who enter the program while logged into the health system’s patient portal, the program works similarly, except that it will auto-populate your insurance and demographic information using information from your patient records, and display available time slots with your physician first. It also lets you scroll and view time slots of other providers in case someone else is available sooner. 

How do they perform?

From over 60,000 visits to Mount Sinai in mid-2023 to mid-2024, approximately 22,000 patients completed digital triage sessions; 80% of those who started using the tool completed their conversation, according to information published by Naved and colleagues in NEJM Catalyst. The program achieved high satisfaction rates (75% of users rated it at least an 8 out of 10). 

The most common complaints were abdominal pain (2,009 cases), weakness/fatigue (952 cases), cough (832 cases), sore or scratchy throat (762 cases) and leg or foot swelling (727 cases). The average user was 36 years old but ranged from parents inquiring on behalf of a newborn to those 100 years old. The program randomly asked users about their plans for care before viewing the results and found 26% were unsure. Some 62% of patients who thought they knew what to do had an intent that was clinically suboptimal, meaning they weren’t seeking a high enough level of care. On the flip side, many other patients were advised to seek less care. 

Among users, 36% were referred to primary care, 27% to telemedicine, 18% to the ER, 7% to urgent care, 4% each to an ambulance or to specialist care, 2% to home care, and 1% each to behavioral health, education or asynchronous care. About 71% used the tool outside of typical business hours, and zero reported safety incidents. Overall, the tool was 88-96% in agreement with recommendations from a group of clinicians given the same information. 

The program has continued to be used by about 2,000 to 3,000 people each month, Naved said. 

What the future could look like

As these tools are more frequently incorporated, look for additional features to emerge. For example, Naved and his colleagues are working on a 

“show your work” type program that would demonstrate to patients in real time how the AI is processing their responses.

“Our biggest source of constructive feedback is coming from patients who get a primary care provider recommendation,” Naved said. “We see feedback like, ‘My mom could have told me I needed to go see a doctor – this isn’t intelligent.’ But what they don’t know is that [the program] actually ruled out emergency care, urgent care, 911, specialty care, and had a lot of intelligence into deciding that a primary care provider is right.”

The new feature would show patients, response by response, how their illness acuity level might go up or down, like on a gauge, so they can see the intelligence at work, he said.

Another feature could incorporate a confirmation-type checker to query patients after they have scheduled an appointment but before they go to an office, to ensure they were referred appropriately, he said.

Story ideas

  • Talk to patients who have used these programs. Why did they use them? What did they think of the responses? Did their condition resolve, using the information suggested by the program?
  • Talk to physicians who have seen patients booking appointments through these programs. Were patients referred appropriately? How is this impacting their practices?
  • Talk to program developers – how are they training the AI? For example, researchers at the University of California San Diego published information about a chatbot they designed that incorporates information from 100 step-by-step medical flowcharts developed by the American Medical Association. 
  • Ask about future features these tools could incorporate. 

Resources

Karen Blum

Karen Blum

Karen Blum is AHCJ’s health beat leader for AI and Patient Safety. She’s a health and science journalist based in the Baltimore area and has written health IT stories for numerous trade publications.