Doximity Tops Independent Stanford-Harvard Study on Clinical AI Safety

Doximity Ask Tops Independent Stanford-Harvard Clinical AI Safety Benchmark, Outperforming Frontier AI Models

Doximity, a leading digital platform serving medical professionals across the United States, has announced that its Doximity Ask clinical artificial intelligence platform achieved the highest ranking in an independent evaluation of clinical AI safety conducted by researchers affiliated with Stanford University and Harvard Medical School. The results position the company’s HIPAA-compliant AI assistant ahead of several leading frontier artificial intelligence models in one of the most comprehensive assessments of AI performance in clinical settings to date.

The findings come from the NOHARM (Numerous Options Harm Assessment for Risk in Medicine) benchmark, an independent framework designed to evaluate how safely artificial intelligence systems respond to realistic medical scenarios. The study was conducted by ARISE, a clinical AI research team led by physicians from Stanford and Harvard Medical Schools, with the objective of assessing the reliability, safety, and clinical appropriateness of AI-generated responses when presented with simulated patient cases.

According to the evaluation, Doximity Ask achieved the highest overall performance in the benchmark’s real-world clinical sample, a portion of the assessment specifically designed to reflect how physicians use AI tools during everyday clinical practice. The study also found that AI systems purpose-built for healthcare consistently outperformed general-purpose frontier AI models across the broader automated evaluation.

Independent Benchmark Evaluates Clinical AI Safety

Artificial intelligence is becoming increasingly integrated into healthcare environments, assisting physicians with documentation, clinical decision support, medical research, workflow automation, and patient communication. As adoption accelerates, ensuring that AI-generated recommendations are accurate, reliable, and safe has become a major focus for healthcare providers, researchers, and regulators.

The NOHARM benchmark was developed to address this challenge by providing an independent framework for measuring clinical AI safety under realistic conditions. Rather than evaluating AI systems using generalized knowledge tests, the benchmark examines how models respond when confronted with simulated patient scenarios that resemble situations encountered in routine medical practice.

Researchers assessed whether AI systems could generate clinically appropriate responses while minimizing the potential for harmful recommendations, unsafe guidance, or inaccurate medical information. The benchmark emphasizes patient safety and risk reduction, making it particularly relevant for healthcare organizations evaluating AI tools for clinical deployment.

ARISE Research Team Led the Independent Study

The NOHARM benchmark was developed and conducted by ARISE, a physician-led clinical AI research initiative that includes experts affiliated with Stanford University and Harvard Medical School.

The research team designed the evaluation to examine how artificial intelligence performs when presented with complex clinical questions and patient cases requiring nuanced medical reasoning. By using standardized evaluation criteria, researchers were able to compare multiple AI systems under consistent testing conditions.

The study focused on both automated benchmark performance and a real-world clinical sample intended to mirror actual physician use of AI systems during patient care.

This dual approach allowed researchers to evaluate not only theoretical model capabilities but also practical clinical usefulness in realistic healthcare settings.

Doximity Ask Ranked First in Real-World Clinical Evaluation

Among all artificial intelligence systems included in the study, Doximity Ask achieved the highest ranking in the benchmark’s real-world clinical sample.

This portion of the evaluation is considered especially significant because it reflects the types of interactions physicians are most likely to have with AI assistants during clinical workflows.

Rather than relying exclusively on abstract test questions, researchers assessed AI responses in scenarios designed to resemble authentic patient cases and physician decision-making processes.

By earning the highest ranking in this category, Doximity demonstrated strong performance in producing clinically relevant responses while maintaining a focus on patient safety.

Purpose-Built Clinical AI Outperformed General-Purpose Models

Beyond Doximity’s individual performance, the NOHARM study identified a broader trend across the healthcare AI landscape.

Researchers found that artificial intelligence systems specifically designed for clinical use consistently outperformed general-purpose frontier AI models throughout the broader automated evaluation.

Purpose-built healthcare AI systems benefit from clinical optimization, specialized medical training, physician oversight, and healthcare-focused development processes.

General-purpose AI models, while capable of answering a wide variety of questions across multiple domains, may lack the specialized safeguards and workflow integration required for complex clinical environments.

The findings suggest that healthcare-specific AI platforms may offer advantages when deployed in medical settings where patient safety and clinical accuracy are critical.

Designed Specifically for Healthcare Workflows

Doximity Ask was developed specifically to support healthcare professionals rather than general consumers.

Unlike consumer-focused AI assistants, the platform has been designed around clinical workflows and the unique requirements of healthcare organizations.

Its capabilities are intended to assist physicians and other clinicians with everyday professional tasks while supporting compliance with healthcare privacy regulations.

Because medical professionals routinely handle sensitive patient information, Doximity built Ask as a HIPAA-compliant platform capable of operating within secure clinical environments.

This healthcare-first design philosophy distinguishes the platform from broader conversational AI systems developed for general public use.

Physician Expertise at the Core of Development

According to Doximity, one of the primary reasons behind the platform’s strong performance is its emphasis on physician involvement throughout AI development.

Rather than relying solely on automated model training, the company has invested extensively in physician-authored evaluation and continuous quality improvement.

This approach reflects the belief that experienced clinicians remain essential in validating AI-generated medical information before it is deployed in healthcare settings.

By combining advanced artificial intelligence with ongoing physician oversight, the company seeks to improve both clinical accuracy and patient safety.

PeerCheck Program Strengthens Clinical Oversight

Central to Doximity’s clinical AI strategy is its PeerCheck™ program.

Through this initiative, more than 11,000 physician experts have participated in reviewing, evaluating, and improving outputs generated by Doximity Ask.

The physician reviewers examine AI responses for:

  • Clinical accuracy
  • Medical appropriateness
  • Patient safety
  • Evidence-based recommendations
  • Overall response quality

Feedback from these experts is incorporated into the ongoing refinement of the AI platform, creating a continuous improvement process supported by practicing clinicians.

The company believes this large-scale physician participation contributes significantly to the platform’s reliability and safety.

Leadership Emphasizes Physician-Guided AI

Dr. Louis-Antoine Mullie, Head of Medical AI at Doximity, said the company’s development philosophy has always centered on the belief that trustworthy healthcare AI must be built in collaboration with physicians rather than independently of them.

He explained that continuous physician review should not be viewed as a competitive advantage but as a fundamental requirement for any artificial intelligence system intended for clinical use.

According to Dr. Mullie, the results of the NOHARM benchmark reinforce the importance of combining advanced AI technologies with rigorous clinical oversight and independent safety evaluations.

He also noted that external assessments conducted by independent research organizations provide valuable validation of AI performance beyond internal company testing.

Broad Adoption Across Healthcare Systems

Doximity stated that its broader Clinical AI Suite, including Doximity Ask, has already been reviewed, approved, and implemented by more than 150 health systems throughout the United States.

These deployments span a wide range of healthcare organizations, including several of the nation’s leading academic medical centers and integrated health systems.

According to the company, the platform is currently used by eight of the country’s top 20 hospitals, demonstrating growing institutional adoption of clinical AI technologies designed specifically for healthcare professionals.

The expansion reflects increasing interest among healthcare providers seeking secure AI solutions capable of improving efficiency without compromising patient privacy or clinical quality.

Supporting Clinical Decision-Making

Artificial intelligence platforms such as Doximity Ask are intended to assist—not replace—healthcare professionals.

Clinical AI can help physicians by supporting information retrieval, summarizing medical literature, assisting documentation, streamlining administrative tasks, and facilitating clinical communication.

However, healthcare organizations generally emphasize that final medical decisions remain the responsibility of qualified clinicians.

Maintaining physician oversight throughout AI-assisted workflows remains a key principle in responsible healthcare AI deployment.

Enterprise Security and Privacy Features

Given the sensitive nature of healthcare information, Doximity has incorporated multiple enterprise-grade security measures into its Clinical AI Suite.

The platform includes:

  • HIPAA compliance
  • End-to-end encryption
  • Role-based access controls
  • Comprehensive audit logging
  • Session isolation
  • Enterprise privacy protections

These features are designed to help healthcare organizations integrate AI into clinical operations while maintaining compliance with healthcare privacy regulations and organizational security standards.

Strong security infrastructure has become increasingly important as hospitals expand digital health technologies across patient care environments.

Growing Importance of Clinical AI Safety

As generative artificial intelligence becomes more widely adopted in medicine, independent evaluations such as NOHARM are expected to play an increasingly important role in helping healthcare organizations compare available technologies.

Healthcare differs from many other AI application areas because inaccurate recommendations can directly affect patient outcomes.

Consequently, safety benchmarking, physician validation, regulatory oversight, and transparent evaluation frameworks are becoming central components of responsible AI deployment.

The NOHARM benchmark contributes to this evolving landscape by providing standardized methods for assessing clinical performance under realistic healthcare scenarios.

Positioning for Continued Healthcare AI Innovation

Doximity’s performance in the independent Stanford-Harvard evaluation underscores the growing importance of healthcare-specific artificial intelligence platforms designed with physician oversight, clinical validation, and enterprise security at their core. By combining advanced AI capabilities with large-scale physician review through its PeerCheck program, the company aims to deliver tools that align more closely with the practical needs of clinicians while prioritizing patient safety.

As hospitals and healthcare systems continue exploring the integration of artificial intelligence into clinical practice, independent benchmarks such as NOHARM are likely to become increasingly influential in evaluating the safety and effectiveness of AI solutions. Doximity’s top ranking in the study highlights the potential advantages of purpose-built clinical AI systems and reinforces the broader industry trend toward combining technological innovation with rigorous medical expertise to support high-quality patient care.

Source link: https://press.doximity.com/

Newsletter Updates

Enter your email address below and subscribe to our newsletter