Independent Validation

AI scribe quality varies widely. Hospitals need a trusted way to measure it.

The Clinical Documentation Quality Initiative provides independent, physician-led validation of AI-generated clinical documentation across realistic clinical scenarios, measuring accuracy, completeness, hallucinations, clinical risk, and physician editing burden.

Validated AI-Scribe benchmark

AccuracyCompletenessHallucination / FabricationClinical Reasoning FidelityMedication & Allergy SafetyFollow-Up & Safety-NettingAttribution & Speaker AccuracyStructure & ReadabilityOverall Clinical Usability

Benchmark Report

Vendor Assessment

Validated

88%

Overall Clinical Usability

3%

Hallucination / Fabrication

92%

Medication & Allergy Safety

Rubric Dimension Scores

Accuracy90%
Completeness86%
Clinical Reasoning Fidelity87%
Follow-Up & Safety-Netting83%
Attribution & Speaker Accuracy89%
Structure & Readability94%

AI-generated notes are not all created equal.

Ambient scribe tools can reduce documentation burden, but health systems need more than vendor claims. Differences in transcription accuracy, medical reasoning, note structure, hallucination control, and clinically significant omissions can directly affect physician trust, patient safety, and procurement decisions.

Doctor reviewing clinical documentation on computer

Variable note quality across vendors

AI scribe tools vary significantly in transcription accuracy, medical reasoning, and note structure.

Risk of hallucinated clinical details

Without independent validation, unsupported or fabricated clinical information may go undetected.

Lack of specialty-specific validation

Generic testing fails to capture the nuances required for different medical specialties.

Independent benchmarking for clinical documentation AI.

We evaluate AI-generated clinical notes against physician-created reference standards using standardized encounters across multiple specialties and complexity levels.

01

Standardized clinical scenarios

Physician-designed encounters covering multiple specialties and complexity levels.

02

AI-generated note submission

Vendors submit documentation outputs for each standardized scenario.

03

Independent physician review

Board-certified physicians evaluate outputs against reference standards.

04

Benchmark report and validation

Comprehensive scorecard with detailed metrics and specialty breakdowns.

What the benchmark measures

Our validated AI-Scribe rubric evaluates every note across nine critical dimensions of clinical documentation quality.

Accuracy

Does the note accurately reflect the encounter and source material?

Completeness

Does the note include the clinically relevant history, examination, assessment, plan, follow-up, and safety-netting present in the source material?

Hallucination / Fabrication

Does the note contain any diagnosis, finding, medication, result, or plan not supported by the source material?

Clinical Reasoning Fidelity

Is the assessment and plan consistent with the actual clinical information from the encounter?

Medication & Allergy Safety

Are medications, doses, allergies, anticoagulation status, contraindications, and instructions documented accurately?

Follow-Up & Safety-Netting

Are follow-up plans, investigations, referrals, return precautions, and urgent warnings accurately documented?

Attribution & Speaker Accuracy

Are symptoms, statements, decisions, and recommendations correctly attributed to the patient, clinician, or caregiver?

Structure & Readability

Is the note organized, concise, clinically readable, and suitable for physician review?

Overall Clinical Usability

Is the note clinically usable after routine physician review and editing?

Tested against realistic clinical complexity.

The benchmark uses physician-designed encounters that include real-world documentation challenges: complex medication histories, changing timelines, comorbidities, ambiguous symptoms, incidental findings, patient concerns, and specialty-specific assessment and plan requirements.

Family MedicineInternal MedicinePsychiatryNeurologyCardiologyOrthopaedicsEmergency MedicinePediatrics
scenario: Internal Medicine
complexity: High
comorbidities: [
"Type 2 Diabetes",
"Hypertension",
"CKD Stage 3"
]
medications: 12
timeline_changes: true

Built for procurement, governance, and clinical trust.

Hospitals and health authorities need objective evidence before deploying AI scribes at scale. Our benchmark helps organizations compare vendors, understand clinical risk, support procurement decisions, and establish ongoing quality assurance.

Vendor comparison

Compare AI scribe vendors using standardized metrics and objective performance data.

Procurement support

Evidence-based documentation for RFPs, vendor selection, and clinical governance.

Ongoing quality monitoring

Track documentation quality over time with periodic re-validation assessments.

Physician-led. Structured. Transparent.

Each benchmark review is based on standardized scoring rubrics, independent clinical reviewers, reference-standard documentation, and structured assessment of hallucinations, omissions, and required physician edits.

Clinical Scenario
AI Output
Physician Review
Scoring
Validation Report
20+

Standardized clinical scenarios

8

Medical specialties covered

100%

Physician-validated assessments

Ongoing Research Study

Development and Pilot Validation of the SCRIBE-AI Rubric for Evaluating AI-Generated Medical Documentation: A Modified Delphi Study

The SCRIBE-AI Rubric is currently being developed and evaluated through an ongoing modified Delphi research study. The study aims to establish expert consensus on the key characteristics of safe, accurate, complete, and clinically useful AI-generated medical documentation.

Evaluate AI documentation quality before deployment.

For health systems, procurement teams, and AI scribe vendors seeking independent evidence of clinical documentation performance.

Contact us:contact@cdqi.ca