The right approach to usability testing healthcare software is a risk-focused, iterative human factors program: formative tests early, summative validation before launch, and post-deployment surveillance after. This sequence aligns with FDA human factors guidance, NIST usability frameworks, and ONC Real World Testing expectations. Start this week with a small formative test on your highest-risk task.
TL;DR:
- FDA guidance requires teams to identify critical tasks, analyze use risks, and report validation results in applicable premarket submissions.
- Formative rounds use four to six representative users; validation often requires at least fifteen per user group, a stable design, and objective success criteria.
- Build interruptions and device variation into realistic clinical scenarios, then record task success, completion time, and use errors for each critical task.
- Certified health IT developers must submit annual ONC Real World Testing plans and results; production tickets, workarounds, and near misses should feed corrective action.
Table of Contents
- Why Healthcare Software Usability Is Different
- Regulatory and Standards Baseline You Must Meet
- Core Methods: Formative Versus Summative Testing
- Designing Tests: Tasks, Scenarios, and Environment
- Selecting Participants and Sample Size
- Metrics and Analysis: What to Capture and How to Interpret It
- Real World Testing and Post-Deployment Surveillance
- Practical Checklist and Sample Test Outline
- Common Pitfalls and When Advisory Expertise Helps
- How We Help With Usability Testing and Real World Testing Plans
- FAQ
- Sources
Why Healthcare Software Usability Is Different
A confusing button in a retail app costs a sale. A confusing medication screen in an EHR costs a dose; that difference changes everything about how we test.
Healthcare software carries safety-critical consequences that consumer products do not. A missed allergy alert, a wrong-patient order, or a misread dosage field can lead to patient harm, not just frustration. This reality forces teams to prioritize critical tasks, the specific actions where a use error could cause serious injury, above general ease-of-use concerns.
The user population also varies more than most product teams expect. Physicians, nurses, pharmacists, and medical assistants bring different training levels, workflows, and mental models to the same screen.
Testing has to reflect the actual workplace, not a quiet office; our User Experience Design Guide for Designers and PMs offers adaptable methods and artifacts that can help structure such realistic usability testing.
- Interruptions are constant, so tasks should be tested with realistic interruptions built in, not assumed away.
- Lighting and device variability matter, since clinicians move between dim exam rooms, bright hallways, and shared workstations.
- Stress and cognitive load are elevated, especially during codes, handoffs, or high patient volume.
- Device mix is wide, spanning desktop carts, tablets, and mobile phones, often within a single shift.
The implication is straightforward: measure use errors in realistic contexts, not just how quickly someone completes a task in a sterile lab setting.
Regulatory and Standards Baseline You Must Meet
Usability testing for healthcare software is not optional polish. It sits inside a documented regulatory expectation, and skipping that documentation creates real submission risk.
The FDA's human factors guidance directs manufacturers to identify critical tasks early, conduct a use-related risk analysis (URRA), and run human factors validation testing with results reported in premarket submissions such as 510(k), PMA, or De Novo filings. This is the backbone of any defensible test program.
Several standards and templates shape how that testing gets structured and reported:
- IEC 62304 governs the software development lifecycle and where usability activities fit within it.
- ISO 14971 provides the risk management framework that critical-task identification depends on.
- NIST's Common Industry Format (CIF), detailed in NIST IR 7804-1, offers standardized templates and EHR-specific use cases for structuring usability reports.
The ONC's Real World Testing program requires certified health IT developers to submit annual plans and results, measuring interoperability and functionality under real production conditions, according to the Real World Testing Resource Guide. This isn't a separate checkbox. RWT measures can and should overlap with your usability evidence.
Together, these requirements change how you sample participants, script scenarios, and archive documentation. Expect to retain URRA records, test protocols, participant demographics, and raw observational data well beyond your launch date.
Core Methods: Formative Versus Summative Testing
Healthcare usability testing splits into two distinct phases, and conflating them is one of the most common mistakes product teams make.
- Run formative testing early and often. Use low-fidelity prototypes, small samples of four to six users, and rapid cycles focused on discovery. The goal is finding problems fast, not proving a design works.
- Reserve summative or validation testing for near-final designs. This phase uses representative users, realistic tasks under real-world-like conditions, and objective pass or fail criteria tied directly to your critical tasks list.
- Set clear decision gates between phases. Keep iterating through formative rounds until critical-task use errors drop to an acceptable level, then freeze the design before moving into validation. Changing the interface after validation testing begins undermines the evidence.
- Choose the method that matches the question. Think-aloud protocols surface why users struggle during formative work. Simulated-use testing in a mock clinical environment bridges the gap before full validation. Actual-use validation, conducted in or close to the real clinical setting, carries the most weight for regulatory submissions. Simulation-based clinical trials, run with realistic patient scenarios and equipment, can reveal problems that bench testing misses entirely.
Formative testing should produce a running list of usability issues ranked by severity and a revised task flow for the next round. Summative testing should produce a validation report: task success and failure rates, observed use errors, root-cause analysis, and a statement on residual risk. That report becomes part of your regulatory submission package and your internal risk file. Treating these two phases as one blurred activity is why many teams arrive at validation with a design that was never stable enough to validate.
Designing Tests: Tasks, Scenarios, and Environment
Good test design starts with your critical tasks list, built from an honest clinical use case. Our guidance on developing a clinical use case for SaaS outlines how to map real workflows before writing a single test script, and a care management workflow playbook can help translate those workflows into testable product features.
Each task needs an objective success criterion defined before testing begins, not interpreted afterward. "Successfully orders the correct medication at the correct dose within four minutes, with no safety-relevant use error" is testable. "User finds it easy" is not.
Environment choice shapes what you learn:
- Lab settings offer control and easy recording but miss real interruptions and ambient noise.
- Simulated clinical rooms, built to resemble an actual exam room or nursing station, add realism without full production risk.
- In-situ actual-use testing captures the truest picture but requires careful planning around live patient safety and workflow disruption.
Data capture logistics deserve attention before the first session. Plan for screen recording, video of physical interactions, and audit log review, and secure IRB or ethics review and informed consent wherever real or simulated patient data is involved.
Pro Tip: Script a two-minute interruption, like a simulated phone call or alert, into at least one task per session. It almost always reveals more than the clean-run scenarios.
Selecting Participants and Sample Size
Recruiting the right participants matters more than recruiting a lot of them. Pull from the actual roles who will use the software: physicians, nurses, pharmacists, or medical assistants, and represent a spread of experience, from new hires to veteran staff.
Sample size depends on your test phase:
- Formative testing typically works with four to six participants per round, enough to surface major issues without over-investing in early prototypes.
- Summative validation testing generally needs a larger, more representative sample, often fifteen or more per distinct user group, sized to detect use errors with reasonable confidence rather than to reach statistical significance in the traditional sense.
- Rare-specialty workflows or vulnerable populations (pediatric, geriatric, or low-digital-literacy users) may need targeted oversampling since a general clinician pool will miss their specific failure modes.
- Define training level explicitly before scheduling, since a participant who received a two-hour onboarding behaves very differently than one using the system for the first time.
Scheduling around clinical shifts is its own project. Build in buffer time, since clinical participants reschedule often, and plan sessions around shift changes rather than against them.
Metrics and Analysis: What to Capture and How to Interpret It
The numbers that matter most are task success rate, time-on-task, and error rate, each tied back to your critical tasks list rather than tracked in the abstract. Granular interaction data, clicks, mouse travel, and pixel distance, adds real value in high-frequency workflows where small friction compounds across hundreds of daily repetitions.

Qualitative data fills the gaps numbers leave behind. Structured interviews, think-aloud commentary, and direct observation explain why an error happened, which matters more than the fact that it happened.
A perioperative clinical decision support simulation trial found that redesigned workflows reduced task time, click counts, and on-screen navigation distance compared to a standard EHR workflow, showing how granular metrics can surface friction that time-on-task alone misses.
The critical analytical step is classification. Not every observed problem carries the same weight:
- Use errors with safety implications (wrong patient, wrong dose, missed critical alert) get escalated immediately and tracked separately from general usability friction.
- Efficiency or convenience issues (extra clicks, unclear labeling) feed the design backlog but rarely block a release on their own.
- Repeated errors across multiple participants signal a systemic design flaw rather than individual user confusion.
Findings then translate into three outputs: design changes that eliminate root causes, training content for issues that can't be designed away entirely, and documented risk controls for your regulatory file.
Real World Testing and Post-Deployment Surveillance
Validation testing before launch is not the finish line. The ONC's Real World Testing requirements call for certified health IT developers to submit testing plans and results annually, with measures and settings published on the Certified Health IT Product List by set deadlines.
Designing RWT measures thoughtfully lets one data collection effort serve two purposes: satisfying your interoperability and functionality obligations while generating genuine usability evidence from live clinical use.
Post-deployment, track production metrics that mirror your validation criteria:
- Help desk tickets and workaround reports often surface use errors that lab testing missed.
- Longitudinal error rates across baseline, early go-live, and steady-state phases reveal whether training closed the gap or whether the design itself needs revision.
- User-reported near-misses deserve the same escalation path as formal adverse events.
When surveillance uncovers a pattern, it should trigger a defined corrective action process, feeding directly back into your next formative testing cycle rather than sitting in a report nobody revisits.
Practical Checklist and Sample Test Outline
A usability program doesn't need to be complicated to be defensible. It needs to follow a consistent sequence and generate evidence at each step.
- Define critical tasks from your clinical use case and risk analysis before writing any test script.
- Run formative pilots with four to six representative users and low-fidelity prototypes.
- Iterate until critical-task use errors drop to an acceptable level, then freeze the design.
- Validate with a larger, representative sample under realistic conditions tied to objective pass or fail criteria.
- Monitor post-deployment using production metrics and your ONC Real World Testing plan.
A sample moderator script for a medication-ordering task might open with a brief scenario ("Your patient has a new penicillin allergy documented; order their next antibiotic dose"), followed by silent observation, then a structured debrief asking what felt unclear and why.
Pro Tip: Keep your moderator script to one page. A long script tempts you to lead the participant instead of watching them struggle.
Common Pitfalls and When Advisory Expertise Helps
The recurring failure pattern we see is teams writing tasks around what the software does rather than what the clinician actually needs to accomplish, then recruiting a convenient participant pool instead of a representative one. Skipping environmental realism, testing in a quiet conference room instead of a simulated clinical setting, is a close second.
Clinical advisory input tends to matter most at three points: defining critical tasks with real clinical judgment, interpreting whether an observed error is a safety signal or noise, and building regulator-facing documentation that holds up under FDA or ONC scrutiny.
— Paul Bergeron MD, MBA
How We Help With Usability Testing and Real World Testing Plans
We work with healthcare SaaS teams through fractional Chief Medical Officer and advisory engagements, assisting with clinical workflow translation into critical-task definitions, interpreting human factors findings, and structuring validation reports and ONC Real World Testing plans.

If your team is preparing for an HF validation submission or an RWT filing, we can review your current test plan and flag gaps before a regulator does. Visit our services page to schedule a consultation.
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
FAQ
What is some good software for usability testing?
There is no single tool that covers every phase; teams generally combine screen and video recording software for session capture with structured templates like NIST's Common Industry Format for reporting results. The right choice depends on whether you're running formative sessions or full validation testing.
What is the most used software in healthcare?
Electronic health record systems are the most widely used category of healthcare software, since nearly every hospital and most outpatient practices rely on one for clinical documentation and ordering. Usability testing standards for these systems, including NIST's CIF templates, were developed specifically because of how central EHRs are to daily clinical work.
What is usability testing for medical devices?
Usability testing for medical devices evaluates whether intended users can operate a device safely and effectively, following the risk-based human factors process described in FDA guidance. It includes identifying critical tasks, running formative studies to refine the design, and completing a validation study whose results get reported in the premarket submission.
What are five types of usability testing?
Common types include formative testing on early prototypes, summative or validation testing on near-final designs, simulated-use testing in a mock clinical environment, actual-use testing in the real clinical setting, and remote usability testing conducted with participants outside a physical lab. Each serves a different purpose depending on how far along the design is and how much realism the evaluation needs.
Does usability testing need to continue after a product launches?
Yes, ongoing monitoring is part of the expectation, not an optional add-on. The ONC's Real World Testing program requires certified health IT developers to submit annual plans and results based on production use, which means usability evidence keeps accumulating well past go-live.
Sources
- Human Factors and Medical Devices (FDA)
- Technical Evaluation, Testing and Validation of the Usability of Electronic Health Records (NIST IR 7804-1)
- Real World Testing Resource Guide (ONC, 2025)
