Machine Intelligence in the Pharmacy
Conditions
Keywords
pharmacy, medication verification, artificial intelligence, machine intelligence
Brief summary
Pharmacists currently perform an independent double-check to identify drug-selection errors before they can reach the patient. However, the use of machine intelligence (MI) to support this cognitive decision-making work by pharmacists does not exist in practice. This research is being conducted to examine the effectiveness of the timing of machine intelligence (MI) advice on to determine if it results in lower task time, increased accuracy, and increased trust in the MI.
Detailed description
Pharmacists currently perform an independent double-check currently to identify drug-selection errors before they can reach the patient. However, the use of machine intelligence (MI) to support this cognitive decision-making work by pharmacists does not exist in practice. Instead, pharmacists rely solely on reference images of the medication which they can compare to the prescription vial contents. Previous research has shown that decision support systems can effectively improve healthcare delivery efficiency and accuracy, while preventing adverse drug events. However, little is known about how MI technologies impact pharmacists' work performance and cognitive demand. To facilitate the long-term symbiotic relationship between the pharmacists and the MI system, proper trust needs to be established. While trust has been identified as the central factor for effective human-machine teaming, issues arise when humans place unjustified trust in automated technologies do not place enough trust in them. Over trust in automation can lead to complacency and automation bias. For instance, the pharmacists may rely on the MI system to the extent that they blindly accept any recommendation by the system. Under trust can result in pharmacist disuse and potential abandonment of the MI system. Furthermore, little is known about the timing of the MI advice on pharmacists' work performance. For example, showing the MI's advice while the pharmacist is performing the medication verification task may yield different results than showing the MI's advice after the pharmacist made their decision. The study investigators have developed a MI system for medication images classification. The objective of this study is to examine the effectiveness of the timing of MI advice to determine if it results in lower task time, increased accuracy, and increased trust in the MI.
Interventions
Participants will complete the medication verification task without any MI help
Participants will receive MI in the form of a pop-up message if their decision differs from the MI's determination.
MI help will be displayed concurrently with the filled and reference images.
Sponsors
Study design
Eligibility
Inclusion criteria
* Licensed pharmacist in the United States * Age 18 years and older at screening * PC/Laptop with Microsoft Windows 10 or Mac (Macbook, iMac) with MacOS with Google Chrome, Edge, Opera, Safari, or Firefox web browser installed on the device * Screen resolution of 1024x968 pixels or more * A laptop integrated webcam or USB webcam is also required for the eye tracking purpose.
Exclusion criteria
* Participated in Wave 1 or Wave 2 * Eyeglasses * Uncorrected cataracts, intraocular implants, glaucoma, or permanently dilated pupil * Require a screen reader/magnifier or other assistive technology to use the computer * Eye movement or alignment abnormalities (lazy eye, strabismus, nystagmus)
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Reaction Time | Throughout the verification task | Difference in task time measured by the number of seconds from starting the task to accepting or rejecting a medication image |
| Decision Accuracy | Throughout the verification task | Difference in detection rate measured by the number of medication verification errors across all participants in the Arm/Group. |
| Trust Change | After every trial in Scenarios 1 and 2 | Participants will complete 100 mock medication verification trials in each of the study arms (i.e., Scenario 1, Scenario 2, and No Help). After each trial in Scenario 1 and Scenario 2, participants will use a visual analog scale (VAS) to respond to the question: How much do you trust the AI advice? The endpoints of the 100-point VAS are 'Not at all' to 'Completely trust'. Participants indicate their level of trust in the MI advice after every trial on a scale from 1-100, with higher scores indicating greater levels of trust. The trust change, as measured by the visual analog scale, will be calculated using the following formula: Trust change (i) = Trust(i) - Trust(i - 1), where i=2, 3, ..., 100. To compute a single, summarized value for the Trust Change variable within a specific scenario, the individual Trust Change scores measured from the trials are averaged. This averaging method provides a comprehensive measure of how trust shifted across the duration of the scenario. |
| Trust | Post-intervention in Scenarios 1 and 2. | Trust will be assessed using the Muir & Moray's (1996) Trust in Automation scale. Scores range from 0 to 100 with higher scores indicating greater levels of trust. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Workload | After completing 100 mock verification trials in each arm | Participants will complete 100 mock medication verification trials in each of the 3 arms. The workload of each arm will be measured by the NASA Task Load Index (TLX). The 5 TLX dimensions assessed are: mental demand, effort, temporal demand, performance, and frustration. For each dimension, participants will indicate their response to a single question. For 4 of the dimensions, the endpoints of the Likert scale are 'very low' and 'very high'. The performance dimension is reverse-scored, and the endpoints are 'perfect' and 'failure'. Participants then complete 10 pairwise comparisons of the dimensions by indicating which dimension they consider to be a more important factor (e.g., effort vs frustration). Each category score multiplied by its respective pairwise comparison count is summed and divided by 10 to get an overall weighted workload score. The result is an overall workload score between 1 and 20, with higher scores indicating higher workload. |
| Usability | After completing 100 mock verification trials in each arm | Participants will complete 100 mock medication verification trials in each of the 3 arms (No MI Help, Scenario 1, and Scenario 2). After completing 100 trials, participants will assess the mock verification interface using the System Usability Scale (SUS). The SUS is comprised of 10 statements that participants indicate their agreement with using a 5-point Likert scale ranging from strongly agree to strongly disagree. Odd-numbered questions have a positive response and even-numbered questions are reverse-scored. Scores are summed and multiplied by 2.5 to get a final SUS score. SUS scores range from 0 to 100 with higher scores indicating greater usability. An average SUS score is considered to be 68. Anything below 50 is Not Acceptable. Scores between 51-70 are considered Marginal, those above 71 are considered Acceptable, and those at 80 or above are indicative of high usability. |
| Cognitive Effort | Throughout the verification task | Participants' eye movements were tracked using a browser-based online eye tracking system. The outcome measure is the difference in cognitive effort as measured by fixation count in the defined areas of interest: fill image, reference image, or MI plot. Higher fixation rates indicate repeated interest in a certain area. |
Countries
United States
Participant flow
Participants by arm
| Arm | Count |
|---|---|
| Pharmacists Licensed pharmacists in the United States with medication dispensing experience who are 18 or older. | 50 |
| Total | 50 |
Withdrawals & dropouts
| Period | Reason | FG000 |
|---|---|---|
| No MI help | Technical issues | 18 |
Baseline characteristics
| Characteristic | Pharmacists |
|---|---|
| Age, Categorical <=18 years | 0 Participants |
| Age, Categorical >=65 years | 0 Participants |
| Age, Categorical Between 18 and 65 years | 50 Participants |
| Age, Continuous | 35.52 years STANDARD_DEVIATION 6.949 |
| Ethnicity (NIH/OMB) Hispanic or Latino | 2 Participants |
| Ethnicity (NIH/OMB) Not Hispanic or Latino | 46 Participants |
| Ethnicity (NIH/OMB) Unknown or Not Reported | 2 Participants |
| Race (NIH/OMB) American Indian or Alaska Native | 0 Participants |
| Race (NIH/OMB) Asian | 6 Participants |
| Race (NIH/OMB) Black or African American | 1 Participants |
| Race (NIH/OMB) More than one race | 2 Participants |
| Race (NIH/OMB) Native Hawaiian or Other Pacific Islander | 0 Participants |
| Race (NIH/OMB) Unknown or Not Reported | 3 Participants |
| Race (NIH/OMB) White | 38 Participants |
| Region of Enrollment United States | 50 participants |
| Sex: Female, Male Female | 34 Participants |
| Sex: Female, Male Male | 14 Participants |
Adverse events
| Event type | EG000 affected / at risk | EG001 affected / at risk | EG002 affected / at risk |
|---|---|---|---|
| deaths Total, all-cause mortality | 0 / 50 | 0 / 50 | 0 / 50 |
| other Total, other adverse events | 0 / 50 | 0 / 50 | 0 / 50 |
| serious Total, serious adverse events | 0 / 50 | 0 / 50 | 0 / 50 |
Outcome results
Decision Accuracy
Difference in detection rate measured by the number of medication verification errors across all participants in the Arm/Group.
Time frame: Throughout the verification task
| Arm | Measure | Value (NUMBER) |
|---|---|---|
| No MI Help | Decision Accuracy | 291 Number of errors |
| Scenario #1 | Decision Accuracy | 238 Number of errors |
| Scenario #2 | Decision Accuracy | 230 Number of errors |
Reaction Time
Difference in task time measured by the number of seconds from starting the task to accepting or rejecting a medication image
Time frame: Throughout the verification task
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| No MI Help | Reaction Time | 3668 millisecond (ms) | Standard Deviation 924 |
| Scenario #1 | Reaction Time | 4727 millisecond (ms) | Standard Deviation 1040 |
| Scenario #2 | Reaction Time | 4510 millisecond (ms) | Standard Deviation 1339 |
Trust
Trust will be assessed using the Muir & Moray's (1996) Trust in Automation scale. Scores range from 0 to 100 with higher scores indicating greater levels of trust.
Time frame: Post-intervention in Scenarios 1 and 2.
Population: This Outcome Measure was pre-specified to be only assessed for Scenarios 1 and 2. No data were collected for this Outcome Measure for the No MI Help scenario.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Scenario #1 | Trust | 53.65 scores on a scale | Standard Deviation 33.12 |
| Scenario #2 | Trust | 60.37 scores on a scale | Standard Deviation 31.95 |
Trust Change
Participants will complete 100 mock medication verification trials in each of the study arms (i.e., Scenario 1, Scenario 2, and No Help). After each trial in Scenario 1 and Scenario 2, participants will use a visual analog scale (VAS) to respond to the question: How much do you trust the AI advice? The endpoints of the 100-point VAS are 'Not at all' to 'Completely trust'. Participants indicate their level of trust in the MI advice after every trial on a scale from 1-100, with higher scores indicating greater levels of trust. The trust change, as measured by the visual analog scale, will be calculated using the following formula: Trust change (i) = Trust(i) - Trust(i - 1), where i=2, 3, ..., 100. To compute a single, summarized value for the Trust Change variable within a specific scenario, the individual Trust Change scores measured from the trials are averaged. This averaging method provides a comprehensive measure of how trust shifted across the duration of the scenario.
Time frame: After every trial in Scenarios 1 and 2
Population: This Outcome Measure was pre-specified to be only assessed for Scenarios 1 and 2. No data were collected for this Outcome Measure for the No MI Help scenario.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Scenario #1 | Trust Change | -0.1715 units on a scale | Standard Deviation 21.1415 |
| Scenario #2 | Trust Change | -0.0552 units on a scale | Standard Deviation 19.3931 |
Cognitive Effort
Participants' eye movements were tracked using a browser-based online eye tracking system. The outcome measure is the difference in cognitive effort as measured by fixation count in the defined areas of interest: fill image, reference image, or MI plot. Higher fixation rates indicate repeated interest in a certain area.
Time frame: Throughout the verification task
Population: Complete eye tracking data was not available for analysis for one participant. The MI plot was pre-specified to be only assessed for Scenarios 1 and 2. No data were collected for the Outcome Measure MI plot for the No MI Help.
| Arm | Measure | Group | Value (MEDIAN) |
|---|---|---|---|
| No MI Help | Cognitive Effort | Area of Interest: Fill Image | 3 Number of fixations |
| No MI Help | Cognitive Effort | Area of Interest: Reference Image | 2 Number of fixations |
| Scenario #1 | Cognitive Effort | Area of Interest: Reference Image | 2 Number of fixations |
| Scenario #1 | Cognitive Effort | Area of Interest: Fill Image | 3 Number of fixations |
| Scenario #1 | Cognitive Effort | Area of Interest: MI plot | 1 Number of fixations |
| Scenario #2 | Cognitive Effort | Area of Interest: Reference Image | 2 Number of fixations |
| Scenario #2 | Cognitive Effort | Area of Interest: MI plot | 2 Number of fixations |
| Scenario #2 | Cognitive Effort | Area of Interest: Fill Image | 3 Number of fixations |
Cognitive Effort
Participants' eye movements were tracked using a browser-based online eye tracking system. The outcome measure is the difference in cognitive effort as measured by the duration of fixations in the defined areas of interest: fill image, reference image, or MI plot. Longer fixation duration indicates a higher cognitive load.
Time frame: Throughout the verification task
Population: Complete eye tracking data was not available for one participant. The MI plot was pre-specified to be only assessed for Scenarios 1 and 2. No data were collected for the Outcome Measure MI plot for the No MI Help.
| Arm | Measure | Group | Value (MEDIAN) |
|---|---|---|---|
| No MI Help | Cognitive Effort | Area of Interest: Fill Image | 619.5 millisecond (ms) |
| No MI Help | Cognitive Effort | Area of Interest: Reference Image | 365.0 millisecond (ms) |
| Scenario #1 | Cognitive Effort | Area of Interest: Reference Image | 399.0 millisecond (ms) |
| Scenario #1 | Cognitive Effort | Area of Interest: Fill Image | 687.5 millisecond (ms) |
| Scenario #1 | Cognitive Effort | Area of Interest: MI Plot | 268.0 millisecond (ms) |
| Scenario #2 | Cognitive Effort | Area of Interest: Reference Image | 419.0 millisecond (ms) |
| Scenario #2 | Cognitive Effort | Area of Interest: MI Plot | 299.0 millisecond (ms) |
| Scenario #2 | Cognitive Effort | Area of Interest: Fill Image | 706.0 millisecond (ms) |
Usability
Participants will complete 100 mock medication verification trials in each of the 3 arms (No MI Help, Scenario 1, and Scenario 2). After completing 100 trials, participants will assess the mock verification interface using the System Usability Scale (SUS). The SUS is comprised of 10 statements that participants indicate their agreement with using a 5-point Likert scale ranging from strongly agree to strongly disagree. Odd-numbered questions have a positive response and even-numbered questions are reverse-scored. Scores are summed and multiplied by 2.5 to get a final SUS score. SUS scores range from 0 to 100 with higher scores indicating greater usability. An average SUS score is considered to be 68. Anything below 50 is Not Acceptable. Scores between 51-70 are considered Marginal, those above 71 are considered Acceptable, and those at 80 or above are indicative of high usability.
Time frame: After completing 100 mock verification trials in each arm
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| No MI Help | Usability | 72.1 score on a scale | Standard Deviation 16.5 |
| Scenario #1 | Usability | 70.4 score on a scale | Standard Deviation 16.4 |
| Scenario #2 | Usability | 73.0 score on a scale | Standard Deviation 15.6 |
Workload
Participants will complete 100 mock medication verification trials in each of the 3 arms. The workload of each arm will be measured by the NASA Task Load Index (TLX). The 5 TLX dimensions assessed are: mental demand, effort, temporal demand, performance, and frustration. For each dimension, participants will indicate their response to a single question. For 4 of the dimensions, the endpoints of the Likert scale are 'very low' and 'very high'. The performance dimension is reverse-scored, and the endpoints are 'perfect' and 'failure'. Participants then complete 10 pairwise comparisons of the dimensions by indicating which dimension they consider to be a more important factor (e.g., effort vs frustration). Each category score multiplied by its respective pairwise comparison count is summed and divided by 10 to get an overall weighted workload score. The result is an overall workload score between 1 and 20, with higher scores indicating higher workload.
Time frame: After completing 100 mock verification trials in each arm
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| No MI Help | Workload | 7.4 score on a scale | Standard Deviation 3.7 |
| Scenario #1 | Workload | 7.5 score on a scale | Standard Deviation 3.8 |
| Scenario #2 | Workload | 6.9 score on a scale | Standard Deviation 3.4 |