Machine Learning, Parkinsonian Disorders
Conditions
Brief summary
The Movement Disorders Society (MDS) Unified Parkinson's Disease Rating Scale (UPDRS) Part III (MDS-UPDRS III) is the primary assessment method for motor symptoms in Parkinson's disease patients. Currently, movement disorder specialists conduct semi-quantitative scoring, which entails limitations such as subjectivity, weak sensitivity, and a limited number of professional physicians. This study, based on machine vision, establishes gold standard labels according to expert scoring. By using machine learning, we develop a machine rating model and compare the model's performance with gold standard rating and general clinical rating to investigate the accuracy of machine vision-based MDS-UPDRS III machine rating.
Interventions
Patients' performance of MDS-UPDRS III will be recorded.
Sponsors
Study design
Eligibility
Inclusion criteria
* Meeting the diagnostic criteria for Parkinsonism established by the International Movement Disorder Society: having bradykinesia, and meeting at least one of the two criteria for resting tremor or muscle rigidity * 20 to 80 years old * Good compliance, voluntarily joining the study, and able to sign an informed consent form or have it signed by a legal representative
Exclusion criteria
* Significant cognitive impairment (MMSE ≤ 23) * Unable to sign written informed consent or unable to complete the trial due to other reasons * Other situations in which the researcher deems the participant unsuitable for this study * Participation in other clinical trials
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Item-level: MAE | 1 day | At the item level, mean absolute error (MAE) between paired AI scores and consensus reference scores, calculated as the average absolute difference across items. Lower values indicate closer agreement. |
| Item-level: P(|Δ|≥2) | 1 day | At the item level, the proportion of paired AI and consensus reference scores with an absolute difference of 2 or more points. Lower values indicate fewer large item-level scoring disagreements. |
| Score-level: ICC | 1 day | Intraclass correlation coefficient (ICC) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. Higher ICC values indicate greater agreement. |
| Score-level: SEM | 1 day | Standard error of measurement (SEM) for AI-derived scores relative to consensus reference scores, calculated separately for each of the four subdomain scores and the total score. SEM quantifies measurement error in the units of the corresponding score, with lower values indicating greater measurement precision. |
| Score-level: SDC | 1 day | Smallest detectable change (SDC), derived from the measurement error and calculated separately for each of the four subdomain scores and the total score. SDC represents the minimum score change required to exceed expected measurement error. Lower values indicate greater measurement precision. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Item-level: Within-one agreement (ACC1; P(|Δ| ≤ 1)) | 1 day | At the item level, ACC1 is defined as the proportion of paired AI and consensus reference scores with an absolute difference of no more than 1 point, i.e., P(\|Δ\| ≤ 1). Higher values indicate closer item-level agreement. |
| Item-level: Exact-error proportions [P(|Δ| = k)] | 1 day | At the item level, exact-error proportions, P(\|Δ\| = k), represent the proportions of paired AI and consensus reference scores with each exact absolute error magnitude k. In particular, P(\|Δ\| = 0) corresponds to exact-match accuracy (ACC), defined as the proportion of items for which the AI score exactly matches the consensus reference score. The remaining values of k characterize the distribution of item-level scoring errors. |
| Item-level: Per-class exact recall | 1 day | At the item level, for each consensus reference score class, the proportion of items for which the AI score exactly matches the consensus reference score among all items belonging to that reference class. Higher values indicate better class-specific exact agreement. |
| Item-level: Confusion matrix | 1 day | At the item level, a cross-tabulation of AI scores against consensus reference scores, showing the number of paired ratings for each combination of reference and AI score categories. Rows represent consensus reference scores and columns represent AI scores. |
| Item-level: Row-normalised confusion matrix | 1 day | At the item level, the confusion matrix normalised within each consensus reference score row so that each row sums to 1, showing the distribution of AI scores conditional on each reference score class. |
| Score-level: Limits of agreement | 1 day | Bland-Altman limits of agreement between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. The limits characterize the range within which most paired differences between AI and reference scores are expected to fall. |
| Score-level: MAE | 1 day | Mean absolute error (MAE) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score as the average absolute difference between paired scores. Lower values indicate smaller scoring errors. |
| Score-level: RMSE | 1 day | Root mean square error (RMSE) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. RMSE gives greater weight to larger scoring errors, with lower values indicating closer agreement. |
| Score-level: Spearman correlation | 1 day | Spearman rank correlation between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. Higher values indicate a stronger monotonic association between AI-derived and reference scores. |
Countries
China