Activities of Daily Living, Aging, Cardiovascular Disease Prevention, Cognitive Dysfunction, Diagnostic Techniques and Procedures, Frailty, Health Services Accessibility, Metabolic Syndrome, Mortality, Musculoskeletal Diseases, Physical Disability, Preventive Health Services
Conditions
Keywords
Longevity, Preventive Health, Health Screening, Biological Age, Artificial Intelligence, Machine Learning, Computer Vision, Sit to Rise, Gait Speed, Pose Estimation, All-Cause Mortality, Life Expectancy Prediction, Healthspan, Preventive Medicine, Functional Decline, Human Performance
Brief summary
This study builds AI models that score diagnostic screening tests, and that predict screening results, clinical judgment, and life expectancy. Longevity Metrics collects a battery of clinical tests on each participant, in whole or in part, and follows every participant for life. The sit-to-rise test and the timed walk are scored by hand today, from a person's count. A model scores the same test from video instead. It also measures what no one can count by eye - speed, asymmetry, steadiness - so one capture yields both the original score and additional measurements, intended to enrich the model and strengthen what it predicts. Every test in a participant's record measures the same body, so the tests are correlated: a test that was performed carries information about one that was not. A model trained across the library learns those relationships and estimates a missing result from the results that are present. Each estimate is checked against records where that part was actually measured, and over decades against death and disease through linkage to the 100-Year Human Aging Study (NCT07563777). The hypothesis is that the full battery can eventually be predicted across modalities with high accuracy using a few short video clips, replacing most in-person screening. That would let preventive screening reach people and places a physical laboratory cannot. How far the input can be reduced is the question this study exists to answer. Every model is a physician-reviewed clinical decision aid until it is cleared by the FDA.
Detailed description
The models serve three aims. First, they automatically score simple physical and cognitive tests that already predict function and in some cases mortality, such as the sit-to-rise test and the timed walk. A model reads richer detail from the same recording than a human scorer can, so it improves on the human score rather than only reproducing it. Second, they predict the parts of a screening a participant did not obtain from the parts that were performed, and increasingly from inexpensive standardized inputs such as a short video. Within a single record, every test is correlated with the others, so each test can both predict the ones that were not performed and serve as the truth against which those predictions are checked. A missing-data engine fills any missing part of a record by leave-one-out across the library. Third, they predict the physician's clinical judgment where no determining measurement exists. Models are developed by milestone freezing with forward validation. No model is validated on records it was trained on. Each model is validated in two stages. It is first validated against a human scorer for measurement accuracy, which gates its use as a clinical decision aid. It is then validated over decades for what it predicts about death and disease. The platform's distinguishing asset is mortality. Every participant is followed for life, so a small library with verified death outcomes answers questions a much larger library without them cannot. The library is the durable asset, studied across geography and time by increasingly capable models. This study is one of four that compound into one system. The 100-Year Human Aging Study (NCT07563777) supplies the clinical data and validates what it means for health, disease, disability, and death. The Human Observatory Study (NCT07646782) does the same with sociodemographic and environmental data, and receives each model's geographic residuals. The Health Ahead Comparative Effectiveness Study (NCT07669168) moves the screening toward increasing automation and mobility while maintaining quality. This study builds the models that make automation, prediction, and broad utilization possible. A physician or licensed provider reviews and signs every result a participant receives.
Interventions
Video and audio-captured physical, cognitive, and observational screening tests scored by frozen, versioned AI models; models automatically score standard tests (sit-to-rise, timed walk, chair stand), predict non-performed results from performed ones, and predict clinical judgment; all outputs are physician-reviewed clinical decision aids, and no model output reaches a participant report without provider confirmation.
Sponsors
Study design
Eligibility
Inclusion criteria
* Age \>= 18 * willing to participate in the study
Exclusion criteria
* Age \< 18 years
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Prediction of Remaining Life Expectancy | Through study completion, up to 100 years | Accuracy of predicting remaining life expectancy from video and library data, against observed mortality through linkage to the 100-Year Human Aging Study. |
| Prediction of Physician Clinical Judgment | At each model freeze, through study completion, up to 100 years | Accuracy of predicting the physician's clinical judgment where no determining measurement exists, against the sealed record. Applies to models that predict a clinical determination rather than a measured value. Where a determining test exists but was not performed, the outcome falls under Primary Outcome 2. |
| Screening-Related Injuries and Adverse Events | Continuously from first screening through study completion, up to 100 years | Participant injuries or adverse events attributed to measurements added under this study, such as the sit-to-rise test. Ascertained from two independent sources so that an event missed by one is still captured: the tester logs any fall or injury at the time of screening, and the participant reports separately on the post-screening form. |
| Measurement Accuracy, Non-Inferiority to Human Scoring. | At each model freeze, through study completion, up to 100 years. | Agreement between the model's value and the reference standard, established as non-inferiority to a qualified human scorer where a human reference exists. Gates deployment as a clinical decision aid. |
| Cross-Prediction of Non-Performed Results and Derived Scores | At each model freeze, through study completion, up to 100 years. | Accuracy of predicting, from video or from the performed part of a record, both the screening results the participant did not obtain, against the measured value; and the derived scores computed from a complete record, including the Longevity Score, biological age, and estimated age at death, against the value the scoring engine produces from the full measured record. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Model Enrichment | Through study completion, up to 100 years. | Whether added parameters improve prediction of 100-Year outcomes over the base measure. Reported once Primary Outcome 3 is estimable. |
| Improvement Over Standard Screening | Through study completion, up to 100 years | Incremental value of video-derived information over standard screening for remaining life expectancy, through the 100-Year linkage. Reported once Primary Outcome 3 is estimable. |
| Incident Chronic Disease Prediction | Through study completion, up to 100 years | Accuracy of predicting new chronic disease onset, through the 100-Year linkage. Reported by predictor set, model and physician, so the prognostic comparison is made on a common outcome. |
| Cause-of-Death Prediction | Through study completion, up to 100 years | Concordance between predicted and actual cause of death, through the 100-Year linkage. |
| Time to Functional Disability | Through study completion, up to 100 years | Accuracy of predicting the timing of functional disability onset, through the 100-Year linkage. |
| Geographic Predictive Transportability | At each model freeze, through study completion, up to 100 years | Change in predictive performance with environmental distance from the validated envelope. Residuals are returned to the Human Observatory Study. |
| Temporal Stability and Drift | Continuously from first model freeze through study completion, up to 100 years | Stability of performance across calendar time and across a participant's repeat screenings. |
| Missing-Data and Completeness Robustness | At each model freeze, through study completion, up to 100 years | Model performance on partial or incomplete input relative to complete input. |
| Rate of Change | Through study completion, up to 100 years | Accuracy of predicting the change in a measure between visits, and of predicting outcomes from that change. Applies to every model with repeat captures. The within-participant change and its pace are tested against death, disease, and functional decline alongside the cross-sectional value. |
| Measurement Accuracy, Superiority to Human Scoring | At each model freeze, through study completion, up to 100 years | Conditional on non-inferiority, whether the model scores more reliably or precisely than a qualified human scorer. |
| Uncertainty Calibration | At each model freeze, through study completion, up to 100 years | Agreement between the model's stated confidence and its observed accuracy. |
| Within-Session Repeatability | At each model freeze, through study completion, up to 100 years. | Where a test is captured twice back to back, the stability of the model's output across the two captures. Not applicable where repeat capture is impractical, or where a second effort measures fatigue because the test is performed to failure. |
| Physician Assessment of Model Output | Continuously from first model read through study completion, up to 100 years | Three physician-recorded fields per result, identical to the structured judgment recorded under the Health Ahead Comparative Effectiveness Study, so that one instrument governs physician review of model output across the platform. Agreement with the model read, 1 to 10, 1 strongly disagree to 10 strongly agree. Safety, the potential for patient harm had the read been acted upon as written, 1 to 10, 1 high potential for serious harm to 10 no potential for patient harm. Override, binary, recording whether the released interpretation differs in any substantive respect from the read. The two rating scales ascend toward the better state. Agreement records what the physician thought of a read; override records what the physician did with it, and the two diverge routinely. |
Countries
United States
Contacts
Longevity Metrics, Inc.