A marker is not a lever
Measurements do two different jobs, and the two get confused constantly. A prognostic marker sorts people by risk. A modifiable target is something which, when changed, changes the outcome. Grip strength is exceptionally good at the first and unevidenced for the second.
The UK Biobank analysis of roughly half a million participants related grip strength to all-cause mortality and to cardiovascular, respiratory and cancer outcomes, and the pooled prospective cohort literature reports the same direction of association in community-dwelling populations. For a measurement that takes seconds and costs almost nothing, that is a remarkable amount of predictive signal.
What none of it establishes is that raising grip strength as such lowers risk. Grip is a cheap proxy for overall condition, and the risk sits with the condition, not with the handshake. This is the single most misread finding in the measurement literature.
Grip strength and the threshold problem
Grip is measured with a hand dynamometer. It takes seconds, needs no individual calibration and costs almost nothing, which is why it appears in cohort studies at a scale no laboratory measure reaches.
The dose-response meta-analysis asked directly whether identifiable handgrip thresholds correspond to mortality risk for all-cause, cancer and cardiovascular death. The relationship it describes is graded across the measured range rather than switching on at a point.
Threshold values still get published, because clinical definitions need a number to operate on. They should be read as operational cut-points chosen for a purpose — and as values that depend on sex, body size and the equipment used — not as biological boundaries that exist independently of the decision to draw them.
Power and strength are different measurements
Strength and power are separate physical quantities and separate measurements. Strength is peak force. Power is force multiplied by the velocity at which that force is expressed. In everyday function the time available is often the binding constraint rather than the force required, which is why the distinction is not merely terminological.
The Mayo Clinic Proceedings analysis put the two head to head as predictors of mortality in middle-aged and older men and women, rather than treating either as a proxy for the other. It reports them as distinct predictors — strength and power are not interchangeable, and results from one do not transfer to the other.
The marker-and-lever caution applies here unchanged. This is prognostic data about how two measurements sort people by risk.
What wearables measure well, and what they do not
The validation literature is unusually clear on which metrics survive scrutiny, and the split is the most useful result in it. The systematic reviews of commercially available devices report steps, heart rate and energy expenditure separately, because the accuracy differs sharply between them: the first two hold up for most purposes, the third does not.
The 2023 validation of the Apple Watch 6, Polar Vantage V and Fitbit Sense measured heart rate and energy expenditure as separate outcomes on the same three devices and found the same split. The 2025 meta-analysis of Apple Watch accuracy across the health metrics it reports describes the same pattern, and the living umbrella review of consumer wearable accuracy summarises it across the systematic reviews as a whole.
The reading follows directly. Heart rate from a wrist device is a measurement, with known conditions under which it deteriorates. Energy expenditure is a model output wearing the typography of a measurement, and it should not be treated as data.
Heart-rate variability: real signal, heavy noise
Heart-rate variability is the variation in the interval between consecutive heartbeats, and it carries real information about autonomic regulation. The review of heart-rate variability and training adaptation in elite endurance athletes is the reference account of both its value and its difficulty: day-to-day noise is large relative to the signal, recording conditions have to be standardised to compare anything, and the same directional change can mean different things in different states.
The measurement question sits on top of that. Because the index is built from beat-to-beat timing, it demands far more of a sensor than a heart-rate average does. The validation study against reference instrumentation reports accuracy that holds under resting conditions and degrades away from them.
The two problems compound rather than cancel. A noisy underlying signal, captured by a sensor whose accuracy depends on conditions, then reduced to a single daily score, is a chain in which small errors travel a long way. This is the most over-interpreted number in consumer fitness, and it is over-interpreted for structural reasons rather than careless ones.
Load monitoring, and the shelf life of a validation study
Training load is monitored in two registers. External load is what was performed — distance, speed, load, repetitions. Internal load is the response it produced — heart rate, perceived exertion, and the recovery markers discussed above. The review of monitoring practice sets out what each family of measures actually captures, and its central point is that they answer different questions.
The distinction is not pedantic. Two people completing an identical session absorb it differently, and two sessions producing identical internal responses can have been very different pieces of work. A measure of one is not evidence about the other, and treating them as interchangeable is how monitoring data gets over-read.
A validation result belongs to the device, the firmware, the population and the activity it was collected in. The wearable umbrella review is maintained as a living document for exactly that reason, and the earlier consumer-tracker review now reads as a snapshot of one generation of hardware. Accuracy is not a property a product category inherits — it is a claim that has to be re-established every time something changes.



