Inertia-1

Inertia-1

An Open Exploration to a Unified Motion Foundation Model

Contact Us

Towards one general motion model

Motion is universal — but the models built for it weren't. Inertia-1 brings the whole landscape under one roof.

Datasets disagree on the basics — sampling rate, window length, sensor modality, body placement, even signal format — and every task gets its own bespoke model. Findings rarely carry from one setup to the next.

Inertia-1 studies the full lifecycle of motion models — data, sensing, objectives, and scale — inside a single, controlled space instead of isolated one-offs.

The payoff: one representation that adapts across placements, devices, and tasks — the same backbone, working far beyond the setting it was trained on.

Beyond benchmarks, Inertia-1 surfaces the choices that decide whether a motion model actually works in the real world.

Learn it on the wrist. Use it anywhere on the body.

A wrist sensor pretrained on accelerometry transfers across the body — head, chest, hip, thigh, knee, shin, and ankle — each showing its own motion signal.

Pretrain once on the wrist, then point the model anywhere. It holds up on body placements — and even sensor types like gyroscope and magnetometer — that it never saw during training. No retraining for each new spot on the body.

Wrist-accelerometer pretraining transfers across sensors and placements.

Add more streams. Get more signal.

Multi-stream fusion improves representation geometry.

Stack on more streams — extra placements, gyroscope, magnetometer — and the learned representation gets both more accurate and cleaner, with activities separating into tighter clusters. The streams are complementary: each one catches something the others miss.

Sensing design is a first-order choice

How you capture motion shapes what a model can do with it. A few practical rules of thumb from the study.

Sampling rate

Pretrained models stay strong even at a low 1 Hz for activity recognition; finer-grained health signals benefit from higher sampling rates.

Pretrained models stay strong even at a low 1 Hz for activity recognition; finer-grained health signals benefit from higher sampling rates.

Window length

30–60 second windows hit the sweet spot across most tasks — long enough to capture context, short enough to stay sharp.

30–60 second windows hit the sweet spot across most tasks — long enough to capture context, short enough to stay sharp.

Keep all three axes

Full triaxial input consistently beats collapsed vector-magnitude summaries — the extra axes carry signal worth keeping.

Full triaxial input consistently beats collapsed vector-magnitude summaries — the extra axes carry signal worth keeping.

Stay in the time domain

Time-domain modeling preserves gait and health cues better than frequency-domain reconstruction.

Time-domain modeling preserves gait and health cues better than frequency-domain reconstruction.

One pipeline, from raw signal to real-world insight

The general representation comes together in three clean steps.

Learn from planetary-scale accelerometry — over 18 million hours across global cohorts — with self-supervision, no labels required.

Adapt the same representation to new placements, devices, and sampling rates with light tuning — or none at all.

Power activity, mobility, and health applications from one backbone — from fitness tracking to clinical screening.

From movement to meaning

The same representation spans the full spectrum of motion understanding.

One general model for human motion

Inertia-1 is a first step toward a unified motion foundation model — and an open invitation to collaborators with motion data, new tasks, or a shared interest in where the field is headed.

Read the paper

Contact Us