CurveCodec

CurveCodecv0.2.0 Skeleton-agnostic animation compression with a learned entropy model

1The University of Hong Kong    2Adobe Research    *Co-corresponding authors

Demo posterScreenshot of the Space: example gallery on the left, three-way viewer (raw / ACL / ours) on the right.
Loading the live demo… Open live demo ↗ 226 example clips · animals, humans, robots, two-person takes · or upload a BVH
Live demo. Pick an example or upload a BVH; the viewer plays the raw clip, ACL and ours on one timeline. Runs on a 2-vCPU server — long clips take a moment. Open full screen ↗
Overview

CurveCodec 2 compresses skeletal animation from any rig into a compact, bit-exact bitstream. Every joint curve is quantized in closed loop against a stated error bound, the encoder chooses per joint which samples not to code, and a small learned entropy model codes what remains. Trained once on 886 hours of motion from 33 datasets, it spends about 2 bits per joint-sample at 0.1 cm, 100× less than raw float32, and transfers without retraining to a species it has never seen.

Footprint
906 h of motion, 775 GB as BVH: 11.9 GB at 0.01 cm, 4.3 GB at 0.1 cm, 1.1 GB at 1 cm
Input
Any skeleton, any rig: per-joint rotations and translations (BVH)
Precision
0.01 to 1 cm, set per clip; every decoded clip is verified against its error bound*
Entropy model
108 K-parameter causal transformer (870 KB); integer inference, bit-exact across platforms
Decode
2.1 s per million joint-samples on one CPU core, 0.76 s on four threads
Total size against error
Held-out test side: 4,472 clips, 20.1 h.
Table
pCurveCodec 2ACL
sizemean errorsizemean error
original9,596 MB as float32 (quaternion + translation per joint, 28 B), error 0
0.01 cm262 MB0.0033 cm702 MB0.0034 cm
0.05 cm136 MB0.0166 cm509 MB0.0166 cm
0.1 cm96 MB0.0332 cm433 MB0.0333 cm
0.3 cm50 MB0.134 cm365 MB0.137 cm
1 cm26 MB0.542 cm374 MB0.560 cm

*Each clip's mean joint error is at most that of the reference codec (ACL) at the same p; clips that fail fall back, and the fallback is counted. Sizes are bits per joint-sample × 342.7 M joint-samples; “original” is the same clips as float32, a quaternion and translation per joint (28 B).

01Questions

Q1Why study compression at first?

  • Motion data is highly redundant. A clip is stored as every joint's local transform at every frame. But when a person performs an action, they rarely attend to how each joint gets from one place to the next: the trajectories are largely produced by a strong prior, the body itself, rather than being what the motion is about. Spending more effort and computation on generating these curves does little for understanding behaviour or action.
  • Compression is a principled way to learn what matters. Embodied intelligence works the same way: an agent decides what to do, and its body, its morphology and its dynamics, decide most of how. A model of motion intelligence should spend its capacity on the decisions, not on the kinematics the embodiment already implies. Compression separates the two: what a good codec must still send is the decision; what it can drop is the body.
  • We want a representation general enough to model the dynamics of motion. If the body is only the prior, the representation should not be tied to one body: it has to generalize across motions and across skeletons. Today's hierarchical skeleton representations bake a specific topology into the data, so every rig needs its own model and little of what is learned on one body carries over to another; they do not scale. CurveCodec 2 treats one model serves any rig, and transfers without retraining to a species it has never seen.

Q2Could it replace ACL?

  • No. Many of our evaluations use ACL: it is the production reference for what an error bound means, and the only tool we could find that truly works on any skeleton, as we aim to. But the two codecs target entirely different goals. ACL is built for the best possible runtime behaviour: it is stateless, keeps the clip compressed in memory and decompresses only the poses a frame needs, and reads any sample at random with as little memory traffic as possible.
  • Matching ACL's precision has a price. Our encoder corrects itself with error feedback through forward kinematics, and our stream is entropy-coded and decoded sequentially, once per clip at load time, without random access. It is still far faster than real time: one CPU core decodes about 8,700 frames per second of a 55-joint skeleton. But it is not a design that chases efficiency above everything else.
02Method

A lossy stage that chooses which samples not to code, and a lossless stage where the network lives.

Two-stage view: shared preprocessing takes the 906-hour corpus from 647 GB raw to 453 GB; ACL then drops w, range-reduces per segment, picks a bit width per joint and packs (32.6 GB); CurveCodec 2 log-maps, quantizes in closed loop, selects RD keys verified through FK, predicts and codes with a learned rANS (11.9 GB plus an 870 KB model).
Where the bytes go. Both codecs share the deterministic preprocessing; ACL then quantizes with per-segment range reduction and fixed bit widths, CurveCodec 2 with a closed-loop quantizer, RD-selected keys, prediction and a learned entropy coder. The 870 KB model is fixed after training and shared by all clips.
Overview of CurveCodec 2. (a) Lossy stage: the clip is folded and log-mapped into one curve per sub-track; a per-clip closed loop grows quantization steps or drops keys, decodes each candidate, measures the subtree shell error through forward kinematics and accepts it only if the error gate holds. (b) Lossless stage: a fixed predictor turns the integers into residuals, a causal transformer reads own-past features and predicts their distribution for rANS; the decoder replays the same model in lockstep and reconstructs the clip with Catmull-Rom interpolation.
Overview. (a) The lossy stage runs once per clip in the encoder: every animated sub-track becomes a curve in the log map, and a closed loop grows quantization steps or drops keys, accepting a move only if the error contract still holds through forward kinematics. (b) The lossless stage turns the integers into residuals with a fixed predictor; a small learned model, the only learned component, outputs the distribution that drives the rANS coder. The decoder replays it in lockstep and reproduces every probability exactly.
  1. e(j,t) = max over axes of |T(t)·δe − T̂(t)·δe|

    Error is measured through the skeleton. The shell error of joint j at sample t is the largest displacement of three virtual vertices δ = 3 cm along its axes, in object space after forward kinematics, so it contains the errors of all its ancestors.

  2. accept a move iff ΔD / ΔR ≤ λ and the contract holds

    The lossy stage trades error for bits. Removing a key or growing a quantization step is accepted only if the error it adds per bit it saves stays under λ, and the decoded subtree still meets its contract; λ rises until the error budget is spent.

  3. bits ≈ Σ −log2 P(r_t | f_≤t), r_t = n_t − n̂_t

    The network is the entropy model. Each integer is predicted from its own past; a small causal transformer gives the distribution of the residual from the curve's own decoded history, and rANS codes it at close to that many bits. The decoder runs the same model in lockstep.

03Analysis
Compression ratio against raw float32
Held-out test side, one row per dataset, sorted by the ratio at 0.01 cm; log scale.
Table
datasethours0.01 cm0.1 cm1 cm
BONES-SEED144.2120×307×795×
Geno 100STYLE22.143.3×84.6×237×
BEAT62.633.7×121×1,109×
Geno ZeroEGGS2.233.5×85.9×403×
MotionPersona39.430.9×114×312×
for_elise (hands)23.730.4×78.9×299×
HiPHI308.729.5×87.2×384×
Motion-X86.729.5×55.6×128×
ZeroEGGS2.229.0×70.3×342×
in-house collection51.128.7×63.0×228×
Xia0.228.2×59.2×137×
wild game assets2.128.1×45.5×88.9×
AMASS27.527.5×51.1×124×
multi-subject2.327.4×59.8×234×
ZeroEGGS 65-joint2.225.5×60.7×288×
Geno InterAct single9.323.9×69.7×391×
Geno InterAct multi9.323.9×69.8×391×
Geno LAFAN14.523.7×47.9×138×
CMU9.123.1×55.8×272×
Geno Motorica6.222.1×41.7×110×
Mixamo0.121.4×36.4×80.7×
kid3.020.6×49.1×164×
InterAct 65-joint12.519.6×53.9×248×
InterAct12.518.8×53.3×308×
100STYLE22.118.6×37.3×112×
PFNN1.118.4×52.1×211×
AnimationGPT26.617.8×30.8×71.3×
Bandai3.915.3×31.9×85.8×
BFA2.214.4×33.9×196×
dog (held out)0.713.9×32.0×121×
Edinburgh0.213.9×29.0×152×
LAFAN14.610.8×18.4×50.4×
HumanAct121.28.7×13.0×24.0×

A1Compression ratio per dataset.

  1. Long, high-rate captures hold the most redundancy. BONES-SEED compresses 120× at 0.01 cm and 795× at 1 cm; BEAT, captured at 120 Hz, reaches 1,109× at 1 cm.
  2. A looser bound pays only where frames are dense. From 0.01 to 1 cm, BEAT's ratio grows 33-fold but HumanAct12's, at 20 Hz, only 2.8-fold: the encoder buys most of its savings by dropping samples, and sparse clips have few to drop.
  3. The same motion costs twice as much on another rig. Retargeted to the Geno skeleton, the same captures compress 2.2–2.3× better than on their original rigs: 100STYLE 18.6× → 43.3×, LAFAN1 10.8× → 23.7×. Part of what a joint curve carries is the body, not the motion.
  4. What the model learns generalizes to bodies it has never seen. The dog was held out of training entirely, yet it compresses 13.9× at 0.01 cm, on par with human captures such as Edinburgh. The entropy model has learned regularities of motion itself, not of the skeletons it was trained on.
Samples kept as keys, by depth in the skeleton
Max-error contract.
Table
joints0.01 cm0.1 cm1 cm
root rotation86.9 %50.6 %16.5 %
depth 1–285.0 %42.9 %13.9 %
depth 3–567.3 %27.9 %8.3 %
depth 6+56.7 %21.5 %7.2 %
root translation85.6 %44.9 %14.1 %
Keys kept, and how short the gaps are
Table
psamples keptgaps ≤ 2 samples
0.005 cm79.1 %92.0 %
0.01 cm72.4 %90.6 %
0.02 cm62.2 %89.0 %
0.05 cm50.5 %83.6 %
0.1 cm39.5 %77.3 %
0.3 cm25.5 %60.4 %
1.0 cm15.5 %37.3 %
The price of a bit at 0.1 cm
Mean-error contract; lower is cheaper.
Table
movedistortion per bit
step growth (1 grid unit)0.17
dead zone0.31
key removal (RD ladder)0.03

A2The hierarchy decides which samples to code.

A joint’s error reaches every descendant, so the root keeps the most samples and the leaves the fewest: 87 against 57 % at 0.01 cm, 51 against 22 % at 0.1 cm. ACL can only strip whole frames whose every joint is linearly interpolable, 0.38 % of the frames at 0.01 cm.

At the margin, leaving a sample out is the cheapest way to spend error: removing a key adds 5 to 10× less distortion per bit saved than coarsening a step. As p loosens the encoder keeps fewer keys, from 79 to 16 % of the samples, yet the gaps stay short: at p ≤ 0.1 cm three quarters of them hide at most two samples.

Depth: 5 development clips, an earlier encoder; keys kept: development set; gaps: curated set.

Bits per joint-sample through ACL’s stages
Table
stagebits / joint-sample
raw (10 floats)320.0
drop scale223.8
drop w191.3
constant / default fold67.7
range reduction + bit widths18.1
Entropy of ACL’s own symbols
Table
modelrelative to ACL
ACL’s own bits100.0 %
entropy, 0th order97.7 %
entropy of first differences83.8 %
entropy of second differences81.1 %
Laplace model80.6 %

A3ACL’s bits are not where the redundancy is.

Almost all of ACL’s 17.7× comes from two steps: folding sub-tracks that do not move (2.84×) and the two-level range reduction with variable bit widths (3.74×).

An entropy coder on top of ACL’s own symbols would save at most 16 to 20 %, even with second-order models. A much smaller stream needs a different decision about what to code, not a better back end.

Curated set of 70 clips.

Removing one component from the final codec
Held-out data.
Table
removedcodec0.01 cm0.1 cm1 cm
predictionmax gate, vs every-sample codec2.3402.9502.990
keysmax gate1.0761.3941.690
keysmean gate1.3171.8132.230
entropy modelmean gate1.0901.1041.074
transformer, as an MLPmean gate1.0521.0481.039
component thinningmean gate0.9941.0061.016

A4Prediction first, then the samples not coded.

Predicting each integer from its own past is the largest single saving at tight bounds, worth 2.3 to 3.0×. Keys come second and grow as the bound loosens: removing them costs 8 % at 0.01 cm and 69 % at 1 cm under the max-error contract, 32 % to 2.2× under the mean-error contract.

The learned entropy model adds 7 to 10 % on top, and the transformer 4 to 5 % over an MLP.

311-clip test subset covering all 33 datasets; a failing clip counts at ACL’s bytes.

The weakest clip against the whole test side
p = 0.01 cm.
Table
bytes relative to ACL
This clip (HumanAct12)192 samples · 20 Hz · 27 joints0.74–0.75×
Whole test side4,472 clips0.373×

A5Where it gains least.

The weakest case is a short, sparse clip: 192 samples at 20 Hz of a 27-joint skeleton. All codecs are visually exact, but CurveCodec 2 saves only a quarter of ACL’s 23.8 KB, against 0.37× on the whole test side.

The per-clip header and the cold start of the entropy model are amortized over few samples, and at 20 Hz few samples can be left out.

HumanAct12, held-out test side.

04Versions and BibTeX
  1. 2.0
    CurveCodec 2 current

    The same curve space, expressed with a more stable structure: comparable to ACL in both mean and worst-case error. Closed-loop quantization and rate–distortion-selected keys are verified through the skeleton, and a small learned entropy model, whose integer inference is bit-exact across platforms, codes what remains. Every decoded clip is checked against its error contract.

    Code ↗Paper soon
  2. 1.0
    CurveCodec · SIGGRAPH Asia 2026

    In this work we first found that curve space is general enough: one learned model codes the joint curves of any skeleton. It reconstructed each curve from sparse anchors with a learned prior and matched ACL's mean error, but its worst-case error was not stable enough and its encoding and decoding were less efficient, so version 2 replaces it.

    PDF ↓DOI ↗

CurveCodec 2

@article{shi2026codec2,
  title   = {CurveCodec 2: Skeleton-Agnostic Animation Compression
             with a Learned Entropy Model},
  author  = {Shi, Mingyi and Lin, Huancheng and Chen, Xuelin and Komura, Taku},
  year    = {2026}
}


@inproceedings{shi2026codec,
  title     = {Neural Codec for Skeletal Animation Compression},
  author    = {Shi, Mingyi and Lin, Huancheng and Chen, Xuelin and Komura, Taku},
  booktitle = {SIGGRAPH Asia 2026 Conference Papers (SA Conference Papers '26)},
  year      = {2026},
  month     = dec,
  address   = {Kuala Lumpur, Malaysia},
  publisher = {ACM},
  isbn      = {979-8-4007-2842-6},
  doi       = {10.1145/3829340.3842192}
}
05Acknowledgements

We thank Nicholas Frechette, the author of ACL, for detailed discussions of ACL's design goals and technical details. We thank Jun Xing, Tianshu Zhang and Zhixin Piao for the very early discussions on motion compression.

The characters of the demo are third-party assets: the animals come from the Truebones Zoo pack (truebones.com), the humans are the Geno character of the ZeroEGGS dataset (Ubisoft La Forge), the robots are Unitree G1, H1 and Go2 models (Unitree Robotics), the low-poly wolf is “Wolf rigged low poly” by 3DHaupt, and the Objaverse-XL examples are the Sketchfab models below, all under CC BY 4.0; the motion clips come from the datasets credited in the paper.

Sketchfab models shown in the demo (CC BY 4.0)
  • “Animated Hovering Flying Hummingbird Loop” by LasquetiSpice (Sketchfab)
  • “sparrow_upload” by faiyaz5yaz (Sketchfab)
  • “Night Sky Down Under” by Miguelangelo Rosario (Sketchfab)
  • “high poly bee modle” by shreebaghel72 (Sketchfab)
  • “Mantis Twitch Walk” by JeffFleetwood (Sketchfab)
  • “Animated Repticect” by DoubelFace (Sketchfab)
  • “Caterpillar Crawl” by michael l. (Sketchfab)
  • “Pangxie” by wsrttys (Sketchfab)
  • “Kraken v2” by lawtrigg (Sketchfab)
  • “Tortuga verde (Chelonia mydas)” by Innoceana (Sketchfab)
  • “Fish Swimming” by geniusrahman155 (Sketchfab)
  • “Cute Fish” by RickStikkelorum (Sketchfab)
  • “Whale666” by a0976623059 (Sketchfab)
  • “Animated Elephant Character” by bgilgen (Sketchfab)
  • “Snowman” by Horizon Studio (Sketchfab)
  • “American Bison” by Damco (Sketchfab)
  • “Cow NPC - Now free to download” by Owlish Media (Sketchfab)
  • “Low Poly wolf” by manoeldarochadeoliveira (Sketchfab)
  • “Fennec Fox Free” by Evil_Katz (Sketchfab)
  • “Roaring Stag ( deepdreamed )” by Miguelangelo Rosario (Sketchfab)
  • “Don't overlook the Hippo” by Miguelangelo Rosario (Sketchfab)
  • “Triceratops occultatum” by Miguelangelo Rosario (Sketchfab)
  • “Cute Sci-Fi Dragon” by hare_ware (Sketchfab)
  • “Flint Maw” by Spinnee (Sketchfab)
  • “Fire Elemental” by InaLaAtzu (Sketchfab)
  • “Monster Plant Enemy” by Jacqueline Sweeney (Sketchfab)
  • “Cactus1” by nathan.connell (Sketchfab)
  • “SkeletonBoss” by kennethcplace (Sketchfab)
  • “Robot Dinosaur Walking - First Mechanics Test” by Instinto Ideal Studio (Sketchfab)
  • “Ezaroid - The Chopping Killer Machine” by evilinvader (Sketchfab)
  • “Low poly mech walking” by WarlockStones (Sketchfab)
  • “Military Drone Low-Poly” by ToporEnterprise (Sketchfab)
  • “Robot Error” by wamala (Sketchfab)