CurveCodecv0.2.0 Skeleton-agnostic animation compression with a learned entropy model
1The University of Hong Kong 2Adobe Research *Co-corresponding authors
CurveCodec 2 compresses skeletal animation from any rig into a compact, bit-exact bitstream. Every joint curve is quantized in closed loop against a stated error bound, the encoder chooses per joint which samples not to code, and a small learned entropy model codes what remains. Trained once on 886 hours of motion from 33 datasets, it spends about 2 bits per joint-sample at 0.1 cm, 100× less than raw float32, and transfers without retraining to a species it has never seen.
- Footprint
- 906 h of motion, 775 GB as BVH: 11.9 GB at 0.01 cm, 4.3 GB at 0.1 cm, 1.1 GB at 1 cm
- Input
- Any skeleton, any rig: per-joint rotations and translations (BVH)
- Precision
- 0.01 to 1 cm, set per clip; every decoded clip is verified against its error bound*
- Entropy model
- 108 K-parameter causal transformer (870 KB); integer inference, bit-exact across platforms
- Decode
- 2.1 s per million joint-samples on one CPU core, 0.76 s on four threads
Table
| p | CurveCodec 2 | ACL | ||
|---|---|---|---|---|
| size | mean error | size | mean error | |
| original | 9,596 MB as float32 (quaternion + translation per joint, 28 B), error 0 | |||
| 0.01 cm | 262 MB | 0.0033 cm | 702 MB | 0.0034 cm |
| 0.05 cm | 136 MB | 0.0166 cm | 509 MB | 0.0166 cm |
| 0.1 cm | 96 MB | 0.0332 cm | 433 MB | 0.0333 cm |
| 0.3 cm | 50 MB | 0.134 cm | 365 MB | 0.137 cm |
| 1 cm | 26 MB | 0.542 cm | 374 MB | 0.560 cm |
*Each clip's mean joint error is at most that of the reference codec (ACL) at the same p; clips that fail fall back, and the fallback is counted. Sizes are bits per joint-sample × 342.7 M joint-samples; “original” is the same clips as float32, a quaternion and translation per joint (28 B).
Q1Why study compression at first?
- Motion data is highly redundant. A clip is stored as every joint's local transform at every frame. But when a person performs an action, they rarely attend to how each joint gets from one place to the next: the trajectories are largely produced by a strong prior, the body itself, rather than being what the motion is about. Spending more effort and computation on generating these curves does little for understanding behaviour or action.
- Compression is a principled way to learn what matters. Embodied intelligence works the same way: an agent decides what to do, and its body, its morphology and its dynamics, decide most of how. A model of motion intelligence should spend its capacity on the decisions, not on the kinematics the embodiment already implies. Compression separates the two: what a good codec must still send is the decision; what it can drop is the body.
- We want a representation general enough to model the dynamics of motion. If the body is only the prior, the representation should not be tied to one body: it has to generalize across motions and across skeletons. Today's hierarchical skeleton representations bake a specific topology into the data, so every rig needs its own model and little of what is learned on one body carries over to another; they do not scale. CurveCodec 2 treats one model serves any rig, and transfers without retraining to a species it has never seen.
Q2Could it replace ACL?
- No. Many of our evaluations use ACL: it is the production reference for what an error bound means, and the only tool we could find that truly works on any skeleton, as we aim to. But the two codecs target entirely different goals. ACL is built for the best possible runtime behaviour: it is stateless, keeps the clip compressed in memory and decompresses only the poses a frame needs, and reads any sample at random with as little memory traffic as possible.
- Matching ACL's precision has a price. Our encoder corrects itself with error feedback through forward kinematics, and our stream is entropy-coded and decoded sequentially, once per clip at load time, without random access. It is still far faster than real time: one CPU core decodes about 8,700 frames per second of a 55-joint skeleton. But it is not a design that chases efficiency above everything else.
First and foremost, ACL is designed to be flexible and minimally invasive. Something easy to integrate into any C++ codebase. To that end I attempted to maintain a balance between decompression speed, memory footprint, latency, and compression speed. First, I think of what sort of codebase uses something like animation compression. By and large, animation decompression today happens on the CPU (and that was even more true 10 years ago when I began the project). Animation often is processed in two distinct steps:
- An animation graph/tree update pass (aka update pass) which advances time, reacts to gameplay changes (e.g. performs graph/tree topology changes through transitions), etc.
- A pose evaluation pass (aka evaluate pass) which samples clips, blends them, and performs any post-process type modifications (e.g. IK).
How data is accessed in both passes is quite different. In the update pass, we typically access two kinds of data:
- Data that is shared between multiple character instances (aka shared data). Data of this nature can comprise of things like graph/tree topology information, constant values a user has authored (e.g. play rate = 0.5), translation tables where an input is used to select among multiple outputs (e.g. for character type X, use sub-graph Foo), etc. This shared data is generally immutable at runtime and can be baked offline.
- Data that is unique per character instance (aka instance data). This would be actual node state, things like the current/previous evaluated type, previous condition values, etc. This instance data is generally read/write at runtime and is either instanced up front in bulk for the whole graph or on demand.
During the update pass, we are latency bound. Shared and instance data is, for the most part, randomly accessed and the processor will not be able to effectively hardware prefetch any meaningful amount. In this pass, it is thus critical to touch as little memory as you can to avoid evicting memory that was expensive to bring in (through a random access cache miss).
The evaluation pass is quite different, it is throughput bound because data is largely accessed in bulk and linearly. Things like decompression and pose blending touch a lot of memory but they do so with a predictable memory access pattern which means the processor can more easily hide latency and prefetch ahead. The nature of the code also means we can fill the CPU pipeline more easily as we perform the same work across many elements in loops. As a result of this, it is less important here to minimize how much memory is touched, but it can still matter.
Crucially, many animation runtimes use the same graph/tree topology structures for both passes (e.g. calling a different virtual function per node, one for each pass) while others use a different representation for the evaluate pass by building something akin to a GPU command buffer during the update pass (effectively linearizing the graph traversal). Which one your runtime uses matters because in the first instance, we will need randomly accessed instance (and perhaps shared) data while in the second form we might not (as the structure is linearized). Animation Blueprints in UE4 and UE5 use the same graph/tree topology while the new Unreal Animation Framework (UAF) is moving towards a command buffer approach.
ACL is designed to perform well in both forms of runtime. To that end, it aims to touch as little memory as possible to avoid evicting precious CPU cache data that might be needed (e.g. topology data). To achieve this, ACL is stateless and close in spirit to texture compression (e.g. BC7) where the compressed format is used as-is in memory without an intermediate unpacking step. It also means that nothing is cached and re-used frame to frame to again minimize the per clip instance memory footprint and avoid polluting the CPU cache. This ensures that once decompression is done, as little memory was evicted from the CPU cache as possible ensuring that the caller can still run at full speed, avoiding the need to reload the data it needed (that might have been evicted). This is especially important on mobile devices and Nintendo Switch where the CPU cache can be quite small and memory latency high.
That being said, we can evaluate the performance of animation compression along a few properties.
Memory footprint is always important because now more than ever, size is king when it comes to performance: the less memory you touch, the faster code tends to execute as memory is often the slowest thing happening. Memory is still precious on mobile/Switch as well even with modern games.
Compression speed is still important. Games today have more animations than ever before (recent games have over 100k clips for example). This means that importing new clips needs to be quick and predictable, cost wise. Bulk re-compression is uncommon, but it does happen. With good caching, this can be mitigated but care must still be taken. Long animations can sometimes take very long to compress. Old UE5 codecs could take up to 45-90 minutes to import/compress a 1-2 minute long cinematic (versus a few seconds with ACL).
Decompression speed is probably the most important part because animation generally runs on the CPU and animation is often one of the slowest parts of the frame on the CPU. It isn't uncommon for a character to use 3-20 clips to produce a final pose depending on how their locomotion is set up (e.g. motion matching versus blend spaces). This also makes decompression latency a critical aspect because gameplay often needs to wait for animation to finish (e.g. to update bounding boxes for spatial queries/ray casts and to update the physics simulation, and to run secondary motion). Games that run at 30 FPS will have more time to produce a frame but it means that it must react to the latest input more quickly as it happened a long time ago by the time the image gets on screen. Meanwhile a 60FPS or 120FPS game could take an extra frame to react to input and that might be fine (depending on the game, twitchy FPS games might need to react within the same frame still).
These points mostly apply to individual characters being animated, largely because they tend to be the most expensive to update. However, there are other things that can be animated that might not need the same speed/latency constraints (e.g. props in the environment, crowds). These might also lend themselves better to GPU evaluation.
ACL was thus meant to be a decent middleground. Something you can drop in and use quickly that would perform well enough to not be something you have to worry about too much afterwards. I think it has been quite successful in that regard (almost no one reaches out to me despite broadly being in use in the industry).
That should broadly answer your #1 point.
For #2, batch compression is a less relevant metric. Compressed animation data is generally cached in editor/offline and so while compression isn't unusual, it tends to occur for individual clips (e.g. on import, after a sync from p4). Batch decompression may be a useful metric, but for niche use cases. In my experience, it is very difficult on the CPU to batch animation decompression. Every single runtime I've seen produces a final pose per character one at a time and doesn't interleave the process (e.g. decompressing every first, then blending, etc). That might be more readily feasible for a fixed function animation runtime, something built for simple behavior (e.g. a crowd with a few nodes). These would be better suited to GPU evaluation as well. Bulk decompression would also require allocating up-front memory for the output pose and that can be expensive. It becomes a cpu/memory tradeoff: it is faster to bulk decompress, but it uses more memory.
For #3, I have seen neural networks that produce good animation fidelity, comparable to ACL, with decent decompression performance on the CPU, enough to be competitive (e.g. comparable to motion matching) and good memory footprint. This can be achieved with minimal latency when evaluated on the CPU. However, due to the model size, they tend to evict the entire CPU cache (sometimes both L1 and L2) which means they might slow down the caller (e.g. the evaluation pass). I haven't seen by how much yet. This might be okay if the animation runtime uses a command buffer approach (as I mentioned earlier). For this reason, I think the best performance measure you could get is of the entire evaluation pass, not just decompression, to capture the full impact on the CPU cache. You would have the same sort of behavior on the GPU with respect to the cache (I imagine).
A lossy stage that chooses which samples not to code, and a lossless stage where the network lives.


-
e(j,t) = max over axes of |T(t)·δe − T̂(t)·δe|
Error is measured through the skeleton. The shell error of joint j at sample t is the largest displacement of three virtual vertices δ = 3 cm along its axes, in object space after forward kinematics, so it contains the errors of all its ancestors.
-
accept a move iff ΔD / ΔR ≤ λ and the contract holds
The lossy stage trades error for bits. Removing a key or growing a quantization step is accepted only if the error it adds per bit it saves stays under λ, and the decoded subtree still meets its contract; λ rises until the error budget is spent.
-
bits ≈ Σ −log2 P(r_t | f_≤t), r_t = n_t − n̂_t
The network is the entropy model. Each integer is predicted from its own past; a small causal transformer gives the distribution of the residual from the curve's own decoded history, and rANS codes it at close to that many bits. The decoder runs the same model in lockstep.
Table
| dataset | hours | 0.01 cm | 0.1 cm | 1 cm |
|---|---|---|---|---|
| BONES-SEED | 144.2 | 120× | 307× | 795× |
| Geno 100STYLE | 22.1 | 43.3× | 84.6× | 237× |
| BEAT | 62.6 | 33.7× | 121× | 1,109× |
| Geno ZeroEGGS | 2.2 | 33.5× | 85.9× | 403× |
| MotionPersona | 39.4 | 30.9× | 114× | 312× |
| for_elise (hands) | 23.7 | 30.4× | 78.9× | 299× |
| HiPHI | 308.7 | 29.5× | 87.2× | 384× |
| Motion-X | 86.7 | 29.5× | 55.6× | 128× |
| ZeroEGGS | 2.2 | 29.0× | 70.3× | 342× |
| in-house collection | 51.1 | 28.7× | 63.0× | 228× |
| Xia | 0.2 | 28.2× | 59.2× | 137× |
| wild game assets | 2.1 | 28.1× | 45.5× | 88.9× |
| AMASS | 27.5 | 27.5× | 51.1× | 124× |
| multi-subject | 2.3 | 27.4× | 59.8× | 234× |
| ZeroEGGS 65-joint | 2.2 | 25.5× | 60.7× | 288× |
| Geno InterAct single | 9.3 | 23.9× | 69.7× | 391× |
| Geno InterAct multi | 9.3 | 23.9× | 69.8× | 391× |
| Geno LAFAN1 | 4.5 | 23.7× | 47.9× | 138× |
| CMU | 9.1 | 23.1× | 55.8× | 272× |
| Geno Motorica | 6.2 | 22.1× | 41.7× | 110× |
| Mixamo | 0.1 | 21.4× | 36.4× | 80.7× |
| kid | 3.0 | 20.6× | 49.1× | 164× |
| InterAct 65-joint | 12.5 | 19.6× | 53.9× | 248× |
| InterAct | 12.5 | 18.8× | 53.3× | 308× |
| 100STYLE | 22.1 | 18.6× | 37.3× | 112× |
| PFNN | 1.1 | 18.4× | 52.1× | 211× |
| AnimationGPT | 26.6 | 17.8× | 30.8× | 71.3× |
| Bandai | 3.9 | 15.3× | 31.9× | 85.8× |
| BFA | 2.2 | 14.4× | 33.9× | 196× |
| dog (held out) | 0.7 | 13.9× | 32.0× | 121× |
| Edinburgh | 0.2 | 13.9× | 29.0× | 152× |
| LAFAN1 | 4.6 | 10.8× | 18.4× | 50.4× |
| HumanAct12 | 1.2 | 8.7× | 13.0× | 24.0× |
A1Compression ratio per dataset.
- Long, high-rate captures hold the most redundancy. BONES-SEED compresses 120× at 0.01 cm and 795× at 1 cm; BEAT, captured at 120 Hz, reaches 1,109× at 1 cm.
- A looser bound pays only where frames are dense. From 0.01 to 1 cm, BEAT's ratio grows 33-fold but HumanAct12's, at 20 Hz, only 2.8-fold: the encoder buys most of its savings by dropping samples, and sparse clips have few to drop.
- The same motion costs twice as much on another rig. Retargeted to the Geno skeleton, the same captures compress 2.2–2.3× better than on their original rigs: 100STYLE 18.6× → 43.3×, LAFAN1 10.8× → 23.7×. Part of what a joint curve carries is the body, not the motion.
- What the model learns generalizes to bodies it has never seen. The dog was held out of training entirely, yet it compresses 13.9× at 0.01 cm, on par with human captures such as Edinburgh. The entropy model has learned regularities of motion itself, not of the skeletons it was trained on.
Table
| joints | 0.01 cm | 0.1 cm | 1 cm |
|---|---|---|---|
| root rotation | 86.9 % | 50.6 % | 16.5 % |
| depth 1–2 | 85.0 % | 42.9 % | 13.9 % |
| depth 3–5 | 67.3 % | 27.9 % | 8.3 % |
| depth 6+ | 56.7 % | 21.5 % | 7.2 % |
| root translation | 85.6 % | 44.9 % | 14.1 % |
Table
| p | samples kept | gaps ≤ 2 samples |
|---|---|---|
| 0.005 cm | 79.1 % | 92.0 % |
| 0.01 cm | 72.4 % | 90.6 % |
| 0.02 cm | 62.2 % | 89.0 % |
| 0.05 cm | 50.5 % | 83.6 % |
| 0.1 cm | 39.5 % | 77.3 % |
| 0.3 cm | 25.5 % | 60.4 % |
| 1.0 cm | 15.5 % | 37.3 % |
Table
| move | distortion per bit |
|---|---|
| step growth (1 grid unit) | 0.17 |
| dead zone | 0.31 |
| key removal (RD ladder) | 0.03 |
A2The hierarchy decides which samples to code.
A joint’s error reaches every descendant, so the root keeps the most samples and the leaves the fewest: 87 against 57 % at 0.01 cm, 51 against 22 % at 0.1 cm. ACL can only strip whole frames whose every joint is linearly interpolable, 0.38 % of the frames at 0.01 cm.
At the margin, leaving a sample out is the cheapest way to spend error: removing a key adds 5 to 10× less distortion per bit saved than coarsening a step. As p loosens the encoder keeps fewer keys, from 79 to 16 % of the samples, yet the gaps stay short: at p ≤ 0.1 cm three quarters of them hide at most two samples.
Depth: 5 development clips, an earlier encoder; keys kept: development set; gaps: curated set.
Table
| stage | bits / joint-sample |
|---|---|
| raw (10 floats) | 320.0 |
| drop scale | 223.8 |
| drop w | 191.3 |
| constant / default fold | 67.7 |
| range reduction + bit widths | 18.1 |
Table
| model | relative to ACL |
|---|---|
| ACL’s own bits | 100.0 % |
| entropy, 0th order | 97.7 % |
| entropy of first differences | 83.8 % |
| entropy of second differences | 81.1 % |
| Laplace model | 80.6 % |
A3ACL’s bits are not where the redundancy is.
Almost all of ACL’s 17.7× comes from two steps: folding sub-tracks that do not move (2.84×) and the two-level range reduction with variable bit widths (3.74×).
An entropy coder on top of ACL’s own symbols would save at most 16 to 20 %, even with second-order models. A much smaller stream needs a different decision about what to code, not a better back end.
Curated set of 70 clips.
Table
| removed | codec | 0.01 cm | 0.1 cm | 1 cm |
|---|---|---|---|---|
| prediction | max gate, vs every-sample codec | 2.340 | 2.950 | 2.990 |
| keys | max gate | 1.076 | 1.394 | 1.690 |
| keys | mean gate | 1.317 | 1.813 | 2.230 |
| entropy model | mean gate | 1.090 | 1.104 | 1.074 |
| transformer, as an MLP | mean gate | 1.052 | 1.048 | 1.039 |
| component thinning | mean gate | 0.994 | 1.006 | 1.016 |
A4Prediction first, then the samples not coded.
Predicting each integer from its own past is the largest single saving at tight bounds, worth 2.3 to 3.0×. Keys come second and grow as the bound loosens: removing them costs 8 % at 0.01 cm and 69 % at 1 cm under the max-error contract, 32 % to 2.2× under the mean-error contract.
The learned entropy model adds 7 to 10 % on top, and the transformer 4 to 5 % over an MLP.
311-clip test subset covering all 33 datasets; a failing clip counts at ACL’s bytes.
Table
| bytes relative to ACL | ||
|---|---|---|
| This clip (HumanAct12) | 192 samples · 20 Hz · 27 joints | 0.74–0.75× |
| Whole test side | 4,472 clips | 0.373× |
A5Where it gains least.
The weakest case is a short, sparse clip: 192 samples at 20 Hz of a 27-joint skeleton. All codecs are visually exact, but CurveCodec 2 saves only a quarter of ACL’s 23.8 KB, against 0.37× on the whole test side.
The per-clip header and the cold start of the entropy model are amortized over few samples, and at 20 Hz few samples can be left out.
HumanAct12, held-out test side.
-
2.0
CurveCodec 2 currentCode ↗Paper soon
The same curve space, expressed with a more stable structure: comparable to ACL in both mean and worst-case error. Closed-loop quantization and rate–distortion-selected keys are verified through the skeleton, and a small learned entropy model, whose integer inference is bit-exact across platforms, codes what remains. Every decoded clip is checked against its error contract.
-
1.0
CurveCodec · SIGGRAPH Asia 2026PDF ↓DOI ↗
In this work we first found that curve space is general enough: one learned model codes the joint curves of any skeleton. It reconstructed each curve from sparse anchors with a learned prior and matched ACL's mean error, but its worst-case error was not stable enough and its encoding and decoding were less efficient, so version 2 replaces it.
CurveCodec 2
@article{shi2026codec2,
title = {CurveCodec 2: Skeleton-Agnostic Animation Compression
with a Learned Entropy Model},
author = {Shi, Mingyi and Lin, Huancheng and Chen, Xuelin and Komura, Taku},
year = {2026}
}
@inproceedings{shi2026codec,
title = {Neural Codec for Skeletal Animation Compression},
author = {Shi, Mingyi and Lin, Huancheng and Chen, Xuelin and Komura, Taku},
booktitle = {SIGGRAPH Asia 2026 Conference Papers (SA Conference Papers '26)},
year = {2026},
month = dec,
address = {Kuala Lumpur, Malaysia},
publisher = {ACM},
isbn = {979-8-4007-2842-6},
doi = {10.1145/3829340.3842192}
}
We thank Nicholas Frechette, the author of ACL, for detailed discussions of ACL's design goals and technical details. We thank Jun Xing, Tianshu Zhang and Zhixin Piao for the very early discussions on motion compression.
The characters of the demo are third-party assets: the animals come from the Truebones Zoo pack (truebones.com), the humans are the Geno character of the ZeroEGGS dataset (Ubisoft La Forge), the robots are Unitree G1, H1 and Go2 models (Unitree Robotics), the low-poly wolf is “Wolf rigged low poly” by 3DHaupt, and the Objaverse-XL examples are the Sketchfab models below, all under CC BY 4.0; the motion clips come from the datasets credited in the paper.
Sketchfab models shown in the demo (CC BY 4.0)
- “Animated Hovering Flying Hummingbird Loop” by LasquetiSpice (Sketchfab)
- “sparrow_upload” by faiyaz5yaz (Sketchfab)
- “Night Sky Down Under” by Miguelangelo Rosario (Sketchfab)
- “high poly bee modle” by shreebaghel72 (Sketchfab)
- “Mantis Twitch Walk” by JeffFleetwood (Sketchfab)
- “Animated Repticect” by DoubelFace (Sketchfab)
- “Caterpillar Crawl” by michael l. (Sketchfab)
- “Pangxie” by wsrttys (Sketchfab)
- “Kraken v2” by lawtrigg (Sketchfab)
- “Tortuga verde (Chelonia mydas)” by Innoceana (Sketchfab)
- “Fish Swimming” by geniusrahman155 (Sketchfab)
- “Cute Fish” by RickStikkelorum (Sketchfab)
- “Whale666” by a0976623059 (Sketchfab)
- “Animated Elephant Character” by bgilgen (Sketchfab)
- “Snowman” by Horizon Studio (Sketchfab)
- “American Bison” by Damco (Sketchfab)
- “Cow NPC - Now free to download” by Owlish Media (Sketchfab)
- “Low Poly wolf” by manoeldarochadeoliveira (Sketchfab)
- “Fennec Fox Free” by Evil_Katz (Sketchfab)
- “Roaring Stag ( deepdreamed )” by Miguelangelo Rosario (Sketchfab)
- “Don't overlook the Hippo” by Miguelangelo Rosario (Sketchfab)
- “Triceratops occultatum” by Miguelangelo Rosario (Sketchfab)
- “Cute Sci-Fi Dragon” by hare_ware (Sketchfab)
- “Flint Maw” by Spinnee (Sketchfab)
- “Fire Elemental” by InaLaAtzu (Sketchfab)
- “Monster Plant Enemy” by Jacqueline Sweeney (Sketchfab)
- “Cactus1” by nathan.connell (Sketchfab)
- “SkeletonBoss” by kennethcplace (Sketchfab)
- “Robot Dinosaur Walking - First Mechanics Test” by Instinto Ideal Studio (Sketchfab)
- “Ezaroid - The Chopping Killer Machine” by evilinvader (Sketchfab)
- “Low poly mech walking” by WarlockStones (Sketchfab)
- “Military Drone Low-Poly” by ToporEnterprise (Sketchfab)
- “Robot Error” by wamala (Sketchfab)