Autodesk
Mila Concordia University

Controllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time Guidance

1 Autodesk Research; 2 Mila – Quebec AI Institute; 3 Concordia University, Montreal, Canada
*Corresponding Author
Neutral dance motion and progressively exaggerated outputs

Our approach produces controllable and plausible motion exaggerations from a neutral reference. Increasing the exaggeration scale progressively amplifies the motion while preserving its underlying action intent. We introduce complementary training-time and inference-time strategies that transfer across diffusion and flow-matching motion generators.

Video Presentation

Abstract

Recent motion generative models synthesize physically plausible character motion, but often overlook the animation principles that professionals use to make characters feel alive, expressive, and engaging. We focus on exaggeration and study how to incorporate it into modern text-to-motion pipelines. We introduce a two-stage framework. At training time, we fine-tune pre-trained motion models on a curated, graded exaggeration dataset. At inference time, we formulate exaggeration with dynamic movement primitives (DMPs) and use that formulation to guide diffusion and flow-matching models without additional training. Across MDM, KIMODO, and HY-Motion, our methods increase expressiveness while preserving the neutral motion's intent and physical plausibility.

What do we aim to solve?

Motion foundation models have made major progress in generating realistic human motion from text. Yet realism alone is not enough for compelling character animation. Professional animators intentionally amplify poses, timing, and trajectories to clarify an action and make it more expressive. Existing generative models do not provide a reliable way to control this exaggeration while keeping the original action recognizable and the resulting motion physically believable.

Baseline and exaggerated motions across multiple motion generators

We therefore ask: Can a generative motion model amplify expression in a controllable manner while preserving action intent and physical plausibility? Our work addresses this question with two complementary routes: learning exaggeration as a text-controlled attribute during fine-tuning, and introducing exaggeration at inference time while keeping the base motion model frozen.

Problem Statement and Preliminaries

Given a neutral motion generated from a text prompt, our objective is to produce a family of motions with increasing expressive intensity. A successful method should satisfy three properties: the exaggeration level should be controllable, the semantic identity of the action should remain unchanged, and the resulting sequence should remain physically plausible.

Overview of training-time and inference-time motion exaggeration

Dynamic movement primitives provide a useful representation for inference-time control. A DMP decomposes movement into an attractor dynamics term and a learned forcing term that describes the characteristic shape of the trajectory. Scaling the forcing term amplifies that trajectory while preserving the interval boundaries. We use the resulting exaggerated motion as a target that guides the sampling process of a pre-trained generative model.

Methodology

Training-Time Adaptation

We introduce ExMoCap, a curated motion-capture dataset containing 26 actions performed at five intensity levels, from neutral to extremely exaggerated. We retarget and augment these performances for each model's motion representation, then fine-tune pre-trained generators using LoRA and prior preservation. Text prompts describe the requested exaggeration level, allowing the adapted model to treat expressive intensity as a native, controllable attribute.

Text-controlled exaggeration after supervised fine-tuning

The training objective balances exaggeration control with preservation of the pre-trained model's broad motion prior. Prior preservation discourages the adapted model from drifting away from plausible human motion or losing alignment with the input prompt.

Inference-Time Guidance

Our training-free approach first generates a neutral reference motion. The sequence is divided into intervals between key poses, and a DMP is fitted to each interval. Scaling the DMP forcing term produces an exaggerated target whose intensity is controlled continuously by a scalar parameter. During diffusion or flow-matching sampling, gradient-based guidance pulls the generated sequence toward this target while the frozen generative prior regularizes the result.

Training-free DMP-guided controllable exaggeration

This formulation requires no additional model training and applies to both diffusion and flow-matching backbones. The guidance strength and exaggeration scale expose a practical trade-off between stronger motion amplification and closer preservation of the neutral action.

Comparative Results

We evaluate controllability, physical amplification, action preservation, and perceptual quality across MDM, KIMODO, and HY-Motion. The experiments compare supervised fine-tuning, training-free DMP guidance, and their combination with the corresponding pre-trained baselines.

Qualitative Comparison Across Backbones

Qualitative results for MDM and HY-Motion

Across diffusion and flow-matching models, our framework increases pose amplitude and characteristic movement while retaining the action and its temporal progression. The examples illustrate that the same control strategy transfers across different model families and motion representations.

Supervised and Training-Free Exaggeration

KIMODO supervised and training-free exaggeration results

Supervised fine-tuning provides the strongest direct control over requested exaggeration levels. Training-free guidance offers continuous control without updating the model and restores plausibility compared with unconstrained DMP amplification. Combining both strategies yields the strongest physical exaggeration, with a larger departure from the neutral action.

Controllability and Physical-Intensity Ablations

Classifier-based exaggeration control ablation

The classifier-based evaluation measures whether generated motions follow the requested exaggeration progression. Increasing LoRA capacity strengthens exaggeration control, while prior preservation protects prompt adherence and the quality of the learned motion distribution.

Physical motion intensity ablation

Root-relative range of motion quantifies the magnitude of physical amplification. The selected operating points balance stronger exaggeration with text-motion alignment and preservation of the source action.

User Study

We conducted a perceptual study with 20 participants across 320 evaluations. Participants identified the most exaggerated and most neutral motions, then assessed whether exaggerated outputs preserved the source action. The highest requested level was selected as the most exaggerated in 81.8% of comparisons, and 76.7% of the action-preservation ratings were positive overall.

Representative questions and findings from the perceptual study

Acknowledgements

This research was supported by Autodesk Research and the Autodesk AI Lab. We thank Noa Kaplan, Larasika Nadela, and Lucie Taglienti for their contributions to the dataset, and Marc Chu, Robert Helms, Helen Lu, Evan Atherton, Shu Ishida, and Jenmy Zhang for valuable discussions and support throughout this work.

If you find this study useful in your work, we kindly ask that you cite it using the following format.

BibTeX

@inproceedings{zamani2026controllable,
  title={Controllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time Guidance},
  author={Zamani, Amirhossein and Rampini, Arianna and Roy, Bruno},
  booktitle={arXiv},
  year={2026},
  url={https://arxiv.org/pdf/2610.12316}
}