FlowMotion: Target-Predictive Conditional Flow Matching for Jitter-Reduced Text-Driven Human Motion Generation

1Universidade Federal do ABC, Santo André, Brazil

FlowMotion for smooth, high-fidelity 3D human motion generation

Abstract

Achieving high-fidelity and temporally smooth 3D human motion generation remains challenging, particularly within resource-constrained environments. We introduce FlowMotion, a novel method leveraging Conditional Flow Matching (CFM). FlowMotion incorporates a training objective within CFM that focuses on more accurately predicting target motion in 3D human motion generation, resulting in enhanced generation fidelity and temporal smoothness while maintaining the fast synthesis times characteristic of flow-matching-based methods. FlowMotion achieves state-of-the-art jitter performance, achieving the best jitter in the KIT dataset, the second-best jitter in the HumanML3D dataset, and a competitive FID value for both datasets. This combination provides robust and natural motion sequences, offering equilibrium between generation quality and temporal naturalness.

BibTeX


@article{CANALESCUBA2025104374,
  title     = {FlowMotion: Target-predictive conditional flow matching for Jitter-Reduced text-driven human motion generation},
  journal   = {Computers & Graphics},
  volume    = {132},
  pages     = {104374},
  year      = {2025},
  issn      = {0097-8493},
  doi       = {https://doi.org/10.1016/j.cag.2025.104374},
  url       = {https://www.sciencedirect.com/science/article/pii/S0097849325002158},
  author    = {Manolo {Canales Cuba} and Vinícius {do Carmo Melício} and João Paulo Gois},
  keywords  = {3D animation, Human motion synthesis, Flow matching},
  abstract  = {Achieving high-fidelity and temporally smooth 3D human motion generation remains challenging, particularly within resource-constrained environments. We introduce FlowMotion, a novel method leveraging Conditional Flow Matching (CFM). FlowMotion incorporates a training objective within CFM that focuses on more accurately predicting target motion in 3D human motion generation, resulting in enhanced generation fidelity and temporal smoothness while maintaining the fast synthesis times characteristic of flow-matching-based methods. FlowMotion achieves state-of-the-art jitter performance, achieving the best jitter in the KIT dataset, the second-best jitter in the HumanML3D dataset, and a competitive FID value for both datasets. This combination provides robust and natural motion sequences, offering equilibrium between generation quality and temporal naturalness.}
}