Project Page · 2026

DynaConTalk

Wavelet-Constrained Diffusion for Long-Form and
Controllable Holistic Co-Speech 3D Motion

Wavelet Motion Space Dynamic Condition Negotiation Matched-Noise Constraints
Paper Examples Code

Abstract

Holistic co-speech animation must coordinate body, hands, face, and root motion over long sequences while remaining editable. DynaConTalk performs diffusion in stationary wavelet transform (SWT) coefficient space, separating coarse posture evolution, gesture strokes, and fast expressive details into aligned frequency bands.

Matched-noise constraint injection re-noises clean history or sparse user constraints to the current diffusion level before injection. This turns long-window continuation and localized editing into in-distribution denoising rather than post-hoc stitching, while dynamic multimodal conditioning negotiates rhythm, acoustic, semantic, text, identity, and activity cues over time.

Method

Structure first

DynaConTalk combines an aligned frequency representation, frame-wise multimodal conditioning, and diffusion-compatible history and editing constraints.

Figure 1 architecture of DynaConTalk, showing multimodal inputs and stationary wavelet motion space, separate face and body diffusion streams, DynaCon conditioning, band-specific heads, inverse SWT, root trajectory, and holistic co-speech motion output.
RepresentationRaw 6D → aligned SWT frequency bands
ConditioningIndependent gates → proposal and consensus
ConstraintsRe-noised history and edits → clean motion

Representation

Stationary Wavelet Motion Space

Frequency-aligned motion decomposition.

SWT separates coarse posture, gesture strokes, and fast expressive detail without changing sequence length. Per-band normalization gives low-amplitude dynamics their own learning scale, while inverse SWT brings every band back into coherent motion.

Conditioning

Dynamic Gating & Consensus

Frame-wise multimodal condition negotiation.

DGN balances rhythm, acoustic, semantic, text, identity, and activity cues frame by frame. A proposal path introduces the useful signal; consensus corrects or reinforces it against the current motion state.

Controllable Generation

Matched-Noise Constraint Injection

Diffusion-compatible history and editing constraints.

Matched-noise injection re-noises history or sparse edits to the current diffusion step before insertion. Time, body-part, and wavelet-band masks localize control while the denoiser preserves the surrounding motion.

DynaConTalk

Paper and code

Paper and code links will be added after the arXiv release.

Paper · Coming soonCode · Coming soonBack to Top

Generated Examples