MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial Geometry

Abstract

Existing 3D head avatar reconstruction methods adopt a two-stage process, relying on tracked FLAME meshes derived from facial landmarks, followed by Gaussian-based rendering. However, the inevitable misalignment between the meshes and target images often degrades rendering quality and suppresses fine details. In this paper, we propose MoGaFace, a novel framework that jointly refines geometry and texture during Gaussian rendering. To mitigate mesh–image misalignment, we introduce a Momentum-Guided Consistent Geometry module, which leverages a momentum-updated expression bank and an expression-aware correction mechanism to enforce temporal and multi-view consistency. In addition, we propose Latent Texture Attention, which encodes compact multi-view cues into head-aware representations, enabling geometry-aware texture refinement via integration into ellipsoidal Gaussian primitives. Extensive experiments show that MoGaFace achieves high-fidelity head avatar reconstruction and significantly improves novel-view synthesis quality, even under inaccurate mesh initialization and unconstrained real-world scenarios.

Overall Framework

Audiodynamic Image

Given multi-view images, MoGaFace initializes FLAME via tracking, refines expressions through a Momentum-Guided Geometry module for consistent and accurate fitting, and embeds 3D Gaussian textures using a Latent Texture Attention module that exploits multi-view texture cues.

Camera-Aware Settings

In the camera-aware setting, we employ 16 camera viewpoints calibrated with NeRSemble, designating the frontal viewpoint as the novel view.

Novel View Synthesis

Self-Reenactment

Cross-identity Reenactment

Camera-Free Settings

In the camera-free setting, we operate on self-recorded videos without calibrated cameras, estimating FLAME meshes from camera poses predicted in a monocular manner. The required camera parameters are obtained using a monocular method [ZBT22].