Existing 3D head avatar reconstruction methods adopt a two-stage process, relying on tracked FLAME meshes derived from facial landmarks, followed by Gaussian-based rendering. However, the inevitable misalignment between the meshes and target images often degrades rendering quality and suppresses fine details. In this paper, we propose MoGaFace, a novel framework that jointly refines geometry and texture during Gaussian rendering. To mitigate mesh–image misalignment, we introduce a Momentum-Guided Consistent Geometry module, which leverages a momentum-updated expression bank and an expression-aware correction mechanism to enforce temporal and multi-view consistency. In addition, we propose Latent Texture Attention, which encodes compact multi-view cues into head-aware representations, enabling geometry-aware texture refinement via integration into ellipsoidal Gaussian primitives. Extensive experiments show that MoGaFace achieves high-fidelity head avatar reconstruction and significantly improves novel-view synthesis quality, even under inaccurate mesh initialization and unconstrained real-world scenarios.