Facial Keypoints as a Generative Modeling Problem

I've been exploring facial landmark detection as an image-conditioned generative modeling problem. Instead of producing coordinates in one pass, the prototype starts with random coordinates and updates them toward a face layout using image features as conditioning.

The recovered project contains two approaches: diffusion with noise prediction and flow matching with velocity prediction. Both operate on landmark coordinates rather than generating image pixels.

The source is available on GitHub: Generative Facial Keypoints. It includes the model backbones, dataset loaders, training configuration, and setup instructions. The published copy replaces private dataset paths, restores the missing diffusion forward return, and defaults to cropped inputs with DINOv2-base. Datasets and trained checkpoints are not included, so it is not a complete reproduction package for the historical results below.

Originally dated January 5, 2026; updated October 6, 2026 with recovered experiment artifacts, a detailed architecture diagram, and a new DiT inference animation. The January 4 validation results and January 8 DiT checkpoint are separate experiments. Implementation details describe the source recovered in October, which includes uncommitted changes, rather than a verified snapshot of either training run.

DiT sampling in action

Actual DiT flow-matching inference on four cropped face inputs, showing generated landmarks moving from Gaussian noise through 25 Euler updates.

October 2026 inference update: this animation was generated from the retained January 8 DiT checkpoint at training step 42,500, using frozen DINOv2-base features, random seed 17, and 25 Euler steps. It shows actual intermediate coordinates for four cropped 300-W sample images, not interpolation between known endpoints. Only generated landmarks are drawn. Playback timing is chosen for presentation and does not represent inference speed; the new run used CPU because the server's GPU runtime was unavailable.

This is flow-matching inference: the model predicts velocity and advances coordinates from noise toward a landmark layout. It is distinct from the diffusion branch's DDIM denoising. The January 8 checkpoint strictly matches the current six-block, 384-wide DiT implementation, and its saved validation summaries show cropped face inputs. These samples illustrate behavior and are not a new held-out accuracy benchmark.

The four inputs were selected at fixed indices from the 300-W loader using the original committed crop implementation: landmark bounding box, 20% margin, clipping to image bounds, then resizing to 224 × 224. Their membership in the checkpoint's original training or validation partition was not verified. These crops use ground-truth annotations, so this is not an end-to-end face-detection demonstration; an application would need a detector or another source of face bounds. The inference script loaded DINOv2-base explicitly from the local cache; this confirms the encoder used for the animation, not the encoder's provenance during the original training run.

Architecture: image features guide coordinate updates

Detailed model architecture comparing the four-layer transformer decoder for diffusion with the six-block DiT-style velocity model for flow matching, including coordinate embeddings, frozen DINO image memory, time conditioning, self-attention, cross-attention, and output heads.

The diagram follows the recovered source: the left branch adds time embeddings to landmark tokens before transformer decoding; the right branch uses time-conditioned normalization and gated residuals inside each DiT-style block. Both attend to projected image patches. The January 4 checkpoint uses the earlier decoder structure, while the January 8 checkpoint used in the animation matches the DiT-style structure. The diagram's full-image padding label describes the current loader; the animation instead uses the original face-crop loader. The January 4 training objective remains uncertain, as discussed with its results below.

The image branch produces spatial patch features with a frozen DINO encoder. The coordinate branch embeds each two-dimensional landmark, adds a learned landmark-identity embedding, and incorporates time. Cross-attention lets coordinate tokens consult the image features, while self-attention lets landmarks interact with one another.

The output is a two-dimensional vector for every landmark. Its meaning depends on the training objective: diffusion predicts noise, while flow matching predicts velocity. A sampler turns those predictions into successive coordinate updates.

The extractor attempts to load a DINOv3 ViT-S/16 checkpoint and falls back to DINOv2-base on loading failure. That fallback changes the conditioning model, so the actual loaded encoder must be recorded with any future experiment. The training code infers feature dimensionality from a dummy forward pass.

Preparing images and landmarks

The dataset loaders combine still images from 300-W and annotated frames from 300-VW, parsing coordinates from .pts files. The dataset organizers provide the background for these benchmarks: 300-W and 300-VW.

The current preprocessing uses the full image, pads it to a square with black pixels, and resizes it to 224 × 224. Coordinates receive the same translation and scaling. An optional landmark-based crop exists, but the default path uses the whole image.

The training collator maps pixel coordinates into approximately [-1, 1]:

normalized_keypoints = (pixel_keypoints / 224.0) * 2.0 - 1.0

Coordinates and images must stay aligned through every transformation. A good first diagnostic is to draw the transformed ground truth on the resized image before training anything.

Diffusion: learn the noise added to a landmark set

The diffusion branch adds Gaussian noise at a randomly chosen timestep:

$$ x_t = \sqrt{\bar\alpha_t},x_{\mathrm{clean}} + \sqrt{1-\bar\alpha_t},\epsilon. $$

The transformer receives the noisy coordinates, timestep, and image features. Its objective is mean squared error between predicted and actual noise.

The recovered diffusion configuration uses 1,000 timesteps, and the process implements a linear beta schedule. Its evaluation path calls a 50-step DDIM sampler rather than the full ancestral sampler. DDIM provides the underlying approach to sampling with fewer iterative steps.

The diffusion backbone is a four-layer transformer decoder with hidden width 256 and four attention heads. However, the current DiffusionTransformer.forward computes pred without returning it. That branch needs repair before it can produce a usable loss. Its presence in the repository is evidence of an implemented design, not a currently working training path.

Flow matching: learn a velocity field

The default configuration selects flow matching with a DiT-style backbone. During training, it pairs Gaussian noise x_0 with ground-truth coordinates x_1 and constructs a straight interpolation:

$$ x_t = (1-t)x_0 + tx_1, \qquad v_{\mathrm{target}} = x_1-x_0. $$

The network predicts that velocity given the intermediate coordinates, time, and image features. This coordinate-space experiment uses the velocity-regression idea described in Flow Matching for Generative Modeling.

Four stages of a straight interpolation from noise to a hand-designed facial landmark layout. This is a schematic, not a trained prediction.

This figure is explanatory: it uses hand-designed landmarks and Gaussian noise. It shows the training interpolation, not a learned sampling trajectory or an experimental result.

At inference, the sampler starts from Gaussian coordinates and applies 25 Euler steps:

x = torch.randn(shape, device=device)
dt = 1.0 / steps
for i in range(steps):
    time_index = torch.full(
        (shape[0],), int((i / steps) * 1000),
        device=device, dtype=torch.long,
    )
    velocity = model(x, time_index, image_features)
    x = x + dt * velocity

This sketch assumes the model is in evaluation mode and sampling runs without gradients. It maps time to a batch of integer indices to reuse sinusoidal time embeddings, matching the approach used for the GIF.

The flow backbone uses six blocks, hidden width 384, and six attention heads. Time modulates the self-attention and MLP branches through AdaLN-style shifts, scales, and gates. Image cross-attention is a separate residual branch without that gate. The final coordinate projection is zero-initialized.

Recovered training configuration and evaluation

These are defaults from the recovered source. They are not independently verified hyperparameters for the January 4 or January 8 checkpoint; saved artifacts and nearby configurations do not establish a complete training manifest.

Setting Recovered default
Default objective Flow matching
Batch size 32
Optimizer Adam
Learning rate 0.0001
Epochs 10
Evaluation/checkpoint interval 250 training steps
Learning-rate schedule Cosine decay with warmup
Warmup 5% of steps, capped at 100
Train/validation split 90% / 10%, random seed 42
Feature extraction during training Frozen encoder, BF16 autocast

Evaluation accumulates squared coordinate error and divides by the number of images. Despite the eval/mse label, this is summed coordinate squared error per image, not mean error per coordinate or a standard facial-landmark normalized mean error. Its scale depends on the number of landmarks.

The visualization logger overlays ground truth and sampled landmarks on the first four validation images. The saved TensorBoard image summaries preserve those overlays, making it possible to inspect the earlier experiment without rerunning inference.

Earlier transformer-decoder validation results

The retained January 4 run contains 246 image summaries and corresponding validation measurements, from step 250 through step 61,500. These results belong to the earlier transformer-decoder checkpoint, not the DiT animation above. The images below are extracted directly from those summaries. Red marks are ground truth; green marks are sampled predictions, following the logger's drawing convention.

Step 250

Saved validation overlays at training step 250, with red ground-truth landmarks and green sampled predictions.

Step 10,000

Saved validation overlays at training step 10000.

Step 61,500

Saved validation overlays at training step 61500, showing sampled landmarks close to many ground-truth positions.

These are the four examples chosen by the existing validation logger, not a curated best-case selection. Many late-run predictions visually approach the annotations, but these examples do not establish accuracy across poses, occlusion, identities, or unseen videos.

Validation coordinate error from the retained January 4 TensorBoard scalar summaries, plotted on a logarithmic vertical axis.

The logged error falls from approximately 1.7705 at step 250 to 0.0370 at step 61,500. These are the trainer's summed coordinate squared errors per image, not normalized mean landmark error. The existing sample-level split also limits what the improvement says about generalization.

There is a provenance mismatch worth preserving: a nearby saved Hydra configuration selects flow matching, but the run's checkpoint contains the earlier 256-wide transformer decoder rather than the current 384-wide DiT backbone. Checkpoint structure identifies the backbone but does not identify its training objective. I therefore label these as results from the retained run, rather than claiming they establish a diffusion-versus-flow comparison. The current source's missing diffusion return also cannot be used to infer that earlier training failed: the saved artifacts show that an earlier implementation did run.

A retained CSV contains episode, reward, and error columns from an earlier experiment. Its schema does not match the current trainer, and its model provenance is unclear. I have excluded those numbers rather than attaching them to diffusion or flow matching.

What needs validation next

The source offers a useful starting point, but several checks come before a credible comparison:

  • Split by video or subject. The current random sample split can place neighboring frames from one video in both partitions. It cannot establish generalization to unseen sequences.
  • Validate the encoder contract. The extractor removes one leading token and assumes the remaining sequence forms a square patch grid. Special/register tokens, feature layout, and the selected checkpoint need explicit checks.
  • Repair and test diffusion. Add the missing return and check the DDIM terminal update: the current schedule ends at diffusion index zero instead of explicitly stepping to the clean-data boundary. Its eta parameter is also ignored in favor of deterministic updates.
  • Establish evaluation units. Report a defined normalized landmark error and compare against a direct coordinate-regression baseline under the same split and preprocessing.
  • Measure sampling variability and cost. Multiple random starts may produce different landmarks. Accuracy, variation, and runtime should be measured together; the configured 25 versus 50 steps alone do not establish a speed or quality advantage.

The appeal of this project is the compact output space: a generative process over structured coordinates, conditioned on a rich image representation. The recovered validation overlays and new DiT sampling animation show that idea in action. The next step is a reproducible comparison with sequence-aware splits, complete training manifests, defined accuracy metrics, and measured inference cost.