EE 428 computer-vision team project · Spring 2026

Deepfake Face Detection

Can a detector trained on one kind of fake face recognize another?

  • PyTorch
  • EfficientNet-B3
  • ResNet-50
  • Xception
  • OpenCV
  • FFT
Average frequency spectra for real and fake faces across two datasets
different generators leave different frequency fingerprints
the short version:
0.9861 → ≤ 0.62
in-domain AUC to cross-dataset AUC
the central research test

does a detector learn manipulation, or one generator's fingerprint?

The project compares fully synthetic StyleGAN3 faces with DFD identity-swap video frames. A model can separate real and fake inside either dataset while relying on artifacts that disappear as soon as the generation process changes.

Our four-person team built parallel model branches; my branch is an end-to-end PyTorch/Colab implementation and the source of the pipeline, cross-dataset experiments, and EDA described here. The final report synthesized team results.

the experiment design

in-domain success followed by a deliberate stress test

01

Two fake families

StyleGAN3 creates whole faces from noise; DFD blends identities into real video frames.

02

Three backbones

EfficientNet-B3, ResNet-50, and Xception run through the same two-stage transfer-learning loop.

03

Cross-dataset evaluation

Each detector is evaluated on a balanced sample from the other manipulation family.

04

Frequency investigation

Mean faces, RGB histograms, and 2D FFT spectra probe whether generator-specific artifacts help explain the model failure.

how it works

the research pipeline from Kaggle to distribution shift

  1. 01

    Acquire + cache

    Download Kaggle data once, persist paths/checkpoints on Drive, and copy balanced training subsets to local Colab storage.

  2. 02

    Extract frames

    Resume-capable OpenCV tooling samples DFD videos without redoing an interrupted 100k-frame job.

  3. 03

    Balance + split

    Cap DFD classes at 10,874 each and use fixed stratified 80/20 splits.

  4. 04

    Augment

    Crop, flip, color-jitter, grayscale, and JPEG-compress train images to simulate social-media degradation.

  5. 05

    Fine-tune in two stages

    Train the new head at 1e-3, then unfreeze the backbone at 2e-5 with AdamW, cosine decay, AMP, and gradient clipping.

  6. 06

    Evaluate in-domain

    Record accuracy, fake-class F1, ROC-AUC, confusion matrices, and training behavior.

  7. 07

    Cross the datasets

    Test the saved detector on the other fake family, then inspect FFT and image statistics for plausible explanations of the collapse.

under the hood

the research infrastructure

PyTorch + timm

Provides custom datasets, the shared classifier head, staged fine-tuning, AMP, and pretrained CNN backbones.

OpenCV

Extracts representative DFD frames with resume logic for interrupted sessions.

scikit-learn

Computes stratified splits and comparable accuracy, F1, AUC, ROC, and confusion outputs.

Colab + Drive

Free GPU compute plus persistent file lists, checkpoints, and JSON histories survive runtime resets.

matplotlib + seaborn

Turn samples, spectra, confidence distributions, and model results into the evidence used in the research story.

datasets

Two very different kinds of fake

StyleGAN3 synthesizes whole faces; DFD manipulates identities inside real video frames. Their artifacts, class balance, and visual statistics are not interchangeable.

StyleGAN3 sample grid
one distribution can look deceptively learnable
training

A resilient Colab pipeline

I built a roughly 50-cell pipeline for environment setup, cached data, frame extraction, loaders, staged fine-tuning, checkpoints, and 13 numbered experiments across EfficientNet-B3, ResNet-50, and Xception.

The design could resume after interrupted Colab sessions instead of losing an expensive run.

generalization

The flattering number was not the answer

EfficientNet-B3 reached 94.4% accuracy and 0.9861 AUC on StyleGAN3. Across datasets, every AUC stayed below 0.62; one SG3-to-DFD configuration produced F1 = 0.0000 because it called every image real.

The result is consistent with a detector relying on generator-specific cues rather than a general concept of manipulation.

Cross-dataset ROC curve
a near-chance curve tells a more important story than 94.4% accuracy
frequency domain

Looking for the fingerprint

I generated paired image grids, averaged faces, RGB histograms, and 2D FFT power spectra. The frequency analysis supports a plausible explanation: StyleGAN upsampling artifacts differ from the signals created by face blending.

FFT power spectrum comparison
frequency space helps explain why transfer failed
my contribution

the complete vertical slice on my branch

I authored the dataset loaders, model factory, training and evaluation scripts, frame extractor, roughly 50-cell experiment notebook, and separate EDA notebook on my unmerged branch.

  • Implemented EfficientNet-B3, ResNet-50, and Xception under one repeatable training interface.
  • Engineered around Colab disconnects with file-list caches, local copies, checkpoint guards, and saved histories.
  • Designed the bidirectional cross-dataset test instead of stopping at the strong in-domain number.
  • Built FFT/mean-face visual analysis to connect the observed failure with a plausible mechanism.
team research structure

parallel branches, one synthesized report

Each teammate developed an independent pipeline on a separate branch. The written report combined the team's model results, but the codebases were not merged; the page therefore separates the project-wide research question from the implementation on my branch.

what happened

results, with context

94.4%

StyleGAN3 accuracy

Excellent performance on familiar generator artifacts.

0.9861

in-domain AUC

The model separated classes well inside that dataset.

≤ 0.62

cross-dataset AUC

Transfer stayed close to chance across every direction.

looking back...

The failed transfer is the project’s most useful result. I would train across more generators and manipulation families, then evaluate domain generalization before treating any in-domain score as evidence of a robust detector.