CSC 487 deep-learning team project · Winter 2026

Deep Purring

Can you estimate a cat’s age from its meow?

  • PyTorch
  • librosa
  • SmallCNN
  • ResNet-18
  • SpecAugment
  • scikit-learn
Log-mel spectrogram of a senior cat vocalization
what the model actually sees when a cat meows
the short version:
74.8%
best held-out-cat accuracy
the audio experiment

a cat is the unit, not a clip

Deep Purring asks whether a vocalization contains enough signal to classify a cat as kitten, adult, or senior. The playful question exposed a serious evaluation problem: recordings from the same animal can make a random clip split look much better than real generalization.

The project is a team research prototype. My work covers the YAMNet baselines, the raw-audio spectrogram pipeline, SmallCNN/ResNet-18 comparison, and the cat-level leakage correction used across team notebooks.

what the pipeline explores

three representations of the same meow

01

Embedding baselines

YAMNet converts audio into embeddings; PCA, scaling, logistic regression, and SVM establish a non-deep-learning reference.

02

Raw-audio CNN

A custom dataset converts two-second clips into 64-bin log-mel spectrograms for a compact SmallCNN.

03

Transfer learning

An ImageNet ResNet-18 is adapted to one-channel spectrograms to test whether a larger visual backbone helps.

04

Inference demo

The notebook accepts a new WAV file, applies the same preprocessing, and returns age-group probabilities.

how it works

from WAV file to an honest held-out-cat score

  1. 01

    Identify the cat

    Keep cat_id attached to every recording so the split can separate animals, not only clips.

  2. 02

    Group split

    GroupShuffleSplit creates train/test sets with no cat appearing on both sides.

  3. 03

    Normalize audio

    Resample to 16 kHz and pad or crop each example to a two-second window.

  4. 04

    Represent the signal

    Generate a 64-bin log-mel spectrogram or use YAMNet embeddings for the baselines.

  5. 05

    Augment train only

    Time and frequency masks add controlled variation without contaminating evaluation.

  6. 06

    Compare models

    Evaluate logistic regression, SVM, SmallCNN, and ResNet-18 with accuracy and macro-F1.

  7. 07

    Repeat across seeds

    Multi-seed results expose the variance hidden by the best seed-42 run.

under the hood

the notebook stack in context

librosa

Loads, resamples, windows, and transforms raw waveforms into log-mel features.

PyTorch

Implements the dataset, augmentations, SmallCNN, adapted ResNet-18, training loop, and inference path.

scikit-learn

Provides grouped splitting, PCA/scaling, classical baselines, and confusion-matrix metrics.

Jupyter

Keeps audio inspection, experiments, visualizations, and demo inference together for a course research workflow.

representation

From sound to image

Each clip is resampled to 16 kHz, normalized to two seconds, and converted into a 64-bin log-mel spectrogram. Custom time and frequency masking adds controlled variation during training.

Log-mel spectrogram
from raw audio to a compact visual representation
evaluation

The split was leaking cats

The first clip-level split let different recordings from the same cat appear in both training and test sets. That made the model look better at recognizing age when it could partly recognize individuals.

I replaced it with cat-level GroupShuffleSplit across the team notebooks so every test animal was truly unseen.

comparison

A small CNN beat transfer learning

On the seed-42 held-out-cat split, SmallCNN reached 74.8% accuracy and 0.67 macro-F1. ResNet-18 reached 49.6% and 0.38.

Across seeds, SmallCNN averaged 64.4 ± 9.1% accuracy, making the dataset’s instability visible rather than hiding it behind the best run.

Feline age class counts
a small, uneven dataset made multi-seed reporting essential
my contribution

the notebooks and evaluation changes I owned

I solely authored the baseline and spectrogram-CNN notebooks and then propagated the cat-level split correction across the team experiments.

  • Built raw WAV → log-mel preprocessing and a custom PyTorch dataset.
  • Compared SmallCNN with one-channel ResNet-18 and classical YAMNet baselines.
  • Caught same-cat leakage in the initial clip split and changed the unit of evaluation.
  • Reported multi-seed mean and variance rather than presenting 74.8% as a stable universal score.
team experiment

parallel models, shared evaluation discipline

Teammates explored other model notebooks and one teammate connected my SmallCNN to a Flask demo branch. The shared value of my contribution was the reusable audio pipeline and the group-split correction, not ownership of every team model.

what happened

results, with context

74.8%

best accuracy

The strongest single grouped split, using SmallCNN.

0.67

seed-42 macro-F1

Macro-averaged across the three age classes on one grouped test split.

64.4 ± 9.1%

multi-seed accuracy

The variance is part of the result, not a footnote.

looking back...

The leakage discovery changed what the score meant. A larger cat-level dataset and repeated grouped cross-validation would matter more than adding a larger model. The technical story here stays centered on the notebooks, models, and evaluation changes I can explain most deeply.