Solo CSC 466 data-mining project · Fall 2025

Student Success Prediction

What predicts student dropout before and after college begins?

  • Python
  • pandas
  • scikit-learn
  • XGBoost
  • Jupyter
  • seaborn
Visualization showing 53.2 percent background and 46.8 percent academic predictive power
the finding that changed the direction of the project
the short version:
53 / 47
background vs. academic predictive lift
the research question

when does useful dropout signal become available?

A model using final-semester performance may predict dropout well but arrive too late to support a student. I treated admission and post-enrollment data as two different decision moments and measured what each stage added.

This was a solo, end-to-end analysis of 4,424 UCI student records, not a deployed intervention system. The proposed two-stage support framework is an interpretation of the evidence, not a product currently making decisions about students.

questions inside the analysis

one dataset, two usable moments

01

Admission view

Sixteen features available before semester grades ask how early a meaningful signal exists.

02

Full view

Twenty-four features add curricular progress and semester performance to measure the later lift.

03

Model comparison

A dummy baseline, logistic regression, random forest, gradient boosting, and XGBoost are compared under the same split and cross-validation protocol.

04

Intervention framing

The 53/47 lift decomposition translates model performance into two possible moments for student support.

how it works

the experiment from raw rows to an interpretable finding

  1. 01

    Audit 4,424 records

    Inspect 36 raw fields and the dropout, enrolled, and graduate target distribution.

  2. 02

    Engineer + group

    Create five performance features, collapse high-cardinality categories, and remove redundant signals.

  3. 03

    Freeze the split

    Use a stratified 80/20 split with 3,539 training and 885 held-out records. Fit preprocessing only on training data.

  4. 04

    Build two feature sets

    Separate admission-time information from the fuller post-enrollment view.

  5. 05

    Compare five models

    Use five-fold stratified CV, then evaluate the selected candidates on the untouched test set.

  6. 06

    Interrogate errors

    Read class-level recall, precision, confusion matrices, ROC behavior, and feature importance instead of relying on accuracy alone.

  7. 07

    Decompose the lift

    Compare both models with the majority baseline to obtain the 53.2% background / 46.8% academic contribution.

under the hood

why these tools and choices

pandas + NumPy

Handled audit, category cleanup, feature engineering, and reproducible train/test preparation.

scikit-learn

Provided preprocessing pipelines, stratified validation, baselines, and comparable metrics across candidate models.

XGBoost

Won on held-out performance with the smallest CV-to-test gap among the leading models.

Jupyter + seaborn

Kept the analysis traceable and turned the model comparison into presentation-ready evidence.

data

Two models for two moments

The UCI dataset contains 4,424 students, 36 raw fields, and three outcomes: dropout, still enrolled, or graduate. I engineered five columns from semester performance and grouped high-cardinality categories before modeling.

The admission model deliberately excludes every semester grade and curricular-unit count. The full model adds those academic signals, so the comparison has a real temporal meaning.

Model progression chart
the baseline, admission view, and full view side by side
modeling

Selecting for generalization

I compared a dummy baseline, logistic regression, random forest, gradient boosting, and XGBoost using a stratified train/test split and five-fold cross-validation.

XGBoost won by a small margin, and its 0.98-point CV-to-test gap was smaller than the runner-up’s 1.68-point gap.

finding

Where the predictive lift came from

The majority baseline scored 49.93%. The admission model rose to 64.18%, and the full model reached 76.72%.

Attributing those gains proportionally gives 53.2% of the total lift to background and admission-time factors and 46.8% to academic performance. Structural context mattered slightly more than grades.

53.2 and 46.8 predictive power split
background conditions carried slightly more signal than academic performance
negative results

What I chose not to keep

SMOTE made performance worse. Additional interaction features looked important but reduced held-out accuracy from 76.72% to 75.82%, suggesting overfitting on the 3,539-record training set.

Keeping those experiments in the final story mattered: the best model was simpler than several versions I tried along the way.

Final model feature importance chart
engineered success-rate features became the strongest academic signals
my contribution

a solo project from question to interpretation

I owned the data audit, feature sets, preprocessing, candidate models, experiments, evaluation, charts, whitepaper, and presentation.

  • Designed lifecycle-specific feature sets rather than one maximally predictive feature soup.
  • Selected for generalization gap as well as headline accuracy.
  • Kept the weaker enrolled-class result visible: the full model recalled only 36% of students still enrolled.
  • Rejected SMOTE and extra interaction features when held-out performance worsened.
what happened

results, with context

76%

dropout recall

The full model correctly identified 217 of 284 dropout cases.

77%

dropout precision

About three quarters of students flagged as dropout cases were correct.

59%

early recall

The admission-only model reached this before semester grades existed.

looking back...

The surprise was that the information available before classes began carried slightly more of the measured predictive lift than semester performance. I would next validate the approach on another institution and add careful explainability before anyone treated the output as an intervention tool. The two-stage framework remains a proposal, not a deployed system.