Three-person CSC 466 applied-ML project · Fall 2025
Potiongram Recommender & Analytics
A fantasy streaming world for exploring recommendations, churn, and user personas.
- Python
- scikit-learn
- pandas
- SVD
- K-Means
- DuckDB

≈3×
Precision@2 improvement
recommend, retain, and understand a fictional streaming audience
Potiongram is a term-long applied-ML project built around a synthetic fantasy streaming service. The shared data supports three separate tasks: recommend content, predict subscriber churn, and turn sparse behavior into interpretable audience personas.
It is an offline course project with no deployed app or real business outcomes. The strongest case-study material is the experiment discipline: ablation, leakage correction, rejected features, and explicit metric tradeoffs.
three models, three different decisions
KNN recommendations
Content, language, genre, and TF-IDF signals are compared under Euclidean and cosine distance with Precision@2 and ranking metrics.
Churn prediction
A temporal snapshot labels near-future cancellations and compares Random Forest, Gradient Boosting, and Logistic Regression.
Behavioral personas
Sparse watch/rating interactions become SVD embeddings and then ten K-Means segments.
Product interpretation
Each metric maps to a decision: what to show, whom to contact, or how to understand a viewing segment.
the offline data stack
Parquet + pandas + DuckDB
Parquet stores the evolving synthetic snapshots, while pandas and in-process DuckDB load, join, and query them without a deployed database.
TF-IDF + KNN
Represents text/content attributes and tests cosine similarity against the weaker Euclidean baseline.
Random Forest
Provides the final recall-heavy churn classifier under the leakage-safe temporal validation.
SVD + K-Means
Compresses a 98.89%-sparse interaction matrix, then discovers balanced behavioral segments.
Custom evaluators
Compute recommender ranking metrics, classification metrics, and clustering silhouette/inertia for each task.
What should an adventurer watch next?
I iterated on the team’s KNN recommender through scaling, feature sets, distance metrics, and evaluation. Cosine similarity plus TF-IDF at k=15 reached Precision@2 of 0.1667 versus 0.055 for the Euclidean baseline.

Who may leave the kingdom?
I built the tracked churn pipeline with a temporal label window, compared three models, and chose Random Forest for recall rather than raw accuracy. A 0.40 threshold better matched the retention use case.
The defensible final temporal-validation F1 was 0.6874, not the higher number from a mismatched chart version. A content-diversity feature reduced F1 by 1.02%, so I left it out of the final model.

Thirty-two thousand users, ten segments
The committed persona outputs use 15-dimensional SVD embeddings and K-Means over 31,693 users. I produced and committed the result artifacts and translated the clusters into interpretable viewing-behavior profiles.
k=10 achieved the best tested silhouette score, 0.528, while the tested DBSCAN settings remained negative.

the slices I can trace most directly
My main contribution spans the iterative KNN ablation, the full tracked churn pipeline, and the analysis and communication of the committed persona outputs. The wider Week 5 recommender comparison is part of the team project, while the personal recommender result highlighted here comes from my earlier KNN experiments.
- Iterated scaling, feature sets, cosine/Euclidean distance, and k for the KNN recommendation study.
- Built the tracked churn labeling, feature, model-comparison, threshold, and visualization pipeline.
- Caught a 26% train vs. 46% validation churn-rate shift and moved to temporal validation.
- Produced persona outputs from SVD/K-Means experiments and kept DBSCAN's negative result in the rationale.
weekly questions inside one shared world
Three teammates worked through seven assignments against a common dataset. Writeups and results were collaborative, while code ownership varied by week; this page uses team language for the overall platform and first person only for verified contributions.
a few more pages from the process



results, with context
Precision@2
About three times the Euclidean KNN baseline.
churn F1
The final leakage-aware temporal validation result.
silhouette score
The strongest tested k=10 persona clustering.
The fun theme made three different analyses feel connected. Because the project was collaborative, this case study stays focused on the KNN, churn, and persona work I can explain most deeply.