Phone Predictor

AboutJune 2025

14· Machine Learning, MLOps

A full MLOps loop around a small classifier: three candidate models retrain on every data update, MLflow picks the winner, and GitHub Actions containerizes and deploys it, with quality gates in CI.

mlopsacademic project

Source ↗

MLOps workflowArchitecture
Data comes in, three candidate models train in parallel, MLflow picks the best one, and GitHub Actions ships it to Railway, monitoring feeding back into the next deploy.
MLflow experimentsTraining
Random Forest, SVM, and XGBoost logged side by side, so the best-performing run is a comparison away, not a guess.
CI/CD runPipeline
A clean GitHub Actions run: pre-build, model retraining and code tests, then build-and-deploy, all green.
UI, before and afterFrontend
The prediction form early on versus the final version, restyled once the pipeline behind it was actually working end to end.
Write-up

Phone Predictor is a course project (Pengembangan Sistem dan Operasi, ITS) that classifies a phone into one of four price tiers, budget, mid-range, high mid-range or flagship, from its hardware specs. Three candidate models (Random Forest, SVM, XGBoost) train on every data update, MLflow tracks them and picks the best performer, and GitHub Actions builds and pushes a Docker image to Railway, wired back to Railway's own monitoring so the loop stays observable. The shipped model lands around 81% accuracy.

Most of the actual engineering was fixing what broke along the way, and the most instructive break was a model that had quietly collapsed to predicting one class no matter the input. Accuracy still looked plausible while the model was useless. The fix was rebalancing the training set with SMOTE and grading it on F1 instead of raw accuracy, nothing to do with the algorithm at all. Alongside that: an MLflow tracking bug traced to a missing mlartifacts/mlruns directory, a GitHub Actions permissions error blocking automated commits, and a SonarQube coverage and duplication failure traced to the test file itself.

What the retraining loop buys is the difference between a model that was accurate once and a model that stays accurate, which is the part most small projects skip. Built with four people, which is also where the CI permissions and code-quality problems came from.

Things to underline
  • Automated the full MLOps loop: GitHub Actions retrains RandomForest/SVM/XGBoost on every data update, MLflow tracks and picks the best one, and it's containerized and deployed to Railway
  • Fixed a model that had collapsed to predicting only one class no matter the input, by rebalancing the training data (SMOTE) and switching the metric it was graded on from raw accuracy to F1 score
  • Diagnosed and fixed two separate CI bugs: an MLflow tracking failure from a missing artifacts directory, and a GitHub Actions permissions error blocking automated commits
  • Gated code quality with SonarQube, resolving a failing coverage/duplication check by refactoring the test suite
Built with
Pythonscikit-learnMLflowDocker