Drug Classification

AboutApril 2025

13· Machine Learning, MLOpsNot recently updated

A machine learning pipeline that predicts a patient's drug prescription from their medical profile, trained, evaluated, serialized and served end to end, so the model answering requests is always the one the published numbers describe.

mlops

Source ↗

Confusion matrixEvaluation
Predicted vs. true drug class on the test set, 95% accuracy overall with most of the error concentrated between drugC and drugX.
Interactive prediction demoDemo
Gradio UI for the deployed model: a 30-year-old male with low blood pressure, normal cholesterol, and a Na/K ratio of 13.5 is classified as drugX in real time.
Write-up

A classification pipeline that predicts a drug prescription (one of five classes) from a patient's medical profile: age, sex, blood pressure, cholesterol and sodium-to-potassium ratio. It reaches 95% accuracy and an 0.86 F1 score on the held-out test set, with only a handful of confusions between the closer drug classes. This is a small dataset and a solved problem, so the shape is worth more than the score.

The trained pipeline is serialized with skops and served through a FastAPI '/predict' endpoint. Choosing skops over raw pickle is a small decision with a real reason: loading a pickle executes arbitrary code, which is an unacceptable property for an artifact meant to be shared. GitHub Actions automates training and evaluation and regenerates the metrics and confusion matrix on every run, which is what keeps the served model and the published numbers the same model.

The confusion matrix is published alongside the accuracy on purpose. Clinical decision support is the wrong setting for 95% to sound like enough, and which classes get confused with which matters more here than the headline figure.

Things to underline
  • Trained a classifier on patient records (age, sex, blood pressure, cholesterol, sodium-to-potassium ratio) reaching 95% accuracy and an 0.86 F1 score
  • Served predictions through a FastAPI '/predict' endpoint, with the model serialized via skops rather than raw pickle
  • Automated training and evaluation through GitHub Actions, regenerating metrics and a confusion matrix on every run
Built with
Pythonscikit-learnFastAPIGradioGitHub Actions
GitHubdarrellathaya/drug-classification
  • Jupyter Notebook 92.3%
  • Python 5.8%
  • Makefile 1.9%
Commits
12
Status
Not recently updated