Posts

[Day 185] Using prefect as my orchestrator for my MLOps project

Image
 Hello :) Today is Day 185! A quick summary of today: creating a flow for model training creating a flow for uploading data to GCS Well, I decided to go with Prefect as my orchestrator. I have little experience using it and I liked it, so went with it.  All the code from today is on my repo . As seen from the pic, I ran jobs many times and have many successes and failures.  The first flow (as they are called in prefect) I created is for model training.  This is an example visualisation of the flow from one of the runs: It contains 4 tasks: read data (not from GCS, I need to make it read from GCS) split data get best params for the model from mlflow train the model, and produce reports on it The reports are quite nice as we can use markdown and python f-strings together to make them: Model report: MLflow info: Deploy instructions: Instead of going to mlflow to get info, I tried to make it a bit easier and clearer on how to get a model, and also showing basic performan...

[Day 184] Mlflow experiment tracking and trying out metaflow

Image
 Hello :) Today is Day 184! A quick summary of today: got various experiments ran and saved in my GCP mlflow server set up and tried out metaflow for MLOps orchestration Today I had to re-create my GCP VM because of a python version issue. So after starting a VM with python 3.10 I could use mlflow's most current logging functionality.  I ran a few experiements and the currently saved are: I will select a model solely based on recall - being able to detect the most fraud cases out of all fraud cases. Based on that, so far (after finding optimal params) the Balanced RF Classifier is the best -> recall being 0.92.  The most recent code and files are in my repo .  Today I did some refactoring to make things look nicer -> added a utils folder with some helper functions that I used a lot throughout notebooks. The current are log_model_performance and get_best_params which log a model run to mlflow, and then load the params of a model from mlflow.  Another t...

[Day 183] Failing to install Kubeflow, and setting up mlflow on GCP

Image
 Hello :) Today is Day 183! Today is actually past the half-way point in this blog journey. 182 days to go! A quick summary of today: setting up mlflow for experiment tracking in GCP Today I started looking at tools I can use for my MLOps zoomcamp project. I used Mage for the data engineering project so now for an orchestrator I want to try another tool. My first idea (which I had on my mind for a while) is Kubeflow. Tldr - it is so hard to set up locally and I could not find a guide that clearly explains how :(  I had this  1hr video in my 'to watch' for a bit and it goes over building an ML pipeline in kubeflow however it does not go over how to install from kubeflow from 0 to what is needed in the video. Then I started looking in kubeflow's documentation - I tried to set it up on GCP but I could not and actually was a bit adamant to use it because it explicitly said that the free credits on GCP do not cover kubeflow hosting costs. So ~ I started looking into local depl...

[Day 182] Learning about feature selection in fraud detection and finding a classifier model with low recall

Image
 Hello :) Today is Day 182! A quick summary of today: learning about IV, WoE, and finding a best model for an imbalanced insurance fraud imbalanced dataset The time has come to start thinking about the project for MLOps zoomcamp.  I was looking around for some interesting dataset related to PD (probability of default) or LGD (loss given default) or EAD (exposure at default), and I found this  notebook. Warning - it is fairly long. But inside I saw something that interested me - it talked about WoE and IV. It says that they are good estimators for evaluating features for fraud and similar classification tasks. This website's definition was the most clear. Weight of Evidence (WoE) It is a technique used in credit scoring and predictive modeling to assess the predictive power of independent variables relative to a dependent variable. Originating from the credit risk world, WoE measures the separation between "good" and "bad" customers. Here, "bad" custom...

[Day 181] Lending club data engineering project - Done

Image
 Hello :) Today is Day 181! A quick summary of today: completed and documented my data engineering project Everything is on my github repo , but below I will provide an overview. A diagram overview of the tech used: Raw Lending Club data from Kaggle Mage is used to orchestrate an end to end process including: extract data using kaggle's API and load it to the Google Cloud Storage (used as a data lake) create tables in BigQuery (used as a data warehouse) run dbt transformation jobs Terraform is used to manage and provision the infrastructure needed for the data pipeline on Google Cloud Platform dbt is used to transform the data into dimension tables, add data tests, and create data documentation Looker is used to create a visualisation dashboard For the dbt documentation, I was using dbt cloud IDE for the development, but to deploy a docs I needed to get its files, so the easiest way was to sync and run dbt locally. Setting up dbt to sync with local files was not hard, and this gave...

[Day 180] From Kaggle to BigQuery dimension tables - an end2end pipeline

Image
 Hello :) Today is Day 180! A quick summary of today: finished data modelling in dbt set up PROD in dbt set up automatic dbt job runs in mage created an end to end pipeline All code from today is on my github repo . 1. Settling down on a data model in dbt I went over a few different today, but I ended up with the above one. Because all my data is coming from 1 source I felt like, in order to avoid redundancy - I just decided to have dim_loans as the main table, and then have dim_borrower which includes just info about the borrower and dim_date just about the loan issue date. I also added data description fields, and some tests. The below pics are taken from the dbt generated documentation: dim_borrower dim_date dim_loans (image is truncated as there are many fields) 2. Setting up a PROD environment for dbt After I finally settled on a data modelling architecture, I created a PROD env to run all the models in a job. Not seen in the pic, but there is an 'API trigger' button which...

[Day 179] Using Docker, Makefile, and starting Data modelling for my Lending club project

Image
 Hello :) Today is Day 179! A quick summary of today: continued working on my Lending club data engineering project Today I added a few cool features, and I learned plenty 1. Introduced docker to the project Here is the Dockerfile I created Today I learned more about the docker folder structure and where to copy what, and where things live. At first I was not sure where things go, in which directory should I point my env vars, and where should I copy files. But then I also found that in Docker desktop I can view the files in a running image, so that is how I figured out what and where. The bash script referenced is here: And my docker-compose.yml (before I added the volumes, the code I was writing in mage was not persisting, so now I know what happens without volumes) 2. Added a Makefile I also made a Makefile (using this for the first time). I saw that adding a Makefile is good in the data eng zoomcamp project advices. And is good for reproducability. These are the options that ca...

[Day 178] Starting 'Lending club data engineering project'

Image
 Hello :) Today is Day 178! A quick summary of today: started my own data engineering project GCP as cloud storage solution terraform for infrastructure management mage for orchestration dbt for modelling maybe more to come The last bit of Module 4 was about creating a spark cluster in GCP - using Dataproc We create a cluster, which creates a VM for it.  On that cluster we can add PySpark (or other types like Spark) jobs where we can submit python scripts (that use PySpark) that we wrote. And we can run them on the cloud And because this cluster is created on GCP, it can directly read data from GCS and also write data to it.  Now ... onto the Lending club data pipeline project It is still early stages but, I had an idea in mind how to combine the different tools I learned from the DataTalksClub data engineering camp. I am taking inspiration from the data engineering zoomcamp where full documentation is encouraged so I will try to create nice graphs, visualisations and exp...

[Day 177] Spark for batch processing

Image
 Hello :) Today is Day 177! A quick summary of today: started Module 5: batch processing from the data eng zoomcamp Spark Spark operations Connecting PySpark to GCS What is batch processing of data? method of executing data processing tasks on a large volume of collected data all at once, rather than in real-time. It is often done at scheduled times (i.e. hourly, daily, weekly, x times per hour or minutes) or when sufficient data is accumulated Technologies used for batch processing: python scripts SQL Spark flink Advantages of batch jobs: easy to manage retry scale Disadvantages: delay in getting fresh data Spark is the most popular batch processing tool, and its variation in PySpark is popular. I have used PySpark during my placement at Lloyds Banking Group back in 2019-2020 so there was not much new info~ The most important bit is that spark works with clusters and in order to utilise spark's power in handling large datasets, we need to partition our data.  Then, I got a r...