Posts

[Day 73] MBR and FUDGE - decoding mechanisms; pre vs post layer normalization

Image
 Hello :) Today is Day 73! A quick summary of today: Covered lecture 6 : Generation algorithms and 9 : Experimental design and human annotation from the CMU 11-711 Advanced NLP course, from which I found out about: Minimum Bayes Risk (MBR) FUDGE decoding Why pre layer normalization is better than post layer normalization in transformers 1) Minimum Bayes Risk decoding for modern generation techniques ( Bertsch and Xie et al., 2023 ) When we get to the generation step of a language model, predicting and outputting the next token in a sequence, we can use methods such as greedy decoding or beam search to select tokens that have a high probability. But MBR which was originally proposed by Bickel and Doksum in 1977, questions whether we actually want to get the highest probability token.  Outputs with low probability tend to be worse than the opposite. But if we just compare the top outputs, it is less clear, as the outputs with the top probabilities might look similar (for example...

[Day 72] Carnegie Mellon University - Advanced NLP Spring 2024 - assignment 1

Image
 Hello :) Today is Day 72! A quick summary of today: Today I found 11-711 Advanced NLP by CMU I am amazed. It is still ongoing and they upload the lecture videos and all the information is on the course website. I had a look at the syllabus and saw that the first few lectures are: Given I have covered CS224N by Stanford, I felt confident and skimmed over the lectures just to check for any new info, and then decided to jump into assignment 1. WOW! Assignment 1 felt great, it was challenging, difficult and made me read through research papers like (RMS layer norm, and Adam) to implement a Llama model from scratch! I uploaded all my assignment code to this github repo , but the main files that needed to be implemented are classifier.py, optimizer.py, rope.py, and llama.py. Below I will go over how I implemented each file (and my struggles along the way).  I am not that familiar with Llama but this was a good exercises to get more familiar with the model - by jumping rig...

[Day 71] Backprop, GELU, Tricking ChatGPT, and Stealing part of an LLM

Image
 Hello :) Today is Day 71! A quick summary of today: Finished up the manual backprop code to make it more clear Read some research papers Gaussian Error Linear Units (GELUs) Using Hallucinations to Bypass RLHF Filters Stealing Part of a Production Language Model Firstly, the manual backprop code cleanup. To finish it up (for now atleast), I decided to get more data. Original was ~32000 names, on kaggle I found with around ~90000 names.  As for the architecture - I settled on the original input > flatten > batch norm > activation > output, I added more even up to 2 hidden layers and batch norm for each, but the training time got longer, and the results were not even better than the simple original one.  Even though I stayed with the original model structure, I tried to modify embedding size, context length, hidden layers, just to play around with them and see the result. At the end the generated names look decent. But I was not able to achieve a much better re...

[Day 70] Testing my backprop knowledge

Image
 Hello :) Today is Day 70! Today, I saw someone on X mentioning backprop and I decided to go back to my 'backprop ninja' code and re-do it to practice, and make sure I know the code. I initially did the 'become a backprop ninja tutorial' by Andrej Karpathy on Day 54 , but since then I have not looked at it in depth, and I felt that I was not satisfied with how I did it. Though I learned a lot about how backprop works, for the code... I did not feel like it was my own. Though I wrote the manual backprop by myself, the rest was pretty much copied from Andrej's code. Well today I wanted to fix that, and fully make the code 'my own' but also it would give me a chance to test my knowledge. Firstly, I started with proper naming for the vars, to make the code more readable. For instance,  original code: updated: The full new version is on this github repo . Besides variable renaming to make the code more readable, I wanted to add something else, so I decided to add...

[Day 69] Training an LLM to generate Harry Potter text

Image
 Hello :) Today is Day 69! A quick summary of today: tried to build upon the built LLM from the book from the days before, and write a training loop in the hopes of generating some Harry Potter text ( kaggle notebook ) Firstly I will provide pictures of the implementation (then share my journey today) The built transformer is based on this configuration Dataset is Harry Potter book 1 (Harry Potter and the Philosopher's Stone) text file from kaggle. Used batch size 16, and a train:valid ratio 9:1 Model architecture code: Multi-head attention GELU activation function Feed forward network Layer normalization Transformer block GPT model (+generate function) Optimizer: AdamW with learning rate 5e-4 and 0.01 weight decay. Ran for 1000 epochs.  Final number of model parameters: 1,622,208 million Now, about my journey today ~ The easy part of was the dataset choice - harry potter. The hard part was choosing hyperparams that would end up in not perfect, but somewhat readable and s...

[Day 68] Build a LLM from scratch chapter 4 - making the GPT-2 architecture

Image
Hello :) Today is Day 68! A quick summary of today:  Covered chapter 4 of Build a LLM from scratch by Sebastian Raschka Below is an overview of the content with not much code. For the full code version of every step - it is on this github repo . This chapter is the 3rd and final step from the 1st state towards a LLM 4.1 Coding a LLM architecture The book will build the smallest version of GPT-2 that has 124m parameters, with the below config. The final architecture is a combination of a few steps, presented below After creating and initializing the model, and the gpt-2 tokenizer, on a batch of 2 sentences: 'Every effort moves you' and 'Every day holds a', the output is: 4.2 Normalizing activations with layer normalization Taking an example without layer norm The mean and var are Apply layer norm, and the data has 0 mean and unit variance. We can put the code in a proper class to be used later for the GPT model 4.3 Implementing a feed forward network with GELU activation...