Thank you very much for this excellent work. I have some questions about the paper, why is the rank of LORA set to 512? Usually, people would choose a smaller rank. Wouldn't using 512 in this work lead to overfitting of the model? Another question is that the 100k dataset was only used for training for 5k times, with a batch size of 16. The data was not fully utilized. Why can we be sure that the current training is sufficient?
Thank you very much for this excellent work. I have some questions about the paper, why is the rank of LORA set to 512? Usually, people would choose a smaller rank. Wouldn't using 512 in this work lead to overfitting of the model? Another question is that the 100k dataset was only used for training for 5k times, with a batch size of 16. The data was not fully utilized. Why can we be sure that the current training is sufficient?