Hi, thank you for your amazing work!
I'm trying to reproduce the training process using a single A100 GPU (40GB memory). I reduced the batch size to 32 to fit the memory constraints. Currently, it takes around 27GB of GPU memory during training, each epoch takes about 18 minutes to complete. I think it may be too long to complete 1000 epochs.
I would like to ask:
- What kind of GPUs were used during your training?
- Approximately how long did the full training take? any way to accelerate?
Knowing these details would help me better plan my training setup. Thanks in advance for your help!
Best regards.
Hi, thank you for your amazing work!
I'm trying to reproduce the training process using a single A100 GPU (40GB memory). I reduced the batch size to 32 to fit the memory constraints. Currently, it takes around 27GB of GPU memory during training, each epoch takes about 18 minutes to complete. I think it may be too long to complete 1000 epochs.
I would like to ask:
Knowing these details would help me better plan my training setup. Thanks in advance for your help!
Best regards.