Hi,
Thank you for your impressive work on Emotion-LLaMAv2. I've been reading the paper and I am trying to reproduce the evaluation results mentioned in the "Experiments" section.
I noticed that the paper describes a specific parsing method to extract emotion labels from the tags for calculating Accuracy and F1 scores. However, I cannot find the corresponding evaluation script (e.g., eval_emotion_llama_v2.py) in the current repository.
Could you please release the evaluation code or point me to the script that handles the parsing and metric calculation?
Thank you for your help!
Hi,
Thank you for your impressive work on Emotion-LLaMAv2. I've been reading the paper and I am trying to reproduce the evaluation results mentioned in the "Experiments" section.
I noticed that the paper describes a specific parsing method to extract emotion labels from the tags for calculating Accuracy and F1 scores. However, I cannot find the corresponding evaluation script (e.g., eval_emotion_llama_v2.py) in the current repository.
Could you please release the evaluation code or point me to the script that handles the parsing and metric calculation?
Thank you for your help!