Hi, fantastic work on the unsupervised text-image alignment!
I have a quick question: have you considered adding the CLIP score between the cropped image and its text into the reward function?
My thinking is that this could help fix potential misalignments that might still occur by directly rewarding semantic similarity.
Thanks for the great project!
Hi, fantastic work on the unsupervised text-image alignment!
I have a quick question: have you considered adding the CLIP score between the cropped image and its text into the reward function?
My thinking is that this could help fix potential misalignments that might still occur by directly rewarding semantic similarity.
Thanks for the great project!