I have a question regarding the dyMEAN benchmark result shown in Table 1 of your ICLR 2025 IgGM paper.
When I test dyMEAN on the RAbD dataset, its average DockQ stays around 0.3. However, your Table 1 reports dyMEAN’s DockQ at only ~0.1, which is a huge gap.
Could you share the core reasons behind this significant performance drop? I wonder if it comes from inconsistent evaluation pipelines, epitope input usage, sampling selection rules.
Thanks a lot for your clarification!
I have a question regarding the dyMEAN benchmark result shown in Table 1 of your ICLR 2025 IgGM paper.
When I test dyMEAN on the RAbD dataset, its average DockQ stays around 0.3. However, your Table 1 reports dyMEAN’s DockQ at only ~0.1, which is a huge gap.
Could you share the core reasons behind this significant performance drop? I wonder if it comes from inconsistent evaluation pipelines, epitope input usage, sampling selection rules.
Thanks a lot for your clarification!