Hi Galaxea Team,
I am studying the G0.5 paper and trying to reproduce the BEHAVIOR-1K experiments.
I noticed that the paper reports BEHAVIOR-1K results for:
- G0.5 (1 epoch post-training): 0.2904
- G0.5 (4 epoch post-training): 0.3136
Could you please clarify:
-
Are the BEHAVIOR-1K post-trained checkpoints released or planned to be released?
-
If not, is there any plan to release:
- the fine-tuning script/config,
- the processed BEHAVIOR-1K training format,
- or the final checkpoint?
-
For reproducing the BEHAVIOR experiments:
- Is the post-training simply supervised fine-tuning with the same next-token prediction objective?
- Were CoT tokens enabled during BEHAVIOR post-training?
- Was the visual memory module enabled during evaluation?
-
Are there any embodiment-specific modifications for R1-Pro in the BEHAVIOR setup?
Thanks!
Hi Galaxea Team,
I am studying the G0.5 paper and trying to reproduce the BEHAVIOR-1K experiments.
I noticed that the paper reports BEHAVIOR-1K results for:
Could you please clarify:
Are the BEHAVIOR-1K post-trained checkpoints released or planned to be released?
If not, is there any plan to release:
For reproducing the BEHAVIOR experiments:
Are there any embodiment-specific modifications for R1-Pro in the BEHAVIOR setup?
Thanks!