I'm trying to reproduce the baseline performance. It appears that the baseline results reported in the paper use the llama-3.1-70B model for the human agent. However, due to limited computing resources, I'm unable to run the llama-3.1-70B model. Instead, I would like to know the performance results when both the human and robot agents use the llama-3.1-8B model as the planner.
If possible, could you share the results for this configuration?
I'm trying to reproduce the baseline performance. It appears that the baseline results reported in the paper use the llama-3.1-70B model for the human agent. However, due to limited computing resources, I'm unable to run the llama-3.1-70B model. Instead, I would like to know the performance results when both the human and robot agents use the llama-3.1-8B model as the planner.
If possible, could you share the results for this configuration?