Description
Evals on arc-challenge for both the 1bit/2bit models are 21.5% which is worse than random chance since almost all questions have 4 options (some have 3/5, assume this averages out). Note that some of the baselines in BitDistiller were also around this accuracy, so it isn't totally unexpected.
Dig into the model and understand why this is happening:
- Understand the data by reading some examples.
- Are there any common mistakes the model is making. Why could this be?
Checklist before resolving issue
Description
Evals on arc-challenge for both the 1bit/2bit models are 21.5% which is worse than random chance since almost all questions have 4 options (some have 3/5, assume this averages out). Note that some of the baselines in BitDistiller were also around this accuracy, so it isn't totally unexpected.
Dig into the model and understand why this is happening:
Checklist before resolving issue
mainand fixed any conflictsdry_run.shon a CUDA device successfully