Hi UniPat team, thank you so much for open-sourcing such impressive models and sharing your research! It’s genuinely exciting to see the direction you’re pushing with UniScientist and the broader UniPat effort, congratulations on the release🎆.
I wanted to suggest a potential benchmark that could be particularly interesting for evaluating UniScientist’s capabilities: CritPt, a frontier physics reasoning benchmark that has recently gained attention for its difficulty and relevance to advanced scientific reasoning.
CritPt has been described as one of the "best proxies of progress towards actual recursive self-improvement, building automated research scientists." Notably, even the most frontier models today still struggle deeply on this benchmark, and the models that currently lead the leaderboard tend to require extremely high (often prohibitive) inference costs to achieve their results. Given UniScientist’s focus on scientific reasoning, benchmarking on CritPt could provide a particularly meaningful signal—both in terms of capability and efficiency.

Source: https://x.com/teortaxesTex/status/2036163570048614437)
For convenience, here are the links to the:
It would be fascinating to see how UniScientist compares on this benchmark, both in terms of raw performance and qualitative reasoning behavior. Even partial results or early experiments could offer useful insights to the community.
Thanks again for your amazing work and openness! Can't wait to see where UniScientist goes next🙏
Hi UniPat team, thank you so much for open-sourcing such impressive models and sharing your research! It’s genuinely exciting to see the direction you’re pushing with UniScientist and the broader UniPat effort, congratulations on the release🎆.
I wanted to suggest a potential benchmark that could be particularly interesting for evaluating UniScientist’s capabilities: CritPt, a frontier physics reasoning benchmark that has recently gained attention for its difficulty and relevance to advanced scientific reasoning.
CritPt has been described as one of the "best proxies of progress towards actual recursive self-improvement, building automated research scientists." Notably, even the most frontier models today still struggle deeply on this benchmark, and the models that currently lead the leaderboard tend to require extremely high (often prohibitive) inference costs to achieve their results. Given UniScientist’s focus on scientific reasoning, benchmarking on CritPt could provide a particularly meaningful signal—both in terms of capability and efficiency.

Source: https://x.com/teortaxesTex/status/2036163570048614437)
For convenience, here are the links to the:
It would be fascinating to see how UniScientist compares on this benchmark, both in terms of raw performance and qualitative reasoning behavior. Even partial results or early experiments could offer useful insights to the community.
Thanks again for your amazing work and openness! Can't wait to see where UniScientist goes next🙏