Skip to content

Suggestion: Benchmark UniScientist on CritPt (frontier physics benchmark) #2

Description

@panademo

Hi UniPat team, thank you so much for open-sourcing such impressive models and sharing your research! It’s genuinely exciting to see the direction you’re pushing with UniScientist and the broader UniPat effort, congratulations on the release🎆.

I wanted to suggest a potential benchmark that could be particularly interesting for evaluating UniScientist’s capabilities: CritPt, a frontier physics reasoning benchmark that has recently gained attention for its difficulty and relevance to advanced scientific reasoning.

CritPt has been described as one of the "best proxies of progress towards actual recursive self-improvement, building automated research scientists." Notably, even the most frontier models today still struggle deeply on this benchmark, and the models that currently lead the leaderboard tend to require extremely high (often prohibitive) inference costs to achieve their results. Given UniScientist’s focus on scientific reasoning, benchmarking on CritPt could provide a particularly meaningful signal—both in terms of capability and efficiency.
Image
Source: https://x.com/teortaxesTex/status/2036163570048614437)

Image Image

For convenience, here are the links to the:

It would be fascinating to see how UniScientist compares on this benchmark, both in terms of raw performance and qualitative reasoning behavior. Even partial results or early experiments could offer useful insights to the community.

Thanks again for your amazing work and openness! Can't wait to see where UniScientist goes next🙏

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions