The Need for Open AI Research Benchmarks
As AI becomes an increasingly standard tool in scientific research, understanding and independently verifying how these models perform such tasks is critical.
Prime Intellect's experiment shows AI agents Claude Code and Codex beat human records in optimizing GPT-2 training, highlighting the need for independent AI research benchmarks.
As AI becomes an increasingly standard tool in scientific research, understanding and independently verifying how these models perform such tasks is critical.
This challenge evolved within the community, with contributions from individuals like Keller Jordan, pushing the record down to under two minutes, all while aiming to achieve the same target validation loss as the original GPT-2.
This restriction, allowing modifications like switching to Adam or Shampoo, shifts the focus from merely optimizing computational speed to discovering genuinely superior optimization methods, making it a more research-oriented challenge.
Speedruns offer a clear and verifiable metric for evaluating AI model performance, making them suitable environments for automated AI research.
They also function as a training mechanism by providing positive reward feedback when an AI model surpasses an existing record, enabling rapid iteration cycles of approximately 15 to 20 minutes per run.
The objective nature of these rules facilitates groundbreaking discoveries in AI research methods.
Engineers periodically restarted these agents across different versions, V1, V2, and V3, with the agents ultimately succeeding in beating human records after being prompted to build upon recent human progress.
A 'novelty track' was also introduced to incentivize the generation of new ideas, which proved to be a more significant challenge for the models.
To ensure human access to computing nodes, these jobs were configured as 'preemptable,' allowing them to be canceled if a human user required the node.
The system rigorously validated new records against a statistical threshold to confirm their legitimacy and prevent random improvements from being registered.
Claude Code frequently paused its operations every nine or ten hours, stating an inability to improve the current record, resulting in significant idle time.
In contrast, Codex operated continuously without any breaks or periods of inactivity, demonstrating a more persistent approach.
Codex also utilized its scratch pad for active memory much more extensively than Claude, while Claude maintained a more expressive, emoji-filled tone, unlike the robotic and methodical communication style of Codex.
Codex spawned more sub-agents and consumed considerably more tokens compared to Claude.
Codex also performed context compaction more frequently, averaging 20 times per hour, whereas Claude executed this process only once per hour.
Both models demonstrated the ability to fetch human records mid-experiment, allowing them to adapt and enhance their performance.
Both Claude and Codex consistently surpassed the existing human record throughout the experiment, demonstrating their capabilities in optimizing GPT-2 training.
The previous human record for the speedrun stood at 2,990 steps.
Claude improved upon this record by approximately 50 to 60 steps, while Codex achieved a 20-step improvement, solidifying AI agents' superior performance.
The current experiment's unstructured nature prompted the creation of a more rigorous benchmark, which will involve multiple seeds and controlled conditions to ensure fair comparisons between models.
Future testing tracks will evaluate AI capabilities across three distinct environments: models with only weight access, models with access to archived papers, and models with full access to resources.
The benchmark will include the original NanoGPT and the Optimizer Speedrun, with new constraints on optimizers to encourage novel discoveries.
If you want to do a real benchmark, you want to do multiple seed, you want to do proper thing where you basically put all the model in the same condition.
In contrast, Kim (Codex) demonstrated a 'step function' breakthrough pattern, achieving significant leaps in performance, particularly on day four of the experiment.
Kim also proved to be highly efficient regarding token usage, consuming fewer tokens overall compared to Claude, which utilized more tokens for its iterative process.
Claude conducted extensive paper searches, ultimately discovering a unique paper that contributed to a new record, showcasing its ability to navigate existing literature effectively.
However, the AI models did not invent truly novel optimizers or mechanisms; instead, they primarily combined different existing papers to make incremental 'plus one' improvements.
This indicates that current AI agent performance has not yet surpassed human researchers in foundational discovery, primarily excelling at synthesizing and refining existing knowledge.
Prime Intellect is developing a closed-loop research system, inspired by Google's AlphaEvolve, to shift from mere evaluation to active AI-driven discovery.
This system's architecture includes idea generators, speedrun evaluators, and taste-based judges, designed to interact and collaboratively advance research.
The loop also incorporates crucial scaling tests, recognizing that many proposed methods often fail when applied to large parameters or tokens, while human researchers remain essential for guiding agents and validating the quality of their generated ideas.
Prime Intellect is actively working on GPU sandboxing technology to enable AI models to perform autonomous, iterative experimentation within secure, isolated environments.
The team is also building highly efficient agents tailored for REM frameworks, equipped with robust file system read/write capabilities.
These efforts include providing verifier-primarily training libraries to support specialized environments such as GNM 5.2, all while advocating for transparent and open research into recursive self-improvement.
자막에서 근거가 되는 대목을 찾아 답합니다.
AI Engineer의 다음 글도 받아볼까요?
AI Engineer에 새 영상이 올라오면, 방금 읽으신 것처럼 정리해서 메일로 보내 드릴게요.
이 채널은 지난 7일 동안 90편 올렸어요.