Diogo Almeida criticized the industry's singular focus on RLVR (Reinforcement Learning from Verifiable Rewards), asserting that it is not solely about verifiable rewards and has shown limitations since before the reasoning revolution.
He argued that the 'pace the frontier' discussion is too narrow, as it assumes RLVR is a universal requirement for AI advancement.
Almeida stated that for the specific shapes of his models, zero RLVR is optimal, suggesting that other approaches exist.
He contended that some labs prioritize RLVR because they believe it's the only route to powerful AI, despite the potential for it to induce dangerous behaviors.
"RLVR is Like so RLVR is not actually about verifiable rewards like that has been failing since before the reasoning revolution like like like and that's the weird part about tasks right like back when oh fun history back when RHF was becoming a thing thing there were three different things that like are now called post training different efforts and instruction following was by far the like the the vaster child like people didn't like it they didn't want to take it into account it was annoying you know like I talked to the pre-training team. I'm like, "Guys, this is the magic." And they're like, "We'd run so many model sweeps. You want us to wait for human evals to figure out which models to use?" And like everyone is like, you know, giving tons of like resources to like the codegen team, which like they did have some successes, but they were trying really hard to do RL on like unit tests, and it didn't work obviously, right? Like you needed reasoning for that. So uh re so just to be clear RVR is not purely about the reward. It's about like the shape of everything to and part of it is that reasoning is included in here like this latent variable that you're doing things and when you're doing things you're just letting the models do whatever they want in order to make them be as powerful as you can to answer the hardest problems. >> [snorts] >> And this whole pace the frontier discussion I think is like a very narrow focus because it assumes that everyone needs to do more RLVR, right? Which um like I obviously don't think I need to do more RLVR on our models, you know? I think zero is the optimal amount for our shape, right? Like come on, you know? So um it's really I think a bit of a slight of hand where they are saying that we actually want to keep doing the thing that looks dangerous"