Scientific benchmark leadership shifts to GPT-6 Astra
GPT-6 Astra, even on its low setting, achieved a score of 54.3 in scientific benchmarks, surpassing Fable 5.1's maximum configuration. This performance comes at less than one-third of Fable 5.1's cost, marking a significant shift in leadership within scientific tasks.
Fable 5.1, while showing substantial improvements over its predecessor, Fable 5, in real-world science applications, could not match Astra's efficiency or output quality. Theo attributes OpenAI's post-training superiority as the likely driver behind these considerable scientific gains.
"Astra on low gets a 54.3 and only costs $11, putting low Astra higher than Max Fable 5.1 at under a third the price for the first benchmark that Anthropic had listed."


