OpenAI Announces o3 — 87.5% on ARC-AGI Sparks AGI Debate
In one sentence o3 achieves 87.5% on ARC-AGI (above the 85% human threshold), solves competition-level math and PhD science problems. Test-time compute scaling at $2,000/task high-compute setting reignites the AGI timeline debate.
OpenAI unveiled their most powerful reasoning model yet, called o3, and the results caused a genuine stir in the AI research community. On the ARC-AGI benchmark — a test specifically designed to measure fluid intelligence and resist memorization — o3 scored 87.5%, which is actually above the average human score of 85%. It can also solve competition-level mathematics problems and answer PhD-level science questions. The catch is that at maximum power, running a single difficult task can cost around two thousand dollars in compute. A smaller, cheaper version called o3-mini was also announced for everyday use. These results triggered a heated debate: does this mean AI has reached human-level general intelligence, or is it still just very sophisticated pattern matching? The ARC-AGI creator himself said the benchmark needs to be updated, as o3 had effectively solved it.
Companies
OpenAI
Tools
o3, o3-mini
Tags
Sources