Same Model, 13.3% to 38.3%

Featured Image
arc-agi-3 benchmarks agent-harness evaluation gpt-5 context-engineering ai-engineering measurement