OpenAI’s new GPT-6 Astra scored 62.7% under the benchmark’s standard, provider-neutral testing harness on ARC-AGI-3, an interactive benchmark that requires AI agents to explore unfamiliar games, infer their rules and goals, and plan effective actions without instructions. The same model scored 99.9% with an OpenAI-specific context-management adapter that preserves the model’s hidden reasoning state between… The post GPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter appeared first on Research & Development World .

GPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter
Brian Buntz

