GPT Astra Scored 99.9%. Here Is the Number Nobody Is Quoting.


Hello Reader,

GPT-6 Astra scored 99.9% on a benchmark nobody prepared it for. Six months ago, the best model scored 7.8% on that same test. Naturally, the internet declared coding dead and Jensen Huang posted that AGI is here.

I went and opened the actual results page from the people who ran that benchmark. There is a second number sitting right there, on the same page. 62.7%.

Same model, same benchmark tasks, two different evaluation setups. Nobody is quoting that second number. Here is where both numbers actually come from, and what it means for your next interview.

Two harnesses, two very different stories

The benchmark is Arc AGI 3. You drop a model into a game it has never seen. Nobody explains the rules. It has to figure out what the game is, what winning looks like, and how to get there. Humans solve 100% of these environments.

Arc Prize tested Astra two different ways. The standard harness is provider-neutral, the same interface for every model, and the model itself decides what to carry forward between steps. Astra's score there was 62.7%, at a cost of $26,098 for that run.

The second harness is OpenAI's own adapter, which preserves what Arc Prize calls opaque reasoning state between requests, letting the model carry its own working memory forward instead of rebuilding it every time. That is where 99.9% comes from, at a cost of $18,817.

This is not a difference in effort settings. Line both harnesses up at the same reasoning level and the gap persists: 62.7% versus 98.6%.

Arc Prize is transparent about what each number actually measures. The standard harness asks how models compare on a level playing field, and their stated position is that real AGI should solve it under those conditions.

The provider-specific harness asks a different question entirely: how well does a model perform when given the context management infrastructure its own creator built specifically for it. Both numbers are real. Only one of them made it into headlines.

Even Arc Prize themselves wrote, directly, "we are not claiming that it is AGI," and described their own benchmark as having "a tightly bounded scope and format" that "does not represent the complexity and openendedness of the real world." The benchmark authors said this themselves.

The part that never makes the headline: cost

Those two Arc AGI runs cost $26,000 and $18,000 for a single benchmark, and that is the cheap end, lower reasoning settings ran up to $49,000 in worst-case scenarios. If your team wants this in production, that is the real meeting you are walking into.

What does one run cost. What happens when it fails halfway through. Who is watching, and what tools did it call. If it touched customer data, where is the audit trail. None of this has one universal answer.

It depends on your compliance requirements, who owns what in your org, and how much risk your business will actually carry. The model got better, and somebody now has to own cost, guardrails, and observability on a system that burns real money every time it runs. That job did not exist three years ago.

What I actually see across 300-plus interviews

70,000 software engineering roles were posted this year on the board I track. Filter to entry level and the pool falls off a cliff. Onboarding timelines are compressing too. Where new hires used to get two to three months to ramp up, students are now telling me they are expected to contribute within the first one to two weeks. If a new model ships every week, nobody wants to spend three months training anyone.

They want the person who can walk into a mess with no runway and figure it out. The candidates who do not get offers are almost never the ones who do not know enough.

I have interviewed people who answered every technical question correctly and still did not get hired, because every answer was textbook-correct and generic. The candidates who get hired bring up the constraint before I ask about it.

They say, "by the way, this is a regulated environment, so here's what I'd need to change." Nobody handed them that. They knew because they have lived it. That is exactly the thing the benchmark explicitly does not measure.

I had two students, identical resumes, same developer background. In mock interviews, one talked about real-world scaling challenges on a big shopping day. The other gave textbook answers. The first one got the Solutions Architect offer at AWS.

What to actually do

Stop collecting certificates. A certification is a closed-ended test with a fixed syllabus, the exact shape of problem this generation of models is already extremely good at. They open some doors, but I never once asked a candidate to show me a certificate across 300-plus interviews. I had a student with 12 AWS certifications and the golden jacket who still could not crack interviews.

Learn enough of the AI stack to talk about it in real terms, not to collect vocabulary. RAG, MCP, tool calls, agents, guardrails, observability, model routing, knowing the definitions is table stakes and getting cheaper to learn every month. The actual skill is knowing which one your specific situation needs, and why the obvious choice is wrong in your environment. Go deep on whichever concept connects to your existing background.

Reframe your experience as the asset it already is. The on-prem engineer who knows exactly how microservices fail under real traffic is worth more right now than someone who has only seen the cloud work properly. That is not legacy baggage. That is the half of the job nobody has automated, and it just got scarcer.

Next time someone tells you coding is dead, show them the Arc AGI results page, the cost, and the ultimate barometer - the number of open jobs. Go crush your next interview, and get that dream job.

FYI - I will be speaking at AI Engineer NYC conference on Oct 13th, on building a second brain for your team and enterprise, using RAG, Memory, and Agents. If you are attending, come say hi.

Keep learning and keep rocking 🚀,

Raj

P.S - If you want to get an AWS Solutions Architect job without coding or learning every AWS service, the 10th cohort for AWS SA Bootcamp is launching on Oct 17th, 12 PM ET (Eastern Time) via live workshop. This program now includes our updated Gen AI - including FDE roles! Please register below:

Here’s what you get when you show up LIVE:

  1. The myths keeping most people stuck - and what actually gets you hired as an SA - I've conducted over 300 SA interviews, so I know what I'm talking about!
  2. How GenAI is reshaping the SA and Gen AI roles including FDE, and the exact AI concepts (RAG, agents, MCP, eval etc.) you need to speak fluently in interviews.
  3. A first look at my new product feature, built to help you practice real-world, interview-relevant hands-on work instead of copy-paste tutorials.
  4. Full bootcamp breakdown for Cohort 10, plus a special offer only for live attendees.
  5. My exclusive Solutions Architect framework to prep you for today's job market! But if you’re not live, you won’t get it. No second chances.

And good news - it already worked for last cohort's students who secured cloud jobs in top companies, including at AWS, Microsoft, Google, JPMorgan, Reddit, and some of them didn't even have cloud experience 💰.

Spots are limited, so don't miss it!

Fast Track To Cloud

Free Cloud Interview Guide to crush your next interview. Plus, real-world answers for cloud interviews, and system design from a top AWS Solutions Architect.

Read more from Fast Track To Cloud

Hello Reader, Next week I will be on stage at AI Engineer (AIE) NYC. Out of all the AI conferences out there, AIE is one of my favorites. The speakers are engineers who have actually shipped AI systems, so the talks cover real-world implementation with all the messy tradeoffs included. You walk out with patterns you can use on Monday morning. Here are the details: Talk: Build a Self-Improving Portable and Secure Second Brain for Your Agents, with a live demo When: Tuesday, October 13th at...

Hello Reader, Six months ago, an interviewer asking about Gen AI was a bonus round. Today it is standard. At AWS, Microsoft, Meta, JP Morgan, Verizon, and most large enterprise technology teams, Gen AI questions are showing up in every SA, FDE, and AI engineer interview regardless of the role's primary focus. The candidates who answer these well are not the ones who have read the most blog posts. They are the ones who can connect the concepts to real architecture decisions and explain the...

Hello Reader, Six months ago, an interviewer asking about Gen AI was a bonus round. Today it is standard. At AWS, Microsoft, Meta, JP Morgan, Verizon, and most large enterprise technology teams, Gen AI questions are showing up in every SA, FDE, and AI engineer interview regardless of the role's primary focus. The candidates who answer these well are not the ones who have read the most blog posts. They are the ones who can connect the concepts to real architecture decisions and explain the...