OpenAI will sell you Astra, but not the system behind its score
This week, TNS covered GPT-6 Astra’s improving ARC-AGI-3 score, and it shows why harness engineering is becoming just as important as model selection. The software surrounding the model can change both what the model accomplishes and how much it costs.
ARC Prize ran GPT-6 Astra through its own standard harness, and the model scored 62.7%. When run through OpenAI’s Provider Adapter, the same model in the same reasoning setting scored 98.6%. The model didn’t change, but the software around it did, adding 36 points to the score. And the better-performing system cost less: $17,332 with OpenAI’s adapter versus $26,098 with ARC Prize’s.
I’ve been arguing the harness matters for months. I didn’t expect the result to be this lopsided.
More caching and a bigger vector database won’t fix what breaks when agent workloads, not human ones, hit your retrieval layer. Join us live to see what actually happens under that pressure and learn:
Why a bolted-together pipeline breaks where a unified retrieval layer doesn’t
Latency stacking, stale context, and relevance drift under concurrent load
Where this wall shows up in production, before you find it yourself
A malicious PR almost turned an AI coding assistant into a wiper. Learn why AI agents need strict external security controls.
WeAreDevelopers is coming to the US to give unsung developers a bigger voice
WeAreDevelopers, the 15,000-person Berlin conference now in its 11th year, is finally coming to the US— San Jose —this September 23-25. In an interview with The New Stack, co-founder Sead Ahmetovic and former GitHub CEO Thomas Dohmke talk about what a developer's day looks like when GitHub was built for humans collaborating with humans, not for orchestrating dozens of agents.
Catch the discussion about the future of software development ahead of the event, and grab a discounted ticket while you're there! Plus, don't miss the opening reception on September 23 at The Tech Interactive — space is limited, so RSVP soon.
Smarter alerting at scale: Live OpenSearch demo on PPL & unified alerting
Most observability platforms force a tradeoff: alert on everything and pay exponentially, or scale back and accept the blind spots. Join us next week as we explore how two OpenSearch capabilities close that gap: no licensing tiers, no ingestion ceiling.
The AI-speed SOC: A peer workshop for security leaders
Alert volumes are climbing, and AI is helping attackers move faster than most SOCs can respond. On September 15, a small group of CISOs and SOC leaders will meet behind closed doors to define what security operations will look like when AI becomes core to its architecture. This room is capped at 25 and is designed for contributors, not for an audience that listens.
WeAreDevelopers welcome reception with The New Stack and Dynatrace
Join us at The Tech Interactive in San Jose as we celebrate the first North American WAD World Congress. You're invited for elevated bites, beverages, and banter from 6–8 PM on September 23 while exploring one of the Bay Area's best tech museums. Mingle with the brightest minds, VIPs, and The New Stack's Editorial team—don't miss out! Space is limited, so get your name on the list soon!
Human review vs. verified pipelines: What catches bugs in the age of AI code
AI agents write code faster than any team can review it, and more review or AI reviewing AI won't fix it. Join us live on September 29 as TNS Host Viktor Farcic and Octopus Deploy's John Bristowe debate what actually catches bugs when volume outpaces review.