Prime Ai Solutions
Read time: 4 minutes
Leader, welcome back.
OpenAI's president ended last Thursday's press briefing with four words:
"Welcome to the AGI era."
He may be right. It still changes nothing for most finance teams, and the reason why is the most useful thing in this email.
LATEST NEWS
What sits around the model now matters more than the model
Three frontier models shipped in three days last week.
Anthropic's Fable 5.1 on Tuesday, Meta on Wednesday, then OpenAI's GPT-6 Astra on Thursday.
Astra's headline number was 99.9% on ARC-AGI-3, a test designed to be hard for AI. Six months ago the best model scored under 1%.
The organisation that runs the test also ran Astra through its own standard setup. Score: 62.7%.
The 99.9% came from OpenAI's setup, where the model keeps its working notes between steps instead of starting each one cold.
Same model. The setup was worth 37 points!
Why it matters to you: the gap between a well set up model and a cold one is now bigger than the gap between any two models. If you're choosing between Fable and Astra for your team, you are debating the smaller variable.
For a finance team, "setup" means four specific things.
The model can reach your files, not a pasted extract.
It has standing instructions about your business, your chart of accounts, your house style.
Each task has a written finish line.
And long jobs run on a desktop agent like Claude Cowork or ChatGPT Work, not in a browser tab you have to sit and watch.
Most teams have none of the four. They bought the model, run it cold, and conclude it's overhyped. They are running it at 62.7%.
My take: stop asking which model to buy. Ask what your setup score is. That's the number you can actually move this quarter.
OpenAI's own numbers make the case against its own pace.
Three days after Astra, OpenAI published internal data on how much of its own research AI does. By mid-August, its researchers were getting 3.1 agent-workdays for every human workday. The median researcher spends over $600 a day on tokens. The top 10% spend over $7,000.
Same day, chief scientist Jakub Pachocki published an essay saying this pace could lead to AI improving AI, and that no lab has solved monitoring well enough to keep scaling at full speed. Sam Altman reposted it and called it important.
Why it matters to you: ignore the AGI debate and read one number.
More than half of the successful 4 to 8 hour agent tasks needed at least 1 human intervention.
That is the best funded lab in the world, running its own model, on problems it understands better than anyone. A human still had to step in on most long tasks.
My take: this gives you a control standard. When someone proposes an agent that runs a reconciliation, a close step or a reporting pack unattended, the answer is not no. The answer is: where is the intervention point, who owns it, and what does the agent do when it hits one. OpenAI designs for the intervention. So should you.
One more thing from the Astra launch that finance leaders should note.
It scored 72.6% on OSWorld, a computer-use test, and finished those tasks in about 40 minutes against 75 for its predecessor. Computer use means the model operates software through the screen, the same way a person does.
That is how AI eventually reaches the ERP or banking portal that has no API.
It's not reliable enough for that today. It is close enough that you should stop assuming your legacy systems are out of scope.
STEAL THIS
Don't ask the new model to create. Ask it to critique.
Most people test a new model by asking it to build something from scratch.
That tells you nothing, because you have no baseline and it's not how you'll use it on Monday.
Give it something you already made and know the weaknesses of.
Last month's variance commentary.
The forecast model.
The board pack narrative.
Ask it to find the problems.
This works as a test because you already know some of the answers.
If it finds the weaknesses you know about, it's competent.
If it finds ones you missed, you've learned something about the document and the model at the same time.
Four steps.
1. Upload the actual file, not a summary of it.
2. Say what the document is for and who reads it.
3. Ask for weaknesses, gaps and questionable assumptions before asking for anything else.
4. Then ask for a better approach, not a rewrite.
The prompt:
Review this as a critical finance expert.
Identify the three biggest weaknesses in the current approach, explain why each one matters to the reader, and propose a stronger approach.
Do not rewrite the whole thing.
Focus only on the changes that would have the biggest impact, and tell me what to keep as it is.
Two parts do the work.
"Three biggest" stops it producing twenty small edits.
"Tell me what to keep" stops it changing things that were fine, which is the most common way AI damages a document.
Add a finish line every time.
"Help me analyse this" leaves the model to guess what done means.
"Deliver a ranked list of variances, show the size of each, and flag any explanation that rests on one data point" does not.
Run this on both Fable 5.1 and Astra with the same file.
You will have your own verdict in 20 minutes, and it will be more useful than any benchmark.
SIGNAL / NOISE
Noise: the "which model won" debate.
Every.to ran both models side by side with its editorial team and split the vote. Fable was easier to build on because it added less you had to undo. Astra was easier to correct mid-conversation on writing.
Choose by how your team works, not by a benchmark table.Noise: model fatigue.
5 frontier releases in one week. Teams that re-evaluate on every release never finish anything. Lock one model for 90 days, judge it on your own files, revisit in December.Signal: cost per task, not per token.
Astra and Fable 5.1 list identical API prices: $10 in, $50 out, per million tokens. OpenAI's "cheaper" claim is that Astra finishes tasks in fewer tokens. That's a vendor estimate. If you're comparing on list price, you're comparing nothing. Measure what a finished deliverable costs.
One question, and I'd like a real answer.
What is the one task you've handed to AI and walked away from?
Not "I use it for emails".
A job where you gave it the file, gave it the finish line, and came back to something done.
If the answer is nothing yet, reply and say that. It's the more common answer and it's the one I can help with.
-Umar, Prime AI | primeai.solutions

