Prime Ai Solutions
Read time: 4 minutes
Leader, welcome back.
Back to normal scheduling before I travel for the weekend.
Bad news after watching Resident Evil: Claude just found a previously unknown biological system inside viruses. Let's hope San Francisco fares better than Raccoon City!
The story that matters more for your budget: Anthropic and OpenAI both cut prices this week.
LATEST NEWS
Three new models, all cheaper!!
On Tuesday, Anthropic launched Claude Opus 5.5.
About 90 minutes later, OpenAI launched GPT-6 Sol and GPT-6 Luna.
Opus 5.5 matches Anthropic's top model, Fable 5.1.
It costs $4/$20 per million input/output tokens.
Anthropic says a typical job costs about 40% less than on Opus 5, because it uses fewer tokens. It also writes more clearly, with less jargon.
Sol costs $2/$10, half the price of Opus 5.5.
Luna costs $0.10/$0.50.
Both are only slightly smarter than the models they replace, and for now they're only in Codex and ChatGPT Work, not the regular ChatGPT app.
Why it matters: both companies are competing on price, because their business customers want to cut AI spend.
Price per token isn't the full picture, though.
YouTuber Nate Herk ran eight real tasks on both. He preferred Opus's output on seven, but Opus took 8h40 and cost $213. Sol took 5h51 and cost about $74. Browser Use found Sol scored higher on its web-browsing test (66.9 vs 59.4) at about 3.5x less cost.
My take: Opus 5.5 is the best all-round model right now. Sol is a step behind.
You don't need to pick one.
Use the strongest model to plan and check work, and a cheaper one for the repetitive steps.
What to measure: cost per completed task. That's model spend plus time, retries, and how often someone had to fix the output.
If you use Claude: Opus 5.5 defaults to medium thinking in the app.
Switch to high for complex analysis.
Claude Code's five-hour usage limits also went up 20%.
The virus discovery
Anthropic says about 950 Claude agents spent 21 hours searching a large DNA database for reverse transcriptases, enzymes that copy RNA into DNA. They then checked what sat next to each one.
They found over 200,000, flagged about 3,500 possible systems, and passed the best 20 to scientists as written reports. One agent spotted an unusual repeating DNA pattern, the same kind of clue that led to CRISPR. Lab tests confirmed it produces small RNAs. Anthropic named it ART. Nobody knows what it does yet.
Why it matters to you: the method transfers.
Humans set the question, AI searched and filtered a huge pile down to a shortlist, and humans checked the shortlist.
You can apply the same approach to a year of supplier contracts ("flag any with auto-renewal or price escalation clauses") or a ledger ("flag transactions that don't match the usual pattern for this supplier").
AI does the first pass.
Your team reviews what it flags.
My take: Dario Amodei says Claude led most of the discovery while humans chose the area and ran the lab tests. Andrew Curran pointed out that lab work still happens at normal speed. The same applies in your business: AI speeds up finding issues, not fixing them.
STEAL THIS
Check if a cheaper model does the job
Most AI spend goes on simple, repeated tasks: coding invoices, categorising expenses, routing queries. A cheaper model often handles these fine.
In June, Lindy's founder Flo Crivello moved all his company's AI work to DeepSeek and saved millions. He basically said you don’t need a genuis to write your email. He later said it didn't work for Lindy's more complex product. Cheap models suit simple tasks, not hard ones.
Here's how to test it.
Pick one high-volume task, such as invoice coding.
Pull 100 random examples from last week.
Write one pass/fail rule.
For invoice coding: "Correct nominal code and VAT treatment."Run all 100 through your current model and one cheaper model. If you use an open model, host it with a US or UK cloud provider so your data stays off the model maker's servers.
Count the passes for each.
Then compare. If your current model passes 92 and the cheap one passes 89 at a fifth of the cost, you're paying five times more for three extra correct answers. Multiply by your monthly volume to see the real money.
If those three answers don't matter, switch. If the failures are easy to spot automatically, such as a missing field, send only those back to the expensive model.
No technical team? Send this to whoever manages your AI tools, or ask your software vendor two questions: which model do you use, and what does each task cost? Either way, you'll have a number for your next contract renewal.
SIGNAL / NOISE
Signal: ChatGPT now works inside Word, on every plan including Free. Install it from Microsoft Marketplace and it opens in a sidebar. Try it on your last board pack: "Write a one-page executive summary for a non-finance board member. Flag any numbers that look inconsistent between sections."
Signal: OpenAI's Astra for Law scored 54% on OpenAI's own legal test, up from 38.7% for the standard model. They published the accuracy rate, which most vendors don't. Ask every AI vendor you pay: "What's your accuracy on tasks like ours, and how did you measure it?"
Noise: launch-day speed demos. Games built in ten minutes don't show how a model handles your reconciliations. Test on your own work.
That's it for this week.
If you know someone who approves AI spend at their company, forward this to them.
The 100-request test is worth running before their next renewal.
- Umar, Prime AI | primeai.solutions

