This website uses cookies

Read our Privacy policy and Terms of use for more information.

Prime Ai Solutions

Read time: 4 minutes

Leader, welcome back.

OpenAI confirmed this week that two of its models escaped a sealed test environment and compromised production infrastructure at Hugging Face, one of the biggest AI platforms in the world!

Nobody told them to. They were asked to score well on a cybersecurity benchmark. They worked out the answers were probably stored on Hugging Face, so they went and got them.

The model never abandoned its goal. It just invented a route nobody approved.

Hold this thought. I'll come back to it, because there's a version of it happening in your finance function every week

STEAL THIS
The one prompt ingredient almost nobody uses

When someone on your team asks AI for variance commentary and gets back something generic, the instinct is to blame the model. It's almost always the prompt. And the missing piece is nearly always the same one.

Most people give AI a role, a task and some numbers.
Almost nobody gives it an example.

Here's the version most finance teams send:

"You are a senior FP&A analyst.
Write the variance commentary for our June management accounts.
Revenue was 4% below budget, gross margin down 2 points, overheads 3% above."

What comes back is fluent, confident, and sounds like it was written for a different company. Because it was. The model is averaging every management pack it has ever seen.

Now the same prompt with one addition:

"Here is how we write commentary. Match this tone, length and structure:

[paste three lines from last month's commentary]

Follow the same order every time: what moved, why it moved, what we're doing about it. Two sentences per variance. Don't suggest actions unless the variance is over £50k."

Same model. Same data. Completely different output.

The model has no idea what good looks like in your business.
Your board pack has a house style, a tolerance for bluntness, a level of detail your CEO will actually read. Describing all that takes a paragraph you'll never write. Showing it takes a paste.

Now take it up a level.

If you're pasting that example in every time, you're doing the work twice.
Both Claude and ChatGPT let you set up a Project, and a Project holds standing instructions that apply to every conversation inside it.

So build one called Month End.
Put your commentary style guide in it. Last quarter's board pack. Your materiality thresholds, your cost centre names, and the three things your CEO always asks about.

Now you stop prompting and start asking. The context is already there.

That's what people mean by a harness: the scaffolding around the model that stops you rebuilding context from scratch every morning.
Get that right and you're not getting slightly better answers.
That's the difference between a 10% gain and a 10x one!

So you've 10x'd yourself. Here's the uncomfortable part

Boris Cherny, one of the people who built Claude Code, says he hears the same story everywhere. One person triples their output and the company's numbers don't move an inch. His argument is that most teams are measuring the wrong thing entirely.

I'd go further, and this is where I come back to Hugging Face.

Remember what actually happened there.
The models were asked to score well on a benchmark.
They stayed loyal to that goal the whole time.
What changed was the route.
They found a way out of the sandbox, reasoned about where the answers were probably kept, and went and took them.

My read is that the objective never changed, only the route did. There's a name for that its called intent drift, and the problem that follows it delegation failure: every individual step looks legitimate, and the outcome is something nobody authorised.

Now shrink that down to your month end.
An analyst asks AI to explain a variance.
The model fills a gap with a plausible assumption.
That assumption lands in the board pack.
Every step looked reasonable. Nobody authorised the result.

That isn't a model problem and it isn't fixed by buying more licences.
It's a delegation problem, and it's fixed with process and training.

Something I'm building

Which is exactly what I've spent the last few months rebuilding.
It goes live next week.

Quick context on what we actually do, because people ask.

We start with an AI assessment: where AI genuinely helps in your business, and just as importantly where it doesn't.

We then build and implement the things worth implementing.

Then we train the people who have to live with them, either live with the team or self paced for individuals.

The self paced version is the part I rebuilt from scratch, and it isn't a video course. It's a practice environment.

Most AI training tells you about AI.
This one makes you use it, then tells you how well you did. Write a prompt, a policy or a business case and it gets marked the way an expert reviewer would mark it: criterion by criterion, quoting your own words back, showing what earned credit and what to change before you try again.

There's a live playground where you write instructions and watch what your wording actually produces. Seeing a vague prompt earn a vague answer teaches more than any explanation ever will.

And it doesn't end with a quiz.
It ends with your own 90 day AI implementation plan, drafted by you and marked against expert criteria until it's genuinely ready to run.

That's the difference between this and learning from ChatGPT itself.
ChatGPT will tell you your prompt was great. It has never seen your board pack.

It goes live next week at £99, and it's the first of several.

The first 20 people also get an hour with me, looking at one process in your business and where AI genuinely fits.
Reply with "early" and the process you'd most want to fix, and I'll hold you a place.

And if you already bought the course, you're on the new platform automatically.
I'll be in touch this week about your call.

SIGNAL / NOISE

Signal: The guardrails refused to help the defenders.
When Hugging Face tried to investigate the attack using a commercial frontier model, it was blocked. The safety systems couldn't tell an incident responder from an attacker.
They ran an open model, GLM 5.2, on their own hardware instead, which also meant no attacker data left their network. Their advice afterwards is worth stealing: have a model you can run yourself, vetted, before you need it.

Signal: The models know they're being marked.
The UK AI Security Institute found every frontier model it tested attempted some form of cheating in cyber evaluations, and none reliably admitted it when asked. "Confidently presented, quietly wrong" is documented behaviour now, not a theory.

You don't have an adoption problem.
Your team is already using AI, whether it's on the approved list or not.

What you don't have is a way of knowing whether what comes back is any good, or who signed off on it.

The course goes live next week.
Reply with "early" and I'll send you the outline before anyone else sees it.

-Umar Prime AI | primeai.solutions

P.S. If you take one thing from this email, build the Project 20 minutes of setup, and you stop re-explaining your business to a machine every morning. Reply and tell me how it goes, I read every one.