AI Safety Needs Governance Too

“AI safety” should never become the governance equivalent of “for your own good.”

There are some words that arrive in a governance meeting with an unfair advantage.

Safety is one of them.

Nobody wants to be the person arguing against safety.

The same is increasingly true of AI safety.

That's probably a good thing. As AI systems move into healthcare, critical infrastructure, financial services, employment and other places where mistakes can have consequences beyond an inconvenient software error, I'd much rather we worried about safety too much than discovered its importance afterwards.

Still, there's something about the phrase that bothers me more than the objective, it is the authority the objective can acquire.

Put "AI Safety" at the top of a PowerPoint slide and almost anything underneath it starts sounding slightly more reasonable.

That's worth being careful about because safety is an objective.

It most certainly isn't a justification.

If I sit here and consider an organisation that's worried about how its employees are using generative AI.

Some employees might paste confidential information into prompts. Others might use unauthorised models, circumvent restrictions or find imaginative ways around controls that nobody anticipated when the policy was written.

Nothing particularly surprising there. Give people a new technology and somebody will eventually do something with it that makes the security team stare silently at a screen for a few seconds.

So the organisation introduces an AI safety control.

Every interaction with its internal AI systems is monitored.

Prompts and responses are retained and analysed.

Employees are assigned behavioural risk scores based on how they interact with those systems. Cross a particular threshold and access to certain AI capabilities is automatically restricted.

The organisation keeps the interaction history for several years so that behavioural patterns can be examined retrospectively.

Let's make the hypothetical control even more uncomfortable.

Suppose it works and it works really quite well.

It catches confidential information being submitted where it shouldn't be. It identifies attempts to circumvent policy. It discovers behaviour conventional security monitoring missed. Incidents that might otherwise have occurred are prevented.

The dashboards look excellent.

Now what?

That's where I think the interesting governance problem begins.

Someone decided what behaviour should increase an employee's risk score.

Someone decided how long their conversations should be retained.

Someone decided which people should have access to those conversations.

Someone decided when a score becomes consequential.

Perhaps the employee can challenge the decision.

Perhaps they can't.

Perhaps a human reviews the restriction before it happens.

Perhaps another AI system quietly does the whole thing.

We could produce excellent evidence that every part of this control operates exactly as designed and still have avoided a rather awkward question.

Should we have designed it that way in the first place?

That's not a control-effectiveness question.

A control can be wonderfully effective and deeply inappropriate at the same time.

Security people should already be familiar with this problem. We could dramatically reduce certain cybersecurity risks by removing everyone's internet access, filling the USB ports with epoxy and confiscating every laptop at the end of the working day.

Effective?

Possibly.

A sensible way to run most organisations?

Probably not.

Governance exists partly because effectiveness isn't the only thing we care about.

We care about proportionality.

Authority. Then privacy. Accountability and eventually consequences.

Who benefits from a decision and who quietly inherits the risk created by it.

An AI safety control could be intrusive, coercive or even Machiavellian (a word I once won a game of Scrabble with) and still genuinely make an AI system safer.

Calling it safety doesn't resolve the contradiction.

It gives us a reason to examine it.

For your own good

None of this started with artificial intelligence.

Governments and organisations have been exercising authority in pursuit of safety for a very long time.

National security.

Fraud prevention.

Public health.

Cybersecurity.

Protection from harm.

The objectives can be perfectly legitimate whilst a particular intervention made in their name isn't.

We've built quite a lot of law, oversight and governance around that uncomfortable fact.

We know that necessity matters.

A good HR department knows that proportionality matters.

A police man knows that authority matters.

So does the ability to challenge consequential decisions.

AI doesn't somehow get an exemption because we're still excited about what it can do.

If anything, the problem becomes more interesting as the controls themselves become intelligent.

Imagine an AI agent responsible for monitoring another AI system.

At first it only observes.

Later we allow it to identify potentially unsafe behaviour.

Then perhaps it can restrict capabilities.

Change configurations.

Quarantine activity.

Initiate remediation.

None of those capabilities is inherently unreasonable. Some could be extraordinarily useful.

But somewhere along that progression we've stopped talking about a passive safety mechanism and started giving one system authority over another.

Eventually that authority may extend to people.

At which point I'm much less interested in what we've called the control than in what we've allowed it to do.

Who governs the system that's governing the AI?

I suspect that question is going to become considerably more important.

A control can work and still be wrong

I've been working recently on a small AI governance assurance proof of concept.

One of the ideas behind it is fairly ordinary:

A claim is not evidence, and evidence is not assurance.

An organisation claims that a control exists.

Evidence is submitted to support the claim.

Someone examines that evidence and determines whether it actually demonstrates what the organisation says it demonstrates.

Nothing revolutionary there. Auditors have been making people's Tuesday afternoons more exciting with variations of this process for quite some time.

But the AI safety question exposes another layer.

Imagine our fictional organisation submits overwhelming evidence for its employee-monitoring control.

The configuration is correct.

The audit trail is complete.

Risk scoring operates consistently.

Retention works exactly as documented.

Access restrictions trigger at precisely the approved threshold.

Every test passes.

As an assurance professional, I might be able to say with considerable confidence:

Yes. This control works exactly as represented.

I still haven't answered whether the control should exist in that form.

That's a different assertion.

And therefore, I think, it requires different evidence.

Evidence of effectiveness tells me the machinery works.

Evidence of legitimacy would have to tell me something about why this particular machinery was justified.

What risk required it?

How serious was that risk?

Were less intrusive controls considered?

Who approved the authority being exercised?

What new risks did the control introduce?

Can someone affected by it challenge the outcome?

When will the original justification be reconsidered?

That last question interests me particularly.

Controls have a habit of becoming permanent.

The incident that justified them fades from organisational memory. The people who designed them leave. Technology changes. The original assumptions disappear into an old risk assessment sitting somewhere in SharePoint.

The control remains.

Still collecting data.

Still making decisions.

Still "for safety."

There's probably a governance lesson hiding in that.

Assurance has to look both ways before crossing the road

Much of AI assurance quite reasonably concentrates on the AI system itself.

Does it behave as intended?

Are known risks controlled?

Is human oversight meaningful?

Are security controls working?

Can consequential outcomes be explained where necessary?

Can the organisation produce evidence for the things it claims to be doing?

We'll be asking those questions for a long time.

I think we'll increasingly need to turn around and examine the machinery we've constructed to make those answers satisfactory.

An AI system can be under-governed.

Its safety controls can be overreaching.

Both can be true at once.

That's what makes "AI safety" such an interesting governance problem. The safer option in one dimension may be the riskier option in another.

More monitoring may reduce misuse whilst increasing privacy risk.

More automated intervention may improve response times whilst reducing contestability.

Longer retention may improve investigations whilst creating an increasingly valuable repository of behavioural information.

Greater control may reduce one form of uncertainty by creating another.

There isn't a universal equation that tells us where the correct balance sits.

I'm rather suspicious of anyone who says there is.

Governance is partly the uncomfortable work of deciding what we're prepared to trade, who gets to make that decision and what evidence would cause us to reconsider it later.

AI safety belongs inside that conversation, not above it.

The phrase should never become a permission slip.

Perhaps that's the part worth remembering as these controls become more capable.

One day we may have AI systems watching AI systems, restricting AI systems and correcting AI systems, all operating faster than a person could reasonably supervise them.

Someone will inevitably describe that architecture as safer.

They may be completely right.

I'd still like to know who decided what safe was allowed to mean.

Hayden Pritchard
Hayden Pritchard

I've spent much of my career helping organizations make difficult decisions about cybersecurity, governance, and risk.

That work has taken me through hospitals, regulated industries, boardrooms, investigations, and more standards documents than I'd care to admit. Along the way I've become increasingly interested in something that doesn't appear in most governance frameworks: how people actually think.

Here, I write essays rather than reports. I explore the ideas that stay with me long after the meeting ends: why frameworks often ask the same questions in different languages, why some human limitations may actually be strengths, and how emerging technologies quietly change the assumptions that regulation depends upon.

Professionally, my work focuses on AI governance, cyber risk, healthcare, and safety-critical systems.

Personally, I'm just trying to understand them a little better than I did yesterday.

https://www.solvingcyber.com
Next
Next

Software Decays. AI Doesn't. (At least, not in the way we think.)