From Volume to Veracity

Hayden Pritchard 

I still remember when “Big Data” seemed to dominate almost every technology conference. 

You could walk through a vendor hall and barely make it past three booths without  encountering a data lake, an analytics platform, or an enormous diagram showing arrows  flowing into a glowing database in the middle. 

Somewhere near the beginning of almost every presentation was the same slide. The Five Vs. 

1. Volume. 

2. Velocity. 

3. Variety. 

4. Veracity. 

5. Value. 

The list changed occasionally, depending on who was speaking and what they happened to  be selling. But those five appeared often enough that most people working in technology  could recite them without much effort. 

At the time, if you had asked me which one mattered most, I would probably have  answered before you finished the question. 

Volume. 

We had discovered that we could collect almost everything. 

Application logs. 

Network traffic. 

Customer transactions. 

Search histories. 

Medical device telemetry. 

Sensor data from machines that nobody had previously thought needed to produce sensor  data. 

The difficulty was finding somewhere to put it all.

Storage was expensive. Processing was slow. Searching across large datasets could  become an afternoon-long exercise in watching progress bars and reconsidering your  career choices. 

The questions were practical. 

Can we store this much information? 

Can we search it? 

Can we move it quickly enough? 

Can we extract anything useful before the data becomes historical evidence of a problem  we should have noticed last Tuesday? 

Those were sensible concerns. In many organizations they remain sensible concerns  today. 

I do not remember Veracity receiving the same attention. 

I am sure it appeared in the presentations. It was right there on the slide. 

We understood that data could be incomplete, inaccurate, duplicated, old, badly  classified, or collected for a purpose nobody could quite remember. Anyone who has  worked with a large asset inventory, customer database, or hospital CMMS export knows  that dirty data did not arrive with artificial intelligence. 

It was already sitting there. 

Waiting. 

Usually in a spreadsheet someone described as the “source of truth” whilst quietly  maintaining three other versions on a shared drive. 

Still, the scale problem felt more urgent. 

First, collect the data. 

Then store it. 

Then work out what to do with it. 

Somewhere in that sequence, perhaps after lunch, someone might ask whether any of it  was trustworthy. 

The assumption seemed to be that more data would eventually produce better answers.

There was a certain logic to that. A larger dataset gave analysts more material to search,  compare, correlate, and investigate. Human beings remained part of the process. We could  question an odd result, examine a source record, or decide that a report did not make  sense. 

Bad data created problems, but it often created problems at human speed. AI changes the speed, scale, and confidence with which those problems can travel. 

I noticed my own attention shifting whilst reading AI governance frameworks over the past  year. 

The documents themselves did not mention the Five Vs in any way that brought back  conference memories. Something else did. 

I caught myself asking a question that would not have occupied much of my thinking  twenty years ago: 

“Can we trust this data enough to let an AI system make decisions with it?” That question has been following me around ever since. 

Imagine an organization has built a model to help identify suspicious financial  transactions. 

The historical dataset contains years of completed investigations. On paper, that sounds  ideal. The model can learn from previously confirmed cases and help analysts recognize  similar patterns. 

Then someone notices that certain customer groups were investigated more frequently  than others. 

Were those groups genuinely associated with greater risk? 

Or were they investigated more often because an old rule, an inherited bias, or a  management decision directed more scrutiny toward them? 

The dataset contains facts about what the organization did. 

That does not automatically mean it contains reliable evidence about what the organization  should do next. 

The distinction matters. 

A model can learn the history perfectly and still reproduce the wrong lesson. Healthcare creates an even less forgiving version of the same problem.

Suppose an AI system helps prioritize patients for review. Its data comes from electronic  health records, diagnostic systems, clinical notes, and previous outcomes. 

The dataset may be enormous. 

It may also contain gaps. 

Some patient populations may have received less consistent care. Certain symptoms may  have been documented differently across hospitals. One department may record a  condition using structured fields while another leaves the same information buried inside  free-text notes. 

The model sees the records it receives. 

It does not see the appointment that never happened, the symptom that was never  entered, or the patient who could not access the service in the first place. 

Volume does not repair those absences. 

It may simply make them harder to notice. 

I have seen similar issues in medical-device environments, long before anyone attached an  AI label to them. 

An asset inventory might contain thousands of devices. That sounds impressive until you  begin asking ordinary questions. 

Which operating system is installed? 

Which software version? 

Who owns the device? 

Is it still supported? 

Where is it connected? 

Some records are excellent. Others contain a device name and little else. Several  apparently different assets turn out to be the same machine recorded under different  naming conventions. 

The organization has plenty of data. 

The practical problem is whether that data supports the decision someone is trying to  make. 

AI does not create that problem. It does, however, give us new ways to automate it.

That is the part I keep returning to. 

A person looking at an incomplete report may notice that something feels wrong. They may  know the hospital, recognize the equipment, or remember that the dataset was assembled  during a rushed discovery exercise five years earlier. 

An AI system has no equivalent institutional discomfort unless someone designs it in. It processes the information provided. 

It may produce an answer that is neatly written, mathematically defensible, and  operationally useless. 

Worse, the answer may look more authoritative because it came from an AI system. 

We tend to associate precision with correctness. A probability score of 87.4 percent feels  more considered than someone saying, “I’m reasonably concerned about this.” 

The decimal places are doing emotional work that the underlying data may not deserve. This is where Veracity stops being something the database team can clean up later. 

Once an AI-supported decision affects a customer, employee, patient, security response,  or business process, uncomfortable questions arrive fairly quickly. 

Who decided the data was suitable? 

Suitable for what? 

Was it collected for the same purpose for which it is now being used? Which populations or situations are underrepresented? 

How old is it? 

What changed after the model was trained? 

What happens when the data no longer reflects reality? 

Who notices? 

And perhaps the most awkward question of all: 

What evidence supports our confidence in the answer? 

Those questions move beyond data engineering. 

They reach risk owners, legal teams, privacy professionals, cybersecurity staff, auditors,  executives, and the people expected to use the system responsibly.

That is governance territory. 

I do not mean that every incorrect database field needs executive oversight. That would  create a governance program capable of producing excellent meeting minutes and very  little else. 

The point is proportionality. 

The greater the consequence of the decision, the stronger the evidence we should expect  around the data supporting it. 

A recommendation about which marketing message to display does not carry the same  weight as a decision affecting employment, credit, clinical care, public safety, or access to  an essential service. 

The data question changes with the consequence. 

So should the governance. 

This is one reason I have become cautious about organizations beginning their AI programs  with model selection. 

Which platform should we use? 

Which vendor has the best capability? 

Should we build or buy? 

Those questions matter, but they may be arriving too early. 

A better first conversation might begin with the information the system will depend upon. Where did it come from? 

Why do we trust it? 

What does it fail to represent? 

The answer is rarely “the data is perfect.” 

That would make me suspicious in a different way. 

The real objective is reasonable confidence, supported by evidence, with enough  monitoring to notice when yesterday’s reasonable confidence starts becoming today’s bad  assumption. 

That is less exciting than a product demonstration.

Nobody fills a conference hall to hear someone announce that their organization has  improved data lineage and assigned an accountable owner. 

Yet those mundane controls often determine whether an AI system remains useful after the  demonstration ends. 

One thing I have gradually come to appreciate is that governance tends to follow whatever  keeps organizations awake at night. 

When storage was expensive, conversations revolved around retention and capacity. 

When breaches dominated the headlines, security and access controls moved closer to  the center. 

Privacy regulation made organizations ask why they were collecting data, how long they  retained it, and whether they had the right to use it. 

AI is pushing trust into the same conversation. 

The world was always on fire. 

We simply have new fuel now. 

Veracity did not suddenly become important because AI arrived. It was always important. AI has made the consequences of ignoring it faster, broader, and more difficult to dismiss. A weak dataset used by one analyst may produce one poor decision. 

The same weakness embedded inside an automated system can produce thousands  before anyone realizes the pattern is there. 

That is the change. 

Not the existence of bad data. 

Its reach. 

I still smile when I see the Five Vs in an old presentation. 

The model was never wrong. 

My attention was elsewhere. 

Back then, Volume felt like the problem standing directly in front of us. We needed storage,  processing power, and tools capable of making large datasets usable. 

Now I find myself lingering on Veracity.

Can we trust the data? 

Can we explain why? 

Would that explanation survive contact with an auditor, a regulator, an affected person, or  simply someone willing to ask an inconvenient second question? 

Ten years from now, another of the Five Vs may move to the center of the discussion. 

Perhaps Value, once organizations become less impressed by merely owning AI and start  asking whether it produced anything worth the cost. 

Technology has a habit of changing which old question suddenly feels urgent. For now, though, I keep returning to Veracity. 

AI did not make it important. 

AI made ignoring it much harder.

Hayden Pritchard
Hayden Pritchard

I've spent much of my career helping organizations make difficult decisions about cybersecurity, governance, and risk.

That work has taken me through hospitals, regulated industries, boardrooms, investigations, and more standards documents than I'd care to admit. Along the way I've become increasingly interested in something that doesn't appear in most governance frameworks: how people actually think.

Here, I write essays rather than reports. I explore the ideas that stay with me long after the meeting ends: why frameworks often ask the same questions in different languages, why some human limitations may actually be strengths, and how emerging technologies quietly change the assumptions that regulation depends upon.

Professionally, my work focuses on AI governance, cyber risk, healthcare, and safety-critical systems.

Personally, I'm just trying to understand them a little better than I did yesterday.

https://www.solvingcyber.com
Previous
Previous

To homelab or not to homelab is not a question

Next
Next

The moment Cybersecurity clicked for me?