From Volume to Veracity
Hayden Pritchard
I still remember when “Big Data” seemed to dominate almost every technology conference.
You could walk through a vendor hall and barely make it past three booths without encountering a data lake, an analytics platform, or an enormous diagram showing arrows flowing into a glowing database in the middle.
Somewhere near the beginning of almost every presentation was the same slide. The Five Vs.
1. Volume.
2. Velocity.
3. Variety.
4. Veracity.
5. Value.
The list changed occasionally, depending on who was speaking and what they happened to be selling. But those five appeared often enough that most people working in technology could recite them without much effort.
At the time, if you had asked me which one mattered most, I would probably have answered before you finished the question.
Volume.
We had discovered that we could collect almost everything.
Application logs.
Network traffic.
Customer transactions.
Search histories.
Medical device telemetry.
Sensor data from machines that nobody had previously thought needed to produce sensor data.
The difficulty was finding somewhere to put it all.
Storage was expensive. Processing was slow. Searching across large datasets could become an afternoon-long exercise in watching progress bars and reconsidering your career choices.
The questions were practical.
Can we store this much information?
Can we search it?
Can we move it quickly enough?
Can we extract anything useful before the data becomes historical evidence of a problem we should have noticed last Tuesday?
Those were sensible concerns. In many organizations they remain sensible concerns today.
I do not remember Veracity receiving the same attention.
I am sure it appeared in the presentations. It was right there on the slide.
We understood that data could be incomplete, inaccurate, duplicated, old, badly classified, or collected for a purpose nobody could quite remember. Anyone who has worked with a large asset inventory, customer database, or hospital CMMS export knows that dirty data did not arrive with artificial intelligence.
It was already sitting there.
Waiting.
Usually in a spreadsheet someone described as the “source of truth” whilst quietly maintaining three other versions on a shared drive.
Still, the scale problem felt more urgent.
First, collect the data.
Then store it.
Then work out what to do with it.
Somewhere in that sequence, perhaps after lunch, someone might ask whether any of it was trustworthy.
The assumption seemed to be that more data would eventually produce better answers.
There was a certain logic to that. A larger dataset gave analysts more material to search, compare, correlate, and investigate. Human beings remained part of the process. We could question an odd result, examine a source record, or decide that a report did not make sense.
Bad data created problems, but it often created problems at human speed. AI changes the speed, scale, and confidence with which those problems can travel.
I noticed my own attention shifting whilst reading AI governance frameworks over the past year.
The documents themselves did not mention the Five Vs in any way that brought back conference memories. Something else did.
I caught myself asking a question that would not have occupied much of my thinking twenty years ago:
“Can we trust this data enough to let an AI system make decisions with it?” That question has been following me around ever since.
Imagine an organization has built a model to help identify suspicious financial transactions.
The historical dataset contains years of completed investigations. On paper, that sounds ideal. The model can learn from previously confirmed cases and help analysts recognize similar patterns.
Then someone notices that certain customer groups were investigated more frequently than others.
Were those groups genuinely associated with greater risk?
Or were they investigated more often because an old rule, an inherited bias, or a management decision directed more scrutiny toward them?
The dataset contains facts about what the organization did.
That does not automatically mean it contains reliable evidence about what the organization should do next.
The distinction matters.
A model can learn the history perfectly and still reproduce the wrong lesson. Healthcare creates an even less forgiving version of the same problem.
Suppose an AI system helps prioritize patients for review. Its data comes from electronic health records, diagnostic systems, clinical notes, and previous outcomes.
The dataset may be enormous.
It may also contain gaps.
Some patient populations may have received less consistent care. Certain symptoms may have been documented differently across hospitals. One department may record a condition using structured fields while another leaves the same information buried inside free-text notes.
The model sees the records it receives.
It does not see the appointment that never happened, the symptom that was never entered, or the patient who could not access the service in the first place.
Volume does not repair those absences.
It may simply make them harder to notice.
I have seen similar issues in medical-device environments, long before anyone attached an AI label to them.
An asset inventory might contain thousands of devices. That sounds impressive until you begin asking ordinary questions.
Which operating system is installed?
Which software version?
Who owns the device?
Is it still supported?
Where is it connected?
Some records are excellent. Others contain a device name and little else. Several apparently different assets turn out to be the same machine recorded under different naming conventions.
The organization has plenty of data.
The practical problem is whether that data supports the decision someone is trying to make.
AI does not create that problem. It does, however, give us new ways to automate it.
That is the part I keep returning to.
A person looking at an incomplete report may notice that something feels wrong. They may know the hospital, recognize the equipment, or remember that the dataset was assembled during a rushed discovery exercise five years earlier.
An AI system has no equivalent institutional discomfort unless someone designs it in. It processes the information provided.
It may produce an answer that is neatly written, mathematically defensible, and operationally useless.
Worse, the answer may look more authoritative because it came from an AI system.
We tend to associate precision with correctness. A probability score of 87.4 percent feels more considered than someone saying, “I’m reasonably concerned about this.”
The decimal places are doing emotional work that the underlying data may not deserve. This is where Veracity stops being something the database team can clean up later.
Once an AI-supported decision affects a customer, employee, patient, security response, or business process, uncomfortable questions arrive fairly quickly.
Who decided the data was suitable?
Suitable for what?
Was it collected for the same purpose for which it is now being used? Which populations or situations are underrepresented?
How old is it?
What changed after the model was trained?
What happens when the data no longer reflects reality?
Who notices?
And perhaps the most awkward question of all:
What evidence supports our confidence in the answer?
Those questions move beyond data engineering.
They reach risk owners, legal teams, privacy professionals, cybersecurity staff, auditors, executives, and the people expected to use the system responsibly.
That is governance territory.
I do not mean that every incorrect database field needs executive oversight. That would create a governance program capable of producing excellent meeting minutes and very little else.
The point is proportionality.
The greater the consequence of the decision, the stronger the evidence we should expect around the data supporting it.
A recommendation about which marketing message to display does not carry the same weight as a decision affecting employment, credit, clinical care, public safety, or access to an essential service.
The data question changes with the consequence.
So should the governance.
This is one reason I have become cautious about organizations beginning their AI programs with model selection.
Which platform should we use?
Which vendor has the best capability?
Should we build or buy?
Those questions matter, but they may be arriving too early.
A better first conversation might begin with the information the system will depend upon. Where did it come from?
Why do we trust it?
What does it fail to represent?
The answer is rarely “the data is perfect.”
That would make me suspicious in a different way.
The real objective is reasonable confidence, supported by evidence, with enough monitoring to notice when yesterday’s reasonable confidence starts becoming today’s bad assumption.
That is less exciting than a product demonstration.
Nobody fills a conference hall to hear someone announce that their organization has improved data lineage and assigned an accountable owner.
Yet those mundane controls often determine whether an AI system remains useful after the demonstration ends.
One thing I have gradually come to appreciate is that governance tends to follow whatever keeps organizations awake at night.
When storage was expensive, conversations revolved around retention and capacity.
When breaches dominated the headlines, security and access controls moved closer to the center.
Privacy regulation made organizations ask why they were collecting data, how long they retained it, and whether they had the right to use it.
AI is pushing trust into the same conversation.
The world was always on fire.
We simply have new fuel now.
Veracity did not suddenly become important because AI arrived. It was always important. AI has made the consequences of ignoring it faster, broader, and more difficult to dismiss. A weak dataset used by one analyst may produce one poor decision.
The same weakness embedded inside an automated system can produce thousands before anyone realizes the pattern is there.
That is the change.
Not the existence of bad data.
Its reach.
I still smile when I see the Five Vs in an old presentation.
The model was never wrong.
My attention was elsewhere.
Back then, Volume felt like the problem standing directly in front of us. We needed storage, processing power, and tools capable of making large datasets usable.
Now I find myself lingering on Veracity.
Can we trust the data?
Can we explain why?
Would that explanation survive contact with an auditor, a regulator, an affected person, or simply someone willing to ask an inconvenient second question?
Ten years from now, another of the Five Vs may move to the center of the discussion.
Perhaps Value, once organizations become less impressed by merely owning AI and start asking whether it produced anything worth the cost.
Technology has a habit of changing which old question suddenly feels urgent. For now, though, I keep returning to Veracity.
AI did not make it important.
AI made ignoring it much harder.