Data to Intelligence

No AI survives
first contact with your data.

Data quality, governance and AI-ready data products — and the process flows they make possible. Because an agent that cannot trust your data cannot be trusted with your business.

Scroll
%

of data and analytics leaders say their data strategy needs a complete overhaul before AI ambitions can succeed

Salesforce, 2025 · n=7,652
%

of AI projects unsupported by AI-ready data will be abandoned, through 2026

Gartner, February 2025
%

of chief data officers are confident their data can support new AI-enabled revenue streams

IBM Institute for Business Value, 2025 · n=1,700
×

more invested in data and analytics foundations by organisations with successful AI initiatives

Gartner, April 2026
The readiness gap

Everyone says they are AI-ready. Their data says otherwise.

The pattern across every credible survey of the last eighteen months is the same: confidence is accelerating faster than readiness. Organisations are deploying agents on top of data estates they would not sign off on for a regulatory return.

The paradox

Ready, and also not ready

88% of data leaders report having data readiness. In the same survey, 43% name data readiness as their single biggest obstacle and the most significant barrier to AI alignment. Both answers came from the same people.

Precisely / Drexel LeBow, January 2026 · 500+ leaders

The evidence

The valuable data is the unusable data

19% of company data is siloed, inaccessible or otherwise unusable — and 70% of data leaders believe their most valuable insights sit inside exactly that 19%. The average enterprise runs 897 applications; only 29% are connected.

Salesforce, November 2025 · n=7,652

The consequence

It shows up in production

89% of organisations with AI in production have already experienced inaccurate or misleading outputs, and 55% report wasting significant resources training models on bad data. Meanwhile 71% of practitioners worry about hallucinated outputs reaching stakeholders.

Salesforce, 2025; dbt Labs, April 2026

Data quality

Quality is not a feeling. It has six dimensions, and each one has an owner.

We measure against the six primary dimensions defined by the DAMA UK working group, applied within the DMBOK2 knowledge areas — and, where maturity needs a defensible score for a board or a regulator, the EDM Association's DCAM. Every dimension is measured per critical data element, and every failure is priced in business terms rather than rated as a severity.

Completeness

Is the data that should be there, there? Measured against what the process needs, not against what the schema allows.

Fails as: blocked automation, manual chase

Uniqueness

Is each real-world thing recorded once? Duplicate customers, assets and vendors are the quiet tax on every downstream number.

Fails as: double counting, wrong entitlements

Timeliness

Is it current enough for the decision it supports? A daily batch is fine for a board report and useless for an agent acting in a queue.

Fails as: agents acting on stale state

Validity

Does it conform to the rules, formats and reference values it is supposed to? Validity is the dimension that automated rules can actually enforce at the point of entry.

Fails as: silent integration rejects

Accuracy

Does it describe the real world correctly? The hardest and most expensive dimension, because it usually requires a source of truth outside the system.

Fails as: confidently wrong outputs

Consistency

Does the same fact agree across systems? Where finance, operations and the CRM each hold a different version, the model will learn all three.

Fails as: irreconcilable reporting

A note on a figure you will see quoted at you: Gartner's "$12.9 million average annual cost of poor data quality" is real, but the underlying research is from 2020. The "$3.1 trillion" figure is an IBM estimate for 2016. We use current, sampled evidence instead — and we measure your own cost rather than borrowing someone else's.

Data flow → process flow

Your data model decides which process you are able to see.

Every process leaves a trace. An order raised, a delivery booked, an invoice matched, a payment cleared — each one is an event, attached to an object, written to a table with a timestamp. That trace is the only honest record of how the organisation actually runs.

The catch is structural. Conventional analysis forces every event into a single "case", but real operations do not work that way: one order becomes three deliveries and two invoices. Flattening that produces convergence and divergence distortions — work counted twice, or not counted at all. Object-centric event data models events against multiple related objects, so the process you analyse is the process that happened.

Which is why data work and process work are the same work. Fix the object model and the process becomes visible. Leave it broken and every downstream artefact inherits the distortion — the dashboard, the process map, the business case, and the context you hand to an agent.

82% of decision-makers believe AI will fail to deliver ROI if it does not understand how the business runs. 85% want to be an agentic enterprise within three years; 76% admit their current processes are holding them back.
Celonis, February 2026 · n=1,649 · vendor-commissioned research, large sample
SOURCE SYSTEMS ERP CRM EAM · IoT Documents OBJECT-CENTRIC EVENT LOG ORD DLV INV PAY events → many objects PROCESS GRAPH cycle time · rework · conformance AGENT · DECISION · AUTOMATION needs data context and process context — both, or neither works

Object-centric event data (OCEL 2.0) as the bridge between the data estate and the operating process.

AI enablement

AI-ready is a property of the data, not a property of the model.

Gartner's definition is usefully strict: data must be representative of the use case — of every pattern, error, outlier and unexpected emergence needed to train or run a model for that specific use. It is a practice, not a state, and it rests on metadata good enough to align, qualify and govern the data underneath.

Which means the work is unglamorous and entirely tractable: semantics, ownership, contracts, lineage, and evidence that any given answer traces back to a governed source.

DataInformationInsightDecisionAction capturedstructuredunderstoodownedexecuted most organisations stop here

Dashboards are an insight artefact. Value is only booked at the last two steps.

  • Semantics before scale. Gartner expects organisations prioritising semantics in AI-ready data to increase GenAI model accuracy by up to 80% and cut costs by up to 60% by 2027. dbt's April 2026 benchmark is more concrete: on modelled data, a leading model answered 84.1% of questions correctly through text-to-SQL and 100% through a semantic layer.
  • Metadata that does something. Active metadata that raises a recertification alert when data drifts, rather than a catalogue that nobody opens after the launch email.
  • Contracts at the boundary. Producers commit to schema, semantics, freshness and quality; consumers build against a promise rather than an observation. Gartner expects half of organisations to have agents translating governance policy into machine-verifiable data contracts by 2030.
  • Grounding, not guessing. Retrieval and knowledge graphs anchored in governed data products, so an answer can be traced to a source — the difference between a model that is useful and one that is merely fluent.
  • Zero trust for machine-generated data. Gartner predicts half of organisations will adopt a zero-trust posture for data governance by 2028: "Organizations can no longer implicitly trust data or assume it was human generated."
The honest caveat. The same dbt benchmark shows a semantic layer scores zero on questions outside its modelled scope, where text-to-SQL manages 70%. The lesson is not "semantic layer wins" — it is that coverage and modelling are the work, and they are never finished.
Australian obligations

Three dates worth putting in the board pack.

Data governance stopped being a hygiene topic when the obligations attached to it acquired penalties and commencement dates. These are current as at August 2026.

10 December 2026

Automated decision transparency

New APP 1.7 requires privacy policies to disclose the use of computer programs that make decisions affecting individuals' rights, including the kinds of personal information those systems use. Complying means first knowing where automated decisions actually happen, and what feeds them — which most organisations do not.

Already in force

APRA CPS 230

Commenced 1 July 2025; transitional relief for pre-existing material service provider contracts expired 1 July 2026. Critical operations need tolerance levels that explicitly include maximum acceptable data loss, and material service provider agreements must specify data ownership, audit access and termination rights.

Since 4 April 2025

SOCI data storage systems

Responsible entities must identify and manage risks to data storage systems holding business critical data inside their Critical Infrastructure Risk Management Program — with impacts to availability, integrity, reliability or confidentiality designated a material risk.

Context worth carrying: 1,205 notifiable data breaches were reported in 2025 — an all-time high, up 8% on 2024. In the first half of the year, human error accounted for 37% of breaches, up from 29%. That is a process and data-handling failure, not a cyber one. The statutory tort of serious invasion of privacy has also been available since 10 June 2025. Privacy Act Tranche 2 reforms have been proposed but not introduced — we will tell you what is law and what is not.

Our method

Six steps from data estate to something an agent can be trusted with.

01

Landscape and lineage

What data exists, where it originates, who touches it on the way through, and what depends on it downstream. Lineage first, because everything else is guesswork without it.

02

Quality baseline

All six dimensions measured per critical data element, with the business impact of each failure priced in dollars rather than rated red, amber or green.

03

Governance operating model

Ownership, stewardship, decision rights and data contracts — designed as an operating model, not a policy document. DMBOK2 for scope, DCAM where maturity needs a defensible score.

04

Data products and semantics

Governed, versioned, documented data products with named consumers, quality SLAs and a shared business meaning — the semantic layer that both people and models query against.

05

AI enablement

Retrieval grounded in those products, agent context assembled from both the data model and the process model, guardrails proportional to autonomy, and an evaluation harness that runs continuously.

06

Observe and sustain

Freshness, distribution, volume, schema and lineage monitored as a matter of course, with active metadata triggering recertification when something drifts. Handed to your team, not retained by ours.

How we work

We fix the foundation in the order the value arrives.

Not every data problem is worth solving. We start from the decisions and processes that carry value, trace back to the data those depend on, and fix that — rather than launching a three-year enterprise data programme that improves everything slightly and nothing measurably.

Gartner's finding is the argument for doing it properly: organisations with successful AI initiatives invest up to four times more of their revenue in data and analytics foundations, and the highest-maturity organisations achieve up to 65% greater business outcomes. The foundation is not overhead. It is the multiplier.

Data quality diagnosticGovernance operating modelAI-readiness assessmentData product designSemantic layer buildPrivacy & ADM readiness
What we commit to
Critical data elements identified and ownedBefore AI use cases proceed
Quality measured across all six dimensionsPer critical data element
Business impact of each quality failure pricedIn dollars, not RAG ratings
Data products shipped with consumers and SLAsEvery product, no exceptions
Automated decision points inventoriedAhead of December 2026
Observability live before handoverAll five pillars

Engagement commitments. Client-specific results provided as references under NDA.

Where to start

Four ways in, depending on how urgent the AI conversation has become.

3–4 weeks

Data Quality & AI-Readiness Diagnostic

Critical data elements identified, six dimensions measured, failures priced, and an honest read on which AI use cases your data can currently support.

Start here
6–8 weeks

Data Governance Operating Model

Ownership, stewardship, decision rights, contracts and the forums that make them real — designed to be run by your people, with a maturity baseline you can re-measure.

Enquire
6–12 weeks

Data Product & Semantic Layer Build

Governed data products for a defined domain, with a semantic layer that both analysts and models query against — and the contracts and tests that keep it honest.

Enquire
3 weeks

Automated Decision Readiness Review

Where automated and AI-assisted decisions currently happen, what data feeds them, what has to be disclosed, and what has to change before 10 December 2026.

Enquire
Common questions

Before you ask

No, and organisations that try usually never ship. You fix the data that the valuable decisions depend on, in the order those decisions arrive. The failure mode is the opposite one: scaling an agent across an estate where nobody can say which number is right.

Not necessarily. Platform choice is downstream of ownership, semantics and quality. We have seen well-architected platforms fail because nobody owned the data in them, and modest platforms work well because the governance was real.

Whichever fits your domain structure and operating model. Data mesh's four principles — domain ownership, data as a product, self-serve platform, federated computational governance — are sound, but they describe an operating model, and adopting them without changing accountability produces the same silos with new labels.

Directly. Event data is the bridge: your data model determines which process you can observe, and your process model determines what context an agent needs. We run the two together rather than as sequential programs.

Yes — DMBOK2 knowledge areas for scope, DCAM where you need a scored and benchmarkable maturity assessment, ISO 8000 where master data exchange is the issue, and the relevant Australian obligations where the driver is regulatory.

It is the emerging governance problem and it is why zero-trust postures for data are being adopted. Provenance marking, human-verified reference sets and recertification triggers all matter. We design for it now rather than retrofitting it after the first model-collapse incident.

Fix the foundation before you scale the agent.

A data readiness diagnostic takes three to four weeks and tells you which AI use cases your data can support today, which need work, and what that work costs.

Sources: Gartner, Salesforce, IBM IBV, Precisely/Drexel LeBow, dbt Labs, Celonis, DAMA, EDM Association, OAIC, APRA. Full source list on request.