Vibe Coding Is Creating a Data Governance Problem

If you are anything like me, you are excited about this AI revolution.

Things that once took hours of research, learning and experimentation can now be explained in plain language, using your own context to help you understand them. No more searching for that video you vaguely remember seeing, only to discover it doesn't quite apply to your problem.

More importantly, we are now seeing real-world examples of businesses using AI to uncover insights from their data that previously would have required teams of analysts and significant amounts of time.

The promise is incredibly compelling:

I understand my business. I have all this data. Now AI can help me put the two together and find answers.

Which brings us to a very real and increasingly urgent problem… the data itself. If your company is like most modern organizations, you are swimming in it.

You probably have multiple databases, warehouses and data lakes, often sitting in departmental silos. At the same time, increasingly capable employees are using AI tools to explore that data, build applications and answer questions that previously would have required specialized technical skills.

That's exciting.

But it also creates a new question:

How do we know everyone — including the AI — is working from the same ‘songbook’ so to speak?

AI Doesn't Fix Ambiguous Data

The early AI success stories can make this look deceptively easy.

Give AI access to your data, write a concise prompt, and suddenly you have deep, previously inaccessible insights about your company. There is some truth to that.

What often gets left out is that many successful early adopters of AI were also early adopters of good data management. They had already done much of the difficult work required to get their data house in order.

They have curated data sources. Their data has been cleaned, organized and described so that an AI model has enough context to interpret it correctly.

More importantly, they have agreed on what their data means.

  • What is revenue?

  • What is profit?

  • When is an account considered overaged?

  • Who counts as an active customer?

Those sound like simple questions until you ask five different departments to define them.

AI doesn't make those disagreements disappear. In fact, AI can make the problem worse because it dramatically increases the number of people who can independently interrogate the data.

Vibe Coding Changes the Scale of the Problem

This is where vibe coding becomes particularly interesting.

We are rapidly removing the technical barriers between people and data.

Someone no longer needs to be a SQL expert, data engineer or BI developer to create an analysis or even build an application against corporate data. AI can increasingly handle technical mechanics.

That is an enormous productivity opportunity, but if ten people can now independently build solutions against five different versions of customer, revenue and profit, we haven't democratized trusted analytics.

We've democratized inconsistency.

The challenge is no longer simply giving people access to data. It is giving people — and AI agents — access to data with enough business context and governance that the answers remain consistent.

This Is What Semantic Models Are For

A modern semantic model sits between your raw data and the people, applications and AI agents consuming it.

It defines the business meaning of the data: relationships, metrics, terminology, calculations and other rules that previously lived inside dashboards, SQL queries, spreadsheets and, far too often, people's heads.

It can also become an important part of the governance layer because there is another problem with simply pointing an AI at corporate data:

Not everyone should be allowed to see everything.

Security, permissions and data access don't disappear because the interface has changed from a dashboard to a conversation with an AI agent. If anything, they become more important.

A useful semantic layer therefore needs to answer two very different questions:

What does this data mean?

and

Who is allowed to see it?

Getting those questions right is becoming foundational to an organization's AI strategy.


Don't Start With the AI Tool

This is where I think organizations need to be careful.

It is tempting to start an AI analytics initiative by choosing an LLM, connecting it to some data and seeing what happens.

That makes for a great demo. It doesn't necessarily make for a great production system.

Before scaling AI access to your data, spend time understanding the semantic layer that will sit between the two.

That doesn't mean embarking on a multi-year data governance project before anyone is allowed to experiment. Quite the opposite. Start small, pick an important business domain and build a semantic model that both humans and AI can understand. Treat the model itself as a product.

That means working through business definitions, relationships, granularity, security, metadata and the inevitable disagreements between how different parts of the organization interpret the same data.

And if this isn't already a core competency within your organization, work with someone who understands semantic modeling — not just the AI platform sitting on top of it.

The important expertise isn't simply knowing how to connect an LLM to a database. It is knowing how to create a trusted business representation of that data so that dashboards, analysts, applications and AI agents can all use the same definitions.

The organizations that succeed with AI won't necessarily be the ones with the best prompts.

They'll be the ones whose AI understands their business well enough to give answers they can actually trust.

Next
Next

Question #2 - Is gold a good investment during inflationary times?