Season 4 · Episode 10

Metadata to Knowledge Graphs: Mapping the Full AI Data Stack

Sanjeev Mohan

Sanjeev Mohan

Principal, SanjMo

in

A former Gartner analyst and author of Designing the AI-Driven Data Foundations: Architecture, Principles, and Practice, Sanjeev has spent decades helping enterprises build the metadata, semantics, context, and data products that every successful AI initiative depends on.

Satyen Sangani

Satyen Sangani

CEO & Co-Founder, Alation

in

As the Co-founder and CEO of Alation, Satyen lives his passion of empowering a curious and rational world by fundamentally improving the way data consumers, creators, and stewards find, understand, and trust data. Industry insiders call him a visionary entrepreneur. Those who meet him call him warm and down-to-earth. His kids call him “Dad.”

Introducing Sanjeev Mohan and the messy middle of the data and AI stack

The AI Radicals producer introduces guest Sanjeev Mohan, former Gartner analyst and author of Designing the AI-Driven Data Foundations, and previews his conversation with Alation CEO Satyen Sangani on context, governance, and data products.

[00:00] Producer: Welcome back to AI Radicals. Our guest has spent his career untangling the messiest layer of the data and AI stack: what data actually means, who owns it, and who gets to trust it. Sanjeev Mohan is a former Gartner analyst and author of Designing the AI-Driven Data Foundations. He spent decades helping companies figure out the plumbing beneath every AI initiative: metadata, semantics, context, and the data products built on top of them.

[00:27] Producer: Satyen and Sanjeev dig into why context layer thinking breaks down once agents enter the picture, why governance now has to cover not just data, but agent behavior itself, and why the next era belongs to data people who think like business owners instead of technicians. If you've ever wondered why an AI system can have perfectly clean data and still give you the wrong answer, this one's for you.

[00:50] Producer (sponsor message): This show is brought to you by Alation. Only 5% of enterprise AI pilots from 2025 delivered ROI in production. Alation's free agentic AI opportunity discovery guide shows you exactly how to find the right processes to automate and build AI that actually sticks. Get it at alation.com/ai-guide.

Why do data foundations still matter in the AI era?

Alation CEO Satyen Sangani welcomes back Sanjeev Mohan, former Gartner analyst, who explains why he wrote Designing the AI-Driven Data Foundations and why he treats AI as a use case sitting on top of the data substrate.

[01:15] Satyen Sangani, CEO, Alation: Sanjeev, welcome back. The last time you joined us, this show was still called Data Radicals. You are, I think, maybe the second two-time guest that we've now had on the show, so you have quite a distinguished role right now. We're now, of course, AI Radicals, and we're talking all about data products and data governance, and now AI agents.

Since then, you've also written a book, Designing the AI-Driven Data Foundations. Why don't we jump right into that and talk about why you wrote the book, what motivated it, and why you, I guess, felt compelled to put it out at this point in time?

[01:46] Sanjeev Mohan, Principal, SanjMo: So, Satyen, I have found myself in many conversations where I feel I represent the data foundation part to this data and AI equation, and I get pushback from people who are saying that, "Wait, why are you talking about data? No one wants to talk about data. It's all about AI." And I'm like, "No, it's not all about AI," because AI, in my mind, is a use case sitting on top of data, and tomorrow it may not be AI. Maybe it'll be robots. Maybe the day after, it'll be quantum. So the technology will keep shifting, but data is a substrate or the foundation.

If we don't get that correct, then the house of cards on top of it falls down. And I felt compelled to demystify, clarify. There's so many topics that we completely conflate almost daily these days: context, semantics, ontology, data and AI governance, and I just felt that I needed to explain how this magic happens.

It's not like you ask a question into a Claude or a chat window, and magically, it gives me all kinds of answers. There's a lot of work that needs to go. The engineering behind it is what I capture in my book.

Metadata, semantics, and context, defined

Satyen Sangani of Alation asks Sanjeev Mohan to define the core vocabulary of AI data foundations. Mohan starts with metadata, which he breaks into business, technical, operational, and social types of metadata, and explains how technical metadata drives data lineage.

[03:05] Sangani: Yeah, it's really ironic, I guess, that a space that's so concerned with semantics has such messy semantics.

And you've done a really amazing job. You had this one blog which I thought was awesome, and it was basically running us through these definitions somewhat carefully and thoughtfully and closely, but only as an astute practitioner plus analyst could do, because you have to look across all of these things.

Tell us what it is that... Give us sort of this: what are the core nouns that we're talking about? Give us these definitions. What is a semantic layer? What is an ontology? What is a knowledge graph? How should somebody think about these things? Let's start with that.

[03:47] Mohan: Okay. So the starting point is metadata. That's the atoms or the molecules that make up everything. Metadata, as you know, is data about data, so it's everything to do with data. For example, if I'm talking about a record that came in to my transactional database, what are the columns? What are the column types? What are the descriptions? How's the data distributed? Who owns that data, and who's allowed to see it? So metadata gets broken down into business metadata, technical metadata, operational metadata, even social metadata. So that's the lowest level. Now, technical metadata is what we use to drive lineage, for example.

So I can see data came from, let's say, SAP, and then it landed through some data integration software into Snowflake, and then somebody did data quality, then it was cataloged, did my lineage. All of that is technical metadata.

How does a semantic layer turn metadata into business meaning?

Sanjeev Mohan explains semantics as the translation of raw technical metadata into business language, covering business glossary terms, business rules, and KPI definitions such as net retention rate (NRR).

Mohan: Businesses, when they have to use the data, they don't speak cust_ID_Q is for a customer ID. They speak business language. So when we start translating the raw metadata into business speak, it's not just for columns and their business glossary, but it's also for: what is the rule for performing a certain job? What is the definition of a NRR, net retention rate, for example? Even who's construed as an employee within a company can be debated between HR, payroll, and sales teams.

So now we're talking about semantics. So semantics is, like the word says, it's a meaning of the data that the businesses can understand. So now they can easily ask a question in their own language, and they can start getting answers. So that's the next layer. Now let's talk about context, because this is where things get very interesting.

What's the difference between semantics and context for AI?

Sanjeev Mohan defines context as the engineered, often real-time, 360-degree view that combines structured and unstructured data for a large language model (LLM), the idea behind what Alation calls an enterprise context layer.

[05:48] Mohan: So semantics do an amazing job of translating technology and business into something that's consumable, but it's still limited to my SAP or Salesforce or all these other systems that I have. What if I want to ask now in natural language, using the power of LLM, I can combine quantitative and qualitative queries into one.

So I can say, "Show me which salespeople did not perform in a certain region for a certain time, for a certain SKU." All deterministic, but give me a performance improvement plan for these people which is very specific to the geography they are from, what customers' buying habits are, what are they tweeting, what is the unstructured data saying, what emails came in, what Slack messages happened.

Now we are talking about giving AI this complete 360-degree view of a customer, a product, or a task. That is context. So context moves into the whole engineering space. So in semantics, we just had to do the mapping and then maintain it in a catalog. But now in context, we are doing a lot of engineering to bring this together.

It's not an easy job. A lot of it is real time. A lot of it can change from person to person for the same question. So those are the three basic must-haves, in my opinion: metadata, semantics, context. Then we get into a little bit, if I may take some liberty, the esoteric world where we now start talking about taxonomy, ontology, and knowledge graph.

Taxonomy vs. ontology vs. knowledge graph: what's the difference?

Sanjeev Mohan distinguishes a taxonomy (a hierarchy), an ontology (defined relationships), and a knowledge graph (an ontology instantiated with real data). Alation compares these structures in its guide to semantic layer vs. ontology.

[07:28] Mohan: Taxonomy is where we start defining hierarchies. That's all it is. It's just a hierarchy. Like the blog you mentioned, I talk about computers. Under that we have laptops. Under that we have Mac, sorry, Apple. We have Dell. We have HP. Under Apple, we have MacBook. We have Mac Air. We have... So that's a hierarchy, which is taxonomy.

Ontology is the super interesting thing. That's where we start defining relationships. Who is allowed to buy? For example, every order must have a customer. Every customer must have an address. And then there's shipping details. So when you start defining the relationship, now you talk about ontologies.

When you take this ontology and you start instantiating it with actual data, Satyen bought an Apple MacBook from Apple Store in Sunnyvale on this day for this price, this is a serial number. When you start creating this linkage of context, we get a knowledge graph.

What's the difference between technical and business metadata?

Satyen Sangani of Alation recaps Sanjeev Mohan's framework, separating technical metadata (lineage, indices, SQL joins) from business metadata (definitions, owners, tags) that typically lives in a data catalog, then moves on to semantic layers and business rules.

[08:33] Sangani: Let's take this back to you. So level one is metadata.

Mohan: Yep.

Sangani: Metadata is data about data.

Mohan: Yep.

[08:40] Sangani: There are two types of metadata. There is technical metadata, which is all about the technology on which these things are implemented. It tells you about things like lineage, like what the actual data pipeline was. It tells you about indices. It tells you about specific joins and expressions inside of SQL. So this is technical metadata. Then you have business metadata, which is literally the superimposed knowledge on top of that specific technical data set.

Mohan: Yes.

Sangani: Like column definitions, table names, table descriptions, owners, tags, just things that happen, and these things mostly happen in catalogs. Right so far?

Mohan: Yep, absolutely.

[09:18] Sangani: Okay. And then there's the next level, which is this kind of semantics. And so this is a hotly debated or discussed subject. There is this idea of semantic layers, which goes all the way back to the BusinessObjects universe. And semantics are superimposed on these data sets. Is that roughly right?

[09:40] Mohan: Yep. And rules, KPIs, all those things.

[09:44] Sangani: And a rule would be... what's a good example of a rule?

[09:48] Mohan: So a rule could be the customer age must be between X and Y. So if somebody enters customer age which is like 400 something, the rule will catch it.

How do data products relate to a semantic layer?

Sanjeev Mohan, who has also written a book on data products, positions data products as orthogonal to the metadata-to-knowledge-graph hierarchy. In his definition, a data product packages data with a user experience, service-level agreements (SLAs), and data contracts.

[10:01] Sangani: Okay. And you wrote a book on data products. How in your taxonomy... I should probably not... How in this language model do data products relate to semantic layers, in your view?

[10:15] Mohan: So data products are how you consume the data. So it's sort of a layer above that. And I won't put it in the same taxonomy, because a data product could be a report, a dashboard, a machine learn... It could be a view. So data product is sort of orthogonal to this. It uses the semantics and ontology to build a data product, but the purpose is different. So the idea is, instead of having people, and now in the future agents, go straight to my semantics and my data, if I have a data product, then I've got not just information about data, but also a user experience.

I have data contracts. There are SLAs involved with that. I know, for example, when was this data last refreshed, and how recent do I expect this data to be? I may even document what is the quality of this data, because not everything can always be 100% data quality in spite of our best efforts. So that kind of layer, we call a data contract.

So the data products with a contract sits orthogonal to this hierarchy of metadata to knowledge graph that we talked about. It uses it. It's a consumer, because it's sitting on top.

Data products as a portable semantic layer

Satyen Sangani explains how Alation defines a data product: a portable semantic layer that bundles business rules, data contracts, dimensions, metric definitions, and evaluations (evals), with apps treated as separate, ephemeral consumers. Alation delivers these through its data products marketplace.

[11:40] Sangani: Yeah. So it's interesting: in Alation land, which obviously is a land unto itself, we've defined data products, but we've defined them differently from how you've defined them.

And I think this is important for Alation people who are listening to this podcast. We've defined data products as a portable semantic layer which contains those rules that you reference. You know, customer age must be between 20 and 100. Equivalently, those rules contain things like those freshness contracts and guarantees.

So we classify those rules under this broader construct of data contracts, and data contracts are part of the data product, as well as these dimensions and metric definitions. And we've additionally done something different, which is that we've also attached evals to this data product.

[12:27] Mohan: Hmm. Very nice.

[12:28] Sangani: Different from your definition, we've basically said, look, the app that lives on top of it, which could be a dashboard, it could be an automation, it could be just simply an app, a real application, like what you might do, that is a distinct thing that can live on top of a data product in our typology.

This is our language model, but it's important to talk about semantic differences, but it's not the same. Our belief in this world is that in AI land, the app is ephemeral and dumb, and the logic needs to live with the data because the rules live with the data, to your point.

And so we've constructed this data product as basically this kind of portable semantic layer that whoever's reading it now can know the rules around how to use this data and therefore can build the app correctly. I think it's an important difference to call out, and I think it's good that we're talking about these things, 'cause when I talk about customers, they're so confused often about what we mean by data product and, you know, whatever.

We've called it what we've called it. You know, we call it Sam, you call it Henry, but these are, I think, important differences to explore.

Is a data product a thing or a philosophy?

Sanjeev Mohan says his and Alation's definitions are close, offers a complex business rule (what counts as a customer), and argues a data product is a philosophy of applying product management practices such as ownership, versioning, and backward compatibility.

[13:30] Mohan: So Satyen, I want to add two things to it, because there is actually... our differences in understanding are not very that apart. So first of all, let me clarify.

I gave a very silly rule of age, because that could be a data quality rule too. So let me clarify a more technical, business-related rule. For example, a rule could be: what is a customer? So somebody may say a customer is somebody who's bought certain product in this range, above a certain minimum threshold, and has had this product for 30 days or 90 days, has not returned the product, is not a renewal of something, so you're not double counting.

Because if we don't do that, then different departments will count different number of customers. So I just want to clarify that rules can get very complex based on company to company. Coming to your second point about semantic layer, calling it, packaging it as a data product. One definition of data product... and in fact, I was alluding to this when I said a data product could be an existing artifact like a report, a dashboard, a machine learning model, or even a view.

It could be a semantic product too. One definition is that data product is a philosophy. It's not a thing. And what data product says is that when you build your artifact, you build it with the product management best practices. For example, there's a data product owner. It has a version number, and it's backward compatible.

It's packaged with SLAs and contracts. So if you are packaging your semantic layer and you're building it using the product management best practices, then that's a data product. Very much so.

Why isn't every dashboard a data product?

Satyen Sangani of Alation argues that a Tableau dashboard or a quick AI-generated artifact usually fails the data product test, because it is built to solve one problem without the reusability and customer intent a product requires.

[15:21] Sangani: When people call a dashboard... let's take an arbitrary Claude artifact that the average person builds, or some Tableau dashboard.

Mohan: Yep.

[15:29] Sangani: When people call that a data product, the reason why I would disagree with that framing is because a product is built with this construct of reusability.

Mohan: Yes.

[15:40] Sangani: And there is an intention that there is a customer and that there's an evolution to that thing. Now, in theory, you're right, a dashboard could be that, and so in some sense the data product could contain that.

It could contain all those things that I'm talking about being a data product. But I think the differences are often that when people build these things, they're not thinking reusable. They're just thinking, "Let me solve my problem."

Mohan: Exactly.

Sangani: And "let me solve my problem" does a thing, but it's only when you've discovered these use cases and started to think about what the rules around the data are... it's really data-centric.

Why should acceptable-use rules live with the data?

Satyen Sangani shares how a chief data officer (CDO) at a major airline restricts customer preference data from marketing campaigns. Sanjeev Mohan cites a reported Google purchase of Spirit Airlines data that excluded customer preference data.

[16:17] Sangani: Let me give you an example. Let me not talk meta. I was talking to a CDO at a very significant airline, and he was like, "Look, people, when they're building these marketing campaigns, should not necessarily be able to use certain data about the customer, like what their preferences are.

And we have a lot of information about the customer, but they shouldn't be using that information in the marketing campaign." Now, the person building the dashboard might just build a dashboard. The person building the product has to think about what are acceptable uses of this overlying thing.

And that's why I distinguish between the app and the data, because the data has to be managed in its own particular way and life cycle and usability, which I think is traveling separately from the app. That's my argument. Does that make sense?

[16:59] Mohan: 100% makes sense. Google just bought actually 40 years of data from Spirit Airlines for $10 million or something like that. You saw that, right? Spirit Airlines went bankrupt, so they bought the data, but they're not allowed to buy any of the customer preferences data. So it's all anonymized, because there's a lot of angst in the market that, "Whoa, wait, all my history, my travel is now owned by Google?" But there's no personal stuff.

Who should own a data product, business or IT?

Sanjeev Mohan recalls data warehouse projects whose teams disbanded and left no owner, a gap that a defined data product lifecycle closes. Satyen Sangani of Alation argues ownership is usually a business and technical partnership.

[17:27] Mohan: And why I actually was compelled to write about data product a couple of years ago was because, in my previous life, when I used to build data warehouses, reports, dashboards, what used to happen was we would bring a team together, they would do work on the data warehouse. After the project finished, the team would disband and go in different directions onto other projects.

Now the end user would run into a problem and would call and say, "Well, I..." There's an error that needs to be fixed. Who is responsible for this? Nobody, because it's a project and the project finished. A product has a life, like a beginning and end of a life, and so it's reusable. There's an owner.

And what you package under that principle is a data product. So I think we are in sync, and I've seen this firsthand. With a data product, business should own the data product, not technical team, in my opinion.

[18:23] Sangani: No, no, I agree. Well, I think there's a combined ownership, because... I mean, we've tried to find in our customers these singular data product owners.

The problem is that there's the business knowledge of the data product and what it needs to do and what it needs to say, but there's also this technical set of requirements: how fresh is the data? What is the actual route, pipeline, and lineage around the particular data set?

And those questions cannot generally be answered by a singular person, and so it tends to be this partnership. That's what we, at least, have seen when people are truly living this methodology. I mean, the reason, by the way, why this is, I think, so important in this world, or at least in our world, is that we're literally...

Solving dashboard sprawl at the data product level

Satyen Sangani previews Alation's intelligence feed, being released at the revAlation conference, and argues that the "Tower of Babel" of thousands of unused dashboards must be solved with reusable data products. Sanjeev Mohan adds that project-era bloat left no one accountable.

[19:03] Sangani: So next week is our revAlation conference. You're gonna come and talk about it. I know this podcast is gonna be aired after that conference, but one of the things that we're releasing is this thing called an intelligence feed, which is basically the same as any Claude artifact, like a dashboard that you create.

You can create these things so easily, so quickly, and it's literally two sentences, and I get the dashboard that would've been as beautiful as anything that I could have created in Tableau, but for two differences. One, it obeys all of the rules of the data underneath it, 'cause it's built on a data product, and two, it has live provenance, that it can be referenced consistently, as opposed to these artifacts that are created as these HTML things that are totally off to the side.

Those differences, I think, are so critical, but you can create these things quickly, and you can just kill them quickly. And people are so frustrated in customers, 'cause they're just like, "I have a thousand dashboards, don't even know who uses them, what uses them, why they exist." And then you have this... it's a Tower of Babel. And the entire problem in this world is that we're trying to fix this Tower of Babel, and I think part of this is you can't solve it at the dashboard level. You've got to solve it at this reusable product level.

That's why I have so much space and religion around this. But...

[20:11] Mohan: And this is my whole point about the team got disbanded and moved on, and all those things just stayed on and on, and there were temporary tables that were created. And now there's a bloat within an organization, and nobody knows who wants these dashboards.

In a product, data product world, you know who owns it, and there's a life. Otherwise, it's just...

Sangani: Yeah.

Should you govern the data or the AI agent?

Sanjeev Mohan raises AI governance as one of the hottest topics as agents go rogue. Satyen Sangani of Alation argues the "stop sign" for an agent belongs with the data itself, where tools such as Alation data governance apply policy.

Mohan: So just moving on from this topic.

Sangani: Yes.

Mohan: You were asking me what is going on, what are the important things I'm hearing. There's a concept 100% related to the whole idea of context, data products, and that is governance.

And governance is hugely critical, especially as we are recording this. We are starting to see how many cases of agents are going rogue, and they're breaking free and doing all kinds of unexplainable things that they were not designed to do. So governance has become one of the hottest topics right now, and it's a step above from data governance.

Data governance we've been doing or not doing or struggling or whatever. Now we are in the realm of AI, which is a very different beast to govern.

[21:20] Sangani: Yeah, I couldn't agree more. And I think the thing about it is... take this example of this airline further, right? You can have agents, and the agents are saying, "Okay, I want to build this fabulous marketing campaign."

And that's what the agent is knowing to do, and it's gonna go do it, and it's gonna do it intelligently and thoughtfully. And damn, these things are powerful, so maybe they're gonna find as much data as they can to get this stuff going. And then all of a sudden it hits this data, which has customer preference data, and all of a sudden it's gonna use it.

And now, without the appropriate governance of the underlying data... Now, is that governance of the agent or is that governance of the data? I would argue it's governance of the data, because the agent shouldn't know what the rules are for how to... It's like, we don't necessarily have to know all of the rules.

We just have to be told them. And so the stop sign has to happen with the data itself, at the intersection where the things are happening. But the point that I would make is that I think this governance is about the context, which is, I think, data products. Data products are a form of context.

And I would argue at the data level itself, these rules are effectively also governance. And you kind of have to have the full chain governed in order to be able to guarantee that the agents are correct.

Why isn't clean data enough to govern AI?

Sanjeev Mohan explains that deterministic applications made clean input sufficient, but nondeterministic AI models forced teams to also govern outputs for harmful content and personally identifiable information (PII).

[22:30] Mohan: So Satyen, I have a slightly different opinion here. So to me...

Sangani: Ooh, I like that.

[22:35] Mohan: Okay. So what happened in data governance was we spent a lot of effort trying to make sure our data was as clean as we could possibly make it, and then passed that clean data to applications, reports, dashboards. All our applications were deterministic in nature until four years ago, which meant that if my input was properly governed and clean, trusted, then my output would be exactly the same every single time, because it's deterministic. It's running the same SQL every time.

But now, when AI models came in, it broke that paradigm, because models now, even with the same data, can produce different outputs at different times of the day. So now, all of a sudden, I was forced to also govern the outcomes, and by outcomes I mean looking for any harmful, hurtful information, any PII stuff, although I protected everything.

Because when you said agents are amazing, they're thoughtful and intelligent and powerful, those are your words. They can also be sometimes a bit too intelligent and powerful. They can go hack into another system to make sure they win and get the task done, whether it's ethical or not. So now with AI, I had to also think about the output.

What is agentic governance?

Sanjeev Mohan defines agentic governance as covering input, output, and runtime. It extends traditional AI governance to the trajectory data and memory state that multi-agent systems generate.

[23:54] Mohan: When agents came into the picture, which was a couple of years later, then it broke that model once again, because agents are reasoning on the input data, coming up with different chain of thought, deciding which is the best outcome to act upon. So a lot of context is now being generated, which was never the case in previous life.

In previous life, data was created in operation systems, end of the story. We've cleansed it, we've transformed it, and we consumed it. Now, I have new data being created. Sometimes we call it trajectory data. Sometimes we talk about memory state. All these things that multi-agent system is doing, how do we govern that?

So now we have a new challenge, because if somebody says, "I want lineage on how my multi-agent system rejected a loan," I need to know what these agents were discussing on their message boards and what messages. So you see, that is, to me, agentic governance. It's input, output, and runtime.

When an AI agent gets it wrong, what's actually broken?

Satyen Sangani of Alation lists the failure points behind a bad agent output, from prompts and outdated tools to an undocumented acceptable-use policy. Sanjeev Mohan shows how swapping one Gemini model version for another can break a fully governed system.

[24:58] Sangani: I would say something slightly different, which is that, if you think of these apps... I think the argument you're making is completely correct, which is that these apps are not deterministic, and therefore you have to observe them over time. And so I agree completely that the output has to be observed.

Mohan: Yes.

[25:16] Sangani: But what I would add to that, or maybe say slightly differently, is that then the question is: what do you do with output? Okay, I see this output. It's gone wrong. Now it could go wrong for multiple reasons.

It could go wrong because you simply did not have a good enough prompt in the model. It could go wrong because you gave it a tool that is old, and really you need to be using the right new version, or the same thing with the model. But it also could go wrong because, using that airline analogy and taking it forward, it could go wrong because that acceptable use policy wasn't well documented.

Is that a problem of the agent, or is that a problem of the data product or the data set that it's using? I would argue it's a problem of the data product, because the agent should have no need... you can't travel all of the rules of the organization in a single agent. So you have to let the work live with the object that it's meant to govern.

And so I would argue that's a governance of the data problem, not a governance of the agent problem, if that makes any sense.

[26:12] Mohan: So Satyen, I think it's a lot more complicated than that, because it could be that I was using Gemini version X, and now there's a new version of Gemini Flash which is cheaper, and I change my model.

Everything remain the same. The rules are the same. The data is properly governed, and suddenly this new model now starts expecting...

Sangani: Yeah.

Mohan: ...a different prompt. And I'm literally just talking about Gemini model version X to Y. I haven't even said I've decided to put DeepSeek instead, and it just broke everything.

So that's why the governance has to be layered across data and AI to do its job.

Why does the "context layer" idea break down?

Satyen Sangani of Alation argues AI governance must cover the whole system, writing corrections back to the data, the data product, the semantics, or the agent. In his view the "context layer" concept assumes a microservices-style discreteness that real AI systems lack.

[26:54] Sangani: Yeah, I think our argument would be the governance has to be of the system as a whole, not of any single layer. And so you have to write the correction back, whether it's down to the data level because the actual data quality was wrong, to the product level because the rule was wrong, to the semantics in the product 'cause the rule was wrong, or to the agent because the prompt was wrong or the model was wrong. And I think it's a really important straw man. There's this idea out there of this context layer, and I think this idea of a layer is designed on this premise of microservices thinking that assumes a discreteness that just simply does not exist.

These are not independent layers. This is a system and an app that you have to think of as an encapsulated whole. And building a layer is a sort of technologist dream, that I can capture all this information and it's sacrosanct. But it's all defined in the function of whatever the application happens to be, and the application is gonna drift constantly.

How does MCP change data sovereignty for AI agents?

Sanjeev Mohan explains how Model Context Protocol (MCP) tool calls take agents outside the system, raising questions about call identity, poisoned outputs, and data sovereignty under the General Data Protection Regulation (GDPR) and the EU AI Act.

[27:48] Mohan: Yeah. For example, one of the big tasks that an agent does is tool calling, and it's calling an MCP server. So now we're basically going out of the system into some third-party system through MCP. What is the identity of that call, and how do I know that I'm not gonna get spurious information from this external system and it's going to poison the output?

Then there's a whole data sovereignty. What is data sovereignty? All this time we've been constrained by saying data must reside in Germany if it is PII, period. But it's no longer just data residency. It is: where is AI operating? What MCP servers is it calling? Who's allowed to see it? How is this identity jumping between geographic regional boundaries?

So, to your point, we have to think through the whole system, and this is the complication that we run into, because it's no longer a well-defined thing. If you do this, then you achieve your GDPR. For EU AI Act...

Why doesn't data have a version number?

Satyen Sangani of Alation says data teams historically worked from a capabilities mindset. Sanjeev Mohan argues data people isolated themselves from Agile and software development lifecycle (SDLC) practices until data products introduced versioning.

[28:59] Sangani: No, I mean, this is, I think, the problem. I think this is the fundamental thinking difference in the space, because historically data people have thought from a very capabilities orientation.

I just do my job, I get the lineage, I do the stewarding, I tag the data, and I just hand it off, and it's all done. But it's not done, because the fundamental problem is you can do all that work and still get the wrong answer.

Mohan: Yeah.

[29:22] Sangani: And that, I think, is a very different style of thinking.

[29:25] Mohan: Yes. So Satyen, it's been my own observation, I could be wrong here, that data people, actually we ended up doing a lot of disservice to ourselves.

Infrastructure, applications, they were moving very rapidly into Agile and SDLC, all kinds of new ways, but the data people for the longest time were in their own island, in a silo. When I wrote my stored procedures using PL/SQL, I had no concept of A/B testing, CI/CD, any of this. We just did it in production, actually, because there was no version control.

Even now, when was the last time you heard data has a version number? It doesn't. A product does. My iPhone has a version number, not my data. But what you're doing with data products gives it that version number. So now we are finally moving into a more structured way of working on data than the time when I started my career.

[30:29] Sangani: I mean, I think we've always tried to impose the structure through governance. On some level, that's the work. But you would do it because of regulatory reasons, and you would sometimes do it because you needed some operational quality or output, but agents are forcing this question, which I think everybody says and knows.

And look, this is not a unique statement to this podcast or anything that we're saying, but it does change, I think, how we think about it.

[30:54] Producer (sponsor message): We're brought to you by Alation, the data intelligence platform trusted by 40% of the Fortune 100. Here's the problem they solve: your catalog helps people find data, but it doesn't actually deliver business outcomes.

Alation's data products marketplace changes that. Teams can build governed, AI-ready data products in a single session using plain English, no code, no SQL. And business users can literally chat with the data to get trusted answers. Governance is baked in, not bolted on. Want to see how it works? Grab the free data sheet at alation.com/data-products.

How will AI change the role of data teams?

Sanjeev Mohan describes AI as an enabler that automates tuning and self-healing. It shifts data practitioners from undifferentiated heavy lifting toward evaluating outputs, monitoring token consumption, and aligning with business outcomes.

[31:32] Sangani: So I guess now, as we think about moving forward... Given the state of today and the state of play of where things are at, and what things are moving forward and what things are not, how do you see the world of data people moving forward, and what role do people have if they have been in the data world, before and after?

And how does this space evolve? I mean, you're such a keen observer. Tell us what changes.

[31:57] Mohan: So AI is a huge enabler in my mind. It can increase people's productivity, data people's productivity, tremendously by doing things like automation of tuning, all the things you see.

So self-healing, all these things. AI is a huge enabler in that. But the data people's job now changes, because instead of doing this undifferentiated heavy lifting, as AWS says, they become more strategic in nature. For example, you mentioned that when you ship your semantic AI as a data product, there's an eval.

So you have to constantly do evaluation of outputs. You have to constantly see what is my token consumption. It doesn't make sense. What if I put something new in there? But if I put something new, will it break my business, like I was saying earlier, because I put a cheaper model? But now, data practitioners' job now becomes more high level, aligned to the business objectives and outcomes.

What should data engineers stop spending time on?

Sanjeev Mohan, recalling his time at Oracle, argues that backups, restores, entity relationship diagrams (ERDs), and parameter files like init.ora should be handled by systems, freeing data people to focus on business use and unstructured data.

[33:06] Mohan: Not anymore about "I need to do a ERD diagram, and then I need to do my backups and restore." Why do I need to spend my time doing backup and restore? Hugely critical, but those things I should just expect from the system. It shouldn't be my primary job. When I started at Oracle... you remember Oracle, there's a configuration file called init.ora, hundreds of parameters.

I knew each one of them, because that was my job, and I took pride in it. Today, I should not even care about init.ora. That's Oracle's job to figure out. My job should be: how is this data being used by my business? What more can I give? Can I get unstructured data into the picture? And if so, what unstructured data is related to my business data, and how do I structure this together so it's not a technical debt?

I don't want to move unstructured data into my object store. I want to leave it where it is, but I want to rethink my architecture so I can build a context layer from all these related sources. So I really think that this is a golden era for data people, because they are liberated from the nitty-gritty low-level stuff into more business strategic direction.

Data teams as the librarian of the organization

Satyen Sangani of Alation argues data teams must solve business problems while building a reusable "library of truth" of structured data products and unstructured collections, becoming the learning organization.

[34:27] Sangani: I completely believe that and agree with that. I think the data teams, on some level, in my mind, one, have to solve business problems. Now, arguably, they always had to solve business problems.

Mohan: Always, yeah.

Sangani: But the work of data, to your point, was so low level that you could be a data engineer or an analytical engineer or whatever that term is that you are, and you basically could be there moving data, doing transforms, and doing all of this stuff that consumes your day, and you didn't have to really be consumed with the why, because the how was so much of your day.

I think that's gonna drift away over time. I guess I also would argue that by solving these problems, you also have to become the learning organization.

These data products, whether they're structured or unstructured, have to basically be the facts of the organization that you manage.

You have to be the librarian of the organization. Now, you can't build that library in its own... I'm not building this castle in the sky. I've got to actually solve business problems while I'm building this library. But I think if I'm a data person, I'm first and foremost solving business problems, and while I do it, building this library of truth. And this library of truth can be unstructured or structured, but it's basically just this collection, in my mind, very simplified, of data products and unstructured collections of things that I basically built so people can reuse them.

Where are enterprises seeing real ROI from AI?

Sanjeev Mohan points to pharmaceutical companies, where drug development takes 10 years, costs $2 billion to $4 billion, and requires roughly 4,000 regulatory filings, as quiet AI return-on-investment (ROI) stories. Enterprises scaling similar programs can explore Alation AIOS.

[35:43] Sangani: I guess, do you see this happening anywhere? Is this happening well anywhere? I mean, where... 'cause we certainly have some customers who are absolutely doing this and many who are on the road, but what's your perspective on where the market lives in this transition?

[35:57] Mohan: So I see this happening. I get evidence when I go to a customer event. For example, at revAlation, one of my biggest goals is to talk to some of these end users and see how they're doing it. When I go to Gartner D&A Summit, I talk to shipping and logistic companies. I talk to pharma. People constantly say, "Well, what's ROI of AI? What has AI done for me?" And I talk to pharma companies, and I see... they don't like to talk about their successes using AI for various political, social reasons, but the volume of work they're doing is phenomenal. To do a new drug approval process, a new drug discovery takes 10 years, and it costs anywhere from $2 to $4 billion.

It requires literally something like 4,000 documents to be filed with FDA and other agencies. If they can use AI to do some of this documentation, for example, and it has to be in English, French, Spanish, all these other languages, and then you use AI to do the translation, you still have to check the work, but my point is that that's a few million here, a few million there.

So I see this happening a lot. In pharma, for example, a lot of data engineers in the past had no clue about the data they were working with. They had no idea about clinical codes. But now what I'm seeing is that if they move up and they're not spending their time just mired in the trenches...

If they move up, then they start learning what is it that matters to the business most. Why is this ICD-10 code more important than something else? So I'm hoping that data practitioners who are listening to this will take it to their hearts that we need to be and think like businesspeople, not as technicians.

And so think, like you said, think about why and how I can make a difference, not like, "Oh, I need to create a script to back up my database." So that's how I see the change.

What will change in data and AI in the next 180 days?

Satyen Sangani previews an upcoming Alation ontologies product and asks Sanjeev Mohan for a 90-to-180-day outlook. Mohan predicts in-house models trained inside the firewall, notes that small language models (SLMs) haven't taken off, and flags rising token and GPU costs.

[38:12] Sangani: For sure. And maybe as we... we're coming to the end of our conversation, although I think I have to get you back on, 'cause we still have to cover this entire knowledge graphs and ontology set of concepts, because we're gonna be launching an ontologies product, and I think it's gonna be super interesting to think through how that gets created and built and managed in this new world.

But rewinding a little bit... I used to ask people about their predictions for the next year, but I feel like that's just completely unfair, 'cause I can't even tell you my roadmap in the next year, but I can tell you my roadmap in the next 90 days, and it's way more certain.

What do you see happening in the next 90 to 180 days in the world of data? What do you think are the big things that are top of mind, and what people ought to be paying attention to and thinking about, and where should people be focusing their attention?

[39:03] Mohan: So, Satyen, I often do my annual prediction reports.

One thing that I predicted a couple of years ago that still has not come to fruition, and I hope at some point it will, is this ability for organizations to have their own models in-house. Even to this day, the only companies that are able to make models are these frontier model companies. End users don't have the... it's not easy for them to do it.

There should be a process by which I take a base model, a pre-trained model, and then I throw everything that's inside my firewall. It all stays inside the company, and it's constantly updating my model. So I'm not relying on Claude that's trained on the whole of internet. I don't care about the whole of internet.

So I don't see this. So small language models still haven't taken off, and I'm starting to see that happen. And then cost is a huge thing, because tokenomics, like token costs have gone up. No, everything's gone up. Even hardware memory is hard to get. GPUs are more expensive. So for AI to be successful, we need a couple of things to happen.

Why is the AI moat shifting to semantics?

Sanjeev Mohan names reliability, cost-effectiveness, and embedding AI into daily work as preconditions for AI success. He argues the moat now sits in the semantics and business logic layer, which Alation addresses with semantic model mastering.

[40:18] Mohan: We need reliability, so no hallucinations, more deterministic answers, and we need cost effective. The last thing that we need, and it's something that I've seen coming out of some work from Alation, is how easy are these models and AI technologies embedded into my day-to-day work? One of the reasons in my mind why catalog, even BI tools, have suffered is because I am doing my work in a certain system of record, certain application, and then I have to go outside that to go manage my data or run a report, then I have to come back here.

So I... The future belongs to AI. In fact, if you look at what's happening with Salesforce, Salesforce has spent a whole decade building out the most amazing user interface, and now what they're saying, "Just use Claude." So use us for the system of record and the business logic. See, this is where business logic, semantics become really important, but the UI goes to the AI tool.

So this is super exciting for me, to see how the consumption layer is getting democratized, but the moat is now shifting to the semantics. Storage is already democratized because it's all open standards: Iceberg for analytical, Parquet files, all that, S3-compatible storage. It's that middleware where you are linking the business logic, the semantics, to me is the most critical piece in this stack.

Closing thoughts: compute, ontologies, and rules that travel with data

Satyen Sangani of Alation asks whether compute will ever democratize given centralization around Databricks and Snowflake. The episode closes with the AI Radicals producers' takeaway on data products, versioning, and agent safety.

[41:51] Sangani: Yeah. And then I think there's an open question as to: does the compute get democratized? That does not seem to be happening at all, on some level. I mean, there seems to be a lot more centralization around the Databricks and the Snowflakes of the world...

Mohan: Yeah.

Sangani: ...than, you know. And also, then there's also this question of: well, does that move up to the semantics, or is that a distinct problem?

Obviously, we'd argue they're distinct problems. But certainly, that's also an interesting set of questions to think through. Sanjeev, I think I could talk to you for, I don't know, hours and hours, so it's always fun. I will see you next week at revAlation, and then hopefully we can actually do maybe a part two of this.

But it's been great to talk to you. You have such great insights at this moment, and I know everybody will appreciate them, and more critically, go buy the book, Designing AI-Driven Foundations, in a store near you.

[42:41] Mohan: Always a pleasure. I'm looking forward to next week.

Thank you so much for having me on this episode.

[42:50] Producer: Here's the line that stuck with us. Sanjeev pointed out that his iPhone has a version number, but his data never did until data products gave it one. That's a bigger shift than it sounds. Once data is packaged like a product with an owner, a contract, and a life cycle, it stops being someone's abandoned side project and starts being something the whole organization can actually rely on, and that discipline is what makes agents safe to use.

An agent can't be handed every rule in the company, but if the rules travel with the data itself, the agent never has to know them to follow them. Thanks for tuning into AI Radicals. See you next time.

[43:25] Producer (sponsor message): A word from our sponsor, Alation. Here's a stat that should give every healthcare data leader pause. 80% of organizations have deployed AI in some form, but only 11% have achieved it at scale.

The bottleneck isn't models or compute, it's the data underneath them. In healthcare, that's especially high stakes. Ungoverned data doesn't just slow down your AI program, it can introduce bias, expose patient information, and produce outputs clinicians can't trust or act on. Alation's new whitepaper, produced with Healthcare Information and Management Systems Society, pulls from real implementations from Children's Hospital of Philadelphia.

It covers what AI-ready healthcare data actually looks like, how to build data products that accelerate clinical AI, how to govern for HIPAA compliance without killing innovation, and a practical roadmap for going from pilot to production. If you're building AI strategy in a health system, this is one to read end to end. It's free at alation.com/ai-healthcare.


Other episodes you might like

  • Cara Tice

    Data Products, Not Platforms: A CDO’s Playbook For The Agentic Era

    Season 4 Episode 9

    Trillions of dollars move through Cara Tice's systems every year. The CDO of Early Warning (the company behind Zelle)…

  • Eugene Wu

    Semantic Coupling & the Interconnected AI Stack

    Season 4 Episode 8

    What if the fix for unreliable AI agents isn't a better prompt? Eugene Wu, Associate Professor at Columbia and…

  • Charlene Li

    The 90-Day AI Roadmap

    Season 4 Episode 7

    Why do so many AI initiatives die in pilot? Charlene Li, New York Times bestselling author of Winning with AI , says…