Season 4 · Episode 6

The Case for Federation in the Age of Agents

Anant Jhingran

Anant Jhingran

Chief Technology Officer, IBM

in

As Chief Technology Officer at IBM Software Anant Jhingran has led IBM's database and information management research, co-founded StepZen, and scaled Apigee through its IPO and Google acquisition, giving him a rare view of what agents change in the layer beneath the application.

Satyen Sangani

Satyen Sangani

CEO & Co-Founder, Alation

in

As the Co-founder and CEO of Alation, Satyen lives his passion of empowering a curious and rational world by fundamentally improving the way data consumers, creators, and stewards find, understand, and trust data. Industry insiders call him a visionary entrepreneur. Those who meet him call him warm and down-to-earth. His kids call him “Dad.”

Anant Jhingran, CTO of IBM Software, on databases, APIs, and enterprise AI

Producer and host Satyen Sangani, CEO of Alation, introduce Anant Jhingran, CTO of IBM Software, and frame the episode around what agentic AI changes in the data layer underneath enterprise applications.

[00:00] Producer: Welcome back to AI Radicals. AI isn't just changing what software can do, it's also challenging some of the basic assumptions behind how we build it. Today's guest is Anant Jhingran, CTO of IBM Software, co-host of the podcast Context Window, and a longtime leader in databases, APIs, and enterprise infrastructure.

Satyen and Anant get into why agents may force us to rethink decades of assumptions about data infrastructure, why federation could make a comeback, and why AI may change just how clean, centralized, and structured enterprise data needs to be. If you're curious about what AI changes when we stop looking at the application and start looking underneath it, this one is for you.

This show is brought to you by Alation. Only 5% of enterprise AI pilots from 2025 delivered ROI in production. Alation's free agentic AI opportunity discovery guide shows you exactly how to find the right processes to automate and build AI that actually sticks. Get it at alation.com/ai-guide.

[01:07] Satyen Sangani, CEO of Alation: Our next guest has spent his career building the invisible infrastructure that powers enterprise software, and now he's helping shape the future of enterprise AI. As CTO of IBM Software, Anant Jhingran is focused on one of AI's biggest challenges: connecting foundation models to the trusted enterprise data they need to deliver real business value.

Before IBM, he co-founded StepZen, led groundbreaking work in databases and information management at IBM, and then helped scale Apigee through its IPO and its acquisition by Google. Few people have spent more time thinking about how data, APIs, and AI come together at enterprise scale. Anant, you've been a longtime friend and, I think, first-time caller. Welcome to the show.

[01:45] Anant Jhingran, CTO of IBM Software: Thank you so much, Satyen. I didn't quite recognize all the words of praise that you said. It was embarrassing. I think I turned a bit red. But no, thank you very much for your kind words. And as you know, we have been chatting on and off, and through some mutual connections, we said, "Let's talk a bit, not just among ourselves, but also to a broader audience."

So for everybody else, I went to Satyen first and said, "Satyen, please come on my podcast," and he said, "No, no, he'll only come on my podcast if I come on his." So this got scheduled first, so here I am, Satyen.

[02:22] Sangani: Yeah, the best kind of quid pro quo. So this is a two-part episode. And so maybe let's just jump right in. You've obviously done an incredible amount of work and research around databases. You founded companies around APIs. Certainly Apigee was a big investment there. And all of that comes to bear in orchestration in this AI moment. Maybe talk a little bit about this moment. Is this the same as many of these other prior revolutions, or is this different?

Is this AI moment different from past technology shifts?

Jhingran's answer to the episode's opening question: the interaction paradigm changes completely with agentic workflows, while the underlying data systems and integration problems stay exactly where they were.

[02:49] Jhingran: It's really, really different. I mean, I feel it in the way I work, and I'm sure you feel it in the way that you work. But the good thing, Satyen, is that the thing that you're trying to do at Alation, or we are trying to do at IBM and others, is that the foundation is still the same.

The interaction changes, right? So we had primarily built stuff for either automated flows or humans who would swivel chair or not swivel chair based on all the fantastic work that you have been doing, Satyen. But in this new world, if it is AI and agents that are doing it, does it represent a different modality and everything else remains the same? Or does something at the base also change? So those are the kinds of questions that we have got to always answer.

But the fact of the matter is that the data systems are the same, the systems of record are the same, the problems of integration are still the same. The problem of which is the authenticated piece of data here is still the same. So problems remain the same, but the interaction paradigms change completely, and because of that, perhaps some of our old layering needs to be rethought.

Why has metadata work moved from side-of-desk to operational?

Sangani, whose company builds metadata management software, describes the collapse of the boundaries between operational estates, data estates, API consumption, and human consumption, and what that does to the status of governance work.

[04:01] Sangani: So I think you mentioned that a lot of these fundamental problems are similar. These problems of provenance, these problems of understanding, these problems of metadata. And the convergence of having AI do things in the way that a human would have done things meant that many of these estates, these operational estates and these data estates, APIs and technical consumption versus human consumption, all of these borders seem like they've collapsed in my mind.

And that fluidity feels like it's forcing both a step function change in speed, but often when that happens, you have this corresponding dynamic where the form of what matters changes too. And so I think a lot of this work, at least what we are seeing, is that a lot of this work around metadata and insight generation and governance used to be side of desk work, and now it's entirely and completely operational, at least for the most progressive organizations. And to me, that shape of that work means that you have to almost rethink what an application is and how it functions.

What the API boom taught us about generational shifts in infrastructure

Jhingran uses the origin of the API economy, which he helped commercialize at Apigee, to make a point about what changes and what persists when a new platform arrives.

[05:09] Jhingran: I'm old enough to have lived through a few changes, right? So if you recall, why did APIs come about? Because people had physical retail, then they went to e-commerce, and then Steve Jobs came out with this, and they said, "We can either create a third different silo, or why not let's build an API that will allow us to cut across everything else."

And that gave the API boom, and that API boom is now lasting for at least fifteen, eighteen years. But the point that I took out from there is that every time there's a generational shift, something changes in the infrastructure also. Okay? All I was pointing out was that the importance of your data, your metadata, your integrations, your "this is the single source of truth," et cetera.

The importance of that still stays, because anything that is exposed that is not correct or good just amplifies the fast agentic reasoning on top of it, presumably in some incorrect directions if you have some mistakes. The importance stays, but obviously the interface, when the predominant things are machine as opposed to humans or predefined, has to change too, right?

Is MCP just a wrapper around APIs?

Jhingran, who spent years building the API management layer at Apigee, questions whether the Model Context Protocol (MCP) is the right interface for agents or simply a convenient shim over the APIs enterprises already have.

[06:32] Jhingran: And today, for example, MCP is the wrappers around APIs, right? And if you poke into it, that's what it is in most of the cases. But is that the best thing, right? Who knows whether that's the best thing. If skills is coming about, what is the role of that? Can you actually... There are people who are now suddenly saying that CLIs are the better way in which some of these skills and mechanisms actually operate.

So I think that the interfaces will change, but the importance of good hygiene around data and everything else doesn't actually go away, irrespective of whether it's humans, pre-integrated workflows, or now agents. That would be my assertion.

Why do CRM and ERP records still hold unreliable data?

Sangani names the oldest unsolved problem in enterprise data, that systems of record depend on humans to populate fields they have no incentive to populate accurately, and asks Jhingran where the money to fix it comes from.

[07:15] Sangani: Yeah, I agree with that. On some level you have these systems of record, like classically Salesforce or whatever system, ERP, workforce management, whatever it is that you want and have. And then you effectively bolt on all of these things onto the CRM in order to be able to gather some of this data.

And historically, it's just been impossible to make this data correct because no salesperson is going to enter the full fidelity of the conversations that they've had. And so now all of a sudden you just get garbage in these fields, and the data is either not correct or not accurately describing reality.

And so you have this world where the operational systems are important and are driving things, but they're not necessarily always accurate or correct. And so the challenge is, these problems have existed and remain, but now the importance of solving them is very, very high, and they're still very hard to solve.

And I guess, how do you see this forcing function evolving in terms of getting this stuff right? Because it's been always side of desk. It's always been expensive. And it's always been something that a CIO doesn't prioritize. And so how does one actually think about where money is put and time is put in order to fix this stuff?

What is the difference between data for AI and AI for data?

This is the framing Jhingran returns to throughout the episode. One direction is making enterprise data usable by models. The other is pointing models at the data problems that resisted automation for decades, which is the premise behind AI-powered data quality.

[08:37] Jhingran: No, that's a really awesome point. So the way I like to think about it is data for AI, which we just talked about, that if the interactions are changing, et cetera, but also AI for data, which is, can you actually leverage AI to get over some of the hurdles that you have had in the past with respect to how do you actually cleanse and standardize and make things better.

So I like to think about it in those two framings. I don't know whether you think about it like that, but let's talk a bit about how can AI actually help us, in my view, with respect to some of these gnarly data integration and data cleansing problems.

Do AI agents need perfectly clean data to work?

Jhingran, drawing on his API background, makes the counterintuitive case that AI consumers are more tolerant of malformed inputs than the automated integrations that preceded them, which changes how prescriptive enterprises need to be at the interface.

[09:26] Jhingran: So in my mind, there are two things. It is actually quite remarkable how the AI that uses the data is fairly tolerant of bad stuff. Okay? So previously, if my API had a /member/ID, and I accidentally typed /member/ID as opposed to /member/ID, then it will just return whatever it is. I don't even know what the code will be. Probably some 400s or something like that.

But today, AI can say, "Okay, I'll type /member/ID," and it'll come back with that code, and it'll reason among itself and say, "Oh, I got it wrong," okay? And therefore, let me readjust it. Obviously, we are paying more to the AI gods with respect to tokens and all that stuff, but they self-correct around stuff that might be bad.

So in some ways, if you believe that the users of all of this stuff are shifting towards AI, you have to be less prescriptive with it being exactly correct. Okay, so that's my view number one. I don't know whether that makes sense to you or not. But then the second part is that for the things that you have to get right, AI does actually help.

How can AI help reconcile structured and unstructured data?

Jhingran describes the perpetual enterprise problem of reconciling structured and unstructured data, and why a shared embedding space may unify records that older techniques could not. He cites Vanja Josifovski, CEO of Kumo, the graph learning company acquired by NVIDIA.

[10:17] Jhingran: And I've been thinking a bit about this in various formats. So number one is that the classic problem of resolving unstructured data and structured data is a perpetual problem, right? I mean, it's a problem that always exists. Is it possible that if you think about an embedding space so that your structured data and unstructured data are lining up in some embedding space, is it possible for you to actually unify some stuff that you couldn't have unified before?

Is it possible for you to think about certain new fusing techniques, correlation techniques, analysis techniques that would allow you to fine-tune a model? So for example, we had somebody on my podcast, Vanja, who was the CEO of a company called Kumo, that got acquired by NVIDIA, and he was talking about how he's taking standard transformer architectures and then applying them to the data problem to solve some of these things associated with it, right?

So I think that the opportunity to leverage AI to get around those very, very laborious and error-prone tasks, so that you can expose to the AI use cases or human use cases something that's much better and much cleaner, and then depend on them to actually work around, not perfection. I think that is, in my mind, some really, really large opportunity.

Why does data quality get harder after the first 85%?

Sangani, whose teams run these pipelines in production, pushes back on how far AI cleanup actually gets you, and points to agent evaluations as the discipline enterprises have not yet built.

[12:05] Sangani: Yeah, I think the AI for data, I love that expression of the problem statement. It is a huge opportunity. It does shift the work, I think. What we've experienced is that you can apply many of these techniques, but you still end up at some level of error. Ultimately it's a function of your evals and what you call good, but it's very hard to get...

Like anything, whether it's web reliability or whatever, you can get to the first 90%, 85% relatively quickly, and then thereafter it just becomes harder and harder. Every order of magnitude becomes more than an order of magnitude harder.

And so now the work shifts to the last set of quality guarantees, and then many of these processes are dependent on very high fidelity if people are willing to automate it. In the human offline process, maybe they're not willing to automate it, but in the agentic world, they have to get extraordinarily high fidelity to trust it in order to make it move forward.

And I think people are struggling with this problem in its essence, because they just don't know how to think about what they can trust and what they can't trust. This evals thing is a very new way of thinking for people.

How far along are enterprises in using LLMs on old data problems?

Sangani sets up the contrast between traditional structured pipelines and embedding-based approaches, then asks Jhingran what IBM sees across its customer base. Jhingran's answer is that most enterprises have a handle on their data estate and a governed data catalog, and almost nothing further.

[13:34] Sangani: I guess going back to some of these new techniques, you mentioned this classic structured and unstructured data dichotomy and applying transformer models to structured data. The way I've always thought about it is that in some ways, LLMs have developed this new pipeline. On some level, your typical structured data pipeline is basically the structuring. It's essentially taking the world of unstructured reality and just superimposing on it boxes of things that we want to count or label.

[13:57] Jhingran: Extract entities, yeah.

[13:58] Sangani: Yeah. And ML features are effectively the same thing. It's like, okay, I'm just taking some things that are unstructured and I'm going to glom on these features and start counting them, or start measuring them in some fundamental way.

But this new kind of LLM, to your point, this embeddings world, is a very different modality of thinking about the world. I guess, where do you see... I mean, IBM sees so many customers and so many people and so many different organizations. Where are people in absorbing this way of thinking about using LLMs to help do some of these classic problems, or deal with some of these classic problems? Are we very early? Are we mid-stage? Where do you feel we exist?

[14:36] Jhingran: No, I think we are very early, right? So I think that what's really going on is that with your help, with IBM's help and others, people have gotten a bit of their handle on their data estate. They've got the integration flows done. They've exposed the right catalog, perhaps through you, et cetera, with all its issues and everything else.

So the quick hits then are actually changing the application paradigm around the piece of work that has already been done, as opposed to using that new paradigm to change and do something slightly different in the data infrastructure, right?

What does it mean to deliver context to AI applications?

Jhingran defines the first wave of enterprise AI work as a context delivery problem, and draws the line that separates a model's general knowledge from a company's own. The distinction between metadata versus context is what determines whether an agent can answer a question about your business.

[15:15] Jhingran: So the very first thing that people are doing, if I may say, is to say, "Okay, can my data estate deliver the context that's needed for these applications to actually work?" I mean, we all know that what Anthropic knows is not what our customers know inside, and therefore being able to deliver that context is really, really important.

So I think that the first sets of hits are happening on this delivery of context. Okay? Which is not fundamentally saying how do you do all the hard problems of AI for data.

How does the agentic paradigm change data infrastructure?

Jhingran explains why pre-wired pipelines fit a world of stable reporting and not a world of chain-of-thought agents, and where the long tail of never-completed integration and catalog work fits in. Alation builds this layer as an agentic data intelligence platform.

[16:02] Jhingran: The second set of hits that are happening is to say that, look, the application paradigm, the agentic paradigm, is a very, very different paradigm. It's try this, try that, go this, chain of thoughts, figure this out, et cetera. Therefore, exactly to your point from 15, 20 minutes back, it changes the interaction paradigm.

And therefore, what infrastructure do we need in order to go, instead of pre-wired, stable flows, to, "Okay, can you please produce this piece of information for me," or, "Can you do something else for me"? And therefore what that means is some of the constructs have to be done which will allow for pipelines to be created on the fly, or metadata to be assembled.

So I see these two happening in a big way, just to summarize: delivering context for these applications, and then seeing whether these applications can be used. You can change some of the infrastructure in support of these applications.

And the idea behind it is that, look, there has always been a long tail of integrations and metadata stuff that has never been done before, right? Can you actually make some of those things happen? In terms of fundamentally thinking about, can I think about the primary data integration and catalog creation and everything else being done through AI, and therefore what really needs to change with respect to either the embedding space, what's really verifiable?

I mean, obviously, AI works very well when things are verifiable as opposed to not verifiable, right? So what does it really mean in that space? And I think that that is something that I haven't... I mean, obviously all of us are thinking about it, but I haven't seen it yet make a big impact in our customers. Probably because we as vendors, or at least us from an IBM perspective, and I haven't followed Alation that much, we haven't made a big deal out of it.

How long does governance take compared with building an agent?

Sangani puts a number on the gap between building an agent and making it trustworthy, then asks Jhingran which infrastructure primitives he expects to change.

[17:49] Sangani: Yeah, I think this context thing is obviously super topical, and everybody's talking about this idea of context and centralization of context and provision of context. And then to your point, I think the second conversation is an interesting conversation, because it's obviously what drives fidelity and all of this governance and context. I mean, you can build an agent in 20 seconds. Getting all the governance and the context right will take you 100x longer.

[18:15] Jhingran: Absolutely.

[18:15] Sangani: But you've thought a lot about this infrastructure question, like what are the primitives? And I actually think you're probably about as deep as anybody into this question, and you talked a little bit about MCP and maybe some of these ephemeral pipelines. What are some of the ideas floating around in your head around what are the areas that could experience change, and how people ought to think about some of these primitives differently than they were before?

Does agentic AI still require centralizing all your data?

The chapter YouTube labels "Federation's comeback." Jhingran separates the two historical reasons enterprises centralized data, cleansing and query efficiency, and argues that enough active metadata may substitute for one of them.

[18:41] Jhingran: So if you really think about it, the fact of the matter is that data is created everywhere, and it's like you're always bailing water to try to get it in the right place, okay? For all the applications that you've got to do. And that's the perpetual paradigm.

Is there a way to think about it slightly differently? Okay. Which is to say that, look, enough metadata, okay, as opposed to enough centralization of the data. What were the reasons for centralizing the data? Obviously, you got an opportunity to clean it, but there's also an opportunity out there, Satyen, to then run these really, really complicated queries, okay, which are not federated across multiple systems, and you can do a table scan as opposed to something else.

And what we have done is typically we have mixed and matched the two. We have said that we want to bring stuff in one place so that we can both standardize it and we can build those various tiers of my warehousing hierarchy, or whatever the new data lake hierarchy. But also I can now optimize and run stuff really, really efficiently.

Agents run short bursts of discovery queries, not one mega query

Jhingran describes the workload shift that undercuts the query efficiency argument for centralization, and why metadata becomes the binding constraint instead.

[19:38] Jhingran: Okay? If the paradigm is shifting towards, no, look, yeah, you've got to run these long-running queries, but in the agentic interactions, really it's short bursts of queries. First let me find this, and let me find this. Oh, this didn't get me there. Let me now find something else.

So instead of, let me run a mega query, fixed report, et cetera, I'm trying to discover my way through to certain things. Then the paradigm changes. And perhaps this concept of centralizing everything in one place for being able to efficiently run queries may not be as important in these new workloads, because they're not really executing something really, really fancy out there. They're just finding their way, hill climbing their way through to the right answer.

So in that case, some of the primitives that you have thought about, which is, let me move stuff, et cetera, don't have to be thought about. More importantly, what has to be thought about is, how do I express what I can do? How do I then efficiently execute on it, give the reliability and everything else?

But in both of those cases, without that metadata, the agents are basically groping around some space trying to figure this out, and therefore, that particular part of it still remains important. Whether that needs to be accompanied with data movement, et cetera, I don't know. Okay? So that's one place where I think that it may be the re-emergence of federation with good, strong metadata space on top. That's one view.

Can Google-style AI techniques clean up messy structured data?

Jhingran traces the arc from Natural Language Toolkit (NLTK) synonym expansion, through Boolean search on Lycos, to Google's decision to absorb user error at scale, and argues the same shift is available now for structured enterprise data.

[20:52] Jhingran: The second view that I have is that the fundamental problem that we always try to solve for is that as much as we think that the data is clean when it ends up in our databases, it isn't. Right? And the question really is, what is the effort?

So for example, just go back. If you even go back to something like NLTK, which was a natural language Python library, it would be very careful about, look, I've got to clarify what my synonyms are in order to find out, right? Anant may be spelled in two different ways, and so therefore this or that, and you do the expansion.

And then Google came about, right? If you all recall, in the Lycos days and other things, you had to do that ANDing and ORing in order to get there. And then Google came about and said, "Nah, screw it. Just tell us what you think you want, and we'll try to figure out what you want." Right? And that worked well. That worked well because Google had that huge amount of data that it could figure out that somebody with a Y-A-N is probably a misspelling of somebody with Y-E-N.

[22:44] Jhingran: So I think the same thing applies now, which is that if you now look at the data fragments and other things that are misspelled or incorrect, nowadays I don't bother doing a spelling correction when I'm in an agentic interaction, Satyen. The only reason I correct spelling is because I don't want it to think that I'm a doof, right? But I don't really need to correct spelling because it figures it out.

So I think there's something really big out there which says, can we apply some of these AI techniques that have been applied on large corpus of data to structured data? Maybe it is based on an embedding space, that Satyen with a Y-E-N and a Y-A-N end up in a similar space and therefore they're probably the same, as opposed to the old stuff. But I think that that space is ripe for a phenomenal amount of innovations. And that's what my friend Vanja is trying to do, to apply some new transformer models for structured data. So those are two examples that come to mind.

Is centralizing metadata different from centralizing data?

Sangani names the conflation that has followed the data management industry for a decade, and asks whether the rise of Iceberg tables at Snowflake and Databricks settles the centralization question or reopens it.

[23:51] Sangani: Yeah, and both are really interesting. In the first case, we've struggled with this federation question for a super long time. Do you have to move the data? And there's always this conflation between the centralization of data and the centralization of knowledge.

We've always talked about centralizing metadata. And look, I know that's obviously talking our own book, but we've always talked about the centralization of metadata and, now, context, whatever you want to call it, versus the centralization of the actual physical data. But even now that's getting super muddy, because now it's like, well, context can also be operational data, so now is that to be centralized or not? There's all of this new same battle, in a new location.

Do you feel like there's an opportunity for... I guess, do you think the compute opportunity in federation is a new opportunity? I mean, with Databricks and Snowflake, you've seen just massive centralization, and now there's this Iceberg table thing that everybody's coming along. And so do you think that maybe we'll start to see... Is that the end of the story, or do you think there might be even yet another story in the world of compute where federation is a more real thing? And I guess, are there things that you've seen where that's been more or less promising?

Will ad hoc agent tasks outnumber pre-built workflows?

Jhingran makes his central prediction: the interesting agent work is the work no one has done before, which means chain-of-thought exploration will eventually outnumber the pre-wired flows enterprise infrastructure was designed around.

[25:12] Jhingran: No, I actually strongly believe that two things are happening. One is that agents are being asked to do increasingly complex tasks. I mean, you know this. When I used to code, the old Cursor joke was that the only keyboard you need is a mic and a Tab button to accept whatever answers come back, right?

But it's not even that case now, right? I mean, I don't even have to accept the answers. Now I'm giving oohs and aahs and all that stuff, and it's actually figuring out and reasoning through what needs to be done. Just in the case of coding, but also in terms of other things. Today, I was trying to do some research on some other topic, and it spawned some agents, and it came back with something, and I said, "I don't agree with it." It's like, whatever it is.

But the point there is that if you think that agents are just going to do the same thing that you're doing, except machines instead of people, it's kind of boring, and I don't think that's going to happen. Yeah, it might happen, but it's just that kind of shift. The real change will be they're doing something different that we haven't done before.

And in that case, my view is that the paradigm of this chain of thought reasoning, which is really where we are going, where AI is really going towards, is so different from predetermined stuff. And to me, it's not at all clear that all the structures that we built to answer repeatedly the same queries, or where someone is so smart that he has thought about a day-long query that we want to optimize, okay? So I think that the paradigm is really, really different.

Now, obviously, the two paradigms will stay together for quite some time. But I'm pretty clear that the dominant paradigm would be more ad hoc agents thinking through some sets of issues. And obviously, once something gets repeated, you want to relegate it and execute it repeatedly, but the number of tasks that will be done in the new way will exceed the number of tasks done the old way.

Do LLMs remove the need for canonical values in a database?

Sangani uses IBM's own name as the example of why databases demanded canonical values, and asks what changes when a model can resolve variants on its own. Alation approaches this problem through semantic model mastering.

[27:25] Sangani: Yeah, it's a really interesting construct. We've done all this work, to your point, to get this last mile fidelity right. IBM is always IBM in the database. It's never International Business Machines, or Intl Business Mach. It's always just this one canonical thing, because by having this one canonical thing, I can count it exactly, and I can identify it exactly, and then I know what to do as a database.

[27:59] Jhingran: Select count star where name equals single quote IBM.

[28:04] Sangani: Exactly.

[28:04] Jhingran: Yeah. No, but then you also unfortunately still have to do what's called invariance on caps. Sometimes IBM will be small ibm, sometimes it'll be big IBM. So you still have to do that kind of manipulation at query time.

[28:19] Sangani: Yeah. And now, to your point, you don't have to. And so then this construct of... And it's not even very expensive for the LLM to figure out that these two things are the same thing, and therefore the level of last mile fidelity you have to have is relatively lower than it used to be.

And then, to your point, these outcome-defined agents can go do a whole bunch of stuff, and therefore they can spin up and spin down databases if they need to. They can build tables and kill tables. All these things that people used to have to do in a fixed way, these LLMs can do in an ephemeral way, and that's super interesting.

I guess it means that now the start point has to be really high fidelity. The implication of all that is that whatever you're starting from has to be very high fidelity, because otherwise something is going to go wrong if your instructions and your inputs are wrong.

Embedding every row, and why data fidelity is still the floor

Jhingran concedes the cost problem with resolving variants at query time, offers embedding every row as an alternative to standard indexing, and closes the section with the sharpest line in the episode about what better models cannot fix.

[29:17] Jhingran: I agree. Now, we are all system builders, so I don't want to undermine the fact that while capital IBM and a small ibm can be reasoned together by the LLM, the fact of the matter is that it is expensive to send the query down to the database to figure out how many different variations of IBM should I ask for from the database, right? So I'll get this big long list of ORs, et cetera, right?

So there are still hard problems to be solved, but I think those hard problems can perhaps be solved in a slightly different way as opposed to just standard indexing. Just embed every row, okay? And when you embed every row, big IBM and small ibm end up in the same space.

But exactly, after that, obviously without that fidelity, everything else is basically garbage. And as much as the agents and AIs change, they're improving on their hallucination. Right? Hallucination was a big problem two years back. It's a smaller problem now because there's enough RL that's happened. But the fact of the matter is, if the data doesn't have high fidelity, any amount of lack of hallucination is not going to solve the problem.

[30:43] Producer: We're brought to you by Alation, the data intelligence platform trusted by forty percent of the Fortune one hundred. Here's the problem they solve: your catalog helps people find data, but it doesn't actually deliver business outcomes. Alation's data products marketplace changes that. Teams can build governed, AI-ready data products in a single session using plain English, no code, no SQL. And business users can literally chat with the data to get trusted answers. Governance is baked in, not bolted on. Want to see how it works? Grab the free data sheet at alation.com/data-products.

What are the three pillars of IBM's data strategy?

Sangani turns from IBM as a database example to IBM as a business. Jhingran lays out three pillars, starting from the observation that technology moves faster than the customer landscape it serves.

[31:25] Sangani: Let's switch gears a little bit. We've been talking so much about IBM as a semantic construct. Why don't we talk about IBM as an actual construct? So you're at IBM. In some ways, it's an institution that's been so central to this change. So many data management and manipulation and centralization technologies have been acquired by IBM over the years.

[31:52] Jhingran: Built and acquired.

[31:52] Sangani: Built and acquired. Totally fair.

[31:57] Jhingran: I did help build a few of the systems.

[31:58] Sangani: And obviously IBM Research has been incredible, all of the thinking that's been done. So there's a lot of heritage here that speaks to this exact moment. So maybe tell us a little bit about this. Everybody kind of feels like they know IBM or has heard of IBM. Tell us about what IBM is today, and tell us about what Arvind is doing today and what you're doing today.

[32:21] Jhingran: One is, it is incredible... I mean, you know this, we know this, that the technology moves faster than the landscape of our customers. Right? So you say, "Oh my God, everything is now AI." It isn't, right? It isn't. Oh, everything is a lakehouse. No, it isn't, right? So there's a huge amount of, look, these things work, let's tinker at the edges of it, as opposed to let's try to make something different.

Pillar one: modernizing the systems customers bet their businesses on

[32:52] Jhingran: And that's why, for example, from a data perspective, we've got a huge Z business. We of course have DB2s, we have DataStage, we have everything else that you would know of, Informix and everything else, that is centered around the fact that we basically run the business for a large part of the data real estate. So that's one.

And the way that is getting modernized is that the core system of records is not changing. But in the case of DB2, for example, we now have something called DB2 Genius, which says that, look, clearly the most painful task about databases is not the excitement that you get out of it, but really managing the damn database, right? And we have always talked about autonomous databases and autonomic databases and all this stuff, but now it can happen, okay? Because AI can actually observe and tweak and modify.

So around that kind of stuff, the basic problem that we are trying to solve is to say, look, how do we ensure that the people that surround it are actually able to deliver phenomenally large value by leveraging AI? Okay, so with something called AI-additional. [Product name unclear in the recording.] So that's one.

Pillar two: new infrastructure for delivering context to AI applications

[34:01] Jhingran: The second one is what we spend a lot of time on, which is how do you think through the new structures? How is data becoming relevant for this new application in the world of AI, right? So obviously, how do you deliver context? How do you build a vector store? How do you collect information from everywhere else? How do you leverage GPUs?

Our customers have GPUs up the wazoo, but GPUs after three years go down in terms of their relevance for AI. But can they be leveraged for databases and everything else? So how do you do some things that we spend a lot of time in, and deliver something for the new AI and applications out there? So that's the second one.

Pillar three: data in motion and the Confluent acquisition

[34:42] Jhingran: And the third mode that we are on is to say that, look, in this new world, we have spent a lot of effort in getting data correct, in the right place, and at rest. But is it the case that data in motion and data at rest... We talked about unstructured data and structured data, but again, the problem of data in motion and data at rest has never been fully satisfied.

And that's why we went and acquired Confluent, and to say, "Look, is it the case?" Confluent has some really, really fantastic pieces of technology. So I would say these are the three pillars. Modernize things that our customers bet their lives on. Second is build new infrastructure for delivering into the AI applications, and third is combine with it. Those are in my mind the three top pillars of our data strategy.

Which AI data investments are enterprises adopting fastest?

Sangani asks Jhingran to rank his own three pillars by customer adoption speed. Jhingran's answer is that modernizing what customers already run pays off first, and that combining streaming with structured data will take longer.

[35:53] Sangani: Where do you see in each of those pillars the greatest degree of adoption and speed, and where do you see more sort of... And obviously you're innovating in all of them, so presume that's the case. But where do you see customers leaning in faster and adoption moving faster? I would imagine that it's in the first plane of taking the thing that you already have and making it available in the AI world.

[36:21] Jhingran: You're absolutely right. In fact, after this, I'm having a call with a large customer who's saying, "Look, how are you AI-ifying your products? And what are the puts and takes, and how do you decide on, okay, these are the two models, but we have somebody else, how does it work," et cetera. And all the basic brass tacks around it. But it's a real payoff for our customers.

On the last category, with respect to streaming and combination, we also have a lot of traction, because the streaming business that Confluent had was north of a billion. And that by itself has a lot of traction. Okay? Now, I agree that combining it with structured data, shifting from Kafka to Flink as a way of processing it, using table flows, combining it with Iceberg tables, that stuff will take time. You're absolutely right. But that has its own center of gravity out there, and that we are just building.

I don't know, you probably don't read IBM's quarterly things, but that part of the business is also doing fantastically well. Okay? Again, centered around our center of gravity around streaming.

Does an AI data strategy have to be cloud-only?

Jhingran describes where IBM finds its sweet spot in the lakehouse market, and why on-premises estates still matter as much as cloud for the enterprises he works with.

[37:46] Jhingran: The middle one is something where our clients are working through. Obviously, data lakehouse is important. Obviously, lakehouse has certain key characteristics, key vendors, key players. But what we are finding is that we are finding a very good sweet spot with respect to hybrid deployments.

Okay? So as much as the fastest way to get started is in the cloud, the fact of the matter is that almost for every client that you work with and every client that we work with, their on-premises real estate is actually equally important. In some cases, even more important. So how do you think through this context and lakehouse and this and that and federation, all that, in that context?

So the niches are all very clear. In the first one, we have an established footprint, and we are saying, "Miss Customer, we are going to make it really, really successful over the next decades." In the last one, streaming is our center of gravity. And in the middle, it really is the hybrid and on-premises stuff. Those are the three vectors.

Why isn't shipping faster with AI a competitive advantage?

Jhingran shifts to the question he says he would not have been working on six months earlier, and states the argument the episode's closing half turns on.

[38:53] Jhingran: The job that I have is, obviously, I understand a bit of databases, so I help out and help people think through it. But very interestingly, I am focusing slightly more attention not just on what this world of data and databases looks like. Which is correct. It's clearly important.

But something that I wouldn't have thought of three months back or six months back, which is, how do we actually build products? And the reason is very simple, and you and I have chatted about this, is that if you just say that AI is going to help us build products faster, then it doesn't actually create a competitive differentiation, okay? Because everybody else is creating products faster with AI, right?

So you have to both do things differently and perhaps do different things. Again, maybe we'll get into some of that conversation when you have part two. So I'm spending a lot of energy beyond all these three pieces that we talked about, and helping really, really smart people out there, but also, how do you take these fantastic sets of technical people that we have at IBM and really create something slightly different? And so getting that thing at scale, not just within the Data Org, but outside Data Org, is where I'm spending a decent amount of energy.

Why IBM's CTO spends his time outside databases

Sangani expects a CTO with Jhingran's database background to spend his days on the future of the database. Jhingran explains why he deliberately works outside his own area of expertise.

[40:23] Sangani: Yeah, it's funny, because I would've expected you to be spending way more time on that first question of, what does the future of databases look like? Even in this pillar one, pillar three modality, what I would expect you to say is, "Oh, I see some convergence between this world of real-time and streaming and this existing operational estate, and these things are going to become ephemeral, and we're just going to build stuff and it's going to happen all dynamically," or something like that. I don't know.

[40:48] Jhingran: No, no, you are a database person.

[40:50] Sangani: But so that's what I would expect. And I guess, do you see any form of... Do you see revolution or evolution in the database world? That's a question on my mind. Does AI change even the shape of this thing, or does the shape of this thing become enveloped by this AI thing and it just becomes another artifact in all the stuff that AI can deal with and then manipulate, or a set of artifacts?

[41:17] Jhingran: No. So very interesting. I might have messed you up with the right question. So you asked me two questions, so let me first... You made a comment about, why am I not spending so much time on that side. I didn't say I'm not, but the fact of the matter is that as a CTO, I want to make sure that I say, "Okay, look, Anant is a data person. That's great, but IBM Software is much bigger, right? So what is it that we need to do in order to scale?"

So I try to consciously get outside my comfort zone and say, "What are the things we can do?" So that's one reason. And the second is, most of the people who are thinking through the data side are my friends and people that I've worked with. So I have a lot of discussion with them, but there are people that I enormously trust, and so I don't have to necessarily just wallow in the weeds out there.

Is AI driving a revolution or an evolution in databases?

Jhingran's direct answer on whether AI breaks the database, including what he expects to happen to standalone vector databases, and the condition under which he would change his mind.

[42:19] Jhingran: But then coming back to the other question that you asked, which is, is it an evolution or a revolution? I haven't seen inklings of a revolution yet. Okay? Now, sometimes small amounts of changes can... But is there a new database that's being built which is fundamentally centered around a vector store as opposed to centered around it? Yeah, of course. But is it then being absorbed into it? I mean, what's going to happen to all the pure play vector players over there? That's going to get absorbed into it. You don't know the answer to that, right?

Databases, operating systems, a few things have just managed to always make themselves relevant for the next phase. And it's likely that that's what's going to happen. But again, things are changing so fast that I see the things that we're doing with respect to innovation as adding to the core constructs, as opposed to something radically different.

The kinds of optimizations you've got to do, the kinds of indexing that you've got to do, the federation versus this, et cetera. None of those has not been thought of before, right? It is just the context that is different. But it's entirely possible that a set of small things will tumble into something remarkably different. It's entirely possible that the shape of the agentic workloads will look so different that just delivering context is not going to be sufficient. You've got to build a completely different data structure in order to support it. But I haven't concluded that yet.

Does AI make taste and distribution the real differentiator?

Sangani cites former Tableau CEO Mark Nelson from an earlier episode, then raises the idea Jhingran floated in their prep call: that many forked, personalized variants may beat one high quality code path. Alation frames this shift as an AI operating system for the enterprise.

[43:51] Sangani: Talk a little bit about this question of how we develop product. On this podcast, we had Mark Nelson, who used to run Tableau amongst other things, and his view was, look, yeah, everybody can create software more easily, which basically just pushes the tension to not the ability to create, but ultimately taste and quality. And he didn't say it, but I also think distribution is a big thing. Is that how you see the world, or do you see the world differently?

One of the things you said when we were doing this prep call that I found really interesting was, you said something to the effect of, "I used to think that creating one really high quality code path was the right answer, but now I'm not a hundred percent sure. Now it may be the case that we should fork things and really develop a whole bunch of personalization, customization," which is a different way of thinking about it than quality and taste, because it's actually just build the thing that somebody wants specifically, period.

How do you keep AI-assisted teams from producing slop?

Jhingran, responsible for engineering practice across IBM Software, describes the adoption gap he sees and the problem of raising quality judgment across thousands of engineers rather than a handful of craftsmen. He references Linux creator Linus Torvalds on accepting AI-generated contributions.

[44:52] Jhingran: No, exactly right. And you're again absolutely right. So just on quality and taste, I think taste is a bit of an overused term now, but I think that's really, really important. I'll just tell you a bit about... So what happens is this. Just like my parents can't use WhatsApp that easily, okay? But a generation, half a generation younger than them, it's really, really easy.

Similarly, when it comes to AI and AI tools, I'm just amazed as to how easy and comfortable it is for people to actually use it, right? So within IBM, we don't find any reluctance for people to actually use the AI tools, though there are some really, really good surveys out there. Lenny and Noam had come up with some survey that half the people are excited about it, another half are the opposite of excited.

But putting that aside, within that context of taste, the thing that I'm focused on, again coming back to the scale, is how do you raise everybody's level, okay? So they're not just slop producers, but are actually working towards and working through taste. And I read a lot about, look, Linus himself came out with how he will accept AI within his code base. And so there's a lot of interesting work to be done out there, where we go from, look, here are the master craftsmen or craftswomen, to how do you go to the mass thing. And I think that's one of the things that I'm really, really focused on.

Product releases were deliberately slow, because customers bet their businesses on them

Jhingran explains the reliability logic that governed enterprise release cycles before AI, and why creation speed was never the binding constraint.

[46:27] Jhingran: But now if you come back to this... So that's how you build it, but what about the structure of the software? Okay? And I think that some things fundamentally are going to be different. You're again absolutely right. Creation is easy.

But businesses, your customers and my customers, don't just depend on the next feature. They depend on the feature to actually bet their business on. Okay? And that's why typically product releases have been slow, because even in the past, if you could build something, making sure that you're going to a customer that is getting 99.999, you're not suddenly going to take it down to 99.9 with that new feature that you have, right?

So the process of product development was deliberately slowed because you wanted to make sure that the receiving ends are actually... And obviously, process also was slow in terms of you really couldn't produce stuff faster.

If AI can debug any branch, do you still need one code base?

Jhingran works through the second-order consequences of forking, including documentation and support, and lands on the reason a single shared code base may no longer be required.

[47:19] Jhingran: So the way to think about this, Satyen, in my mind, is what changes, okay, if AI is there? How do you make sure that the end-to-end stuff is actually faster, okay, and delivers? And therefore you've got to think through not just code, you've got to think through documentation, you've got to think through support.

And the question really is, in the past, support was good because everybody was trained on the common code base. But if you're now going to branch, how do you do support? But no, I can point my stuff to Claude, and Claude can actually figure out what's going wrong. Therefore, the reasons to have one code base may not exactly exist right now, because I can actually solve some problems. So there are some really, really fascinating things which say that some of our long-held beliefs may not work with respect to the structure of the software that we built.

[48:26] Sangani: Are you seeing this done in practice? Are you seeing people who have done branching successfully and are maintaining multiple variants of the same thing? And are there examples that you...

[48:35] Jhingran: No. So within IBM, I would say that it's very nascent. I'm pushing hard at it because we have trained ourselves to, a feature for one is a feature for all, right? That's the way we have trained ourselves. So it's a different muscle. But in my way, there are two or three really fast-paced products I'm building this way. And hopefully we can show that this can be done that way.

What stays shared when you fork a product into variants?

Sangani presses on what a product even is once it forks. Jhingran names the shared primitives worth staffing with the best engineers, and the two of them test the argument against Palantir.

[49:04] Sangani: I guess then the question becomes, what is common? Because the definition of product on some level is this reusable thing, and so once you fork, the question is, what are the shared primitives?

[49:17] Jhingran: So for me it's... Obviously. So the shared primitives are, in my mind, the thing that you put your best engineers on. Is it an entity graph? Is it a way of optimizing stuff? Is it an economic model? Right? So that becomes the shared stuff, okay? And then you create clean APIs around it, and then try to stamp out certain features off there.

But the shared stuff doesn't go away, because without that, you're not really hanging your hat on anything, right? You have to hang your hat on something differentiated, and that differentiator has to have your best engineers, okay? But you don't want to then not be able to deliver what the customer wants on the surface. So it's a tough thing, and we are working through it. And if you guys have figured it out, it will just be very, very awesome to exchange some of the best practices.

[50:12] Sangani: Yeah, I mean, Palantir has figured it out on some level.

[50:15] Jhingran: Yeah. No, but with Palantir it's an FDE all the time, right?

[50:22] Sangani: This is not quite the forward deployed engineer model. I'm still talking about engineers building it.

[50:24] Jhingran: Yes. And so it is really code as a free asset. But Palantir is a phenomenally well doing company, but they kind of keep peddling their knowledge graph and FDEs, and really it looks like the company's built on those two constructs. I'm not exactly sure whether that's the case.

[50:48] Sangani: Yeah. Brilliant. Okay. Well, we're going to pause here, and I guess pause because you're going to turn the tables and subject me to all sorts of questions that I won't be able to answer well come the next time. But like we always... It's always fun to talk to you, and the time goes by super fast, and it did, of course, this time. And so thank you for taking the time.

[51:06] Jhingran: No, really appreciate this. So see you soon on the other side, and thank you very much.

Closing takeaways: agents explore, so infrastructure has to adapt

The producer's close, summarizing both threads of the conversation: what agentic workloads demand of data infrastructure, and what AI changes about how software gets built and delivered.

[51:16] Producer: We've spent decades designing enterprise systems around predictable workflows. Centralize the data, clean it, structure it, and build software that behaves the same way every time. But agents don't work that way. They explore, adapt, and reason their way toward an outcome, and that demands a different kind of infrastructure.

The same is true for building software. As Anant points out, anyone can use AI to ship faster. The real opportunity is different. Keep the core technology that makes your product valuable, but tailor what you deliver to each customer. AI may not replace the foundations we've built, but it could give us a very different way to build on top of them.

Thanks for tuning in to AI Radicals, and look out for Satyen's episode of Context Window with Anant. See you next time.

This episode was brought to you by Alation. AI can quietly break, and it's not always obvious why: the data, the context, or the agent itself. Join hundreds of data and AI executives, governance leaders, and practitioners to learn how to catch it before it costs you at revAlation this fall. It's Alation's global event series, Chicago on September 17th, London on September 30th, and Sydney on October 8th, bringing together data and AI business leaders to talk about what's actually working in production, not just in the demo. If you're trying to move from "we bought some AI tools" to measurable business results, this is where that conversation is happening. Find your city and register at alation.com/revalation.

Other episodes you might like

  • Brian Solis

    Infinite: Why AI Business Reinvention Beats Automation

    Season 4 Episode 5

    Is your company using AI to become a faster version of yesterday? Brian Solis and Dave Wright of ServiceNow, co-authors…

  • Nathalie Berdat

    AI Governance in Public Media

    Season 4 Episode 4

    As Data Director of Product at the BBC, Nathalie Berdat leads digital transformation at one of the world's most trusted…

  • Francois Ajenstat

    Is Business Intelligence Truly Dead?

    Season 4 Episode 3

    What if every employee could work like a 10x analyst? Francois Ajenstat, Founder and CEO of Golden Analytics and former…