Chris Harrison: Why Large Language Models Fail at Enterprise Data
About this episode
Chris Harrison, CEO of EmergeGen AI, explains why large language models fall apart on enterprise data and how his team fixes it with an approach they call super ontology. The idea is a unified semantic data layer, built from ontologies and knowledge graphs, that gives AI agents the memory and context that general models lack.
Harrison walks through a case with a large commercial real estate firm whose New York office writes 40,000 lease proposals a year using 2,000 staff. His system read a 300-page proposal and filled a 95-metric template in about 10 seconds with full precision. He also covers zero-copy architecture, where the model moves to the data and never sees, trains on, or stores it, so banks and government agencies stay in control. He argues that small, domain-specific models are the future, noting that his own model has 7 billion parameters and runs at a fraction of the cost.
In this conversation
- Building a unified semantic data layer with ontologies and knowledge graphs
- How their model read and filled out a 300-page commercial lease proposal in seconds
- Why AI should put humans "on the loop" instead of in it
- Zero-copy architecture: working with banks and government without ever seeing or storing the data
- Why small, domain-specific language models may be the future
- The energy and cost problems with giant models, and how quantum computing could change AI
Full transcript
Welcome back to another episode. Today we're joined by Chris Harrison from EmergeGen AI. Chris, it's great to have you here.
thank you very much for having me on the on the show. Thank you.
Chris, could you tell us a bit about what's the problem that you guys are solving at EmergeGen AI?
Yeah, so the so the overarching problem is for people to be successful creating AI agents, the first thing you need to do is create a unified semantic data layer, meaning you need you need one layer of of data for the whole organization that takes all the historical information that sits behind typically be sits behind people's firewalls and silo data sets, which has been accumulating for years.
And you take all that information, we run it through our s our through our model, and it creates a unified data layer, which now is basically chat, but for the company's data. Once that's done, now you can build very secure intelligent AI agents on top of that to handle workflows efficiently and effectively that generate a positive ROI. So that's what we do.
Mm-hmm. And for for other people that are using LLMs like me, I sometimes struggle with the context. Basically AI not knowing what I previously talked about in the earlier sessions. do you guys solve that?
No.
Yeah, so there's a lot of problems with large language models, depending on what you want to do. Obviously, they've been very helpful writing emails and writing, you know, a paper or doing research. That they're excellent on. But when it comes to enterprise data, they're not very good. Right? One of the problems that you just mentioned is memory. They don't have memory. They don't have context. They really don't have any intelligence. All they're doing is guessing the next word.
Based off of thousands of examples of people that asked a similar prompt. Right? That's all they do. What our model does, we build something called an ontology, which is the roadmap for the company of how all of their data relates to each other, like a person would have in their brain. And then we create a knowledge graph. And that knowledge graph is what stores all of the most relevant data, the most up-to-date data.
So that you always that data is always at the the the fingertips of any of the employees, right? So it always has the memory. So that's the whole idea. So you know, what we see is a lot of people in enterprises using something like OpenClaw, building their own AI agents. But the AA AI agents don't have the intelligence. It's just like bringing in a new person to a company, sitting them down and say, hey, go do this, this, and this. And they're saying, Well, what are you talking about? I have no idea what you're talking about.
Right. Agents are only as intelligent like people, the data and the information, the quality and the volume of what you s you you give them to train on and what they what they can understand. So that's what's missing. And in large language models can't do it today. They won't be able to do it in the future. So these people that think Claude's going to solve everything or open AI have no idea how these company how these large land models are built and how they're trained and their real limitations.
on what they can do.
I saw this super ontology concept on your your site. where did that idea actually come from? Because I'm not familiar with it and I would love to know what it is about.
Yeah.
It's kind of it's kind of our tech team's play on words. So the the value that we bring to the market is experience. So the senior tech four senior tech people on our tech team, of which are there's twenty, have individually twenty-five to thirty years each in AI and machine learning, and then all of the underlying technology within our model, which would include things like knowledge graphs, ontologies.
Taxonomy, semantic layers, data extraction, OCR, NLP, computer vision. Because it's it's a it's an ensemble of technologies that we've these this team has accumulated over decades, right? That works together in concert that allows us to do what we do. So the underlying tech is very similar to the approach Palantir has pioneered using ontology and knowledge graphs. We call it super ontology because basically
Some companies have an ontology and have a good understanding, other companies don't. So in the absence of the company having an ontology, our our model will build an ontology for them, or it will make their existing ontology more accurate, because there's typically errors in people's ontologies anyway.
Mm-hmm. I also saw on your website that you have some great testimonials and some impressive claims like giving someone twenty five X ROI, which which is huge, right? who do you think benefits most from what you guys do?
Well, it it's very interesting. So we've probably had over a hundred discussions with different enterprises across every industry. Everyone struggles basically with the same problem. So if you think about the world in data, enterprises, eighty to ninety percent of enterprise data is in an unstructured format. So for the last fifteen, twenty years, people have been trying to solve that. How do you get that unstructured data safely, securely, accurately?
into a structured format that you can create automated like data pipelines, automated workflows with using AI. That's been the big challenge. And obviously Palantir has been very successful at it in a very specific part of the market. We're created more of a middle market Palantir to handle the masses. So the use cases are every company is struggling with this. No matter how big, how sophisticated, how big their tech teams are, they still haven't solved this.
Right. They've solved parts of it, but they haven't solved the difficult information. So if you take a look, here's an example. We had one of the largest commercial real estate companies in the US and actually globally. Their New York location creates 40,000 lease proposals a year to win 4,000 deals. Right? How did they do it? They have 2,000 people that cut and paste old lease proposals to make new lease proposals.
So what we showed them is by giving us one lease proposal, which is a very complicated 300 page document that has tables and graphs and pictures of architectural drawings and car parks and all kinds of information. And we put that into the system and they gave us 95 metrics that the people manually are using to create these lease proposals. Put them in, we we hit the go button. And in like 10 seconds, it had read and understood that whole 300 page document.
extracted all the relevant information and filled out that template. Now the thing that was impressive was it did it with a hundred percent precision. So not much better than the accuracy of humans, right? And the first thing the CEO said to me was, Holy crap, I can get rid of a thousand people. Because right now they have two two thousand people doing this. That was the first thing out of his mouth, right? But what what we what what
Wow.
You need what people need to understand is this isn't just to replace people. That's not what AI is. AI is a tool to leverage people. So what we said was that's one way of looking at it. But really the way you should be thinking about this is you're only you're you're only creating 40,000 a year. If you use our system, when a new RFP comes in, it'll run through, generate 90%, 95% of that RFP, then the humans only doing the last five percent, right?
So instead of doing 40,000 lease proposals, maybe you could do 150. Right? The other problem is you're losing 90% of your effort is generating zero revenue. So how about if now we take all of the winners? Well what we'll do is take the last 10 years of proposals, load them into the system, create this queryable database of your old belief, and you'll see these are the ones we won, these are the ones we lost. And now we can run models and analytics to improve your win ratio.
Mm-hmm.
So if you can move here from 10 to 20% success, you'll double revenue with no additional cost. Right? So that's the leverage, that's the value of AI and how to use it properly. And that's just one simple function within a very large organization. Right? There's contracts, there's all kinds of other unstructured data or other things that could be automated that they have people doing manually. So it reduces errors.
Hmm.
Yeah.
It increases efficiency, reduces costs, and drives revenue. Because look at no no CEO cares about just reduce you can't reduce cost and win. You have to keep growing. You have to drive revenue. That's the key to AI.
Mm-hmm. So it's still human in the loop, but you're basically the tasks that are for humans are the most higher leverage ones. Is that what you're saying?
Yeah, so we Yeah, we would say instead of the human in the loop, we would say the human on the loop. So we want to take the human out of the loop and put them on the loop. So yeah, you need a lot less people, but the idea is AI can do ninety to ninety-five percent of the heavy lifting of the a lot of the manual work that gets done that does not need to be done by humans. And that's in every business, every industry, whether it's healthcare, manufacturing, insurance, asset management, banking.
logistics, real estate, it doesn't matter. They're all they're all in the same boat.
How does this play in for example, federal government and banking, which which must be much more, how should I put it, much more secure and you cannot like
Yeah, I know what you mean. Y y it's confidential information. And and that and that would that would be healthcare has HIPAA, you know, the banks have very confidential information. So all of them, all the industries that we deal with, all have confidential information. So how do we how do we solve that problem? Right? Which is around security. So the other big problem with large language models is when you give it information, you have no idea where that information is going.
Yes.
It's completely going into a black box. You don't know if they're storing it, training on it. It's going into the internet. You have no idea. Now their their latest trick has been like, you sign this agreement, we'll give you a private instance. And we we pledge to you that we won't do that. But you as the in the enterprise have no way of validating what they're doing with your data. That's why big enterprises are never gonna give their confidential information to any large language model, right?
Mm.
So what we do is we take our we take our technology, our model sits in a container, we put it into the client's environment behind their firewall within their data sale. So whether it's in Snowflake or Databricks or Mongo, we don't we're at completely agnostic to the database or the cloud provider, or they want to use the LLMs eventually, right? So we don't do that. So what we do is we install there. So we use something called zero copy.
So our model actually moves to the data versus the data moving to the model. So we never see, train, or store data. So that's how we do business with government. And the US government is in the whether it's in the Department of War or whether it's in you know Treasury or Social Security, have hundreds of millions of documents that need to be put into queryable databases. And and we're we're pitching deals all the time.
Mm-hmm.
to try to help automate what they've done. Now the scary part is a lot of it's still in paper, which is kind of mind-boggling if you think about where we are today with technology. But but there's huge opportunities there in government and that's across federal government, state government, you know, even the public, you know, the the local government, all of them need a huge upgrade to become more efficient with using people's tax dollars.
Mm.
Mm-hmm.
Mm-hmm. How do you guys approach go to market strategy when it comes to pitching these kind of systems and AI layers?
So we've the b we we probably want to have eighty percent of our revenue driven through partnerships and twenty percent direct. So the partners we've already integrated with is you know, Collibra, Data IQ, Snowflake, Salesforce, and now we're integrating with Oracle and soon to be Databricks and then IBM. So all those platforms are saying, why would they use a startup's technology? Like why do they need that?
Because those companies are very good at structured data. So they they work on the 10 to 20 percent of an enterprise's structured data. That's what they can use and run through their systems. Unstructured data does not work easily in their systems. Now, the the more generic unstructured data that can work with some of those systems. And there's lots of unstructured structured data tools that are just extraction tools. But to create this unified data layer across the organization.
The only real competitor we have is Palantir. And Palantir does not focus on the part of the market that we do. Right. So that's the nice thing for us, is there really is no other competitor doing what we're doing. so the opportunity set is it's it's unlimited. But that's we're we're using those partners. We use Carahsoft, who's the largest, you know, federal government in the US.
prime contractors. So we use prime contractors there. We use Beat, we use C D C D G. So we use a lot of prime contracting partners in the US. And then we use these big database companies in the US.
Hmm. It must be very early still, right? Like the AI adoption into companies. It's still all getting started. So it's
Yeah, it's i i i it started with people throwing information in a chat a few years ago. and it's evolved since then and the individuals are all using it. Enterprise use is still nascent. they're trying it, they're dipping their toe in the water. But the biggest problem they're having is the security issue around large language models. That's one. The other thing is hallucinations.
And then if they're still making up information 20% of the time, it's not like they're making up information. Really, what it's doing, it's just making the best guess based off of your prompt. But the problem is if you ask the same question three different ways, you're gonna get three different answers. And in an enterprise, if you need an earnings number, you need that number. There's one number. There's not three numbers, there's one number. And that and that's the that's the problem for enterprises. They need exact information.
Ha ha.
So really the future is small language models that are like our model runs one hundredth the cost of what it costs to run a pr you know a token for you know the tokens and you think about token usage. That's the other big problem. So I just talking to one of my friends this weekend and his company is forcing him to use more AI, right, which is more token usage. So the the CTO of Uber said that their company had blown through their token usage for the year.
by April. So the other problem with people writing, doing this stuff is they don't know how to efficiently use you know the the the AI agents, right? So w we build and manage them for people because people don't know how to do them efficiently, safely, and and how to get the max and then they're not secure. So another thing that we're we've doing is we're we've created something called quantum key distribution.
Mm.
which allows our agents to communicate securely using both classical and quantum computing. So it makes them unhackable today, but also make them unhackable in the future.
Mm-hmm. Where do you think the AI industry is heading? Like what what's the what's the next thing?
I think the next thing is is is the small language models. I think there's going to be people walking away from the large language models just because of the cost, because of the cost of energy. Like if you look in the United States, I think forty percent of the databases that are trying to get built are not getting built, and states are pushing back because it's driving up consumer rates for electricity. So I think what you're going to see is people having to be more efficient.
with what they're doing with AI. And that's going to be driven to small like our model has seven billion parameters. That's going to be the future, not one trillion dollar parameter, one trillion parameter models, which just uses tremendously more energy. And so I think that's going to be, you know, one of the things we're going to see is the move towards domain specific smaller language models like what we do that are much more efficient, much more effective, much more accurate.
So I think that's one thing we'll see. I think the other thing we're gonna see is the adoption of quantum is gonna come faster than people anticipate.
Mm-hmm.
And and when you plug quantum into AI, you get results that are mind boggling.
What w what does it mean? Like what what kind of results are you talking about? Because I'm not very familiar with quantum. So could you
You need yeah, so i in simple terms, this is the way it was explained to me because I'm not a quantum physics person either, but the way it was explained to me is classical computing uses chips and one plus one plus one is how you get to Moore's law, right? Quantum uses qubits, one plus one becomes exponential, right? So the every additional qubit you get exponential capability.
Uh-huh.
with the technology. But the real key is if you think about a one of the like a very complex maze, and if you try to solve it with a classical computer, what it'll do is it'll go through each potential opportunity to get to the end one at a time. What a quantum computer will do is it'll go every single way at the exact same time. So for certain functions, quantum computer can handle things that would take
a classical computer thousands of years to calculate, a quantum computer can do in minutes. So it's not for everything. So what you're going to see is hybrids. So you know you can see like Nvidia's computer where it has both you know qubits, has both quantum and classical computing built into the same system because there's certain things classical computing will always be better at, but there's going to be a lot more functions that quantum computing is much better at.
Hm.
And and it's gonna allow us to crack things like cancer, it's gonna crack things like climate change, it's gonna be able to crack things that the today's computers can't even possibly attack. That's the good news. Bad news is it will break all the encryption that we currently have in place. So classical computing would take years to break through like the encryption eight-year bank. A quantum computer can break that in minutes. Right. So
Mm.
The bad actors is are working. So there's a lot of movement on encryption and how to protect how to protect or how create quantum encryption basically so that you protect it against when quantum becomes cause quantum itself, what they call quantum supremacy, is depending on who you want to talk to, it's three years away, five years away, ten years away. It just depends on what your definition is. It's kind of like AGI.
But it's we're seeing people using quantum computing today, including ourselves. So I I think that's the next big leg and quantum computing will be dramatically more impactful than AI alone. By combining those two things, it's going to be we're gonna see changes that people didn't even understand.
Mm-hmm.
Hm.
So
Mm-hmm. So it will speed up the learning curve, right? A lot. Mm.
A yes. Yes. Yeah. It'll help with things like cancer research, drug research, things like that that would have taken years, well can be done much faster now. And large language models are the beginning of this, but when quantum gets in when when when large quantum computers are actually able to handle things properly and accurately, it will it's gonna change forever.
what what we can do with AI.
Mm.
Hmm. Okay, Chris. Well wrapping this up, where where can people find you or if they want to check EmergeGen AI?
They can they can either contact us through the website or they can contact me directly. It's Chris at emergen.ai is my email. and we're happy to walk people through what we're doing and and how we can be useful to their organization.
Okay. Well thank you, Chris.
Thank you very much.