Stop Renting Generic Intelligence for Your Business

Manos Koukoumidis on Specialized AI, Agents, Post-training, and the Cost of Generic Models.


Subscribe: AppleSpotify OvercastPocket CastsYouTube •  AntennaPodPodcast AddictAmazon •  RSS.

Manos Koukoumidis, CEO of Oumi, returns to the podcast to make the case that most enterprises are wildly inefficient, using bloated generic AI models for tasks that demand specialized ones. He argues that companies should stop renting intelligence from frontier labs and start owning, customizing, and compounding it. Manos explains how Oumi’s platform automates the entire lifecycle of building, deploying, and improving custom models. The conversation spans open weights risks, AI sovereignty, reinforcement fine-tuning, agent specialization, and whether the AI industry’s economics are sustainable.

Subscribe to the Gradient Flow Newsletter

Interview highlights – key sections from the video version:

Jump to transcript



Related content:


Support our work by subscribing to our newsletter📩


Transcript

Below is a polished and edited transcript.

Ben Lorica. All right, so today we have my friend Manos Koukoumidis, who has been on the podcast before. He’s the CEO of Oumi, which you can find at oumi.ai. I guess it stands for Open Universal Machine Intelligence, but the tagline, if you go to the website, is “The right model meets the biggest model. Stop renting generic intelligence, build and deploy the best specialized model for your task, and own it fully.” And with that, Manos, welcome back to the podcast.

Manos Koukoumidis. Thank you for having me, Ben. It’s really wonderful to be back.

Ben Lorica. All right. As you know, I’m a big fan of open weights, and I think we’re at the point now where, when we last talked, maybe the gap between open weights and the leading frontier models was on the order of 12 months. Maybe now it’s more like six months. It seems like the models continue to get better, but obviously the rate of progress has slowed down, so that has allowed the open weights models to catch up.

So first question to you, Manos: What is the state of open weights models in the enterprise? It seems to me that with open weights, your motivations are many, but one of them obviously is cost. The second is control. On the cost side, it seems like you only care about cost when you’re using it a lot, right?

Manos Koukoumidis. Yeah.

Ben Lorica. I have this contradictory information about enterprise adoption of AI. On the one hand, everyone is obviously excited and wants to do it, but we’re in a tech, Silicon Valley bubble here in the Bay Area and in Seattle. When you talk to people on the East Coast, for example, yes, they’re adopting AI, but they tend to do it more slowly. When you’re still at the beginning, you’re more likely to say, “Oh, let’s just use Claude or an OpenAI model, because we just want to understand the capabilities,” right?

Manos Koukoumidis. Yes.

Ben Lorica. What’s your sense in terms of interest in open weights models now?

Manos Koukoumidis. Both in terms of my interest and, as you said, what’s happening right now in the enterprise with open models versus closed ones, I think you said it perfectly, Ben. When enterprises right now say, “Hey, we have a new use case,” for them to start with Claude, GPT, or Gemini, which I was building at Google, makes a lot of sense. It’s the easiest thing to prototype with.

It’s usually when enterprises get more serious, when they are a little bit further along the maturity curve of using generative AI, that they have use cases in production and they have seen that things sometimes fail. They say, “Hey, if we just prompt the model, and then we change the prompt, it’s just not maintainable. We can’t keep fixing whatever we see in production that way.” They start really caring about improving quality beyond just a prototype or an MVP. They really care about quality, differentiation, and doing better than competitors.

The other case is when, as you mentioned before, they start hitting scale. When they hit scale and they get large bills, they say, “Okay, we can’t sustain this. We need something better.”

And the last one, which you also mentioned, is especially true for regulated industries, but I would say in the last couple of months it’s not just regulated industries. Even other enterprises have concerns about AI sovereignty, about controlling their own AI, because they realize, “Hey, AI is going to be the most critical component that powers my business. If I’m renting this, if I’m beholden to OpenAI, Anthropic, Google, or anyone else to consume it, I’m going to be in trouble one day. That’s not a viable strategy for me.”

For all these three reasons — quality and differentiation, cost when you hit large scale, and realizing that this is critical for success and needs to be controlled internally — I’ve seen, over the last two years, and at an increasing rate, enterprises embracing open models, especially customized, specialized open models.

Ben Lorica. Let’s tackle these one by one. On customization: I assume that if you’re Anthropic, OpenAI, or Google, you would provide some tools for enterprises to customize Claude, a GPT model, or Gemini. Is that correct?

Manos Koukoumidis. Yeah. Google provides some tools. Actually, some of the tools for customizing Gemini were built by my team at Google, and by some of my now co-founders at Oumi. Anthropic, on the other hand, provides none.

Ben Lorica. Okay.

Manos Koukoumidis. When I asked one of their representatives at AWS re:Invent a little over a year ago, I told them, “Hey, my model doesn’t work as well. Should I fine-tune this?” They said, “No. If you change the prompt and you don’t get the results you want, there’s nothing you can do,” which is extremely inaccurate, but it is self-serving for Anthropic.

And OpenAI announced recently that I think in October, or sometime soon, they are going to be deprecating their tuning APIs, which is unfortunate, because that was the way any enterprise using OpenAI could specialize it. But I think the reason makes sense for them. For consolidation, it’s also to their advantage if everybody is consuming not their specialized models, but the big generic models they build, because that can help make those models even better.

That’s why I often tell enterprises: Don’t just rent your intelligence. You should own it and let it compound for yourself. For every interaction your customers have with your business and with the AI that powers your business, your AI should be getting better. That should be your core asset — not OpenAI’s, not Anthropic’s, not Google’s.

Ben Lorica. As far as customization, or post-training in general, early on it was supervised fine-tuning, which seems easy for people to understand: create labeled examples of prompts and desired outputs. Very early on, Manos, there were services where you could just upload this data, go away for lunch, and come back with a fine-tuned model. More recently, people have started thinking about or talking about reinforcement fine-tuning.

Manos Koukoumidis. Yeah.

Ben Lorica. The tools there are still somewhat advanced. I guess there are startups trying to democratize it, but it’s not quite there yet. Do you provide both?

Manos Koukoumidis. Yeah, we do provide both. It’s exactly as you described it. Until recently, most people would do supervised fine-tuning, full fine-tuning, or parameter-efficient fine-tuning like LoRA. Now reinforcement fine-tuning has become even more popular, I would say.

Ben Lorica. Stop there for a second. Who does this? Do enterprises actually use that? Are they aware of what reinforcement fine-tuning is? Or do you have to educate them?

Manos Koukoumidis. To some extent, you definitely need to educate them, even though there has been more and more noise about RL. Everybody knows RL is a thing, even though they may not deeply understand how it works or where they should use supervised fine-tuning versus RL, because there are pros and cons. It’s a little complicated which one you should use and when.

But I think the main problem, because you said we need to educate them, is that most enterprise leaders — and I think this is changing fast — still don’t realize how irrational what they are currently doing in production is.

As I tell people, pretty much every company right now is becoming an AI company. AI is at the core of their business one way or another. But the vast majority of them are doing it wrong. Why? Because they are just using the biggest generic model, or even if they use a mini model or Sonnet instead of Opus, or Flash instead of the big Gemini, it doesn’t matter. They are still using a big generic model that does a thousand things, as opposed to the one thing they care about.

They pay the cost and the price, not to mention that they are again renting it, not owning it. They are not improving it and compounding it for themselves; they are doing it for those other companies. It is wildly inefficient, to say the least.

Now people say, “Oh, we need more data centers.” It’s because we are using models that are 10 to 100 times bigger than they need to be for the task that most enterprises are solving. Why? Because we’re using generic tools, not tools specialized for the job.

That’s the main education that most enterprise leaders are currently missing: you shouldn’t be using the generic tool. If you’re going to do surgery, you should not use a huge Swiss Army knife to cut somebody open. You should use a scalpel.

Ben Lorica. Another way to translate this is: if I’m going to do sentiment analysis or classification, I don’t need a model that can speak 200 languages. I just need a model that’s good at this one task.

Manos Koukoumidis. Exactly. I can’t overstate how wildly inefficient this is. Sometimes I tell people it’s the equivalent of us commuting to work every day in a semi-truck, as opposed to using a sedan or a Vespa. We’d have a fuel crisis the next day. That’s what’s happening now in AI.

The only reason this is happening is what you mentioned at the beginning: enterprises started by consuming these big models. The alternative — using an open model, evaluating it robustly so you know it’s at least as good, if not better, and then specializing it so you get the quality you need — was prohibitively hard for them. It was hard in terms of both the amount of effort and the expertise required.

That’s why we’re now in a wildly irrational state in the industry. But I think the natural rationalization of the market and the industry is coming at an increasing pace. It’s only a matter of time.

Ben Lorica. Once you start realizing how expensive these things are and how slow they are, it seems like that’s a natural nudge to go down the customization route. And then once you get down the customization route, I’m skeptical that a typical enterprise has even heard of reinforcement fine-tuning.

Manos Koukoumidis. Yeah.

Ben Lorica. So then it’s your job to educate them, but set realistic expectations. How many of these conversations have you been having recently?

Manos Koukoumidis. Numerous. It’s hard to count them. Definitely dozens and dozens — more than 100 over the last two years since we started. I had such discussions when I was at Google, starting with the efforts called PaLM, before they were handed off to DeepMind and rebranded as Gemini.

There have been many more discussions lately. What I’ve seen is that education — not necessarily about what reinforcement fine-tuning is or how well it works, but even just the idea that there is an alternative, that you could own open models, especially specializing and customizing them to your needs, and get better results, lower cost, full control, and faster results — is becoming easier. More and more people are starting to realize this, and the amount of time we need to spend educating leaders is a lot less.

In some cases now, more and more people realize the problem and come to us, as opposed to us reaching out to them and saying, “Hey, listen to us. There’s a better way.”

Even now, sometimes we get large organizations — I won’t name the companies, but some of the biggest organizations that exist — telling us, “Our CTO, CIO, or CFO thinks that for us to stay at the frontier, we should consume these labs’ models.” Quite often I then have the discussion with them: “Do you really want to depend on these frontier labs for what is most critical for you? Do you think Anthropic, OpenAI, or Google knows better how to do risk analysis or analyze stocks — if you’re a bank, for example — than you do, given that this is your core business, you have institutional knowledge, and you have in-house experts?”

Then they say, “Okay, I understand what you’re saying. This is specifically the area where we should own our own intelligence and not just use what they give us.”

Ben Lorica. By the way, these two leading labs, Anthropic and OpenAI, are starting to hire people from these verticals, including finance, so they can sell their solutions to banks and things like that. I don’t know exactly what they’re offering — Claude for work, or agents tuned for those specific tasks. I would imagine that if you’re one of these AI leaders, you go, “Okay, maybe we start there, because we just want to know.”

The psychology, Manos, is that these are the leading models. Maybe they present some sort of upper bound for what’s possible. Once I know that upper bound, maybe then I can go to customization. But on the other hand, if I can’t customize those models, maybe that’s a false upper bound, because I’m throwing a general-purpose model against a domain-specific task. I don’t care what harness is around it; it’s not going to be as good as a specialized one.

Manos Koukoumidis. Exactly. I can’t tell you how many times I hear, “We’re building agents, multi-agent orchestration.” I’m like, “Okay, but what about the core of the agent? Is the core model that is powering all of this the most reliable and accurate model it can be?”

Yes, you can build extra harnesses around it. All of that is good. You should have it. But you should start by putting specialized intelligence — the specialized model — at the core.

To add a little comment on what you said before, I think most organizations right now may still be a little brainwashed or intimidated by these large frontier labs, thinking, “There is no way we can compete.” As you said, they’re hiring finance people, and Anthropic released maybe a dozen finance agents.

But what I tell enterprises is that they still have much more expertise than those companies. They have the data. They operate the business. The algorithms that are already out there, compared to what Anthropic or anyone else is using internally, don’t have a huge delta. You have the data, you operate the business, and you can put that to your advantage. Build a specialized model, and then when it’s in production, with self-learning RL and these approaches, you can have it continuously improve and compound for you. Then no frontier lab is going to be able to catch up with you, especially for that domain and those specialized problems.

It’s not about having more AI than your competitors. It’s about picking your bets. What is the core AI, the core thing that differentiates your business from anybody else? Solve those specific problems with your own specialized AI.

Ben Lorica. I don’t know to what extent people are aware that, in many cases, these agents — setting aside coding agents — if you go into finance, these finance agents are going to call financial calculators and financial tools that you built. Even though these labs keep talking about AGI, these models are not going to do the sophisticated financial calculations and simulations that you already spent years building. They’re just going to call those tools anyway.

Manos Koukoumidis. Exactly. I can’t name the exact examples, but there are examples from classical scenarios like document classification, other types of classification, information or entity extraction from documents, and slightly more complex RPA scenarios, where you have models clicking on different applications and automating workflows.

Time and again, we work with customers where the models, as they came out of the frontier labs, were just not good enough to go to production. The moment they specialized the model for the specific problem they were solving, they saw dramatic quality improvements — in many cases not just one or two percentage points, but 10, 20, or 30 points of improvement, at a fraction of the cost. They fully own it.

The reason is simple. When we ask something like AGI from those models — no matter what task I give you, for whatever domain, by just a prompt description, you’re supposed to understand it, automate it, understand where to click in the app, and automate this process — that is 100,000 times, or even more, harder than saying, “Let me teach you exactly how to do this automation, and then you go do it.”

So whether AGI comes tomorrow or who knows when, specialized intelligence can give you the benefits today.

Ben Lorica. The reality, Manos, is that knowledge workers use tools. I don’t care what you do. I gave the example in finance: you’re using all sorts of financial calculators and models. I’m sure designers have their own specialized tools as well. These models will end up calling those specialized tools. So why do I need a super-expensive model?

Manos Koukoumidis. I’m with you. Quite likely, the model that will be best at calling the specific tools, knowing how to best call them, organize them, and orchestrate them, is going to be something you teach specifically how to do this, as opposed to a generic model that you just prompt and hope it will be able to use all the tools in the same way the best human expert would.

Ben Lorica. Obviously there’s no free lunch. There are also downsides. Downside number one is that if you try to do this on your own and not use a platform like Oumi, then you have to grab the model and maybe do some basic security checks, just in case the model is phoning home to China. I don’t know what it’s doing, even though it’s open weights. So there’s governance and supply-chain risk. And if you’re going to do this on your own, you have to deploy it, which means you have to manage inference infrastructure. Another reason people start with these frontier labs and cloud providers is because they can provide all of these things.

Manos Koukoumidis. Exactly. When we talk to very small startups all the way to some of the largest organizations you can imagine — top 10 banks, for example — it’s often the same story, although it has been changing very fast. They say, “For us to build, operate, deploy, and figure out all the different steps is too hard.”

Even just building a model: there are many platforms where you can train a model, but what they don’t do is the rest of what an engineer needs to do. You need to evaluate the model. You need to build your evaluations, build your test sets, analyze the evaluations, figure out how to curate the right data, figure out the right training algorithms and hyperparameters, train, evaluate, and repeat the process.

Usually, what most of these platforms do is assume you’ve done everything. You give them the data, choose the right algorithms, choose the right hyperparameters, and then they just run the algorithm for you, which is, in many ways, the easiest part of the whole process.

That’s why enterprises were stuck. Doing all the steps an engineer needs to take — not just pushing the training button, but all the steps — and then deploying it would take prohibitively long. Quite often, large organizations tell me, “We don’t have people who know how to do this,” or, “Maybe we only have a few, but we have dozens or hundreds of use cases. Why are we going to do this? It’s not going to scale.”

As you said, the ability to build and maintain this is prohibitively hard unless you have the right platform. That’s why we built Oumi: to have a platform that helps enterprises, without exaggeration, build, own, and compound their intelligence in production 100 times faster and easier.

Ben Lorica. Another risk that I’ve been worrying about more lately is supply. As you know, open weights is a business decision that can be reversed. Meta could decide, “We’re not going to do Llama or open weights anymore.” Thank God for Google and Gemma, but who knows? Down the road, maybe they change their mind as well. Some of the Chinese open weights providers have backed off, or signaled that maybe they may not do as much of that in the future. Is this something you worry about?

Manos Koukoumidis. It’s definitely something I worry about. I’m still optimistic that Mark won’t go back on “open is the path forward” — I think that was the title of his blog post maybe two years ago — and that they will continue to release open models. I would surely hope so, because it aligns better with the ethos of Meta.

But this is definitely a risk. Open models may be taken away at any time. Things like Llama may be taken away at any time. Even if they’re not taken down by the government, maybe one of these frontier labs decides one day, “Actually, I’m competing with you in coding, financial services, or legal AI. I’m sorry, you can’t compete with me using my model.”

Ben Lorica. The Cursor clause.

Manos Koukoumidis. Exactly. I think that’s exactly why Cursor built its own models. It’s not just that they built their own model and it was better than GPT or Claude at 10 times lower cost. It was also so they could get independence, because it was a huge risk for them to know that anytime Anthropic or OpenAI could say, “No, you can’t compete with us using our own models.” They own the models and control the terms.

I think Cursor did absolutely the right thing there, and now there are more companies like Intercom following the same example. Harvey announced that they’re doing the same thing.

Ben Lorica. The supply of open weights — is that something?

Manos Koukoumidis. Let me go back to the supply of open weights. The point I was trying to make is that this risk exists for closed models, but also for open ones. Open models could be taken away at any time. If anything, what it tells me is that, as a leader of a large organization, I would make sure I can take the best open model that works for me, specialize it, put it in production in this compounding, self-improving way, so that even if one day the next version of the model is taken away, I don’t care for the use cases that matter most to me.

I have taken my own model, specialized it, and had it compound so it can stay ahead. It is at the frontier — actually beyond the frontier of the big labs — for the specific tasks. I have my sovereignty and independence. If anything, it’s an alarm for organizations to get that AI independence as soon as possible.

Ben Lorica. To our listeners, I’m hearing about early initiatives and conversations around the possibility of training open weights models on a regular cadence, away from companies and labs — more like consortiums and much more decentralized. Some of them seem determined to do that for a variety of reasons. One is research; they don’t want to be beholden to the labs. But also, as I said, they want a steady supply of open weights models.

I think Google seems committed because they’ve made the distinction between Gemma and Gemini. Gemma seems like something I’m not sure they would take away. What’s your sense?

Manos Koukoumidis. I would like to think not. Going back to the supply risk of open models, I think between Google, Meta, and other companies like Nvidia, there are always going to be open models. There are a lot of tailwinds and reasons why many large organizations around the world need open models to exist as a counterweight to the big frontier labs.

That’s why I’m a little hesitant to say that one day we won’t have open models that are competitive. There are many tailwinds and many reasons why companies — from Nvidia to Meta, hardware companies and software companies — need this technology to exist and be accessible, so that one day they can’t be locked out of it. If OpenAI, Anthropic, or Google dominate and then decide they can lock some other company out of this technology, that’s a problem.

So I think it’s unlikely that we will have no supply. Open weights supply is there. Some companies may come in and out, but I think there’s always going to be something there.

Ben Lorica. Since you talk to many regular companies, when you go into a company, what’s the tell that this is going to be a longer engagement and you need more forward-deployed engineers, versus one where you think, “This is going to be okay; we’re going to get these guys up and running very quickly”?

Manos Koukoumidis. This is related to the capabilities of the Oumi platform and what we’re building. Just to give a quick summary, what we’ve built is, you could say, the first AI engineer for AI, or the first platform that acts like an operating system for you to build and compound your intelligence. It’s the first system that can fully automate the process of building, deploying, and compounding a model in production.

Even just for building, when you ask whether we need forward-deployed engineers and how much effort is required, there are many capabilities we have already automated really well. Some are still coming and are not fully automated yet.

For example, right now if somebody comes and says—

Ben Lorica. “I want to do the following, but my data is all over the place.”

Manos Koukoumidis. Exactly. In some cases they may have the data; in some cases they may not. Let’s take the simplest case where they don’t have the data, but they say, “We’re a bank. We’re classifying customer support tickets based on these 10 categories and by urgency — low, medium, high. We’re using Sonnet, but it’s expensive,” or another bank might be worried because they are using GPT Mini and it’s getting deprecated.

Ben Lorica. Or, “We’re using Sonnet. It works until it doesn’t, because they made an update. We tried to adjust the prompt.” It’s like whack-a-mole.

Manos Koukoumidis. Exactly. These things happen many times. They can come to me and say, “Build a model that classifies customer support tickets based on category,” and they can even copy and paste their GPT or Claude system prompt. They just describe it, and Oumi automates the whole process of an engineer.

It will ask questions to better understand the problem and scope it. Then it will say, “The first step, like a good AI engineer, is to establish your baseline. You need evaluations for this problem. Here are the right metrics. You need evaluations. I’m going to build them for you. You need test sets. If you don’t have any, I will build them for you and make sure they’re thorough.”

Then it evaluates the baseline — closed models, but also the open model family and size that it thinks are good for the specific task. It won’t just give you scores, like 87% versus 93%; it will also go deep and tell you the most common failure modes and the patterns in which this model fails in production.

Then it will do the next step of an engineer: build a training set for you to improve the quality. It will create a comprehensive training set that addresses these failure patterns, then tell you the right algorithm and hyperparameters to use, build the training recipe, train the model, evaluate it, and then say, “We did an iteration. We trained the model. Here is where I still see potential for improvement. Do you want me to keep going?” If you say yes, it will do the loop again to keep improving quality.

All the steps I described — defining the right evaluation metrics, building evaluations, building test sets, curating the right data, figuring out the right training algorithm, running experiments and training, and then evaluating again — are the reason AI engineers would typically take weeks or most likely months, even for the simplest use cases.

Now, with no exaggeration, what I described can take at most a couple minutes of human time in Oumi. It will take hours on the clock, because training and evaluations still take time. But in terms of human effort, it’s minutes. The reason this works is because it’s the equivalent of Cursor or Claude Code for AI development, not software development. You start with a prompt, it automates all the steps for you, but like Cursor or Claude Code, you still have full visibility and transparency into all the steps. You can go in and change them the same way you can change the code you get from coding agents.

If you ask me, automating this is perhaps easier than automating software development.

Ben Lorica. Do most of the teams you talk to have strong opinions about model families? Do they say, “Manos, we want you to use the Kimi family,” or do they just defer to you and the platform?

Manos Koukoumidis. Most of them are not opinionated, except in some cases where they may say, “We don’t want Chinese models.”

Ben Lorica. What percentage?

Manos Koukoumidis. I don’t have good numbers, but I would say perhaps 30% or less. Even for some of them, when they say that, I ask, “If you fine-tune those models and you’re hosting them yourself—”

Ben Lorica. Hosting it yourself. There’s no security issue?

Manos Koukoumidis. Exactly. It operates within your trust infrastructure. There’s no outbound connection. It only classifies information. It’s not an open-ended model that was aligned to do all sorts of things. We realign it to classify only, or to extract information only, while giving you high quality.

Then I ask, “Would you care?” They might say, “Okay, in this case it would be fine.” There are a few organizations that say no, it’s bank policy or company policy that they can’t use them. But the reality is that between Qwen, which is a strong model family, and Gemma, which is very good too, the moment you start specializing those models to your task, even if one is a little better or worse than the other, you recoup this. It usually doesn’t make a huge difference.

Ben Lorica. What about agents? What’s your story around agents?

Manos Koukoumidis. Agents are definitely very useful.

Ben Lorica. Everyone you talk to is probably crazy about agents, right?

Manos Koukoumidis. The word “agents” — I don’t know how many times I’ve heard it. So many times that I often try to avoid using the word myself, because I hear it so much.

Ben Lorica. You’re getting allergic to it now.

Manos Koukoumidis. Yeah.

Ben Lorica. You described a process that allows me to customize a model. But most people probably don’t think in terms of models anymore.

Manos Koukoumidis. Yes.

Ben Lorica. They’re thinking about an agent doing a sequence of things. So you’re not really specializing a model; you’re specializing a model for an agent, as you described.

Manos Koukoumidis. Exactly. That model could be doing all sorts of things. It could be classifying information or extracting information, and then using that information to call the right tool. A company we’re working with right now has a scenario where they have specific documents and want to extract information so they can make a query into their database based on the information just extracted. A similar case happened with one of the RPA scenarios with another company.

All of those qualify as agents. They reason, they call tools, they do all these things.

Ben Lorica. Reasoning models — is that table stakes at this point? The starting model can reason.

Manos Koukoumidis. Yes.

Ben Lorica. No one is using a non-reasoning model anymore?

Manos Koukoumidis. Reasoning is very useful. There could be some cases, for example, if you’re building a classification or extraction model, where you may ask, “If I make my model spend more time thinking, and when I train it I align it so it can spend all these deeper thinking and reasoning steps, would it help?” Usually we experiment to see whether it helps or not. But for more complex scenarios, it typically helps.

Ben Lorica. Agents obviously come with a bunch of buzzwords: protocols, MCP, harness engineering, and now the latest one people are throwing around is sandbox. Do you encounter these terms when you engage with companies?

Manos Koukoumidis. Definitely. These are components of an agent. You need components for agents to be able to call tools that may be exposed on MCP. Sometimes you need to execute whatever the agent wants to do — if it’s code or things like that — in a sandbox for security and other reasons.

All these things are very common. If anything, it makes the work that we do, and that other companies do, harder, because now it’s not just a model in isolation. It’s becoming a more complex system, and you need to be able to train it and evaluate it as a more complex system.

Ben Lorica. Do you have a solution or offering around things like memory?

Manos Koukoumidis. We don’t do anything for memory right now.

Ben Lorica. Okay.

Manos Koukoumidis. We don’t do anything when it comes to memory and orchestration. Memory is something we may address in some ways in the future, but not right now.

Ben Lorica. As you describe it, you’re training and customizing a model for an agent, and then I’m still responsible for the harness.

Manos Koukoumidis. Exactly. You can take it and put it inside any other harness framework you like, which may have memory or may call other agents.

The biggest gap we saw was that people took these generic things and created agents out of them, sometimes even agents with multiple steps. Then they asked, “Why is this not reliable?” It’s because they never put in the effort to make the model reliable — to teach it how to do the specific thing.

That was the biggest gap we were seeing. It was clear to us that unless we solve this holistically, organizations would say, “Only if you solve all the steps can I use this. If you only solve the training but not the evaluation or data synthesis, it’s still too hard for my team.”

Ben Lorica. When you’re customizing the agent, you’re actually customizing the agent as it exists within the harness and as it’s doing the work.

Manos Koukoumidis. Yes. We are able to train the agent in terms of how we test it, evaluate it, and train it in its ability to call specific tools and perform actions. That’s why we built all this infrastructure. We can do not just a classification or extraction model, but even agents that call tools, and we can synthesize the right data, evaluate them, and train them.

Ben Lorica. What if I don’t have an agent and I go to you and say, “Manos, this is what I want to do,” and I basically describe an agent to you? Can you help me? Or do I have to have an agent already, and then you help me customize the model for the agent?

Manos Koukoumidis. It could be either. There are cases where we work with customers that have a use case that has not been put in production yet. They come to us and say, “We need help. We want to get started on this journey of building our own specialized intelligence. We have some people, but we’re worried we don’t have a scalable way to do this. We like what you’ve built because it enables us to build that intelligence, as opposed to you coming in as an FDE and delivering it to us. We think this is an important capability and we don’t want to outsource it. We like that you can teach us how to use the platform and tools to build it ourselves.”

Sometimes we may lean in and help organizations a bit more. Others are able to get going completely on their own. Typically, whether we help them a little in the beginning or not, after the first or second model, they can go on their own.

Ben Lorica. Are your helpers called forward-deployed engineers or engineers?

Manos Koukoumidis. One of the two, yeah. By the way, we’re talking about agents and harnesses. Pretty much what we did with Oumi is build an AI development harness and agent. If somebody looks at the platform, it’s a whole platform that gives you access to training, evaluation, and data curation tools. We have built our own agent that can access those tools and, based on what you tell it, build a model that does X, evaluate the model on Y, and automate all the steps by calling the right tools.

Another way to think of Oumi is as the first AI development harness and agent.

Ben Lorica. You help me with customization. Are you going to help me with deployment?

Manos Koukoumidis. Yes. The short answer is yes. When we launched our open source, we had many organization leaders say, “This is still hard for us.” I could see what they were saying. That’s why we built this automation that makes it fully automated to build the model.

Quite often, someone says, “The platform looks like magic. It’s amazing. It would be great if you could also help me deploy the model.” That’s what we launched a couple of weeks ago: one-click deployment.

The next thing we’re working on, which is coming soon, is how you can also monitor and improve the model in production — not just with a few manual steps, but fully automatically. Somebody can do it now in minutes, but it takes a few manual steps. Soon, we’ll be able to do it fully automatically: build, deploy, and compound, all fully automatically.

Ben Lorica. In terms of optimization, obviously you’re optimizing for accuracy and reliability. What about latency and cost?

Manos Koukoumidis. Both. Quite often, for most enterprise tasks, you don’t even need a very big model. Maybe a 32-billion-parameter model, or even a 4-billion-parameter model.

Ben Lorica. I just had breakfast with someone in healthcare. They’re helping people build healthcare AI, so they have their own specialized models. He said some of the models are for things that are so specific they are in the hundreds of millions of parameters.

Manos Koukoumidis. Yes. With one of the companies we worked with — not to share more details, and I’ll change the scenario slightly — they needed to classify different line items in an invoice correctly, based on the type of line item. We built something like a 200- or 400-million-parameter model that they could put on device, and it was as good as an almost 4-billion-parameter model or a much bigger model, because the task was so simple. They could even put it on device and completely eliminate cloud costs.

That’s what I mentioned before. What’s happening now in the industry, with most enterprises using these big models for such simple tasks, is wildly irrational. They are using models that are 10 times, 100 times, and sometimes even 1,000 times bigger than they need to be.

Ben Lorica. I honestly don’t know why we still don’t know the actual cost of inference.

Manos Koukoumidis. Yeah.

Ben Lorica. I think if OpenAI and Anthropic actually charged us the cost of inference, we would use these models much less. The $20 or $100 a month subscription is completely subsidized. I create a few images in OpenAI and maybe I’ve blown through my $20 monthly budget already.

There’s no transparency in terms of the true cost of inference. If you start hosting these open weights models on your own, you might have a better sense of how much this is costing you, because you’re paying for the compute, energy, and all of that yourself. That’s another way to learn more about what’s truly happening.

Manos Koukoumidis. Very likely your bill would immediately go down by at least 2x, and most likely more.

Ben Lorica. You’re basically using a model that can speak 100 languages and was trained on Reddit and Wikipedia for things it knows that it doesn’t need to know.

Manos Koukoumidis. Yes.

Ben Lorica. Question about multimodality. Obviously all of the models are multimodal. Are you sensing that companies actually need multimodal?

Manos Koukoumidis. Multimodal is important. When people think of multimodality, the first thing that may come to mind is images, then speech or video. But the vast majority of enterprise use cases right now are text. Images often come next as a common scenario, and of course speech in some use cases. It largely depends on the scenario, but text is still the dominant use case.

Ben Lorica. And text includes code.

Manos Koukoumidis. Yes, text includes code in my mind.

Ben Lorica. You alluded to the cost savings being huge. If I’m using Claude or OpenAI, I get the bill, so I know, “Oh my God, this is expensive.” If I move to you, are you charging by tokens? Is that how they find out how much they’re saving?

Manos Koukoumidis. If you are using your own fully fine-tuned model, it has to be placed on your own GPU, and then you get charged per GPU hour, or more precisely, per GPU second that the GPU is being used.

To make it easier for people, we enable autoscaling. They can say, “I may not use the model for a couple hours. I’m still experimenting.” The model becomes cold on its own so you don’t get charged. Then as you start using it, it warms up. If you use it more in production, it autoscales. You get charged per hour based on usage, and you can decide whether you want to shut it down, allocate more GPUs, allocate fewer GPUs, or let it autoscale.

Ben Lorica. Are there examples on your platform of custom models that you can even run on CPU?

Manos Koukoumidis. Yes. Those hundred-million-parameter models were running on device, on iPhone devices. These small models can definitely run locally. Most people now have laptops or workstations that often have some GPU, and there are a lot of models you can run. Especially if you care about a specific scenario, you can easily develop a model that you can run locally.

Ben Lorica. Do you think there will be a breakthrough where, because right now scale is still the holy grail — if you listen to OpenAI and Anthropic, they think this is what’s going to take them to AGI — maybe we don’t need such big models? What you’re alluding to is already happening in enterprise workflows. But as far as the people heading toward AGI, can there be a surprise looming where someone is building something that says, “Hey, I can be as good as Opus, and I’m actually not that big”?

Manos Koukoumidis. I’m going to focus on a word you used, which is “surprise.” It could happen. But if I’m the leader of my organization, a strategy that relies on surprises is not a strategy. At least I would be worried if the future of my enterprise depended on that.

Enterprises need to get on the right AI strategy. A surprise could happen, and it would be great if AI improves in different ways. But leaders need to plan and make sure they future-proof their businesses now.

Ben Lorica. For you and your team, with custom models, there are different model providers and different model families. They have different release cadences. There are model updates happening every two weeks. Even in open weights models, do you keep in touch with all the open weights releases, kick the tires right away, and then add them to the mix of your offerings?

Manos Koukoumidis. Exactly. We integrate them very fast, because people often want to use and try the latest. The good thing, going back to building and deploying your own AI, is that we make it very easy to automatically switch from the model you have in production to one that could be better, using the same consistent evaluation. The same process you used to decide to put the original model in production makes the transition seamless.

It’s just: look at the results as a human, trust them, click the button, and do the switch. In many ways, I would argue that if you rely on the system, it is lower maintenance than consuming a model from OpenAI. One of the banks we work with is doing some risk analysis on communications using GPT Mini, and OpenAI is deprecating it, I think in October. Now that’s a maintenance problem for them. They need to go back, retune everything, and retest it. They don’t have a platform that makes it easy for them to do that. Plus, they are asked to switch to another model OpenAI has released, which is twice as expensive. They are being forced to do all this maintenance and transition, and pay twice as much.

What I’m describing is something better, cheaper, and automatically maintainable.

Ben Lorica. Can you explain to our listeners, in plain and accessible language, how this compounding happens?

Manos Koukoumidis. The best way to put it, without going into nomenclature like RL, is that when you operate the model in production, it interacts with different customers for different use cases and examples. Sometimes the model gets it right, and when it does, you can say, “Good job. Do more of that.” Sometimes it gets it wrong, and then you say, “This didn’t work out very well. Don’t do that exact thing next time. Try something a little different.”

With every interaction, every turn in which your customers use your model in production, it becomes better and smarter for you — not for somebody else, but for you, because you own the model.

Ben Lorica. Does Oumi offer OpenCL?

Manos Koukoumidis. We don’t offer OpenCL.

Ben Lorica. Is this something you hear people asking for? They might say, “You’re already taking care of the model classification. You might as well offer this thing.”

Manos Koukoumidis. I haven’t gotten a request for OpenCL yet. Maybe the time will come, but not yet.

Ben Lorica. What trends in AI are you paying attention to?

Manos Koukoumidis. I try to pay attention to everything, from what’s happening at the frontier labs to open weights and everything else. But the biggest change, trend, and shift I’m seeing — which I think is good news, not just for us, because that’s the strategy we built Oumi on, but because I think it’s a better future for enterprises and humanity overall — is that enterprises are increasingly moving from generic intelligence that they rent toward specialized intelligence that they own.

There is news now: Cursor, then Intercom with Fin, and yesterday Harvey. As more examples come out, more companies will say, “Hey, why are these companies doing this? Should we do it as well?” I can see that growing. Some people have realized and noticed this trend; others maybe not as much.

To give you an idea, we’re getting outreach from analysts and even investors saying, “We’re seeing the market move in your direction.” We can see that as well. That’s good news not just for us, but for the whole industry.

Ben Lorica. This is why OpenAI and Anthropic should go public as quickly as possible, because there is going to be pricing pressure. People are realizing that if we depend on these two providers, it’s going to be expensive, especially as we roll out more and more AI across the organization.

Manos Koukoumidis. As they go public, there is going to be more scrutiny of their economics and how everything works.

Ben Lorica. The pricing pressure is going to get worse for them.

Manos Koukoumidis. I agree with you. They’re going to have to do something, because what is happening now is not sustainable. If it’s not sustainable, something will have to change, or some company is going to have to break and pop.

Ben Lorica. On a scale of one to 10, one being less worried and 10 being more worried, how worried are you about the AI bubble?

Manos Koukoumidis. I would say I don’t think AI is a bubble, but I do think some specific companies may become bubbles. Overall, AI is an extremely powerful technology. People who put it to use rationally and build their business on top of it rationally will be very successful.

Ben Lorica. I agree. But the financial markets drive a lot of what’s happening. I’m increasingly worried, mainly because of the data center buildout, how expensive it is, and how debt-laden the announcements are. Oracle, for example, seems to have bet the whole company on AI data centers. To some extent, the SpaceX IPO is a data center, but in space.

Manos Koukoumidis. Let me tell you why I’m bullish both about data centers and AI overall. I mentioned before that enterprises are finally rationalizing their strategy, moving from generic intelligence to specialized intelligence, where they can get more value for much lower cost.

I believe in Jevons paradox, which says that if you make something cheaper and get more value at a lower price, you don’t end up with less usage. In the end, you get more usage, more revenue, and more of everything. Because we will be able to harvest more value out of AI at a cheaper price, more and more businesses will want to use AI for more problems, and we will need more data centers. We’re going to get a lot more value. The way we’re doing it now has an opportunity to become 10 to 100 times more efficient.

Ben Lorica. To me, the litmus test is: if this AI data center business is such a great business, look at CoreWeave and Nebius. It goes back to what we talked about earlier, which is that if you actually charge me what it costs to do this inference, I’m probably going to pay more, in which case I’m probably going to do less. Even inference is subsidized.

Manos Koukoumidis. Yes.

Ben Lorica. The only winner seems to be Nvidia.

Manos Koukoumidis. Definitely the picks and shovels of the AI industry.

Ben Lorica. The AI data center business is basically taking Nvidia and reselling it.

Manos Koukoumidis. There’s a lot of discussion here about the circular deals and funding that Nvidia may or may not have made.

Ben Lorica. You’re incurring all this debt, and the economics are unclear. The shelf life of these data centers is also uncertain. GPUs get outdated in a matter of years.

Manos Koukoumidis. The thing that makes me more concerned is less data centers and more the business model of some closed-model providers: how capital-intensive it is, how their margins are negative, and how they are negative because it was hard for people to use a solution that was 10 to 100 times better and more efficient.

The moment people realize there are companies like Oumi — and right now we are the only ones that do this problem in this way — and that there is an easier, better way to do this, I think a lot of enterprise usage is going to follow that path. We’re already seeing that.

The more enterprises realize how easy it is to transition to the better path, the fewer people will use frontier models for enterprise AI. For consumer AI, generic models make sense because you need a model that can do everything. But for enterprise AI, I see companies moving away from that.

I’m less worried about the data centers because I think we’ll find ways to put them to use for specialized AI and more efficient use of AI, in line with Jevons paradox. I’m more worried about the closed-model providers.

Ben Lorica. Let’s close with this. You just said something that is contradicted by your blog posts, which is that for consumer AI, generic models make sense — but maybe not all the time. Didn’t you do a study that checked Google’s AI summaries and found they’re not as good as we think?

Manos Koukoumidis. We did the case analysis with The New York Times, and we found that even with the latest Gemini 3 at the time, AI Overviews were fully trustworthy — meaning both the answer and all the claims it makes are correct — only in roughly one out of three cases. I think it was 39%.

For people deep into AI, they may rationalize it and say, “This is a hard problem. I understand why it’s not better.” But the problem is that most people understand the Google brand, and they don’t stop and think, “This is generative AI, and it is not as good.” They don’t realize that this is not coming from Google’s highly curated and accurate knowledge graph. They don’t understand what they’re consuming and the risks that it’s not as trustworthy as they thought it was. I think that’s the main problem.

Ben Lorica. By the way, the citations are bad too. When you click on “I got this from here,” it’s wrong. It’s not that.

Manos Koukoumidis. Yeah. I don’t remember all the details, but what we found was that when you get an answer and it has a couple of citations, only in about one out of three cases were all the citations correct. In most cases, at least one of the citations was not accurate.

The problem is that you have something that says, “Here’s the answer, here’s the proof.” Who is going to check? Especially when somebody says, “Here’s the proof,” you may be biased to trust them. For a user to follow the citation and double-check whether the claim is grounded in the citation is too hard. Nobody does that.

Ben Lorica. Last question in closing: What can I do with the open source Oumi, and what does the commercial offering provide that is not in the open source?

Manos Koukoumidis. The open source is a library that enables researchers and deep AI scientists to do the core operations they would like to do for research: anything from pre-training to post-training, evaluation, running evaluations with full control, and data synthesis.

Ben Lorica. So it’s meant for researchers.

Manos Koukoumidis. Exactly. Mostly researchers and very deep AI scientists who want to spend weeks and months doing frontier research or advancing something to get the last percentile improvements.

But as I mentioned before, what we hear from most enterprises is, “This is still hard for us. It will still take us a lot of time to build our own AI models.” That’s why we built the enterprise platform. The enterprise platform runs on top of our open source. We built the open source to enable anyone, including Oumi, to build their own products, and that’s what we did. We built it to enable ourselves.

When we improve the enterprise platform, roughly 50% to 60% of the improvements still trickle down to the open source platform, which we make better every day.

The difference is three things. One is that it does all the resource management: here are your models, datasets, evaluations, and so on. Two, it gives you compute if you don’t have your own. And for me, the most important one is the AI development harness and agent I mentioned before. It automates and makes the entire process of an engineer so easy. What takes weeks or months using the open source is reduced to literally minutes.

That’s the main benefit. It democratizes the development of specialized AI for anyone. I tell people it’s like Cursor or Claude Code: it makes the deepest experts dramatically more efficient, but it also democratizes it for people who were not able to do it before.

Ben Lorica. We last spoke in May 2025, and there has been tremendous progress. Thanks for coming back on the podcast.

Manos Koukoumidis. Absolutely, Ben. Thank you very much for having me. Always wonderful to chat.