Denise Teng on Enterprise AI Adoption, RL Bottlenecks, Agent Infrastructure, and the AI Bubble.
Subscribe: Apple • Spotify • Overcast • Pocket Casts • YouTube • AntennaPod • Podcast Addict • Amazon • RSS.
Ben Lorica speaks with Denise Teng, partner at Gradient Ventures, about the real state of enterprise AI adoption, why security and data privacy still slow deployment, and why open weights may become increasingly important for regulated industries. They also discuss reinforcement learning, agent infrastructure, AI memory, inference costs, vertical AI startups, multimodal infrastructure, and whether the current AI boom is showing signs of a bubble.
Interview highlights – key sections from the video version:
-
-
- Enterprise AI Adoption Reality Check
- Why Open Source and Open Weights Matter to Enterprises
- Open-Weights Economics, Supplier Risk, and Price Pressure
- Custom Models, Post-Training, and the Future of Smaller Models
- Why Reinforcement Learning Still Feels Just Out of Reach
- The New Infrastructure Stack for AI Agents
- Agent Memory, Context Layers, and Institutional Knowledge
- Vertical AI and Domain-Specific Agents
- Inference Infrastructure and the Hidden Cost of AI
- Token Spend, ROI Tracking, and the Chief Token Officer
- AI Bubble Risks, Security Threats, and Investor Anxiety
- AI Data Centers, Fragile Economics, and Tougher Fundraising
- Multimodal Models, Robotics, and Emerging AI Infrastructure
- Enterprise AI Safety, Compliance, and Monitoring
-
Related content:
- A video version of this conversation is available on our YouTube channel.
- 12 GW announced. 5 GW under construction. What happens next?
- Here’s my uncomfortable bet on OpenAI and Anthropic
- Does AI Actually Make Developers More Productive? The Evidence, For and Against
- Helen Gu → The Hidden Failure Modes of AI Agents
- Zhou Yu → Why Your AI Agent Isn’t Ready to Ship (And How to Know When It Is)
- Hamza Tahir → AI Agents Are Implemented, Not Adopted
Support our work by subscribing to our newsletter📩
Transcript
Below is a polished and edited transcript.
Ben Lorica. All right. Today we have Denise Teng, partner at Gradient, which you can find at gradient.com. Its tagline is “the seed fund designed for founders in AI.” Prior to Gradient, Denise spent six years as a product manager on the Meta AI platform team, where she built infrastructure tools for machine learning workflows. With that, Denise, welcome to the podcast.
Denise Teng. Thank you so much, Ben. I’m really, really excited to be here. As you mentioned, I currently work at Gradient as a partner and focus on early-stage AI investments. As a firm, we have focused on B2B for 10 years at this point, and we support a lot of founders across dev tools, infrastructure, and agents for horizontal use cases and vertical industries. I’m happy to share more background and context about the firm and myself, but overall, I’m very excited to be here today.
Ben Lorica. All right. Let’s start with the two topics I’m super interested in, which are enterprise and AI. I just had breakfast with someone from one of the largest banks in the U.S., who is pretty high up at this bank and is from New York. We were having breakfast in San Francisco, and I said, “Give me a reality check about AI adoption in your bank and in banks in general.”
He was telling me that a lot of these banks tend to move much more slowly than we think. They are very conservative in terms of how they use these products. Sometimes I feel, Denise, that we are in this Bay Area bubble where we think things are happening so fast and people are using OpenAI or Claude at work. But the reality is there are a lot of basics that enterprises have to overcome, including having their data ready for AI and agents.
Based on your conversations with enterprise leaders, forget people in tech, but people in non-tech companies: what is the real level of adoption?
Denise Teng. I would say the real level of adoption for non-tech enterprises is still relatively early, for all the reasons you just mentioned. I think what these companies care about most is using AI to increase productivity. That is the return. But the really important thing they care about is that their institutional knowledge and data have to stay within the enterprise itself. Banks, insurance, healthcare, you name it. A lot of the time, privacy is top of mind, and internal knowledge is the moat for these companies.
If you think about enabling AI, you want to make sure you apply the proper guardrails from a security, data privacy, and monitoring perspective before those enterprises feel comfortable using AI. And if you talk about security, privacy, and compliance-related topics, it can become very complex. It is not just a policy being written and then somehow you can monitor all the AI. It is about deployment: making sure data does not go to closed-source model companies, making sure the data stays, perhaps on-prem or in a private cloud. Those are not easy implementations in any way.
So again, I agree with you. Security, privacy, and data are the core reasons why enterprise adoption is relatively slow. But given the advancements in the tooling that enables enterprises to do private and secure deployments, perhaps adoption rates will increase very quickly later this year and next year, especially as the open-source ecosystem gets stronger over time. I can definitely share more thoughts there, but in general, adoption is relatively slow. Even with open source and deployment-support tooling in the market, I do see the timeline accelerating later this year and next.
Ben Lorica. How important are open source and open weights to these companies?
Denise Teng. When I think about open source and open weights, I consider them another important way to let everyone access AI without letting the data go into closed-source model companies. Historically, let’s say two or three years ago, people debated open source versus closed source much more. Two or three years ago, open-source models’ performance might have been 10 to 12 months behind the frontier closed-source model companies. But recently, I would say this year, the gap has shrunk a lot. This year, the gap is more like six months, or even three months.
Three months after the closed-source model companies announce their latest frontier model, very quickly open-source models also have something on par available in the market. Meanwhile, if you are able to find the right inference provider to host those open-source models in a very efficient and easy way, or if the enterprise is able to buy its own GPU accelerators and spin up the whole inference-serving stack, they can access pretty good AI model performance through the open-source route while keeping everything in a private cloud or on-prem. Open weights play a very, very important role in enabling enterprises to access AI in a more secure and private way.
Ben Lorica. I confess I’m starting to use open-weights models more, partly because the leading-edge models are getting expensive. You do one thing and then look at it and go, “What? That cost me $10?” So there is both a control and a price incentive.
My worry with open weights is that the list of suppliers is unpredictable, because at a basic level, open weights are a business decision, and as you know, business decisions can be reversed. Look at Meta, which backed away from open weights. Some of the Chinese providers are signaling, “Maybe we’re not going to do as many open weights moving forward.”
On the other hand, you have Google Gemma, which clearly is not as good as Gemini, but is good enough for a lot of routine work. I think there may still be price pressure on proprietary model providers because companies can offload a lot of routine work to smaller open-weights models. At least I hope Google, and maybe a few others, will remain committed to releasing them on a regular cadence.
Denise Teng. Indeed. And I also want to highlight that beyond relying on open-source model providers to provide constantly updated open-weights models, it is also critical — and we see this all the time — for the ecosystem and enterprises themselves to produce more custom models through open-source models. In that case, they can bridge the performance gap even further.
That is why there is another wave, another category of tooling and platforms in the market, to enable post-training and reinforcement learning. Even though an open-weights model may not be able to tackle a very sophisticated and complex task directly, if enterprises are able to find the right platform or the right team that provides RL as a service, for example, they can post-train and use reinforcement learning to customize their models through the open model. They can produce many smaller custom models after running that post-training process, and in that way they can bridge the performance gap even further to support complex tasks.
We also believe that in the future, the direction is more and more smaller models. Small-model proliferation will happen, and hopefully it will cover as many complex small tasks as possible, instead of always calling the biggest and more expensive model, whether from an open-source or closed-source model provider. I think reinforcement learning as its own era or its own category will bridge the gap you just mentioned.
Ben Lorica. I’m a big reinforcement learning fan. In fact, every six months I write the same article, which is “reinforcement learning is finally here; the tools are finally easier to use.” But I always feel like it is just around the corner. It seems like this reinforcement fine-tuning is extra motivation to make it even more accessible, but I still sense that in reality, it is still a bit advanced. Am I wrong?
Denise Teng. I think you are right in the sense that people feel RL is more ready because RL libraries and the recipes for how to train RL models are getting more mature. However, if you look at the end-to-end RL pipeline, the real bottleneck is actually the data. You still want the RL pipeline to train or post-train a model against some pattern, and those data patterns are often not even available for enterprises, for a few reasons.
One reason is that enterprises do not really have cohesive data patterns that agents or reinforcement learning pipelines can easily learn from. Think about enterprise data. Sometimes it sits in internal dashboards or internal data warehouses. Sometimes the data comes from the CRM or some other SaaS tooling that they purchased. But if you think about enterprise use cases, a lot of the time employees have to access many different tools in order to complete one workflow. All the tools and all the actions you are taking sit independently in different entities.
There is no universal way to generate agent-learnable or post-training-learnable traces. The data is the real bottleneck. That is why we always feel like RL is there, but it is actually not there in production.
Ben Lorica. Is there also a UX issue? In supervised fine-tuning, the UX is simple. I am going to create some labeled examples of prompt and desired output, prompt and desired output. I do not know, maybe a few thousand. Then there are so many services where I can just upload this thing, go away for lunch, come back, and it is fine-tuned. With reinforcement fine-tuning, it seems like there is still a UX issue. Do you think so?
Denise Teng. Not necessarily a UX issue. I think again—
Ben Lorica. How do you get the domain expert comfortable doing it? Supervised fine-tuning, as I described it, a non-programmer can do now, as long as they can come up with the labeled examples. Reinforcement fine-tuning: do you think it is at the point where a non-programmer can do it?
Denise Teng. It is still early. Again, the data processing and collection piece involves a lot of heavy lifting. A lot of the time, you actually need to generate additional synthetic data in order to bridge the data gap. In that case, it is not a UX layer where another tool can easily help non-technical people bridge that gap.
Ben Lorica. But why not? You could say, “Generate some synthetic data based on this description.”
Denise Teng. It depends on the use case. A lot of enterprise productivity use cases are about asking your agents to do the job for you. Those synthetic data sets or synthetic environments are essentially another software program. You want to synthetically generate another internal dashboard. You want to synthetically generate another terminal environment and also synthetically generate another potential MCP or API call.
Those are not necessarily prompt-and-generation pairs. They become different program environments or software environments as synthetic data. That is why it is not easily reproducible by a human writing another spreadsheet of mock data or creating golden sets for prompt-and-generation pairs.
Ben Lorica. Speaking of agents and platforms, since you came from ML platforms at Meta, people are also starting to realize, “Hey, if 20% of our workforce is agents, that means we may need to revisit our infrastructure for agents.” Identity management, access control, even databases. People are building databases that are agent-friendly. What else have you seen out there? Are people starting to think that in a world of agents, we should rethink our infrastructure?
Denise Teng. Yes. One of the more visible areas of traction in the market for agent infrastructure is this idea of the sandbox. Traditionally, we would have talked about virtual machines, but nowadays the new term is sandbox.
Ben Lorica. Everyone uses this term now, right?
Denise Teng. Everyone uses this term now. Basically, the longer agents can run on behalf of you and me, usually the better the results will be, because that means the agent can achieve a very complex and long-horizon task. But at the same time, you do not want your agents to sit on your laptop and use your laptop environment to do things, because as humans, we need to walk away. We need to stop the session.
That is why a lot of agents need to run in the background. Because the background needs to sit in the cloud in some way and also have proper isolation, so individual agents do not interfere with each other’s tasks, sandbox runtime environments have become a new primitive for running agents. That is one runtime-layer infrastructure area that became important after agents became more popular.
And of course, the things you mentioned: agent identity. How does the security team monitor what agents can access or not? Does the agent directly derive access from credentials?
Ben Lorica. What credentials should the agent have? If these 20 agents are Denise’s agents, are they using Denise’s credentials?
Denise Teng. Exactly. People can debate that, and I am pretty sure people will have different thoughts. A lot of people think agents should derive maybe partial access from the human, but maybe the agent should be given additional sub-access that is only 10% of what Denise as a human can access. Many different use cases can happen. I think agent credentialing and identity in general is a really massive and unsolved problem from a security tooling perspective. It is also very promising, because it is obviously another blocker for enterprises adopting AI agents.
Ben Lorica. There are also things that humans take for granted, like institutional knowledge and teams. What is the equivalent for agents? One of the complaints about MCP is that every time I use this thing, it is going to burn so many tokens because it is going to load all of the tools and go through the tools. But the reality is that when humans assign a project to another human, they also impart institutional knowledge: “Hey, Denise, build this pipeline, and by the way, you may want to look at Ben’s pipeline because Ben built a similar pipeline a couple of months ago.” Why do I have to burn all these extra tokens to start from first principles when there was prior knowledge? I think people are beginning to build solutions to reduce these inefficiencies, right?
Denise Teng. Yes. We did not talk about the data layer for agent infrastructure, which I think is really relevant to what you just mentioned. That is why a lot of people are talking about another category under AI agent infrastructure: AI memory, agent memory, or the agent context layer. No matter how you market the term, they are trying to solve exactly the problem you described.
You want agents to have not just individual memory in terms of what you historically did, what prompts you issued, and what skills and files are relevant to Denise. They also need to know, in an organizational setting, how Denise’s AI memory and Ben’s AI memory collaborate with each other. This can be a data problem as well, because there is a large amount of knowledge, and you need to associate relationships between that knowledge.
Does that become something similar to semantic search? Maybe. Does it also connect to a graph, like an organizational knowledge graph? I think it also has that graph flavor in the data solution. Many AI memory companies tackle this problem by providing additional memory connectivity, building a graph around it, making sure it is semantically connected, and making sure it is searchable. Again, the moment an agent works on a task, you want it to fetch the proper memory on time and in a low-latency way.
You can consider this a new data layer that is natively built for agents. That is another pretty big category, and it is still an unsolved problem in the market.
Ben Lorica. The other frustration for me, Denise, being in the Bay Area, is that a lot of people here are engineers, and engineers build things for other engineers. The problem is their engineer friends will never pay for anything. I keep telling people that they might want to go deep into some domain and build in a specific domain.
For example, coding agents are super popular, built by engineers for engineers, but that is a very domain-specific example. Surely there are many other examples. One example I keep pointing people to is accounting. In accounting, you basically have the tax code and GAAP, generally accepted accounting principles. Maybe that is enough to build a really productive agent, because a lot of the knowledge is in a few concrete sources.
What is your sense? Do you also feel like people are ignoring these domain-specific solutions? Everyone is building horizontal here in the Bay Area.
Denise Teng. Other than coding, there are still quite a few vertical AI companies.
Ben Lorica. Which ones have impressed you?
Denise Teng. In terms of the domains where we see traction in the market, I think legal is one. Many larger startups have already appeared in the market.
Ben Lorica. Lawyers are not happy with the solutions, but I like that people are trying and building.
Denise Teng. People are definitely trying and building. Legal and healthcare are common. In New York, quite a few vertical AI agent companies also focus on finance-related use cases. In legal, Harvey and Legora are two big names. Accounting has Accordance. Healthcare has Abridge. In finance, there are also a few here in New York, like Hebbia and Rogo. There are quite a few vertical AI agent companies that are at a later stage. By the way, none of these are our portfolio companies.
Ben Lorica. What approach are many of these companies taking? Are they going very narrow, into very specific use cases or tasks? How do they go about building for a specific domain?
Denise Teng. These companies have all been around for at least two years, so I am pretty sure that at the very beginning they were more high-level: something similar to a ChatGPT experience on top of internal domain-specific knowledge. But nowadays, most likely, they have all moved into agent workflows as well, to automate a lot of the deep research work, maybe for a legal case or deposition.
In healthcare, it may be more related to getting the scribe and the proper insurance data or patient history data. A lot of them have turned into agent-first workflows to automate complex analysis and help the human see the final analysis outcome. That is my understanding. I personally have not tried any of these products, so I cannot share too much on the detailed use cases here.
Ben Lorica. What are other pieces of infrastructure that have caught your eye?
Denise Teng. Another piece of infrastructure that is interesting, although it sounds like a cliché, is inference. It is still really top of mind if you think about the amount of adoption.
Ben Lorica. Super important, right?
Denise Teng. Super important. Demand is skyrocketing, but supply is very constrained by hardware, data centers, and the number of chips. How do you really use the existing inference stack to optimize as much as possible? I think that still has a massive market opportunity, given that demand is very, very high — higher than all historical software adoption. Inference, all the way down to hosting, serving the model, optimizing the inference engine and kernel, and all the way up to routing to the model of choice: I think these companies or use cases are still very promising in the market.
Ben Lorica. By the way, with inference, I feel like we do not really have any clarity on the true cost of inference, because inference is heavily subsidized. If they were to actually charge us the cost of inference, I think we would be shocked. The $20 a month or $100 a month that we are paying to subscribe to one of these chat providers — we are blowing past that within the first week of the month. It just seems like there is no clarity in the market in terms of how much this is costing.
Already people are pulling back on the more expensive models, and those are probably still somewhat subsidized. If we were to actually pay the actual cost, I think we would be much more judicious in how we use these things.
Denise Teng. People would be more careful, right? You are not going to token-max every subscription you have. That makes sense. But again, optimization-wise, inference still has a lot of room to improve. Hopefully the price will be more stable, and hopefully the inference providers that can optimize more can still increase their margins.
Ben Lorica. My joke is that the CFO is now the CTO: the Chief Token Officer.
Denise Teng. Token Officer. Too much token-maxing. CFOs are having a hard time recently.
Ben Lorica. That is going to be an important part of companies. Basically, I need a dashboard. I need fine-grained attribution for, “Hey, this team is blowing all these tokens, but where is the ROI? On the other hand, this team has clear ROI. Let’s move all of the token spend over to this team.”
We need solutions where people can track all of these things somehow. I do not know how sophisticated dashboards are for tracking token consumption within large companies. I do not know if you have heard of people trying to build something like this.
Denise Teng. We definitely see quite a few very early-stage companies trying to provide that visibility for leadership teams so they can justify the ROI even more.
Ben Lorica. Within a large company, people are using many different providers.
Denise Teng. Many different providers, yes.
Ben Lorica. You cannot just rely on, “Hey, this is our OpenAI bill.” Plus, even if you have that, you want fine-grained attribution down to the person level.
Denise Teng. Down to the person level. Also, how do you define the return? You can see the usage, but how do you define the return? Everyone’s personal OKR is usually a little more arbitrary versus a clear revenue number. In that case, how do you really complete this ROI metric? That itself is quite challenging too.
It is not just usage, but also what the value really is. That is very hard to define. Quite a few startups and vendors want to tackle this problem. But I also want to call out that fundamentally, the way agents are consuming tokens might not be optimal. That is why the cost is really high.
I am hoping the optimization happens first, so token-maxing will not be such a huge problem for CFOs, because perhaps the bill will be more reasonable than before. But of course, visibility is always helpful to let the leadership team understand overall adoption, usage, and ROI.
Ben Lorica. We are clearly in an AI bubble. So two questions. First, on a scale of one to 10, one being least worried and 10 being most worried, how worried are you? And secondly, what are the key things that worry you the most? What are the key indicators that worry you the most?
Denise Teng. I would say I am more of a six. Slightly worried, but not too worried, because at the end of the day, I see AI as — again, it sounds like a cliché — a generational technology. It is changing our lives entirely. We are still early in language models. We have not even talked about other modalities, in terms of multimodal, not to mention physical and world models.
There are way more use cases AI can unlock for us professionally and personally. I am very bullish on AI in general. That is why, when it comes to the current bubble, I am not too worried. Eventually, once the technology and the supply side catch up, and once the model research breakthroughs continue to happen, usage will continue to grow. The market is unlimited for AI. That is the core reason why I am more of a six versus a nine.
What bothers me or worries me most as an investor is the number of companies in the market. Everyone is building AI, so that gives me a lot of anxiety in terms of, “There are too many companies I need to look into.”
The other thing I worry about is how we make sure AI is very helpful, but also behaves relatively safely and securely. This goes back to the initial discussion we had. We see many AI models doing so well, even in security use cases. More and more cyberattacks are happening in the industry. Cyberattacks can happen to companies, so they lose their data. They can also happen on a personal level.
What if you and I constantly get a phone call, and it is a voice AI telling you that your kids or family member has an issue? What if it is another phishing email that looks so personalized, and you click the link? I do not want those things to continue to happen. That is another reason why I think AI advancement is great and AI development is great, but applying a proper safety net and security net around it is extra critical nowadays.
Ben Lorica. I am more of an eight or nine.
Denise Teng. You are more of an eight or nine.
Ben Lorica. Yes. My main worry right now is the AI data center buildout, which is a bit of a contradiction in my mind because, on one hand, I want more of them. But the economics are starting to worry me: the amount of debt these companies are taking on. Look at Oracle, for example. Very worrisome.
Fundamentally, I am not even sure that the AI data center is a great business for many of these companies at the level of what they are paying for. If it were a great business, look at CoreWeave or Nebius. Those are pure AI data center clouds. Again, it goes back to what we talked about earlier: the true cost of inference is not being reflected.
It would be a good business if you could charge the actual cost of inference. But if you do, people will cut back on usage. I think that tension around the data center buildout might be the thing that tips us over. Plus, it is infrastructure that is outdated within three years. The GPUs move to another generation. It is not like fiber optics back in the dot-com era.
Going back to your point about there being too many companies, one thing that seems to be happening, Denise, which might be a good sign or a bad sign depending on your perspective, is that the people who raised seed rounds maybe a year ago are now trying to raise another round. It seems like they need to show a lot more progress than they would have two years ago.
Denise Teng. That is true.
Ben Lorica. Is that a good sign or a worrisome sign?
Denise Teng. From a market adoption and traction perspective, it is a good sign. It really means AI is so helpful, and therefore customers and buyers are excited to quickly buy usable AI products. It is a challenging sign—
Ben Lorica. But before, you could just go back and raise another round. It was easy. Now it seems like the metrics are a lot more demanding. Again, the question is: is the fact that there is more scrutiny with subsequent fundraising rounds a good sign, or is it a sign that the bubble is about to pop?
Denise Teng. It is more that there are many competitors working on similar things in the market, and they all have pretty good traction. So later-stage investors usually want to pick the winner. They want to pick the winner among 10 different competitors.
If the obvious winner is doing so well, they want the second potential winner to do something similar to the first winner. If the second winner already reached $8 million for their Series A, of course later-stage investors will want to see the second or third players also reach maybe $6 million or $5 million, something closer to the first winner’s performance. Basically, adoption is higher, but given that a lot of markets might be winner-take-all, the traction bar is more demanding than before for later-stage investors.
Ben Lorica. As we wind down, you brought up multimodal world models and physical foundation models for physical AI. Is there anything else in that much more R&D stage that people should pay attention to?
Denise Teng. These are all still on the model side, the intelligence side. But of course, how you bake that intelligence into another modality of AI, let’s say robotics, can be very interesting and important as well. That is why there is a massive amount of robotics markets, companies, and VCs in the ecosystem pushing the frontier of AI. That is another big category beyond language model agents.
Ben Lorica. As far as multimodal models, I think most of the frontier models are multimodal, and I think a lot of enterprise data is multimodal. But the existing data infrastructure tends to revolve around structured data warehouses, lakehouses, and things like that. It seems like there are infrastructure opportunities around multimodal, right?
Denise Teng. Yes. Databases for multimodal data formats, databases that can store not just files and PDFs, but also images or videos, and store them in a very cheap way, with an index on top of them to enable fast retrieval. That seems to me to be a very obvious gap in the category. I have not really seen a lot of interesting vendors in that category, but yes, you are right. Infrastructure for multimodal — serving voice faster, serving image or video faster — can be a very interesting infrastructure opportunity too.
Ben Lorica. You mentioned safety earlier, but it was in the context of foundation models. Based on your conversations with enterprise AI leaders, does safety and compliance come up? In other words, at this point in the cycle within enterprises, is the push, “Let’s get these agents built”? Particularly if they are outward-facing, that is one issue. But if they are internal-facing, is the attitude, “Let’s just build them fast and worry about legal, compliance, and safety later”? Is that what you are hearing, or are enterprises already very concerned about compliance, safety, and all these issues?
Denise Teng. There are different degrees of concern, of course. If it is a more high-stakes industry, they care about those things even during the POC phase. For a software tech company or enterprise, they might have had a head of AI, so everyone was trying to quickly get some POCs done and did not immediately think about compliance until something was really in production.
Nowadays, given that more adoption is already in production for many tech enterprises, policy and compliance have become more important than they were a year or two ago. Going back to the monitoring use case we talked about, of course enterprises will want to monitor their employees’ AI traffic and compare that not just against ROI, but also against company policy, and see whether there is any potential issue there. That is why monitoring is another important tooling category we keep hearing about this year for enterprises.
Ben Lorica. With that, thank you, Denise.
Denise Teng. Thank you so much, Ben. This was really fun. Thank you so much again for having me.

