What Comes After Transformers?

Zuzanna Stamirowska on Latent Reasoning, Continual Learning, Small Reasoning Models, and Long-Horizon Agents.


Subscribe: AppleSpotify OvercastPocket CastsYouTube •  AntennaPodPodcast AddictAmazon •  RSS.

Ben Lorica speaks with Zuzanna Stamirowska, CEO of Pathway, about the company’s push beyond transformers toward models that reason in latent space, learn from experience, and maintain state over long time horizons. They discuss Pathway’s 150-million-parameter BDH model, its ARC-AGI results, continual learning, small reasoning models, dramatically lower inference costs, and why long-horizon agents may require a fundamentally different AI architecture. They also explore enterprise deployment, AI infrastructure, open weights, and whether future reasoning systems can deliver far more intelligence per watt.

Subscribe to the Gradient Flow Newsletter

Interview highlights – key sections from the video version:

Jump to transcript



Related content:


Support our work by subscribing to our newsletter📩


Transcript

Below is a polished and edited transcript.

Ben Lorica. All right. Today we have Zuzanna Stamirowska, CEO of Pathway, which you can find at pathway.com. The taglines are “Intelligence with no ceiling,” and they are a frontier AI lab building architectures and models that autonomously reason, learn, and evolve. And with that, Zuzanna, welcome to the podcast.

Zuzanna Stamirowska. Hello, Ben. Thank you so much for having me.

Ben Lorica. These are big, big phrases. “Intelligence with no ceiling”—I guess aspirationally that’s what every AI lab aspires to do. But since our audience is more pragmatic and practically oriented, let’s start there. Zuzanna, what specific problem were you trying to solve when you started the company?

Zuzanna Stamirowska. Absolutely. I guess for those phrases, you need to run after something, and that definitely makes sense. As a scientist, there’s this beauty in discovery and understanding. And, funnily enough, intelligence is actually a tool for understanding. So cracking intelligence is this kind of joker that you play on science and on uncovering the mysteries of the universe.

Especially when I was younger, I used to hang out a lot with physicists. That’s something that stayed with me and, actually, with part of our team, because we do have physicists on the team.

But very concretely, what prompted us to start Pathway and to start working on post-transformer architectures was that we were looking forward to reasoning becoming a thing. We have a very strong background in dealing with the dimension of time in machine learning—any sort of online learning. And we saw RAG, context, and memory, and we’re talking about 2023 and 2024, as bottlenecks. They were not efficient or sufficient to unlock AI as we would like—AI with no ceiling, if you wish.

Specifically, we’re talking about continual learning, memory, and reasoning first. And as we moved through it, we realized that reasoning shouldn’t just be linguistic. Verbalizing things in chain of thought is not the optimal way to do reasoning. With chain of thought, we’re essentially gluing it on top of a transformer, and transformers are naturally ill-suited to dealing with time and learning from experience.

Those were the problems that prompted us to start this work. A bit of background: our CTO was the first person to apply attention to speech recognition. He was at Mila and then Google Brain. And Lukasz Kaiser, one of the co-authors of Transformers, was the first check-in at Pathway.

Ben Lorica. Oh, nice. So you have—

Zuzanna Stamirowska. We have this background. I’m saying this because those experiences shaped us. I’m a complex systems scientist, so I worked on emergence and, generally speaking, on analyzing how small particles interact—how you get this link from micro to macro phenomena, and how that works. Sometimes, by defining the rules of local interactions, you can find properties that stay the same regardless of the scale of time or the size of systems.

There are also interesting things that our chief scientific officer, Adrian, applied that are linked to understanding what an operator is, what state is, and trying to model a system that evolves over time.

So the very big topic for us was time. Time is naturally linked with memory. It allows you to accumulate experience. And as we move to reasoning, you need the notion of time for thinking.

And “intelligence with no ceiling,” which may be a funny phrase—ultimately, one would like to get to extrapolation.

Ben Lorica. All right. So we—

Zuzanna Stamirowska. Think you need memory and reasoning that hopefully can be freed from language. We’ll see.

Ben Lorica. Right. All super interesting. On the reasoning front, you folks make the point that chain of thought is inefficient for a variety of reasons. But one advantage, I guess, is on the UX side: people can see what’s happening.

In the models that you’ve proposed, you’ve gotten rid of that. The reason chain of thought is inefficient is that you’re using text to do the reasoning and problem solving, right?

Zuzanna Stamirowska. And you do it sequentially, token by token.

Ben Lorica. Yeah.

Zuzanna Stamirowska. You don’t play like a chess player who has many possibilities in their head and may not even know why they choose one over another, because they have certain shortcuts in their thinking.

Ben Lorica. Right.

Zuzanna Stamirowska. But go on with your question.

Ben Lorica. The advantage is that there’s some sort of paper trail.

Zuzanna Stamirowska. Yes.

Ben Lorica. So then—

Zuzanna Stamirowska. It’s being argued, right.

Ben Lorica. Right. So on that point, what’s a replacement for the paper trail?

Zuzanna Stamirowska. The way I would take it is that there has been a lot of discussion. We were very happy because we were probably, publicly, among the most advanced—or the most extreme—in terms of latent reasoning out there, at least in terms of what’s proven.

What we see with Astra is that there was a big debate in a piece in The Information about the fact that Astra starts to have some elements of latent thinking. The entire piece was about: “Okay, but we won’t know what the agents are doing because we won’t have the paper trail of their thoughts.”

I actually discussed this with some people from Constellation Lab who are focused on safety.

Topic number one is: to what extent do we really trust the chain of thought of models in terms of actually explaining what they’re doing? There is a human-readable element, and then messages may be encoded in your commas. So that’s one issue. I don’t think many researchers trust that chain of thought is some kind of magic bullet for supervising your models.

Ben Lorica. The average user looks at it and gets some sort of reassurance. But maybe, to your point, there’s no basis for that reassurance.

Zuzanna Stamirowska. Yes. So already the situation is not that good.

The second safety argument comes almost from legal doctrine. Humanity took a while to understand that people shouldn’t be punished for their thoughts. Should we punish a model for its thoughts? Do we actually care?

We punish or control actions. Safety should be applied at tool calling, at the harness level, at the level at which models can interact with the world.

Latent thinking has huge benefits. It can help us solve constraint-satisfaction problems natively, like what we’re showing. We’re solving constraint-satisfaction problems and math with tiny, tiny models. And then there’s cost efficiency. There are just insane cost gains coming from latent reasoning.

Ben Lorica. And this is where your approach definitely shines: cost efficiency.

But going back to the paper trail and lineage, another way to look at it is that with a paper trail, you can see how the model is thinking, which may also allow you to debug it—or at least understand the failure mode. You can see the chain of thought and say, “Oh yeah, this is where it made a mistake.”

Zuzanna Stamirowska. Perhaps.

Ben Lorica. Yeah.

Zuzanna Stamirowska. I guess it depends on the problem and how you can measure the final output.

Ben Lorica. I’m just giving the traditional approach.

Zuzanna Stamirowska. This is something that might make sense, depending again on the use case. Can you measure the output?

I honestly don’t know how debugging would work here for us. What I can say, though, is that with BDH-CQ—the model that we use to show the benchmark on ARC-AGI-1—this is a modern sequence model. It’s a model capable of working with language.

That was actually the point we wanted to make: you can add chain of thought on top of it if you really want to. So I’d say it’s more of a design decision if there is demand and usefulness from having chain of thought on top of latent thinking.

But that will probably be just part of the thinking. Latent thinking still won’t be traceable in the same way as the chain-of-thought part.

Ben Lorica. You used the phrase “continual learning,” which is a big term. It’s kind of like “world model”: people mean different things when they use the term.

Broadly speaking, aspirationally, the goal is for your model on day 500 to be better than your model on day one, right?

Zuzanna Stamirowska. I define it by the model, right? I define it by the model, not the harness. Not the fact that you have memory, but that your model is actually better.

Ben Lorica. Right. But as it stands, when people use the phrase continual learning, the learning can show up in various places. As you mentioned, it could be in context, in memory, in the harness, in model weights, or in the overall system design.

In your case, continual learning means what? Because you’re not updating model weights, right?

Zuzanna Stamirowska. We would be looking at the weights.

Ben Lorica. Oh, you are. Okay.

Zuzanna Stamirowska. Ultimately, we would be looking at the weights.

First of all, step one is simply having long context. This is like having very, very long context. Imagine a state-space model: you just have a lot of state. That’s step one.

But ultimately, we’re actually looking at modifying the weights. There’s the concept that originates from Geoff Hinton of fast weights and slow weights. In fact, this is something where we published the theory—the first theory of it, let’s say—in a paper called BDH, Dragon Hatchling.

There we formally managed to prove the link between the transformer and Hebbian learning, models of the brain. We showed that through synaptic plasticity—some local interactions—we can boil everything down to four equations and show that reasoning can emerge from those four steps happening on synaptic connections between neurons.

What we saw experimentally is that this sort of architecture allows us to avoid catastrophic forgetting as we move from task one to task two.

With transformers, the traditional problem with continual learning is catastrophic forgetting. Let’s say you speak Spanish and then start to learn French: you forget your Spanish. We actually see this as we compare against a GPT baseline. We see this catastrophic forgetting of the previous language.

Whereas our BDH architecture keeps acquiring new skills and avoids this problem of catastrophic forgetting.

But these are early results. Before it’s brought to market as a full product, I’d put it later. For now, it seems that the biggest differentiation and immediate value will come from long context and latent reasoning. Then this element will kick in later.

Ben Lorica. But if you fast-forward and your approach catches on and takes off, listeners to this podcast might be familiar with things like supervised fine-tuning or some form of reinforcement learning. In both cases, there’s some data or examples involved.

In your notion of adaptation or learning, how much data is needed, and how do you represent that data?

In supervised fine-tuning, for example, here are some examples, and then I update the model. It’s not happening in real time. Here’s a bunch of examples, then there’s a batch process and an update.

Zuzanna Stamirowska. No. Already in the benchmark for ARC-AGI-1, we’re having some sort of in-context learning. It’s latent reasoning with in-context learning.

Ben Lorica. Translate in-context learning for civilians.

Zuzanna Stamirowska. You look at examples and figure it out.

Ben Lorica. Just like a human being.

Zuzanna Stamirowska. At least, yeah. You try to get there.

I know from Lukasz Kaiser, for example—the co-author of Transformers, who was at OpenAI for the longest time—that one of the most exciting things, at least for him, would be to find a way of learning that is much more data-efficient.

If you think about a human, how many times do you need to burn your fingers touching an oven to understand that you shouldn’t touch it? Probably once. Hopefully never.

Whereas models—transformers—have to see so many examples to learn. The paradigm of scale is still very much there.

Having an architecture that is more efficient, more data-efficient, and hopefully dramatically more data-efficient is a bit of a holy grail.

Ben Lorica. So how far—

Zuzanna Stamirowska. Are we there? I don’t think I can disclose that right now.

But generally speaking, as we started with “intelligence with no ceiling,” of course this is something that you want to get.

It becomes a bit easier when you get to reasoning, because when you optimize for reasoning, you’re building something slightly different from an LLM.

An LLM, understood through its first use cases as a chatbot, would be evaluated by asking: Does it answer my questions correctly? First, you would have intuitions from a search engine. You’d look at summarization. These are probably the hopes you would have for any model.

Those things need to be a given. A basic floor needs to be respected there.

If you optimize for reasoning, what you really want are systems that are capable of gaining autonomy and maybe learning to behave autonomously in a specific environment, context, or task—or ideally adjusting to whatever that is.

Then you don’t need so much general knowledge. You want something that is smart enough that it can quickly create shortcuts for reasoning—more like reasoning frameworks—from the experiences or data that flow in.

Ben Lorica. But the bottom line is, you still have to provide it a few examples or demonstrations.

Zuzanna Stamirowska. Yes. The question is how many.

Ben Lorica. And then it learns in real time, right away?

Zuzanna Stamirowska. At inference, yeah.

Ben Lorica. And then when you show it future examples, it starts applying what it’s learned.

But can it also unlearn once the situation changes? Let’s say I apply this to the stock market.

Zuzanna Stamirowska. Yeah, and you have a regime change and you want it to detect: “This is a regime change. Doing what I was doing yesterday is dumb. I should understand it’s a regime change.”

I don’t have tests that I can disclose for that.

But this is actually a property of live systems. They are more adaptive than batch, static ones. You can override.

If I were applying something like this to investment, my use case would be spotting regime changes. It wouldn’t just be investing literally, because if you can spot a regime change very early, that’s where you get your gains.

There’s a certain property of live systems where I think you can spot regime changes faster because you see trends instead of thresholds, et cetera.

Ben Lorica. Just to give our audience some baselines here: the latest BDH model is not open source, right?

Zuzanna Stamirowska. No.

Ben Lorica. You’re building a company. Previous generations were more open, right?

Zuzanna Stamirowska. Correct. The BDH architecture, the paper we published back in October, and the main version that allows you to reproduce the results from the paper are available on GitHub.

Ben Lorica. So it’s not open source, and it’s actually not big relative to gigantic LLMs. We’re not even at a billion parameters, right?

Zuzanna Stamirowska. Exactly.

One thing you said is that to get to some sort of reasoning, you could learn by having more examples of the task, and over time you would get better at this.

The other path would be to scale data and have recipes for different problems. This is the transformer type of scaling. You have so many recipes that ultimately you end up finding a recipe for the problem—but then you have to scale.

One thing we worked very hard on was trying to uncouple reasoning power from the size of models.

Of course, language skills help to a certain level. But how much can we squeeze out natively from this weird reasoning fabric?

Imagine you have a virgin brain. How much can it do even at a very small size?

That’s what we’ve shown in the ARC-AGI-1 result, where we are close to—sorry—30% accuracy, at less than one-tenth of a cent. It’s completely all the way to the left, such that you almost don’t even need to move the y-axis in terms of cost.

This is purely latent reasoning on a model that has 150 million parameters, and the cost reported is the actual compute.

Ben Lorica. 150 million?

Zuzanna Stamirowska. 150 million, sorry.

This is purely compute cost. Many prices are a bit arbitrary, whereas the point we want to make is: with this paradigm, we’re shifting the cost so much to the left.

Then we can go up from there. It gives us a new scaling possibility. It opens a new quadrant in terms of cost to accuracy that we’ll be able to gain.

Zooming out, I believe—and I think many people believe by now—that AI will simply be forced to find ways to get much more intelligence from every watt.

Ben Lorica. I see. Does that mean you can do inference using CPUs, not GPUs?

Zuzanna Stamirowska. I didn’t say that. Right now, we’re very traditional: GPUs. Likely any GPUs.

But when we published this benchmark, everybody doing anything on edge started knocking on our doors, because this is what you need for autonomy.

As you read on our website: autonomous systems. You need reasoning for autonomy. You don’t need to know the queens of England. You need to learn patterns of thinking and behaving and, hopefully, even be able to get them dynamically yourself while being thrown into a new situation.

This is the generalization capability.

Ben Lorica. It seems like this would fit into how many companies have come to realize: look, we don’t actually need massive foundation models. We can do a lot of customization and specialization on open-weights models because, for many tasks, we frankly don’t need models that speak 20 languages.

So it seems this could fit into that paradigm, where I could take one of these latent-space models and develop custom ones for different parts of my company.

Zuzanna Stamirowska. Yeah. Something that I’m starting to see is everything around security, where you actually need to deploy close to hardware.

What you want are something I call SRMs: small reasoning models. You’re used to having SLMs, right? But imagine a small reasoning model. You get a different thing: you get reasoning power.

This can be a local AI. And, funnily enough, this is something I see emerging as a market signal that we’re getting.

Ben Lorica. One more baseline on the model you folks just announced. It’s not open source. It’s not big. Is it multimodal?

Zuzanna Stamirowska. Not yet, not exactly. ARC-AGI is visual, so many asterisks apply. It kind of can be, and ultimately somewhat will be, but right now it’s primarily—

Ben Lorica. For our listeners who don’t follow the inside baseball of ARC-AGI, why should a person care about ARC-AGI, and why is that benchmark important?

Zuzanna Stamirowska. I think the easiest, most direct answer is that this is the most renowned benchmark for measuring how far we are from AGI.

It was created by François Chollet and the ARC Foundation, and its goal is pretty much to measure some sort of autonomy and adaptation of the models.

François Chollet wanted to show where LLMs fail and how much more needs to be done.

There are by now three different versions of this benchmark: ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3. They report accuracy on specific tasks and the costs.

Most of the models that are reported actually report the cost of API calls. That can be a bit arbitrary.

Since the case we’re making is a scientific case for a paradigm—not a case for a product—we report the seconds on a GPU of specific characteristics and the cost we’re actually paying for it.

In ARC-AGI-1, these are visual puzzles. They’re fairly easy for a human, and they were very difficult for standard LLMs.

On ARC-AGI, we’ve been seeing progress in how AI, with increasing sizes and different techniques, has been getting better and better.

Ben Lorica. So again, this is—

Zuzanna Stamirowska. The most renowned benchmark where you ask: are we far or are we not far from AGI?

Ben Lorica. For people who don’t follow ARC-AGI, what kinds of models are in there? There are LLMs. Are there world models?

Zuzanna Stamirowska. I’m not sure there’s a world model there, honestly.

Ben Lorica. And I know there are—

Zuzanna Stamirowska. Some approaches. There are some research approaches using latent thinking.

Ben Lorica. Okay.

Zuzanna Stamirowska. And then—

Ben Lorica. Your specific model in the results is not necessarily the most accurate, but it was the cheapest for its level of accuracy. Is that a fair assessment?

Zuzanna Stamirowska. That is exactly a fair assessment.

But this level of accuracy is still above DeepSeek, as we know it. When it was launched, it was around the level of Luna, which was launched at the same time, while having dramatically lower costs.

As I said, the point we wanted to make was the case for latent thinking as an overall approach.

As you scale the models up and train this—this is a 150-million-parameter model—as you scale up, and potentially even have chain of thought, there’s so much you can do with this. This curve goes up, and of course we’re working on it.

The truth is ARC-AGI-1 is one measure where you can show performance. Labs internally use very different benchmarks to measure progress.

A good way to measure the amount of intelligence you’re getting—and this is maybe culturally an OpenAI way—is going through different math benchmarks.

Ben Lorica. So—

Zuzanna Stamirowska. These are the newest kinds of reports we’re getting from Astra and all the—

Ben Lorica. I’ll try to map what I understand into plain language.

It seems like your model can learn new relationships from examples at runtime.

Zuzanna Stamirowska. Yeah.

Ben Lorica. It might even be able to extend some of the learned rules.

Zuzanna Stamirowska. This is finding reasoning shortcuts. Yeah.

Ben Lorica. So where does it struggle? Extrapolation?

Zuzanna Stamirowska. We still haven’t gotten there, but we haven’t scaled it.

We’ve run experiments, but we haven’t trained the scaled model yet. We ran experiments for the scaling laws at 600 billion parameters, but we still haven’t trained the fully scaled model.

What we’ve been doing until now is stabilizing training recipes. ARC-AGI-1 is one thing, but for such a model to be able to claim the post-transformer crown, it needs to speak language and behave reliably, exactly as well as—or better than—the transformer on language, and then have all the other properties like constraint satisfaction, math, et cetera.

There’s a lot of moat still to be built and a lot of depth to be extracted from this scientific direction we’re digging into.

This is what we’ve been stabilizing before investing a lot into training a huge model.

Ben Lorica. So the main thing here, at a high level, is that at 150 million parameters, you can already see some nascent capabilities that show promise.

The expectation is that if you scale to a larger model, the parts it’s struggling with will hopefully go away. Although there’s no theorem that guarantees—

Zuzanna Stamirowska. Scaling laws from a transformer, since it is a—

Ben Lorica. Oh yeah, that’s true.

Zuzanna Stamirowska. And chain of thought. There’s some scaling behavior, but—

Ben Lorica. You’re a different architecture.

Zuzanna Stamirowska. Post-transformer. This naming was actually very important.

Unlike, for example, energy-based models or approaches that are fully different—in vision, for example—we actually—

Ben Lorica. Yann LeCun.

Zuzanna Stamirowska. Yes.

In Yann’s case, as far as I understand it—and I hope we actually agree very much on this—the case is for finding an internal representation, which I would call a latent space. The question is: how do you get to it?

We started from something that works. As we started this research program, we went from the transformer and wanted to find a bridge to the brain—understood as brain models and Hebbian learning.

We wanted to understand this at a very deep theoretical level, and then know how to navigate it, how to select and prioritize experiments, et cetera.

Ben Lorica. The model as you have it right now, the 150-million-parameter model, is already good at copying patterns from examples, right?

Zuzanna Stamirowska. I’d say with those models, we were even surprised by how much native capability they have, just out of the box.

Ben Lorica. But we saw—

Zuzanna Stamirowska. Incredible results in state tracking, for example.

Ben Lorica. Right. But it still struggles with turning those examples into some sort of reusable business logic.

Zuzanna Stamirowska. Still, yes. But this is something that comes with scale.

Ben Lorica. Is it memorizing, or actually learning a rule?

Zuzanna Stamirowska. Learning a rule.

Ben Lorica. Okay.

Zuzanna Stamirowska. This one is learning a rule. That’s fine.

The full approach to memory—what memory should be and how memory is understood in the market right now—is interesting because we see a lot of actors around memory, and usually they operate, let’s call it, at the harness level.

Ben Lorica. You’re not a fan of memory, I take it?

Zuzanna Stamirowska. Why?

Ben Lorica. For some reason, I thought I read somewhere that you think people are over-relying on memory.

Zuzanna Stamirowska. No, I don’t think so. It depends what you mean by memory.

Ben Lorica. My bad.

Zuzanna Stamirowska. Maybe something I said could have been understood that way.

There’s memory as in, I hope, our memory after this conversation. Maybe we were inspired. We internalized somehow what we heard.

And there is the sort of memory where you write something in a notebook and store it on a shelf or in a database, and maybe compress it somehow because somebody takes a number of notebooks and—

Ben Lorica. This is an architecture that, as it gets better, you can imagine sliding into the stack people are building now, where there’s memory and context storage.

Zuzanna Stamirowska. Yeah. It actually internalizes context.

Funnily enough, as this model goes out, it’s probably first replacing some harnesses and making harness building more generalized and easier because you have memory that’s internalized.

Ben Lorica. Bad news for all the harness startups.

Zuzanna Stamirowska. What can I say?

Ben Lorica. By the way, Zuzanna, there are three different things.

There’s memory, which is what I just did. There’s context: here’s some information, answer my question. And then there’s knowledge: here’s everything our company knows.

For this, you—

Zuzanna Stamirowska. You’ll still need a harness, right?

I don’t think models—I mean, you need to be connected to the outside world. It doesn’t make sense to put all the knowledge in the model, perhaps. You want a system of record.

Ben Lorica. That’s what Anthropic and OpenAI want to do: put everything in the model.

Zuzanna Stamirowska. Yeah, but I don’t think you need to put everything in the model.

Ben Lorica. I think people are realizing this because it’s so expensive to run these models.

Zuzanna Stamirowska. Because it’s so expensive. Imagine it wasn’t so expensive. Imagine you had a radical mixture of experts where just a small portion of your brain fires up for a specific thing.

But I still think for general knowledge you may want to call out, and you don’t always need to internalize everything.

The second very big question is the system of record—the actual books, right? Why have it in the model?

Ben Lorica. ARC-AGI, as you described it, has visual puzzles. What about language?

Zuzanna Stamirowska. Language is, in fact, the number-one thing.

As I said, we started on the more traditional path, looking at what works, and followed the breadcrumb trail toward models of the brain—or some properties that we know the brain has.

It actually scales like transformers on language.

Ben Lorica. Okay.

Zuzanna Stamirowska. What we want is dramatically long context, state tracking, and latent reasoning.

Language is fine. This is why we’re not investing in it so much right now, because it’s relatively the less differentiated and easier part of the story for us.

It’s more about our research and company strategy: where to dig and where to run faster.

Ben Lorica. And—

Zuzanna Stamirowska. Which parts are the most difficult.

But as I said, we run a suite of benchmarks. There are basic linguistic benchmarks, and there’s a floor we have to meet there.

Then there are things that are about reasoning capabilities: puzzle solving and all of that.

Ben Lorica. I think it makes sense that inference using these types of models is cheaper.

I’m not sure if this is possible to answer, but as you scale, what about training costs? Is it going to be as expensive as scaling—

Zuzanna Stamirowska. Definitely not as expensive as scaling LLMs.

Somehow, by definition, because this race is a race for Vikings. We have to travel light.

But as we were just discussing, there’s the data-efficiency element and the capability of learning from—

Ben Lorica. Don’t use data from Reddit. Is that what you’re saying?

Zuzanna Stamirowska. At the end of the day, it really depends on what value you deliver to customers.

If you target the enterprise, maybe all the subculture from Reddit and the weird threads where people are trying to mimic the sounds of machines—there was a thread like this—you don’t necessarily need.

Or how—

Ben Lorica. How many languages do you need to be fluent in?

Zuzanna Stamirowska. Yeah, right. Sorry.

Ben Lorica. How many languages does your model need to be fluent in?

Zuzanna Stamirowska. Exactly. These are things that maybe you don’t need to optimize for.

The second thing is deployment: how AI will ultimately be deployed. I’m talking beyond chatbot use cases.

In cybersecurity, you probably can use very small models that are protectors at specific endpoints.

Ben Lorica. The types of models you’re training, Zuzanna—as you mentioned, the training isn’t going to be as expensive as training the types of models people are more familiar with.

Does that mean more groups can train the types of models you’re training? In other words, professors can train them? You know how professors have no compute.

Zuzanna Stamirowska. I know, unfortunately. And I know the GPU market very well by now. I’ve run very intense due diligence on that.

We’re looking at something that can be called—you’ve probably heard so many different terms—sovereign AI, but perhaps corporate sovereign AI: you should have your own AI.

And I think the entire story between Anthropic and OpenAI right now—

Ben Lorica. Yeah, I think people—

Zuzanna Stamirowska. Right. What happens to your prompts? How safe really are you? Can you have your own model?

One of the approaches is: let’s take open-source weights, or open weights, and fine-tune them. And this is, of course, a very good business.

With models that can train with experience, we wouldn’t need that fine-tuning in the same way.

And for compute, something I can give you as an early estimate: we predict that one rack of GB300s could potentially be considered a mainframe to run a model that would be up to one trillion parameters.

Ben Lorica. One of the things people have come to realize now is: okay, there are GPUs, but increasingly one of the bottlenecks is memory and networking. How do these models fit into this?

Zuzanna Stamirowska. If you have a model that can be fully trained on one rack, you kind of resolve the problem of interconnect. There you go. That’s your win.

And you have so much addressable memory on this—

Ben Lorica. One rack. Right.

Zuzanna Stamirowska. Right.

Also, this architecture is very sparse, and we don’t hit the same limits as transformers. But we do need a lot of memory.

So in GB300s, or the new Vera Rubin, and equivalents from AMD, memory is actually the thing that grows.

Right now, as I look at the market and where and how you can get capacity, the less reliant you are on interconnect, the better it is for you.

Generally speaking, you want to find a place where the only two things that can fail are electricity and internet.

Ben Lorica. Right. And hopefully—

Zuzanna Stamirowska. Of course, not even then.

Ben Lorica. Right. What’s your sense, Zuzanna?

Obviously, for your company, the next step is to train a slightly larger, more capable model, keep doing that, and then at what point do you say, “Okay, we need revenue and a product”?

Zuzanna Stamirowska. Yeah, you know—

Ben Lorica. Yeah.

Zuzanna Stamirowska. It always depends.

We just announced our round at a $500 million valuation, and the purpose of this round is to prove the architecture and prove the technology.

The ARC-AGI-1 result is public now, and that’s step number one in us literally delivering what we promised.

Ben Lorica. I think I read somewhere that some people also validated what you did, right?

Zuzanna Stamirowska. Yes, indeed. It has external validation. It was reproduced externally by folks from NYU.

Ben Lorica. We’re not talking Theranos here.

Zuzanna Stamirowska. Lukasz Kaiser, the co-author of Transformers, literally took it for a spin. So did researchers from NYU and other groups.

Ben Lorica. So the next step is you’re going to train a slightly larger, more capable model, and then you’re just going to go up and up the—

Speaker 2. Step.

Zuzanna Stamirowska. No. The next steps will be us showing generalization in terms of capabilities.

Of course, it’s somewhat bigger, irrespective of size. But as you start to show things involving language, of course you need to go above 100 or 150 million parameters.

Maybe a good analogy is to go back to the names of the models. Our models are named after dragons.

Ben Lorica. Okay.

Zuzanna Stamirowska. A dragon is a mythical creature that can blow out flames, has claws, and has these magical properties.

Our first paper, and even the architecture, is called BDH. It stands for Dragon Hatchling.

The point is: it’s a dragon, but it’s a hatchling. It’s tiny, but it shows those kinds of capabilities.

Our job right now is to show the flames.

Ben Lorica. I see.

Zuzanna Stamirowska. Puffs of smoke, you know. It has claws. And then grow it.

Ben Lorica. You’ve gone from science fiction to poetry, and then at some point it’ll be nonfiction.

Zuzanna Stamirowska. I believe that as you actually have the flames and you have the benchmarks—the benchmarks are benchmarks, right?—if you see something doing math, doing crosswords, doing ARC-AGI, and having the same results on language, you see generality across the spectrum at a small size.

Then, sure, as you move up, you expect scaling.

Probably some things will surprise us as we scale, because we tried the scaling laws. There are questions like: how long do you keep the model training at a specific scale to measure other things?

We’re excited to see things that will probably surprise us at scale.

Ben Lorica. What’s your expectation for cadence? Is it every six months there will be a new—

Zuzanna Stamirowska. Hopefully way, way faster, but I can’t commit to that.

The only comment I can give you is that I like to say the years in AI are hamster years, not even dog years.

This is the rate at which it moves, and this is the rate at which we have to move as well.

Ben Lorica. At this point, you’re functioning almost purely as a research outfit, right?

Zuzanna Stamirowska. I’d say a bit more than that.

We do have partnerships with AWS for distribution. We also work out deployment scenarios and all of that.

Ben Lorica. Distribution to whom? Who’s using it?

Zuzanna Stamirowska. We have design partners, and we work with people on defining and scoping use cases.

Actually, to whoever is listening here, specifically right now we’re looking at any use case that would require reasoning over long horizons.

Ben Lorica. Oh, there it is. You just uttered the buzziest term right now: long-horizon agents. Think, think, think—

Zuzanna Stamirowska. Think. Long horizon.

You really need to understand what happens. Maybe it doesn’t need to know every language or all kinds of general knowledge. Can we scope it?

But we’re looking at very long context. You understand the state and what was happening throughout the problem. You keep it in mind, and you need powerful reasoning on top of that.

This is literally what we’re hunting for.

This is something we take for training and testing, and we discuss it with a number of cool design partners.

Ben Lorica. Long-horizon agents are the thing right now.

Zuzanna Stamirowska. That’s a very difficult thing, because you have to have state and memory.

Ben Lorica. People are using all sorts of tools, even formal verification like Lean.

Zuzanna Stamirowska. I would imagine, because you want to patch it up somehow.

Ben Lorica. Yeah. But maybe with this sort of approach, if you can scale it, these models will end up not having to be that big.

Zuzanna Stamirowska. This is something I believe, but it’s also something I’d like to actually test.

Ben Lorica. I mean, they don’t need to be AGI. They don’t have to be so big to solve domain-specific problems.

Zuzanna Stamirowska. Exactly.

You can start with small SRMs. The intuition is that they can be very small.

If you have a problem that looks like a puzzle, you can have an SRM on it if you can formalize it that way.

But specifically, what we’re looking for right now and discussing with people is reasoning over long horizons. Same for data: how do you benchmark this? How do you train it specifically?

Of course, we work with some very cool people on that.

Ben Lorica. I don’t want to be a Debbie Downer, but I think over the long term some sort of openness in models will prevail. Do you think—

Zuzanna Stamirowska. So?

Ben Lorica. Yeah, I think so.

Look at how much people love open-weights models. Part of it is that you have some control. You can customize it. You’re not beholden to someone else’s timeline. You know how OpenAI and Anthropic deprecate models.

And then there’s your IP and your control.

Right now we’re kind of like in the year 2000, when Solaris was clearly much more capable than Linux, and Oracle was much more capable than Postgres. But you know what? A few years later, a lot more people were using Linux than Solaris.

And right now, we’re not even that far apart with open-weights models.

Zuzanna Stamirowska. I wonder, because there are two things to look at.

There’s the topic of the models themselves, and then the need. What’s the real need for open-weights models?

Ben Lorica. But my point is—

Zuzanna Stamirowska. Is it that you really want to see the weights, or is it that you want control?

Ben Lorica. It’s control.

Zuzanna Stamirowska. What if you had a model that you could deploy on your own instance?

Ben Lorica. Well, that’s not—

Zuzanna Stamirowska. Necessarily open weights, yeah.

But you have it on your own, maybe in escrow. The question is: how do you deploy it?

And frankly, this is what we work on with AWS. Can you find some sort of escrow—

Ben Lorica. The challenge is that if Linux didn’t exist, obviously we’d all still be using Solaris.

But if an open alternative becomes comparable, then it’s game over.

In your case, right now you’re the best at what you do. But if what you do becomes clearly good, people will try to replicate it, just like DeepSeek and other Chinese companies try to replicate things.

Zuzanna Stamirowska. The question remains about capital. How do you make money over the long term? How do you make money out of open weights?

With traditional open source, unless you do hosting—from what I’m hearing, you pretty much do hosting. You have your GPU infrastructure, you host it for people, and they have their models and do whatever they want.

So does—

Ben Lorica. Does anyone care whether anyone is making money on Linux?

Zuzanna Stamirowska. Well, I imagine the people who built companies around it do.

Ben Lorica. The Linux Foundation basically supported the developers over time.

But this is more complicated, obviously, because you have to retrain the models constantly, and then you have data and so on and so forth.

I just think it’s such an important part of the stack that if there’s an open alternative—whatever “open” means, open weights—people will gravitate toward the open alternative.

So for you, the hope is that you can fly under the radar for as long as possible.

Zuzanna Stamirowska. Generally speaking, in the R&D sense, yes.

I’ve had discussions with different investors about the outlook for open weights.

Would you invest—just to make money, not to have something dominant, but to make money—in an open-weights model?

I understand the game of saying, “We don’t want to use foreign open-weights models. We need to have a domestic one,” or making the market slightly more difficult for some folks who want to IPO.

And of course—

Ben Lorica. Yeah. Look at Nemotron from NVIDIA. It’s not as good right now.

Zuzanna Stamirowska. It’s an ecosystem game, right?

Ben Lorica. Yeah. Nemotron from NVIDIA isn’t as good right now, but over time it could get better. And it’s not like NVIDIA needs to make money on Nemotron.

Zuzanna Stamirowska. Right. That’s right.

Ben Lorica. So what if Apple starts doing latent-space models?

Zuzanna Stamirowska. On-device. I think they—

Ben Lorica. Build it into the OS.

Zuzanna Stamirowska. The point is, it would actually make a lot of sense for them. So I strongly encourage them to have a chat with me.

Ben Lorica. Yeah.

Zuzanna Stamirowska. At least the big story is that they have one of the largest deployed bases of powerful chips. You can—

Ben Lorica. Just put the latent-space models in the OS.

Zuzanna Stamirowska. Everywhere, right?

And the same applies, honestly, to different hardware producers. We have a lot of discussions with those folks.

Ben Lorica. And with that, thank you, Zuzanna. Again, go to their website at pathway.com.

Zuzanna Stamirowska. Thank you.