Podcasts
Paul, Weiss Waking Up With AI
Cost and Efficiency: The Engineering Behind China's Open-Weight AI Models
In this episode, Katherine Forrest and Scott Caravello explore the rise of Chinese open weight AI models, breaking down the innovative engineering techniques behind their efficiency and the significant cost advantages driving their rapid adoption among developers worldwide.
Episode Speakers
Episode Transcript
Katherine Forrest: Hello everyone, and welcome to today's episode of Waking Up with AI. I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello. Katherine, I need to do a quick call back to last week's intro when you again discussed the hats in my closet. And you raised the question if I knew about the book The Man with Many Hats. And this morning I was just minding my own business and out of the blue I got a text from my sister that just says “CAPS for sale,” which confused me. And then it turned out that she was listening to the episode and wanted to let me know the name of the book.
Katherine Forrest: Caps for sale!
Scott Caravello: Caps for sale.
Katherine Forrest: You know, actually, Amy, my wife, had told me about that right afterwards. She sort of was like holding up a post it because she sometimes, you know, we have desks that face each other and what she held says “caps for sale.” But of course I didn't see it. I was concentrating so much on the caps behind you when you're in your video.
Scott Caravello: Exactly. But so my family's in on it now, so thanks.
Katherine Forrest: All right, OK, OK. It's all in the family there for the Caravello's. Shout out to the Caravello family, who spawned such a brilliant young man. All right, so you're also suffering with something completely hideous right now, which is the smoke accumulation in New York City.
Scott Caravello: Oh my gosh, it's awful. So, you know, starting in advance to everyone if my voice is a little off. I bravely went outside for a few minutes this morning, and I'm suffering the effects of the smog from the wildfires. It's better in Maine, right?
Katherine Forrest: Yeah, well, Maine—it actually moved South. It was here. But then I was back in New York City for a few days and it was in New York City. And now when, when I flew back here yesterday, it had actually moved, pretty much, pretty much, moved away, not completely, but not like what you guys have. So, boy, it's a big deal. It's a big deal. And, you know, my heart goes out to anybody who's got sort of breathing issues where, you know, they're having to sort of grapple with this. But, in any event, we'll try to make today a low impact day, right? Not like last week when we put people through 30 minutes of three different huge AI events.
Scott Caravello: I was exhausted after that one, so…
Katherine Forrest: Were you?
Scott Caravello: This, this will be a bit more focused.
Katherine Forrest: Yeah, today's going to be a little bit more focused, so you want to give us a little intro to what we're going to be doing?
Scott Caravello: Yeah, absolutely. So, I mean, broadly, we're going to talk about Chinese open weight models. And Katherine, I know that you're going to discuss our purposeful use of the term “open weight models,” but those have gained a lot of traction lately thanks to how capable the recent releases are and at a lower cost compared to a lot of the closed source frontier models that we talked about a lot in the podcast and that you see in the headlines.
Katherine Forrest: Yeah. Well, let me just sort of like pause on that “open weight” and what it means, because I think usually people throw around the phrase open source models for a lot of the Chinese models. And in fact, there's a difference between an open source model and an open weight model. And you can be one and not the other. You can be an open weight model and not be an open source model, although if you're truly an open source model, you're going to be also an open weight model. So the difference is that in an open weight model, which is the, that's what the primary Chinese high capability models are, that have been released in the United States. They released the weights of the model, the parameters, and remember that the parameters are the relationships between the data inside the neural network. And so you've got all of this data that's sort of being ingested into the neural network and it's getting related to one another, to different pieces of data. And the relationships are changing as more data comes in. And the weights of the parameters are they tell you a lot about how, a huge amount, about how the model can reason and the way in which it has sort of considered the data. And I'm putting that sort of like in just English language versus engineering language versus open source. What open weight does not tell you is it doesn't tell you exactly what the training method was. It doesn't tell you, for instance, how many layers the neural network has. It doesn't tell you about the guard rails that the model developer might have imposed, or any output filters. It tells you about the weights. So, it gives you a lot of information, but it doesn't give you all of the secret sauce of the model. So, there really is a true difference between an open source model where you get everything and you can really look at it and study it from soup to nuts, versus an open weight where you're just getting sort of one piece of it.
Scott Caravello: Yeah, absolutely. But also the open weight piece and having access to the waste within the model is what can also make it so useful for folks because then they're able to adjust and fine tune the models when they're, you know, hosting it in their own environment to make it more useful for their specific use case.
Katherine Forrest: Yeah. And so the reason we really want to talk about these Chinese models is because they are becoming models that a lot of US and ex US developers have been writing on. They have been used more and more as the base for tool development and additional fine-tuned model development. And recently, we saw an uptick in adoption of open weight, Chinese open weight models. And one of the places that you can go to look at adoption rates is something called OpenRouter. And that's all one word. And for those who aren't familiar with OpenRouter, it's basically a place that you can go, on the web, to essentially you can create an account and you can access a variety of models for your needs. In some ways it's akin to a cursor, but it's not a cursor. It's not, it doesn't have the functionality of a cursor. It's sort of different from a cursor, but it's got some relationship to what cursor does. It's got many models inside of it and it routes your task to the appropriate model or you can choose a model. But what you get from that is you actually get a lot of information about what models things are going to. And so the Chinese models that are open weight that we're going to be talking about, the usage had gone from what was virtually 0 in 2024 to nearly 30% of the usage in open router in recent weeks. So we're seeing a real uptick in the use of these Chinese open weight models.
Scott Caravello: Yeah. And, just to add to that, I mean, OpenRouter has a lot of traffic passing through it. So, it is actually providing a pretty reliable measure of the models that developers are actually using. But it's also, you know, even though the usage went from nearly zero in late 2024, it's not like all this came out of nowhere. And I know Katherine, you had covered way back in January 2025 AI so-called Sputnik moment when DeepSeek released its R1 reasoning model under an open license. And I think back then you had called it a wakeup call for the industry.
Katherine Forrest: Right. And actually, you know, the press was all over that and they were talking about it as a wakeup call and there were issues with whether or not certain stock for certain companies was were overvalued because maybe there would be sort of real drops in what the needs were for compute. And, anyway, let's just go back to that and remind everybody that within days of this R1 release by DeepSeek, which became the most downloaded free app in the US App Store basically for a little bit. And it actually erased, once it was released and a lot of the press had come out about its efficiencies and it's lower cost or inference or and training that, allegedly, the claim that it made was that it only cost $6 million, I think was the number to actually train the model versus a billion dollars that roughly a trillion dollars in the US tech stock market or the tech stocks in the stock market. There was a loss in value and the shock wasn't the performance of the model. It was that there were claims by DeepSeek that the model was being trained at a far lower cost because of, you know, we won't go into a lot of this, but we've, and we've mentioned it on other episodes, but the architecture of the Chinese models, and this is true for the models that we're going to talk about today, has to do with not just paying attention like the transformer architecture to all the data that's going into the model, but using things like, for instance, something called sparse attention where the model actually will pick and choose the data that it actually looks at. And so that is one of the ways for reducing the cost.
Scott Caravello: Right. And so that is still very much the story today. It was a story back in 2025 and the story today that these models are cheaper to use and that's driving the adoption rate. Let's say whatever you're using AI for doesn't need that kind of frontier model firepower you're getting from the closed source models. Maybe it's just for document summaries, or running a customer support bot, or some very routine code generation where a model that's maybe 90% as capable as the top of the line frontier that's available at a fraction of the price then becomes much more attractive.
Katherine Forrest: Right. So we've got two things going. We have claims about the efficiencies of the Chinese models on the training on the front end and then claims about the inference and the compute necessary to make the models perform sort of on the back end. And these efficiencies are actually real. I mean, I think people were wondering exactly whether or not the numbers were as significant as DeepSeek had said when it first came out with R1. But now we're several generations in and it's quite clear that these Chinese open weight models do bring some efficiencies to the table. You know, Andreessen Horowitz, Martin Casado from Andreessen, actually estimated that there's roughly an 80% chance that startups that are actually pitching “open source stacks.” So, there's the use of the word open source, but open source stacks are already running on Chinese open weight models. And so that's a pretty remarkable statistic. I haven't checked that out myself. And that's just sort of one data point. But what it really means is that there are, you know, people who are watching this industry closely, who are saying that there are tech stacks in the United States that are being widely used by startups that are developing on top of these Chinese open weight models.
Scott Caravello: So, with all that said, maybe we can hop into some of the innovative engineering that is behind this efficiency increase.
Katherine Forrest: Yeah, let's go ahead and do that. I've mentioned one piece, but there's more than that.
Scott Caravello: Yeah. And so I guess the other one I just call out to start is Mixture of Experts, which is used within American frontier models as well. So, I don't mean to make it sound like it's only used in the Chinese open weight models, but the analogy I'd sort of used there is a hospital because in the mixture of experts architecture, each task is routed to specialized sub networks that are most appropriate to handle the task. So, like in a hospital, you're routed to the right specialist. If your brain's the issue, you see a neurologist, but if it's your heart, you see a cardiologist advancement, but again, also part of American models. And Katherine, I know, you know, you mentioned the sparse attention, which is also used in American models. I think Opus 4.8, for example, is both sparse attention and mixture of experts. But then we can quickly go into multi head latent attention, which is actually something that DeepSeek itself pioneered. And so think about it this way, right? A model works through a long document, and while it's doing that, it has to hold on to everything that it's already seen and looking at that document in kind of like a working memory, but that memory balloons quite quickly. And So what they do is they have this technique, the multi head latent attention that compresses what it's worked through, like jotting shorthand notes instead of a word for word transcript. And so that way it can hold on to all the information it needs to give you the output you want without lugging around the full weight of everything that is processed.
Katherine Forrest: Right, exactly. And so here are some examples like Alibaba makes the Qwen, Q-W-E-N, family of models and they use something called grouped query attention, which is actually an innovation of Google's that Alibaba has adopted. And inside the model now, the Qwen models, there's like a lot of little readers that are all looking at the text at once and the text being the data that's being sort of pulled into the model. And normally each piece of data gets its own private stack of reference notes. And you know, you know, you have with grouped query attention, they are sharing a stack of reference notes instead of duplicating them. So there's sort of they're getting more from less if you will. And then there's something else that's called, and it's on the other end of things, which is called multi token prediction. And that's really on the output side. And that comes from DeepSeek and you can see it and it's V3 model and most models are writing one word at a time, like, you know, typing a letter. They might have, in terms of what they have reasoned, they might have a whole lot to say, but they're going to be doing it in a linear fashion. But with this DeepSeek version of things, the multi token prediction, I was making sure I was getting it right. They actually draft several words ahead. They sort of jump ahead in a single move. And so rather than doing next word prediction, so to speak, they're doing multiple word prediction and jumping over words, which speeds things up. And it does that, you know, really quite considerably.
Scott Caravello: So then to end this lightning round of innovative engineering techniques, we'll take it back towards training with the Muon clip technique from Moon Shot AI. And so that's a training technique that does two things at once. It helps the model learn more from less data, and it also keeps the very long training process from going off the rails part way through. That's because training a big model is an unstable process, right? It takes so much to do, and things can spiral and the training run itself fails. So this technique actually helps the model both learn efficiently and finish the job reliably.
Katherine Forrest: Right. And so let's just talk about a few of the models that are using some of these techniques because there's a lot of innovation that's happening right now in these Chinese open weight models. And I was actually just talking to a client this morning about one of them. And one of them that's drawn a lot of attention in Silicon Valley, but really all over the world is something called GLM—G, as in George-L-M, as in Man 5.2. And it was released just a couple of weeks ago, June 16th by Zhipu AI, which is also known as Z, capital Z-dot-little, A-I, which grew out of Shanghai University's knowledge and engineering group. And that is Shanghai University is one of, in my view, this is the Katherine Forest view, one of the most sort of advanced computer science and AI engineering groups in the world. So it's an extraordinary model.
Scott Caravello: Totally. And so for just a little bit of detail on the GLM 5.2, it has 744 billion parameters, uses that sparse attention mechanism that we discussed for efficiency and it's really focused on long running agentic work. So it's not just answering a question and stopping, but instead it's reasoning between steps, it's using tools, and then it carries out these multi step tasks from start to finish, all while rivaling the capability of leading US frontier models. But what I think might be interesting to talk about with GLM 5.2 Katherine is its 1,000,000 token context window.
Katherine Forrest: Right. And before we get there, I want to talk about the what the phrase “long running agentic work” actually means just to sort of, you know, refresh our listeners. And so when you're talking about autonomous agentic AI, one of the things that people are really looking at now and measuring is how long, literally like duration, how long can an AI agent perform before it needs human input, human direction. And you know, you can actually require the agent to come back and get human direction at certain points, at certain tasks and other things. But sometimes you just want the model to continue to work. So when we're talking about long running agentic work, we're talking about the ability to keep on going without human intervention. And that is now something that is expanded from what was earlier minutes to now days and even weeks. But the context window that you mentioned is, and we've talked about this before, it's essentially, you know, when you're entering a query, you're entering sort of your little prompt in the prompt little area, that little area that you're typing into, for instance, is considered to be the context window. And what the measure of a context window is, is how much text can a model actually hold in that, what is really a working window, at once? And so if you're talking about a million tokens, it's roughly 750,000 words. So what does that mean? That means that you're not only, for instance, typing in instructions, would you please do this XY or Z, but you can actually upload a variety and a huge amount of documentation and of data. So for instance, if you've got a legal use case like in merger and acquisition data room that's being used for due diligence, you might be able to, in a single prompt, if you will, with that big a context window, be able to upload and have the model actually grapple with a huge amount of data from that data room. So, what's notable is that the leading closed source models, which are things like, you know, the Anthropic models, the Open AI models, et cetera, like Fable 5 and Opus 4.8, they have the same sized context window as the GLM 5.2. So. the leading open weight models, at least you know this, the GLM 5.2 is on par with the American closed sourced frontier models, you know, when you're looking at this particular dimension. So, you know, earlier actually Alibaba had had a model, that's the Qwen model, it actually hit a 1,000,000 token mark as well. So it's not just the GLM, it's the several of them.
Scott Caravello: That's actually a very good segue to sort of briefly touch on some of the other popular open weight models like Qwen. So that's Alibaba's Qwen family, which is now at the 3.7 version. And that's actually the most downloaded model family on Hugging Face, which is the Public Library where you can go on the Internet and developers download and share open models and that has over 700 million downloads worldwide. And that model uses efficient mixture of experts, architecture and excels at multilingual tasks.
Katherine Forrest: Right. And then after we've gone through the Qwen 3.7, we can look now at, you know, the DeepSeek V4 preview. So remember we talked about the R1 and now we're talking about Deep Seeks V4 preview, which you know, we've discussed it before, so we're not going to rehash it, but it does remain the asserted cost efficiency leader and is the first model to actually reason. While conducting multi-step tool use during workflows. So being able to actually use tools and conduct reasoning during a workflow is sort of a simultaneous utilization of its abilities. It's pretty impressive.
Scott Caravello: And so I know we're sort of running short on time, but I think that really to sort of compliment this technical part of the discussion, just add a little bit more context on the cost savings beyond just the fact that we're talking about these incredible engineering techniques that are helping the models run more efficiently. Because there's actually a second component to that. And that's just the nature of that. These are downloadable models that can be hosted within a business or other entities like own environment. And so there you're only paying really for the cost of hosting that model and for the actual compute use. There's no extra margin baked in for a model provider to recoup running the model on your behalf. So that's complementing the efficiency gains that people are getting from the architectural choices that the model developers have used in creating and offering the models.
Katherine Forrest: I think you've got some stats on some of the potential costs that follow right on what you just said.
Scott Caravello: Yeah, for sure. So just level set really quickly and get down to brass tacks. Opus 4.8, for example, cost roughly $5 per million input tokens and $25 for output tokens. So we can just use that as a baseline. The flagship Alibaba Qwen model, which comes in as potentially one of the cheapest among his Chinese peers, is around $0.40 per million tokens for input and $2.40 for output. Moon Shot AI’s Kimi. K 2.6 is $0.95 for input and $4.00 for output, again per million tokens. And K 2.5 is $0.60 for input and $3 for output. We've talked about DeepSeek before, preview a while ago, but there we're at $1.74 for input and $3.48 for output. And then finally, the GLM 5.2 which has gained so much attention as of late is roughly $1.40 per million input tokens and $4.40 per million output tokens.
Scott Caravello: So really…
Katherine Forrest: Right. All of these are per million output tokens, right? So even when we're talking about clawed, it's not $25 per token, that's $25 per million output tokens.
Scott Caravello: But so these are landing at about 1/3 to a 12th of the input price in 1/6 to 1/10 of the output price for the closed source Frontier model. So that that's significant savings.
Katherine Forrest: It really is.
Katherine Forrest: And so we're going to see, I think a lot as the pricing models are starting to change, which is a whole different sort of episode that you and I are going to have to do is the pricing models are starting to change. We're going to see I think some developers really sort of comparing some pros and cons of these open weight models and certain capabilities that they have compared to the American or other models in different parts of the world. But one thing that is interesting is we are having something that you could think of as and the press has talked about it like this a little bit as like a silicon curtain coming down where we've got this difference now and a geopolitical overlay on top of all of this. Because we have the USA that's in a self-stated sort of race for super intelligence against primarily China. There are lots of other places in the world where there are models being developed and worked on. And certainly with you know, when you've got Mistral in France and you've got a number of models in the UAE etcetera, etcetera all over the world, but you've got the US in this self-stated race for super intelligence. But we also have, at the same time, capabilities that are developing and being used in different parts of the world while there's this race going on. And there are increasing numbers of questions about whether or not the US might pull a model, whether China's actually giving X China access, you know, everybody access to the best of the Chinese models, or whether or not what's leaving China are in fact not quite the best of the best. And so there are going to be some developments that we're going to see about how the geopolitics of this whole thing plays out. But I think, Scott, that's about all we've got time for today. You know, I wish you luck with the air quality down there. And if you want to, you guys can just jump on a plane and Portland's only like a 35 minutes flight.
Scott Caravello: Thank you. Yeah, I actually dug out an old box of N95s from the COVID days, unopened, you know, pristine condition. So maybe walking around the streets today, the throwback to COVID era with a mask on. We'll see.
Katherine Forrest: Oh, you know what, I just was watching yesterday. I was watching The Four Seasons and I don't know if you've seen that show it's, you know, on…
Scott Caravello: Yeah, of course.
Katherine Forrest: Yeah. OK, so my sister is a first AD and she did, and actually has a producer credit on Italy. She just did the two episodes in Italy for this past Four Seasons and she's going to be working on the next one. So. But anyway, you got to watch The Four Seasons. So yesterday I was watching The Four Seasons and I'm not yet at her, at Bellamy's, Italy episodes yet. That comes next. But there's a whole thing on Thanksgiving during COVID in The Four Seasons. If you haven't seen this season, have you seen this season?
Scott Caravello: I have.
Katherine Forrest: Ok. And this was like, “Oh my gosh…we all lived through that.” We were all like double masking up. And like, I mean, it was wild to actually watch it and think #1 how recently all that happened and boy, what a strange, strange time that was. So pulling out that N95, boy, I feel for you. I haven't had to do that.
Scott Caravello: To be totally honest, I actually did it a couple weeks ago because the building where my gym is was like the epicenter of the Legionnaires outbreak on the Upper East Side.
Katherine Forrest: Woah. Wait hold on, so you like went into the gym and used your N95?
Scott Caravello: No, no, I, I they said the gym was fine so I wore it on the way to the gym, walk in the gym and then took it off. But…
Katherine Forrest: Wait, hold on, your building is like the epicenter of the Legionnaires?
Scott Caravello: The gym building. The building next to it had a cooling tower that was spewing off the Legionnaires.
Katherine Forrest: OK, you can go use my gym all right if you need to.
Scott Caravello: Crazy days. Crazy days.
Katherine Forrest: Crazy days. All right, folks. Well, that's all for today. I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello, don't forget to like and subscribe.
Katherine Forrest: We'll see y'all next week.