Liang Wengfeng's Long-Form Conversation: Why Does DeepSeek Insist on Open Source and Restraint? (Full Transcript of the Investor Exchange Meeting)
At an investor exchange meeting, Liang Wengfeng discussed the company's vision, open source and restraint, the AGI roadmap, commercialization pace, c…

梁文锋
Welcome, investors. When we first started this company, our original intention was not to think about how much money I would ultimately make, or whether we would go to the capital markets, go public, or anything like that. So we did not start with that intention. The first few dozen people who joined us never thought that way at all. If they had, they would not have come.
So overall, we were doing this with a great deal of goodwill toward the world, and we felt that this was useful to humanity, something beyond money. Of course, later on, after we found that the benefits of this thing were very large, there were other temptations too; that is another matter. But our original intention, our vision, and the vision we have maintained to this day were not based on maximizing commercial interests.
I think this point is quite important. About twenty years ago, the person I admired most in management was Jack Welch, the former CEO of GE. Looking back now, most of what he said may have been wrong, but the most important thing he got right was this: the most important thing for a company is its vision. What does it take to manage a large company? Not your rules and regulations, but your vision. What is vision?
Vision is not a slogan on the wall; vision is how you do things, not how you talk about them, that is, how you actually operate. Anyway, I forgot Jack Welch’s exact words, but that is roughly what he meant. So how do we manage so many people and organize them? In fact, we do not organize in the conventional sense; we are vision-driven, organized around a vision. We do not have organization in the traditional sense. That has both advantages and disadvantages.
In the future, we will try to capitalize on our strengths and avoid our weaknesses, but that is our distinctive feature. We are not doing this in a way that says, “I need to achieve some KPI, there is no assessment,” only a vision. This vision is not even written down; it was never written out as any formal document. This vision exists in the way we do things and in our attitude toward the world.
Maybe each person in our company understands this vision differently, and maybe each person’s vision is different too, but at a broad level they are consistent.
I think it still comes down to having a great deal of goodwill toward the world and wanting to do something. That is what we use to bring ourselves together. Next I’ll talk first, and after I’m done everyone can ask questions. I may continue around this vision as I talk about the rest. This vision is real, not made up; we truly think this way and truly act this way. Otherwise you cannot explain many of the things we do. Why are we so committed to open source?
Because this vision itself requires open source. Without this vision, you cannot organize people. For example, also open-sources, but 's open source is different from ours. **'s open source feels somewhat forced; they feel it is not their original intent. But for us, this is our original intent. And on open source, we were very clear from the very beginning.
First, the vision; second, we believe that to make AI commercially successful, open source has benefits. That may sound a bit contradictory, a bit counterintuitive, because historically open source and commercialization have been at odds. But I think AI is different from before.
Historically, a software company’s market might have been only a few billion dollars a year; if it went open source, it would disappear, maybe leaving only a few tens or hundreds of millions of dollars. But AI is big enough that, in the end, it may account for ten percent of human society’s GDP, for example. That is actually a very large number.
One person cannot monopolize this thing; you cannot monopolize it. You have to share it with others, otherwise you definitely won’t survive. This is different from open-sourcing a piece of software before, because that software market was not nearly as large. But AI is simply too big. If we want to monopolize the benefits, then we are certainly bound to be left behind by history. I think this is primarily an objective law, a historical perspective.
It is not that if I do not open source, I can monopolize the market; that does not conform to objective reality in theory. You will definitely encounter a lot of resistance, and there will definitely be other ways to prevent you from achieving that goal. In this situation, I think you do not necessarily need to follow traditional commercial thinking. You need a mechanism to ensure that the benefits you can obtain are limited; only then is it possible to succeed. Restraint is necessary, I think restraint is necessary.
If we want to make AI succeed in our hands, first of all, I think restraint is needed. We cannot think that a certain percentage of human GDP belongs to me, or that a certain percentage of China’s GDP belongs to me. The more you think that way, the less likely you are to succeed.
So from the very beginning we felt that restraint was necessary. The more restrained you are, the more likely you are to make it. This is a commercial consideration, of course, but it is also a macro-level consideration. I think this is intuitive, at least it is intuitive to me, or at least this is truly what I think. We do not have many other advantages; we are not especially capable, we are not richer than others, and our people are not better than those at other companies either. Actually, not at all.
Think about it: when we founded this company two years ago, we did not have much money, we did not have many chips, we did not have much visibility, and we did not have much influence. We were just a group of very ordinary people.
We really were just a group of ordinary people. If the preferred narrative is that a group of ordinary people did extraordinary things, rather than a group of geniuses doing extraordinary things, that is very much related to our restraint, and is in the same vein as our restraint and our vision. So, will open source and commercialization conflict? I think in AI, if you are not restrained, you won’t get off the ground.
Open source is part of restraint, and our restraint is not only reflected in open source; it is also reflected in many other aspects. But overall, we do not need to dwell on open source or restraint. The more restrained you are, the easier it may be to succeed, or at least so far that has been borne out; so far it is explainable.
Otherwise, there is no way to explain how we could succeed: we do not have any special weapons, our starting point was very low, our resources were very limited, and our people were actually just a random group of ordinary people. I myself am just a university graduate, not from the very top school. This restraint is also part of our vision. AI is too big; the benefits are too large.
We are very restrained. As long as we can make it succeed, the final benefits will be very large. Even if you only take a small slice, the benefits are already very large, so there is no need at all to think about which part of the benefits to take, or how to take them. I think there is no need to think about that at all, because the benefits are large enough already. As long as you take a tiny bit, that is more than enough. So when we previously said that we only earn a reasonable profit, we are only looking at your willingness, not at maximizing profits—that is different.
This is not our API pricing. Our API pricing is based on what we consider a reasonable profit: roughly enough to buy a batch of equipment from the market and recover the cost in ten months. I think that is a reasonable profit.
Under current circumstances, considering risk, upfront investment, and so on, if a server is depreciated financially over three or five years, but commercially we think about ten months to recoup the cost, we think that is enough. OK, we think that is enough. That is the logic behind our current API pricing. Our V3.2 Flash and the others are all priced so that the equipment cost is recouped in ten months.
That is our standard. It is not actually profit maximization; if it were profit maximization, we should set the price higher. Because in this price range, user demand is inelastic: if I raise the price by half again, or even double it, the change in token consumption would not be that different. If I make it twice as expensive, my total revenue is close to doubling. Wait a second, let me check. Ah, that’s great.
Let me tell you a story: one of our models. At first, we were worried about too much demand, so we set the price relatively high in the beginning, and everyone on the team was not very happy. Later I lowered the price again, down to one quarter, and everyone was very happy. I think that reflects what we really think.
What I said earlier about the vision is that we still want this thing to be useful to people, not to make the most money, but to let everyone use it affordably while still earning a reasonable profit. I think, in terms of our company’s other people’s thinking, when we lowered the price then, many people in the company group were cheering; everyone was very happy. Because that is the purpose of all the effort and care we put into making this model good.
The purpose is to make it very cheap and very effective, so everyone can use it fully. We felt very happy about that. That is our motivation, that is our vision, and that is the consensus within our company about why we can come together and do this. I think this is a relatively special point, because for our competitors, lowering prices definitely is not a good thing; they certainly would not be cheering.
Because your revenue, your ARR, if you cut it in half, then ARR drops by half. Right, that is one difference between us and others. We feel that is enough. From within the company, recouping costs in ten months is already very satisfying commercially. Externally, we also feel this price makes people happier and more willing to use it; it is a win-win, for the company and for society and everyone.
I think, OK, someone just left a message on the screen saying that recovering costs in ten months is too high a profit. It is true that there is still room to cut prices; there is indeed still room to cut prices. There is also room for optimization in the model, so overall there is still quite a lot of room to lower costs. But this cost, recouping in ten months, is something we can achieve ourselves; other companies cannot. For example, Alibaba or Tencent are not our optimized setup; their costs should be several times higher than this.
There is still a lot of optimization work here. Just now you asked why we do not continue lowering prices. It is because there is no elasticity. That is, if I lower the price again, demand will not increase further, or if I lower it again, demand will only increase a little. Because at this price everyone can afford it, and everyone feels the price is satisfactory; they won’t stop using it because it is too expensive.
So lowering the price would not bring the company more revenue, and it would not create more value for society either, because everyone is already satisfied with this price. If the price gets even lower, people’s happiness in society would not increase much either. Right, OK, but on this question just now, when it comes to pricing, we definitely do not start from
maximizing company revenue or profit. That is not the starting point. This is part of our restraint, because in the short term, if you price a bit higher, you may earn a bit more; but in the long run, it is not easy to say. Because I think restraint is a strategy. For me, restraint is a strategy. Sometimes you can give up some things in exchange for more of other things.
Open source is actually the same; it can also be seen as pressure on us, or as a concession on our part. First, this concession, from within the company, makes us very happy; everyone is happy, employees feel a strong sense of accomplishment, and we become more cohesive because of it. And this concession is beneficial to society; society is happy too, and other peers, or ordinary people, will all be happy.
So I understand this kind of restraint as something that, in the long run, can increase our probability of succeeding in AGI. When considering something, I have no doubt that AGI will have very large commercial value. On that basis, what I prioritize is not how I can take a bigger share, how I can take more of the pie; what I prioritize is how I can increase the probability that I can actually make it. This restraint may also be reflected in many other aspects.
For example, last year during Spring Festival, we suddenly had many users, but we did not pursue keeping those users, or monetizing them, or trying to seize commercial interests from them. We did not try to take users or make money; instead, we worked very hard to serve those users well.
We would not have thoughts like, “I’m going to make the next super app,” or “I’m going to compete with someone,” or “I’m going to build the next ByteDance or the next Tencent.” Absolutely not. We could do that, but we did not. My understanding is that this is also part of restraint. Don’t think that you have to monetize everything; once you have users, it seems like you can make the next ByteDance, and then you just eat that up.
I think that is commercially feasible; it is possible. If last year we had spent a lot of money to fight ByteDance for users, that would also have been a possible strategy. But we chose a very restrained approach: I won’t fight you for that, because what comes later may be the watermelon, and what comes before may all be just sesame seeds.
I should not grab every sesame seed. Of course, some sesame seeds may be quite large, but I think the AI opportunities later are not small compared with the earlier ones. Looking at it now, not pushing on the consumer side last year may have been the right choice. Because it is clear that there really are bigger watermelons later, and what came before really were just small sesame seeds. If I had a lot of money last year and made this thing very big, what would have been the benefit? You would not have gained much.
These are my real thoughts, because I think the AGI opportunity later should be very large; the AGI opportunity later will always be very large.
I don’t even need to think about whether I will occupy a position in it, or what my business model will be in it; we really do not need to think about that. As long as there is such a huge business opportunity, you will definitely find a way. As for the sesame seeds in front, we will pick them up too, but we will just pick up a little along the way; we will not stop and make them an important thing. So last year’s consumer daily active users and such, I think, may just have been a small matter.
But we also picked them up, and we maintained user usage at a relatively low cost, because they may be useful later. Although right now we do not know what these users are useful for, it is currently a pure cost expense, but it may be useful later. Since it is something we can take along the way, we will take it along the way. Including this year, it is very likely that we also have an opportunity for ARR revenue in API or AI.
That is, if this demand can continue to expand, and if it continues to expand, and GPUs can be bought in greater quantities, then reaching several hundred million dollars in ARR is very possible. If AI can reach one billion dollars, then basically the cash flow of my company may be able to break even, enough to cover my R&D expenses and all my expenses.
So that is also possible, but we have not treated it as a priority. We will do it, but I think it is an important matter. It is not our first priority, and it is not what we are truly focused on today. The bigger opportunity should still be later; the earlier opportunities, including last year’s consumer side and this year’s B端, I think we need to do them, and do them well, but that is not our goal.
Or rather, most people in our company do not think this is a very important matter, and do not think it is equally important compared with AGI. Let me say a bit more about open source, since many of the questions before were about open source. First, I think we will open source, and our strongest models may also be open sourced. Because I cannot see any good reason for closed source, I cannot see any inevitable benefit.
ByteDance’s model is closed source; what benefits does it have? I cannot see any benefits. Even if a model is open source and you tell everyone everything, the barrier is still very high. It is still very hard for others to actually use it. For them to use it is difficult; secondly, to use it while also keeping costs very low is also extremely difficult. It is not easy.
It is not as if I open source and they can easily deploy it at the same cost as me. There is still a lot of work to do here. Although the principles are all understood, not every company is willing, or has the willingness and ability, to organize people and resources to reach that goal. I am also very used to this. They may simply not be good at doing this, because the resistance is too great. It is hard for them to control these costs; there are many managerial and physical constraints.
This is also an advantage of startups, because if a startup is too small, you do not have the power to do this; if you are a large company, it is hard to organize.
Every situation has its difficulties, so this belongs to the sweet spot of a company of our scale. If we were bigger, maybe we would not have other problems; if we were smaller, our strength would be insufficient. So as for open source, I think we should set pricing. At present, we should recognize that we will not force people.
Because with the pricing model, I also will not charge a very high fee; I will probably still charge based on recouping costs in ten months. Based on recouping costs in ten months, that already allows us to beat competitors... By charging based on recouping costs in ten months, we can already make third-party independent deployments unprofitable. Third parties cannot do it; they cannot achieve this cost, they definitely cannot. So open source will not affect my revenue. Of course, if I wanted to earn a hundred times profit, then open source would be...
主持人
I heard you, but the video seems to have dropped, boss.
Maybe a phone call just came in.
梁文锋
So for open source, I think it has no impact on our business model. The premise is that we only earn six times profit; recouping costs in ten months corresponds to about six times profit. If we only earn six times profit, then open source will not have much impact. But if you want to earn a hundred times profit, then open source will indeed affect your ability to earn a hundred times profit, because third parties will deploy it, and they may have a cost that is twenty times lower than yours, so it is cheaper than yours.
Is this model sustainable in the long run? I think it is. Under our vision, I think open source is sustainable in the long run, or at least that’s what we plan to do. You could say that, with restraint, there is also long-term benefit. This strategy gives us more opportunities ahead of technology, and the probability that we can achieve AGI is greater. We are more composed.
Think about it: we don’t even need to work overtime, because it just isn’t that hard. But for other companies, it may be very hard, because they think about too many things. In fact, it’s not hard at all, really not hard. At the beginning it was a ... From the outside, it may look like we chose a very hard mode: we do research, we go after the hardest things, as if it were a hard mode.
But in fact, we have given up a lot elsewhere, which still leaves us very powerful, and lets us do things very easily. So my judgment about open source is that it is sustainable. There is no conflict between open source and commercial paid offerings, provided that there is no conflict under a six-times profit margin. Six-times profit sounds high, but it really isn’t. Because given how efficient AI is now, a reasonable profit today may be about that much.
In the future it may drop, for example, to four times or three times. I think that already ... it really can’t go any lower. But even then, there will still be very large profits. Just looking at selling API, I don’t think selling API is that attractive. But on this point, you can say there is no conflict; I haven’t seen any conflict.
And I’m not worried at all about others deploying our model and competing with us. We actually hope they can deploy it. We try our best to provide help to the open source community and assist everyone in deploying our model. I’m not worried that he will take this business from me, because the market is big enough. I only worry that he can’t deploy it, that some details are not done right, that the results become worse, or that his costs become relatively high.
Right, there is no conflict here. Then last year, when I was being asked about the To B business, one of the more common questions was: if I open source our C端, then will C端 conflict with my C端 again?
Because I don’t have a traffic advantage, or rather, Tencent itself has a lot of traffic. If it deploys our open-source model, it will take over all the C端 users, and then it will steal all my C端 users. But in fact that won’t happen; there are many reasons for that. And the question is: is the open-source model we provide the same as the model we deploy ourselves? Yes, it is the same.
We won’t open source a worse model and then use a better one ourselves when we deploy it. We won’t do that; it’s the same. That also shows there is actually no conflict. Last year, for the whole year, in C端 I basically just open sourced, and I didn’t see any conflict in C端 services. I really didn’t see any conflict. So that’s the open-source part.
Then below, there is also the company’s long-term vision. I think our goal should be AGI. Everyone’s definition of AI may not be the same, but that does not stop us from taking AGI as our goal. From the technical route, the roadmap to AGI is actually quite clear.
With the current generation of AI technology, if you can describe a problem very clearly and give it complete context and instructions, it already surpasses humans. But there is a definition, a premise: you give it complete context, you give it complete instructions. And that premise is very hard to achieve.
For example, in a meeting today, we actually have a very long and very strong context ahead of us—maybe everyone has decades of context—and that is something AI does not have. What AI can have now is that, within a limited context, it can do better than humans. But it still cannot replace humans. What is still missing is continuous learning. Because humans can also keep learning.
You hire an employee, and he may spend two months getting familiar with the company environment and his job. After two months, he can get started. He can do many things. He can understand what you say; for example, if you say, call Xiao Wang over, he knows who Xiao Wang is. But for AI, because it lacks that context, it has not gone through those two months of learning.
If you tell it, bring Xiao Wang over, you have to tell it who Xiao Wang is, what position he holds, where he is, how to find him, and what to pay attention to when finding him. You have to give AI all the context. In that case, AI can do it, but you can’t possibly give it all the context, and that’s unrealistic. So AI cannot replace your employees.
But if AI has continuous learning ability, like your employee, and learns at the company for two months, then
it can replace everyone under heaven, so we are still one step away from the next stage: learning to learn. We can understand AI development as a staircase. The staircase we climbed last year was CoT, chain-of-thought. Because we discovered that through chain-of-thought, intelligence can reach a higher level. By letting it think for itself, we can raise the ceiling and let AI do more things. So we crossed another staircase.
This year’s staircase is Agent, because we found that using Agent, even with many things to do, it can still do them; its capability range becomes larger and its intelligence ceiling higher. Why is it a staircase? Because every step after that is built on the previous foundation. Agent needs CoT, and CoT also needs the earlier staircase, which is the language model, so none of the steps are wasted.
So the development of AI, the direction of intelligence, is traceable. This year’s staircase is Agent, but the Agent staircase will also eventually be completed. Once it has solved all the problems it possibly can, it still cannot replace your employees, but it has already reached the limit of its capability.
It’s like CoT: once CoT reaches its ceiling, it has already surpassed the most top-tier humans, in solving Olympiad math problems and writing programs. But it still stops there; that technology has not reached AGI. So you see, the direction of AI intelligence is traceable.
Then after Agent, we think the problem that should be solved is continuous learning—how to let the model keep learning continuously, rather than requiring a very strong training setup. It should be able, like a human, to undergo relatively long-term continuous learning. This issue, together with task completion and so on, is the same thing; they are related and solve the same problem.
Right now, standing at Agent, what we can see is the next bottleneck: continuous learning. The next problem to solve is how to achieve continuous learning. That is visible, and relatively clear. It is the obstacle in front of us; you have to cross it, and there must be a way to cross it, but it takes time. After continuous learning, we may reach a singularity.
That singularity is when, once the model can continuously learn, it can already do everything humans can do. It can develop its own versions, conduct research on its own, and then develop its next version, developing more advanced artificial intelligence models. So it will reach a singularity and be able to iterate itself. But this singularity is not really a singularity; it is also a gradual process.
This process may also be a relatively long gradual change; it is not a sudden change. But by habit, we all think it may be a singularity. Because long ago, the prophets thought there would be a singularity here, but in fact it is not a singularity; it is a continuous process. And after this step is completed, I think that is when embodied intelligence comes.
This is our speculation. We think the timetable should be: first solve learning to learn, then reach that singularity of self-iteration, and only then embodied intelligence. After embodied intelligence, it enters the real world and can do housework for you and take care of you in old age. We think this is a relatively ideal roadmap, but everyone has a different view, and there is no right or wrong. We just think this roadmap is the easiest.
This roadmap is easy because you have to do very little new work at each step. With this roadmap, we don’t need to work overtime. But if the roadmap is reversed, for example, if embodied intelligence has to be achieved first, then doing it oneself would be very exhausting; it would be very hard work. We don’t want such a roadmap; we want to do it a bit more easily.
If we first solve continuous learning, then solve that singularity of self-iteration, and then solve embodied intelligence, the road will be very easy. Because later on, you can use earlier technologies to help develop later technologies. After the singularity, embodied intelligence doesn’t need to be done by people, it doesn’t need us to do it; the model itself can emerge. So that is the answer to what our long-term goal is.
I told him, this is our long-term goal, which is what we call AGI. Let’s come back to reality. Last year’s most important reality was that everyone had to do Chatbot and compete for C端 traffic. This year’s reality is that everyone has to compete for To B revenue and get involved here, because if you don’t get involved, you’re not even at the table, right?
But we don’t think that is an important thing, or rather, inside our company, what we truly care about is the AGI roadmap I just mentioned and how the next technological breakthrough will happen. But one very strange thing is that the thing you most want to get is often exactly the thing you can’t get. The things you don’t care about that much, on the other hand, are often quite easy to obtain.
There is a strategic advantage here: in our minds we are thinking about AGI, and what we are doing is AGI. Then when we do applications, or C端 and B端, we don’t need to put much thought into those things at all; in fact, very little effort is enough. I think that when you stand at a higher technical level and work on relatively lower-level technology, there is a kind of dimensionality reduction attack.
At least in last year’s C端, we saw that this was indeed the case. We did not spend much effort on C端, and at one point we even didn’t want to maintain those users, but the users could not be chased away. Because they really could not be chased away, they all stayed in the end. But a little ... And this year, the B端 revenue now looks relatively optimistic in terms of growth. I think this number may also be pretty good compared with peers, I guess.
But we did not spend a lot of effort on this. We basically did not
make this happen deliberately; it was something we did on the side. When building internet intelligence and moving toward an online launch, the AGI step is a step I must take. On the road to AGI, I have to pass through this step. Then I provide all these technologies to everyone via API; I am not doing any extra work. We are still doing AI; this is a byproduct.
I only need to have a few people maintain this API, and there is not even any customer service; no sales are needed; nothing is needed, and users will come on their own.
Or rather, the users considered to be C端—both C端 and B端—are byproducts of our journey toward AGI. They are all intermediate outputs and do not conflict with my AGI work. I am not doing C端, or B端, for the sake of doing them. Rather, we are doing AGI for the sake of AGI, and it just happens to produce these things, so I take them and use them for commercialization. This is different from other companies.
Other companies do this for the sake of this—either to serve C端 users or to serve B端 users, and then they build the model. But for us, that was not the original intention; our original intention is still to pursue AGI. I think, to some extent, this is a kind of dimensionality reduction attack. AGI is a bigger vision, and that vision can bring together more excellent people; it has stronger cohesion.
So I have an advantage in organization, and I use that advantage to ... It is a kind of dimensionality reduction attack. But if you are a commercial company and your vision is to serve C端 users well, then that is another story. It has other advantages, with strengths in product, user service, and traffic, but not in technology. The current favorable situation is that model technology is the most important.
You need to make the model good, and the rest ... Right, and this can explain our previous development path. We really chose AGI, and I never thought about making a lot of users. When we suddenly became popular during Spring Festival last year, that was not in our script at all; we had never thought about that. We just wanted to make the technology good.
But back then I found that, compared with organizations that are fully commercialized and product-oriented, our organization actually had extra advantages in talent and organization. That was also very magical. At that time, everyone in C端 was fighting tooth and nail, and in the end they were pulled away by someone who had not gone to compete. That also really shows that what I said earlier makes sense: there are indeed advantages in talent and organization.
This talent advantage does not mean my people are smarter than theirs, but rather how I organize these talents, how I motivate them, and how they cooperate. That is an advantage. Because bringing smart people together does not automatically mean they will cooperate naturally, or naturally have the passion to run toward a goal and complete it, so you need a vision. What our previous experience taught me is that the AGI vision is very powerful.
OK, that’s the question. Then the next question is, what is the importance of core interests?
I said earlier that in many areas we need to be very restrained. But what are our core interests? Actually, we have only one core interest: our biggest core interest is to maintain the stability of the team. This is our biggest core interest, and you could even say it is the only core interest.
As long as I can maintain team stability, I will definitely succeed, definitely achieve AGI, that’s how simple it is. As long as everyone stays, and we can keep going, then I will definitely ... basically there is not much risk. It’s just that sooner or later there will be setbacks; if there are setbacks and everyone still stays, then I can continue. Money is definitely not a problem, resources are not a problem, and all the other factors are easy to obtain.
For us, there is only one core interest, only one thing we cannot compromise on: we must maintain team stability. This is also the very big challenge we face, or rather, I think it is the biggest risk. Of course, with our recent financing, this risk has been greatly reduced. Because everyone’s options are still relatively large, and the amounts are still relatively substantial.
From the perspective of team stability, as long as some of the most important employees and the oldest employees remain stable, the others are not likely to leave. Even if others have a little less equity or a little less income, they won’t leave. Because they are not coming for the money alone; everyone hopes to work in an environment where AGI can actually be achieved. So for talent, it is still attractive.
Historically, our talent turnover has been relatively low. Compared with peers, our talent turnover is always relatively low. But this is still our biggest challenge, and can be said to be the only challenge. The rest are all matters of time; at most they would make us six months or a year late, but they would not make us unable to do it. We definitely are not short of money, and we definitely are not short of resources. In fact, none of these are lacking.
So many of the things we are doing now are to maintain team stability. Aside from that, I think we can give up on everything else; we can restrain ourselves on everything else. We have always been very restrained, and we do not want to become adversaries of any large or small internet company. I hope I can empower them, or help everyone do this, and hope I can assist everyone in doing this.
This is also part of the commercial meaning we talked about earlier. The premise is that everyone should not ...
Under that premise, we are very willing to assist and help anyone, even our competitors, including Alibaba, Zhipu, and Moonshot AI, to do better. Because we do not lose anything; we were open source to begin with, and open source also does not try to make the boundaries as clear as possible. As for how to do it, we hope you can reproduce it; if you can’t reproduce it, say so, and I’ll tell you how to reproduce it. That is originally part of open source, and it won’t change just because you are a competitor.
Of course, if it is a partner, I will do more. But at the level of big interests, there is no conflict.
When dealing with the outside world, our attitude is: we only do the main line of AGI. That is what I just said about GPT, CoT, Agent, and so on—we only do the main line. The AI field is very broad, and there are many things that we think are not on this main line, such as 3D and video generation. I think they may not have much to do with the main line of intelligence, so we will not do them.
There are also some things, such as world models. I think they currently do not have much to do with the upper limit of intelligence either, so we will not do them. But if others do them, we are also very happy to help. Whether we have time is one thing, but there is no conflict in interests. We also hope that these AI technologies can be used in all kinds of production environments, can improve social productivity, and can help all industries improve productivity.
We are very motivated to do this. Whether I have time, whether I have enough people, or whether our teammates themselves are interested—that is another matter. But there is no conflict in interests; we hope to achieve this goal. And we believe this has no conflict with business at all; the benefits I should get have not been reduced at all.
I think that, in the attitude we held before, we actually didn’t lose anything because of it. We didn’t lose anything because I open-sourced things, or because of our goodwill, or because I helped other people. For example, last year’s C-end users — we still have quite a lot of C-end users now, and they are still relatively stable. This year’s B-end, I think, is also relatively optimistic.
I have not hurt my business interests because of our goodwill at all; absolutely not. On the contrary, it may even have added points. This seems counterintuitive, but it really is like that. Or, if we think about it the other way around, even if we violated it, would it necessarily mean I could get more? No. There is a question: how should we understand that world models have nothing to do with raising the upper limit of AI intelligence? What I’m saying is that, at this stage, this is our judgment.
We do have a roadmap; this is our own AI roadmap, not the only roadmap. From our understanding and judgment, what matters most right now is doing AI training well. Doing AI training well does not require world models, and does not even require multimodality. Because if you narrow the scope of AI training a bit, without multimodality, you just won’t be able to do some tasks, but that does not affect the validity of the algorithm. Multimodality will ultimately still need to be done.
What matters now is training, and then the next step is to solve the problem of continual learning, and then the next step after that is getting it to ask questions on its own. But this roadmap does not include world models, and does not include video generation. When video generation first came out, it was very hot, as if it was something that had to be done; if you didn’t do it, you didn’t seem like an AI company.
So I found that very strange. Actually, if you think about it carefully, it has nothing to do with the roadmap to intelligence.
In fact, you also saw that after Sora came out at the beginning of video generation, everyone did it — big companies and small companies all did it. But small companies later cut it all off. It has nothing to do with the upper limit of intelligence. But commercially, it is a good business; commercially, it is a good business. But it has nothing to do with intelligence. We would not do it just because it is a good business; we would only do it if it is on the intelligence roadmap.
Video generation is relatively clear, so I’ll use it as an example. As for world models, the meaning of world models is not so clear, because many things can be said to be world models. From our judgment, world models and intelligence are not yet the most important things at this stage. What is most important is AI training, and then after AI training, how to solve continual learning.
This is our company’s judgment. Of course, every company has a different judgment. What I just said is that, for our company, the most important issue is staff stability. From another dimension, what do we lack? What is the gap between us and the United States? Actually, there is only one gap: resources. We don’t have that many cards; the number of our cards is still relatively small.
We currently have about 20,000 H-equivalent compute. Most of this only just arrived; it came in the last one or two months, and there may still be many machines that have not arrived yet. Our total compute was relatively small last year; this year we are expanding compute very aggressively. We are now at about 20,000 H-equivalent, and over the next few months we will also buy a large batch of machines, basically all NVIDIA.
How many cards do we need? Right now, of course, the more the better. Within the range we can afford, the more cards the better, without question. So our strategy now is: within a reasonable price, buy as many cards as we can. If this financing is used up, however many cards I can buy, I will buy that many. The speed of spending is not in the plan; rather, as long as the price is reasonable, I will buy however many there are.
So if I spend all the money within half a year, I think that would be a good thing. If I spend all the money within half a year, then that would be very happy, very ideal. In fact, it is very difficult to spend so much money; you can’t buy that many cards. It’s hard to buy them, and the prices are high. You also can’t buy them at an extremely high price; you still have to make sure the price is reasonable.
If I can spend all the money within half a year, that may be the most ideal outcome. Because turning my money into NVIDIA cards is definitely better than leaving it in the bank. If it stays in the bank, I’m probably only getting around 2%, it seems. But buying NVIDIA cards is a ten-month social cost.
So of course, you buy as many as you can. If you buy cards first, then later I’ll have plenty of room. By providing services or doing whatever, I can always have cash flow. If I have cash flow, I can survive; I don’t need to keep a lot of cash on my balance sheet. So what we worry about is only not being able to buy that many cards.
If we can turn all the money into cards, then we will without hesitation turn all the money into cards, and in doing so we are willing to pay a certain premium. We are willing to pay a certain premium to turn it into cards, because that is too cost-effective. Even after paying the premium, in fact, it is still hard for us to achieve this goal. So objectively speaking, if I can spend 20 billion this year, then our procurement team will have performed super well.
The gap between us and the United States is mainly in resources, and the gap in people is not very big. There is almost no gap in people, because it’s basically the same batch of people, probably Chinese people. When Chinese people go abroad, some stay in China, some stay abroad, and some go abroad; it is not that the smart people go abroad. No. It is actually quite random.
The smartest people may not be more than half going abroad, and slightly less than half staying in China. China does not lack talent, and our base is large; every year we have so many new people coming in. Talent is not the bottleneck; resources are the biggest bottleneck. Resources first affect talent cultivation, because with less compute, we have fewer opportunities to do experiments, so our talent as a whole is behind the United States. The talent gap, in essence, is also due to the compute gap.
At the largest models right now, we actually can’t afford to train them. Even if we spent the full 50 billion, we still couldn’t train them. Even if we could assemble them, we couldn’t afford to use them. The largest model today has about 800B activations; domestically, we are still at the tens of B scale. The largest domestic model may only have tens of B activations, so we are off by an order of magnitude.
If I want to train a model as large as AI, I would probably need 50,000 GB300s, or Huawei 950, 200,000 cards. And that is only for training; it doesn’t even include research. So the biggest gap between us and the United States is in resources. Our current resources, plus the resources we will have within this year and over the next few months, including the big resources that are coming very soon, are only enough for us to do more experiments at the 10B activation scale.
Because within the tens of B activation scale, there are still many experiments I need to do, and many things that need to be figured out. We are still quite far from being able to train an 800B model; there is still a lot of time, and not that many cards. So the difference between us and the United States, I think, is a difference in resources.
We may think that all the differences we see — including differences in talent, differences in model capability, and differences in applications — can all be attributed to differences in compute resources.
Compute resources are, on one hand, because cards simply cannot be bought domestically; on the other hand, our capital investment is smaller than that of the United States. We have much less capital invested. In terms of capital, the share occupied by talent salaries here is very low. You see salaries like 100 million US dollars, but when you calculate it, talent salaries are still only a small proportion; the main cost is compute. This issue is basically unsolvable at the moment, because Huawei’s output is also limited.
Because if I want to train 800B, I need 200,000 of Huawei’s latest cards; and that is only for training, without even considering research. So right now we simply won’t consider competing with the United States at such a large scale. What we are losing now is still at the scale we can afford to train and use — the tens of B activation scale — and we need to do that well first.
Then when we have more resources in the next step, we can bring it up to the 150B, 156B, or 250B activation scale. So between us and the United States there is currently a gap that seems very hard to close. If you insist on training such a large model, it can be trained, but you can’t do sufficient research. That is, before training, you can’t do sufficient research.
Another issue that everyone is quite concerned about is: in all large-model competition, where will the final gap be reflected? That is, when large models finally diverge, where will the difference show up? I think that difference may ultimately not be very large. The final differences should be in three aspects: cost, time, and user experience. Beyond that, there may not be much difference.
Cost is easy to understand: if you provide the same service, with the same quality, at what cost can you provide it? The same service, to take BYD batteries as an example, under the same level of technology, can other companies provide it at that price? I think that is a difficult thing. It is not so easy to achieve; it is definitely a moat. So cost is definitely one difference, and I think cost may be the number one distinction.
Then the second is time: when can you do it? If you are a few months earlier or a few months later, it is different. The third may be experience; user experience is still somewhat different. There is still some user stickiness and user barriers in the middle, but that may not be fundamental. Fundamentally, it is still the first: cost. The second is time. Whether you are the first to do it or later to do it, whether you achieve the same thing faster or slower.
After the long-term commercialization path and product line are further enriched, how should pricing be set? We think that what is most worth doing right now, and where it is most worth spending effort, is still to do AGI. Right now, what needs to be done is to push AGI further forward, meaning to push up the lower bound of intelligence and move forward. In this stage, that should be more cost-effective than doing more product lines and considering more commercialization paths, or rather, it is a path with greater returns.
I think in the future for some time, and in any period in the past, it may have been the same. That is, if we
We spent a lot of time thinking about whether enriching our products has any use? What was the commercialization discussion half a year ago? It must have been: I need to do advertising, then e-commerce, then embed e-commerce in the product, and then integrate heavily with local services or something like that. It definitely wouldn’t work, because the changes are too fast.
If your product is ahead, or in our current stage, and you spend a lot of time thinking about commercialization, commercialization paths, the product line’s lifecycle is very short. I don’t think we are at that point yet. That is our judgment, or at least all the previous experience supports this judgment. In the past three years, at any time, if you came to me and talked about commercialization paths and product lines, it was a waste of time, because you cannot predict or foresee the future.
What you can foresee is very little. Because I’m looking at the time, I don’t know whether everyone wants a mid-session break? Do you want to eat something first? Or should we keep talking? If everyone has no objection, I’ll continue. In doing global-leading AGI research and commercializing at the appropriate time. I think we have always been doing commercialization; I think we have always been doing commercialization, just not with commercialization as the goal.
We have AGI as the goal, but we have always been doing commercialization, which is why we have C-end users and B-end revenue. From historical experience, this strategy has been successful. And I think the point at which we completely pivot to commercialization is still very far away. So the biggest thing is still the extension of technology, and then doing the next generation of technology, to solve current problems more.
These return ratios, in my personal view, in the foreseeable future right now, at any time focusing on products is too early. So this is also part of our restraint. I hope these business opportunities can be done by others. We hope these business opportunities, and how to use this AI, can be done together by the whole society and all of our partners, sharing the returns together, rather than me wanting to monopolize them; that is impossible.
And we don’t have that much energy, and our organization doesn’t have that many people to do this. For partners, actually, our financing was carefully selected. The specific proposal on how to cooperate is one thing, but first of all, I think the interests are relatively aligned — the ones whose interests are most aligned with ours, who are least hostile to us, or who most want us to succeed.
Not everyone wants us to succeed, because we still harm the interests of many other people. The process and decision mechanism for major corporate strategy, technology, and business decisions. Our company as a whole is built on consensus. I’m not saying I decide everything by myself; rather, I need to seek consensus. My authority and influence within the company are built on consensus.
For example, if I want to do something, I will definitely first look at what our consensus is and whether everyone wants to do it. Then I may have some guidance or inclination, but the role of that guidance is limited; it is very limited. It still has to be built on consensus.
This decision mechanism is actually a mechanism for seeking consensus. It is not that I can push through any one thing; it must be consensus, and only then can I push it through, and then I will push it. At present, most of the main energy is basically on DeepSeek.
Moderator
We’ll take a five-minute break.
Investor
I’ll write quickly.
Moderator
Everyone can turn on your mics and give me some feedback.
Investor
No problem on our side, maybe let’s just take a five-minute break first.
Liang Wenfeng
The final gap in model performance across companies should be a comprehensive one. When comparing model performance, it definitely has to be compared at the same cost; that is what is meaningful. Because when you compare two cars, you compare cars in the same price range. The difference between a model that is well done and one that is poorly done should not be in any specific link; it should be overall. Is Anthropic, which is now ahead of OpenAI, a long-term thing?
I don’t think so. This is definitely episodic. OpenAI and Google will probably continue alternating in the future; they should alternate in rising. In fact, Anthropic’s advantage in Code Agent is not that big right now; it’s not like it is crushing OpenAI.
Maybe half of the people in our company, in their daily work, think OpenAI is better. In fact, Anthropic had a first-mover advantage, but that first-mover advantage should disappear very quickly; it is not an advantage it can hold onto for the long term. All three of these are very strong; among these three, the one with the highest efficiency is the one that spends the least cost, the least money, the least burn.
When global AI is in a dominant division of labor, the role Chinese companies are very likely to play is still the largest producer. By common sense, our capacity is the largest, including chips — our chip capacity may be the largest — and we have the most electricity, so our AI will very likely be one of the three bodies.
Chinese people will make this product the cheapest, and then in terms of performance, after all, for foreign goods, many Chinese-made and American-made products now do not differ that much. In the future AI may also be like this, but AI made in China may be cheaper. This cheapness may be systematically lower, just like services provided by China in other industries may be cheaper.
When I do things, the way I usually think is: what should I do right now to get the highest return? If I think the highest return right now is to do products, then I’ll do products; if I think the highest return right now is to realize AGI first, then I’ll go do AGI first. Obviously, I think doing products right now is not the highest-return thing. If it is a matter of adapting domestic chips, domestic chips actually have a historical opportunity right now.
Because before, there was a difficult problem in domestic chip adaptation, called poor ecosystem. You bought the card, but you couldn’t use it; it didn’t have the NVIDIA ecosystem. So NVIDIA’s moat is very strong.
But that is changing. NVIDIA CUDA’s moat is being rapidly dismantled, and there may be three reasons why it is being dismantled so quickly. One is that now we have AI, and after having AI, it is much easier to build this ecosystem than before, because AI can write code.
I can use AI to build this ecosystem, and then build an ecosystem exactly the same as NVIDIA’s. First, because of AI; second, because of some new technologies. For example, our company came out with a technology called TileLang, which is a high-level language.
Using this high-level language to write CUDA operators can quickly write out NVIDIA’s entire ecosystem, and then combined with AI, there seems to be no obstacle. But it still hasn’t been completed; it hasn’t been finished yet. However, this technical route seems to have no obstacle.
Also, because CUDA, NVIDIA evolved out of gaming cards, so in many places, the design of gaming cards and the setup of gaming cards are part of the same lineage. CUDA is compatible with gaming cards. In the past, because AI computing was a very small field, smaller than the gaming-card market, this was reasonable. But now the market for compute cards is already larger than gaming cards, so there is no reason the two still need to be coupled.
The trend now is that they will no longer be coupled in the future. Then dedicated chips — whether Huawei’s or NVIDIA’s own — will all be dedicated chips in the future, and none of them will be the old stuff anymore. In this context, the role of the original NVIDIA ecosystem is greatly reduced. Because for a dedicated chip, it has nothing to do with CUDA; it is not bound to CUDA.
Or rather, when this chip was being designed, it had already taken into account how to build this ecosystem. It’s a bit complicated, but anyway, domestic AI chip substitution now has a historic opportunity. We believe that within the next year, we will see one thing validated: the ecosystem for domestic chips is completely fine.
What was previously believed to be a problem — that it couldn’t be used, that it wasn’t easy to use — I think within a year we will reverse that perception, or use facts to reverse it. The hardware and ecosystem of domestic AI chips are both fine; the only problem is insufficient production capacity. There are no barriers to domestic card adaptation, and Nvidia cannot stop it.
If we were in a normal business environment and I could buy Nvidia cards, then domestic substitution would be relatively difficult; but when Nvidia cards cannot be bought, everyone is forced to do domestic chips. In this context, adapting domestic cards has no barriers at all. Building an ecosystem for domestic cards that is the same as Nvidia’s, or even better than Nvidia’s, I think there are no barriers, but it still takes time.
Right now we mainly cooperate with Huawei. Huawei does the adaptation itself, but we will participate in this ecosystem ourselves and get deeply involved in Huawei’s side. Huawei’s problem is still insufficient capacity. Huawei gives us roughly 16,000 cards of capacity, while internet giants may have 100,000-plus; we have just over 10,000, and I think that ratio is also relatively…… but this may already be all the capacity Huawei has.
So we can’t really count on Huawei to train the even larger model later on, or to train a model with hundreds of B activated parameters; sometimes people say that could happen this year. But next year or the year after, maybe there will be a chance. For Huawei card adaptation, our main work is to make its high-level language compiler work well, and to make TileLang work well. Once TileLang is done well, the problem may be solved naturally.
This is somewhat complicated to explain, but we are doing it, and once it’s done, explaining it will be much clearer. You can understand it this way: when V3 was trained, it still used Nvidia cards, but it no longer used Nvidia’s ecosystem.
V3 uses Nvidia cards, but not Nvidia’s ecosystem. Instead, we first write a high-level compiler called TileLang, and then based on the TileLang ecosystem we complete everything else, and we are almost no longer dependent on Nvidia’s ecosystem. As long as I take this whole set of things and redo the process on Huawei cards, then it’s done.
I think this may be a historic mission, namely, it can completely reverse everyone’s previous perception that domestic card ecosystems are bad. Right now, the timing, the geography, and the people are basically all in place; the only thing missing is time. I think within a year, many people should have it, or rather, everyone will understand that this problem has been solved, and what remains is a capacity issue. I am relatively optimistic about domestic computing power. On this point, I think Nvidia is digging its own grave.
Huawei’s supernode, Huawei’s 950 supernode, can completely replace Nvidia’s GB200 and GB300 in performance and price. The price will definitely be higher, but the increase is limited. If the price is 50% higher, 100% higher, 100% higher is fine, 200% higher is fine. For example, if it’s 100% higher, I think it can already be considered a replacement in terms of price.
It can also be a replacement in tasks; all the tasks GB300 can do, Huawei supernodes can do too, latency and everything are the same. The only cost is that four Huawei cards equal one Nvidia card, while being two years behind. Four to one is understandable. Being two years behind means that four Huawei 950s can match one GB300. Two years behind means two years behind in time.
Huawei 950 supernodes are shipping in this year’s Q3 or Q4, while Nvidia GB200 is from Q3 two years ago — a two-year gap. Nvidia may already have a new generation this Q3. So the gap between us and the U.S. in chips, I think, will no longer be an ecosystem gap in the future, but on chips it is four times plus two years.
There is also the question of whether we will vertically integrate upstream. I hope not. So there is a question: will we vertically integrate upstream into applications? We hope not to do that; I hope other people will do that. I don’t want to eat everything. I just want to eat one piece, the piece I’m best at, or the piece we ourselves think is the most core, and then the piece that is directly related to users
I think for many of our industry partners, they care about similar things more than we do…… this should be someone else’s significance; it shouldn’t be interpreted by me. Will we build large-scale clusters ourselves in the future? I think building large-scale clusters ourselves is definitely necessary; we have always been doing this ourselves, and all our clusters are self-built.
But whether we will develop chips ourselves in the future, I think, depends on how large the returns are, depends on how large the returns are. Tesla, nurturing…… this sort of thing. If you are operating a power plant, you don’t necessarily need to make the generators yourself, right? The power-generation equipment can be made by others; as long as the price is reasonable, why make it yourself? So I hope we won’t need to do chips. I hope we can buy chips at a reasonable price, so I don’t have to make chips myself.
I think this is very likely how it will be: recently, the profits on Nvidia chips may not necessarily be…… although…… we hope to only do one piece. I think AI is a very big thing; it doesn’t require me to…… I just do one piece. If we focus, and I believe the business value here is already big enough — if the AI era will produce many trillion-dollar companies, I think we are one of them.
We’ve always done one small piece of it, and being one of them is already no problem; there’s no need for me to…… I’m not brimming with confidence that I have to do everything. I think the most likely thing is that it’s hard to change, and if you really want to do something else, you’ll still have more ideas, and that will…… that’s our attitude.
For example, at least in the To B and To C businesses, from what we can see now, the ones that really want to build a closed loop in To C, in fact we do better; the ones that really want to build a closed loop in To B, maybe they still don’t do better than us. The more you want…… and also, To B, I can say a bit more.
The upper limit of the To B business should still be demand. Under the current generation of AGI and AI technology, To B demand should be limited. It will grow rapidly, but it is not something infinite; in the end, it is still constrained by demand, not computing power. Under the current technical conditions, how much revenue there can be ultimately still depends on demand.
Because think about it this way: if I can recoup my costs in ten months, then if I have a destination, I would definitely buy one piece; it’s because there aren’t that many things. Right, demand should keep getting bigger, and if technology keeps making breakthroughs, that demand will keep getting bigger.
Multimodal layout, we’ve always been doing it. For products, it is very important; for C-end user products, it is very important. But for the upper limit of intelligence, it is a component; it is not the main line itself. But as a component, we will definitely do multimodal, and we are doing it. We should release related models — our V4 and the subsequent versions of V4 will support native multimodality.
But for multimodality, for intelligence, it is a component; we don’t treat it as intelligence itself, and its main line is like search. Search is also a component; multimodality can be too, and in our understanding it is also a component. Scaling — we believe in Scaling. The larger the scale, the better the effect, and the more capabilities it can unlock.
What actually stops us from Scaling is computing power; it’s not that we don’t want to Scale, it’s that we don’t have that much computing power to do this Scaling. We haven’t touched the upper limit yet. We train models this large not because I think this size is enough, but because I happen to have this much resource.
I calculate based on my resources: how large a model can I afford and train? That’s how it’s calculated, not because this model is enough. At present, model returns are still very obvious. We still haven’t had the chance to hit the wall of Scaling; we are still very far from that.
When Silicon Valley talks about Scaling reaching its limit, that is for Silicon Valley; for Chinese people, we are still very far from that. We simply haven’t scaled to that level yet. This Scaling includes data Scaling, model-size Scaling, and then training cost. We are still relatively far from exploring the upper limit of this Scaling; there just isn’t that much computing power.
But we will also spare no effort to push the upper limit of this Scaling.
Including after we finish financing, we will have more computing power, and may train larger models. The sooner we can buy it within that time, the sooner we buy it; if we can spend all the effort within half a year, that would be best, but in reality it can’t be done. The core capability of the next-generation model, I think, must have continuous learning ability; only then can it be called a next-generation model. Before that, what we can do is lower cost, and make the effect better, and make the speed faster.
But to have a big breakthrough, it should have continuous learning. And then there is another question: why do we seem to care so much about the model’s computational efficiency? Because I really found that not everyone cares that much about the efficiency of this model. For a commercial company, there is no incentive to pursue model efficiency, because if model efficiency is a bit higher…… so there isn’t that much incentive to pursue model efficiency to be very high.
So several startup companies won’t themselves…… you haven’t heard them say low cost is what they pursue, because that doesn’t align with their interests. If costs are low, what are you still making money on? If costs are low, you won’t be able to collect much money.
On the contrary, this is part of our vision. Or rather, our vision — our teammates care about this cost, because our teammates are all ordinary people, and they know that using this all costs money. They can empathize: others have to pay to use it, so if it’s a bit cheaper, others will be more willing to accept it. So many of our teammates still hope that we can bring costs down further.
But if you look at it from a business perspective, you wouldn’t think that way. From a business perspective……
From a cost or business perspective, this is not the top priority. Whether for startups or for big companies, the service cost is really not high. But I think we hope it is relatively light; I hope it is affordable, especially in the context of China’s computing-power shortage, affordable and usable on domestic cards. I think low cost is first and foremost a result.
Our models have indeed always been moving in a lower-cost direction in terms of architecture, and that is related to our vision. We also have many algorithmic methods, and costs can still go down further. Another reason costs go down is that the lower the cost, the larger the model I can train, and the larger model I can afford. Under the same amount of computing power, if my computational efficiency is higher, I can afford a larger model.
For big companies, they may not think this way. For big companies, resources can be added; the problem can be solved by adding resources. But we will prioritize cost efficiency. The value of data-type models — data is a relatively broad category; data should almost equal half of the model.
Why do I think that if I wanted that, or if AI could account for 20% of GDP, and then if I wanted to take 5% of that, it would absolutely not work? Because I would definitely be beaten by someone else; if someone else says they only want 1%, then they will definitely beat me. If my goal is to take 5% of AI, of all human GDP, then theoretically the math still works out.
You can see OpenAI, it seems like the math works for them; theoretically there is no problem. But they have a problem: they will be beaten by another person who is willing to take only 1%. Because that other person says, I do such a good job, but I only need 1% of global GDP, then they will beat him.
At that point, if another person comes along and says I only need 0.1%, then that person will beat the people before him again.
From a macro perspective, no matter where that percentage comes from, there is no difference between each other; it’s all the same. The one who takes more will be beaten by the one who takes less. You don’t even have to actually take more — if your vision is to take more, then you will be beaten by the person whose vision is to take less. In fact, nobody has taken any money; it’s just a vision. If your vision is to take more, then you lose first, and you will face greater difficulties. That’s just how the world is.
OpenAI originally thought it could really monopolize the world, but in reality it will encounter many, many challengers. It will face challenges, and it won’t be so easy. The U.S. will have faced challenges, and in the future it may also face challenges from China, because Chinese people are willing to take less, and can still provide you with this service. In China, there will also be people willing to take a little less.
But in the end, there will be a balance here, because if you take too little, the company’s business logic won’t hold, and it won’t survive. So if you take too little, you can’t survive; if you take too much, you will be beaten by those who take less. So for us, it’s not about taking the most profit, or maximizing return in pricing, but only earning a reasonable return. That’s an explanation.
I believe in this matter. I’m not finding reasons for it, because there’s no need to find reasons. This is just how I do it; if I do it this way, there must be reasons. Those reasons may not be very conventional, but I think the company itself is not conventional. Our company’s management actually has two lines: one from top to bottom, and one from bottom to top. Bottom-up means each person decides what they want to do and does it themselves; nobody manages them, and there is no KPI.
Top-down means doing the important work; if we need to collectively do something, everyone in the company needs to coordinate. For example, if we want to release V4, then we need division of labor, and everyone needs to do a part. That is top-down, and what we call top-down here is “doing the important work.” Generally, we hope that the “important work” does not take up half of an employee’s time — it should not exceed half.
They still have half of their time unassigned; they can do whatever they want. This is a research scope, allowing them to explore on their own, according to what they think is important, with no prior requirements. As long as the company can support it, and the company’s computing power can support them doing it, or if they don’t need computing power and only need to do very little, then they don’t even need to coordinate. So this is the organizational approach we have now.
Some people think we are top-down, some people think we are bottom-up; I think both are true. My criterion is that “important work” should preferably not exceed half.
We usually don’t work overtime much either. There are two reasons for overtime. First, research needs a relatively relaxed environment. If you pressure people too tightly, they can’t do research. Since you need your own interest and you need to think about these problems normally, it has to be in a relatively relaxed environment for exploration to be possible. That’s a need from research culture. Second, we are very focused.
Being very focused means that the things we need to do are very few. Then I don’t have that many things to do, so I don’t need to work overtime. This is consistent with the restraint I mentioned earlier. Because I restrain myself, many times I just don’t do things. Then with fewer things to do, the work assigned to each person is less. You can see that many of our products are not very complete, and we haven’t gone to fix them. That is also part of our culture.
OK, because there are many questions, I’ve mostly skimmed through them. If everyone still has questions, please ask.
Host
Please feel free to speak up and exchange ideas, investors. Let me remind everyone that Liang Wenfeng mentioned quite a few sensitive pieces of information just now, so please do not share any numbers or situations externally, including card counts and such. Also please do not screen-record or share externally. Thank you very much, and if you have questions, please feel free to speak up.
Investor
Yang-ge, could you share more about the timeline for when continuous learning might bring breakthroughs? Also, if continuous learning is achieved, what other architectural or algorithmic innovations, and other key factors, are still needed?
梁文锋
Fewer people; research is needed. Right now the whole world is studying this problem. Or rather, for investors, what investors see most now is AGENT; but for us researchers, what we see more now is learning, and how to solve the problem of learning. In fact, learning may not be a technology; it is a problem. How to solve this problem may involve many kinds of technologies. It is not a single technology; it is not one thing, it will be many things.
Or rather, AGI is composed of many things. It needs the model, and it also needs many other things. In fact, it is also an engineering and algorithm problem. This problem is quite specialized, but there are many methods and also a lot of research.
Investor
Mr. Liang, thank you. Thank you very much for today’s opportunity. First, I really want to respond with gratitude. What you said at the beginning moved me deeply and also gave us a lot of inspiration. You mentioned that this team is carrying the greatest goodwill, hoping to make some contribution to the development of this industry and of human intelligence in this society.
And within that, it carries this sense of mission and vision. I think that is very similar to the corporate culture of the company we serve, which is “to cultivate oneself and help others.” I deeply understand why you lead the team to do open source. I can vividly imagine it like creating a little bird paradise ecosystem with a banyan tree — benefiting all things without contention, and in that way it will be accepted by everyone, all things living together in symbiosis and coexistence, and ultimately it will be everywhere.
So, through this investment, we also want to express our recognition of, support for, and respect toward this mission and vision. At the same time, we also hope that in the future of this industry, in some of the areas we are good at, we can contribute some strength. On this point, I also want to continue asking for your guidance and discussion. For example, in the co-building of the future ecosystem, now that it has been open-sourced, how many partners, talents, and teams in this industry can fairly well reproduce some of our currently open-sourced models and results?
In the next step, when we want this ecosystem to further develop, in which areas do you feel we need more high-quality talent to connect to our models and reproduce them? Or is it that GPU computing power is relatively scarce right now? In the future, will everyone move toward a model-matrix approach? For example, as we make the foundation models of large models better and better, partners and teams across industries go on to build vertical industry models within the model matrix, or some application models.
How is development in this area right now? In two years, three years, what kind of ecosystem do you think it will grow into? That is my first question for you. Second, you also shared a lot of observations with many partners just now about AI hardware. For example, global AI giants may each be making investments at the hundred-billion-dollar scale, while China currently seems to have some shortcomings in hardware and computing power.
How long do you think it will take to solve this, and support the development of our AI, so that the shortcomings in computing power and hardware do not hold AGI back? Do you think this is something Chinese people will eventually achieve for sure, that we can definitely make it happen, and that it is just a matter of time and capital investment? But at the same time, it may be a two-sided process.
On one hand, progress in models will improve model intelligence, causing the consumption of hardware and computing power for single tasks or certain intelligent single agents to gradually decrease, no longer requiring such large computational power, because model progress will make it think more efficiently. I don't know if I understand that correctly. On the other hand, progress in hardware technology will make the computational energy efficiency of compute more powerful. Will this be a path where both sides move toward each other?
At this point in time, if we use the current 960 or H200, and make a hundred-billion-dollar-scale investment in computing power, you just mentioned that you amortize it over three years. What do you think its actual lifecycle is? Is the technology iteration four or five years? Or to put it more bluntly, will it be that a compute center is built today with current cards and a 10,000-card cluster, but three years later it is actually relatively no longer very advanced compute?
Will there be a situation where something is under construction and not enough now, but three years later it becomes relatively less high-quality compute with excess capacity? I don't know whether such a phenomenon will exist. Please advise on these two questions. Thank you. Thank you.
梁文锋
The first question is about the ecosystem. What we feel now is that perhaps every company faces the problem of not having enough talent. But I think this talent shortage will be temporary. In every industry at the early stage of development, there is not enough talent. Including when websites were first being made, there were very few people making websites, and talent was very scarce. Later, when the internet needed server-side development, talent was also very scarce.
But this kind of talent shortage is solved very quickly, in just two or three years, because a large number of people will be trained. The shortage of AI talent is also temporary, and we have already seen it greatly eased. Because there really is no shortage of AI people. Every company will quickly train people; training people is very fast. So, overall, in the AI industry, whether it is the ecosystem, model companies, or anything else, talent is not scarce. Talent scarcity is definitely a short-term phenomenon.
Historically, there has never been a long-term shortage of any particular kind of person. I still remember that more than ten years ago people said pilots were in short supply. The training cycle for pilots is long, but that too was quickly resolved. So everyone doesn’t need to worry about talent shortages. Also, there are a bit too many companies making models in China right now, still too many. In the US there may be just three companies. In China there are too many foundation model efforts. In the end, it will definitely not require so many companies to build foundation models; it will definitely converge.
So resources are also relatively dispersed, and to some extent relatively wasted. Just before
everyone had to do the same thing, but in the US only three companies need to do it, and resources are concentrated in just those three. In China, resources are spread out very widely, and each company gets much less. I think this will definitely converge, it will definitely happen, but it takes time; in the end it will definitely converge. There is no need for so many companies, because right now everyone may feel that the profit margin for doing this is very high, so they must do it themselves.
But when they find that this may not be such a high-margin business, they may stop doing it. Recently it definitely has not been that high-margin. I don't believe there is such a high margin, because that does not conform to objective laws. That means we are at a stage where if there is a very high profit margin, that definitely does not conform to objective laws. We should have a reasonable profit. So this is the current state of the industry, and I think it will definitely converge.
That part where everyone is building large models, not to mention one company monopolizing everything and saying "I want to take all the profit," that definitely won't work. If each company only takes a reasonable profit, then in fact there is no need for so many companies to build large models. In the end, China having three or four companies competing would already be enough; competition would already be very sufficient, and the pricing would absolutely be enough to fight a price war.
For large models, it may not even be two big companies and two small companies; maybe that is already quite enough. As for the ecosystem, I don't have too many ideas. We hope to support more people, but we don't have that much energy. We do have this intention, and there will be no conflict of interest, but whether we actually do it is another matter. But at least there is no conflict of interest here; we hope for cooperation and a win-win outcome.
First, I absolutely do not believe that large model companies can take away most of the profit; that is impossible, because there are so many large model companies. The gap does not need to be that big now. There are only two things that create a gap: one is time, the other is cost. So it is not like any one company can have excess profits; I don't think there can be excess profits. Those who control costs well earn a bit more, and those who control costs poorly earn a bit less, and that is all. Did that answer it? Was the first question answered?
投资人
Everyone can... you believe that in the future there will definitely be many people who can... in the future it will actually... everyone’s data applications and such will iterate in a cycle... can you hear me now? Thank you. Right, thank you for your answer, and I also deeply understand and respect your ecosystem strategic positioning within the industry. For example, on the data side, publicly available data—I'm sure model companies already have channels to obtain it, so that method should not be a problem; it's just a matter of time and cost.
Then later, for example, when we really get to AGI, one possible imagined or ideal state is that the model can self-iterate and self-learn, meaning it trains itself. On this point, regarding the current data, do you think simulated data can be used, or is real data still of the highest quality?
If real data is still needed, would that limit AI intelligence to human history... at that level, because it depends on real data that humans truly had in the past? Or can this upper bound be broken through using simulated data, synthetic data, generated data, and other methods, allowing model capabilities to surpass all of humanity's past real...
梁文锋
I think it can surpass it. I think there are two kinds of surpassing, for example Go, AlphaGo made a move that humans had never seen before. That is to say, it definitely surpassed humans within a certain range. But it may also have an upper bound, and it may also have limitations. But we cannot see those limitations now.
We believe, so broadly speaking, that it can surpass what humans already have, and the knowledge we can already express.
投资人
Then later on, will this rely on real data or simulated data? Will it work?
梁文锋
There are many methods; it is not impossible.
投资人
Okay, thank you. I’ve also taken up your time, and I’d like to continue asking about the AI Infra issue just now.
梁文锋
What was the second question? A bit...
投资人
Okay, I’ll repeat it briefly and quickly. I wanted to ask about AI Infra. In the future, we believe computing power is now being invested at the hundred-billion-dollar scale by everyone. On this point, we may believe that Chinese people will eventually achieve their mission on hardware, and one day there may be highly efficient compute, but in practice, it is still currently a bottleneck. So in the future, will the two sides move downward toward each other?
On one hand, after model capabilities iterate, it actually shifts from brute-force compute to clever compute, so the requirements and consumption of compute per unit model or task will gradually decrease at the margin. On the other hand, if hardware such as card capabilities improve, iteration speeds will become faster and single-card efficiency will improve. What kind of phenomenon is this in the process?
Will it be that now you build a 10,000-card cluster and buy H200 or 960, but after two or three years it becomes relatively less high-quality compute and relatively becomes an obsolete device?
梁文锋
NVIDIA cards can basically be depreciated over five years. Huawei cards should be depreciated over at most three years. Huawei 950 is still pretty good to use this year; I think it is still okay to use it next year, but using it after that I think it may really be too power-hungry. The lifecycle of Huawei cards is definitely shorter, because they are already two years behind NVIDIA to begin with. But I think the gap is not that big. If B200 can be bought now, I think it is all worthwhile.
If it were Tencent, and Alibaba could buy it, then it depends on the scale; if it can be bought at a reasonable price, it is definitely worthwhile. But when calculating costs, believe me, you can't buy it.
投资人
Understood.
梁文锋
Our compute is lagging behind; that is a fact. This fact is alleviated through three aspects. The first aspect is that we accept model lag; we can only use smaller models than theirs, and how much smaller is a training issue. We need to accept a certain degree of model lag, as well as smaller model sizes. This lag has one advantage: lag means you have more time. You have a certain amount of technology, and then you can use clever methods
投资人
Thank you.
梁文锋
So our gap with the US may be that we are 12 months behind the US, perhaps 12 to 18 months behind, or 6 to 12 months behind. Simply put, we are two years behind the US, and then we do this with only one-twentieth of the compute. That narrative is: one to two years behind, but using only one-twentieth of their compute.
Then in the future, we need to rewrite that narrative: we use a fraction of their compute, but shorten the time further, down to 6 months, 3 months. I think that is a goal. And we may even surpass them in some areas. But when there is still an order-of-magnitude gap in overall compute, comprehensive surpassing is unrealistic; however, in some key areas where we make trade-offs, surpassing in some places may be possible.
投资人
Understood, thank you, Mr. Liang. Full of confidence, let's work together. We’ll leave the time to other partners, thank you. Thanks for the sharing just now. I have two quick technical questions. In the technical roadmap just discussed, it was mentioned that the core problem we need to solve at this stage is continual learning, which is currently a hot topic in overseas research, called RecursiveImprovement. I’d like to ask, from a technical standpoint, what is the biggest difficulty right now?
From your perspective, when can this be solved? That is the first question. The second question is that you just mentioned solving continual learning first, and then doing intelligence. I also want to understand the technical root behind your saying this. Does it mean that after solving continual learning, DeepSeek will also go on to do general intelligence? Please help explain these two questions.
梁文锋
Technical issues are actually a bit hard to explain... The difficulty is that we still haven't found a very workable method. The whole world still hasn't found a good method, and everyone is still exploring. So we are still in the exploration stage, meaning we don't know who will find the next method to solve this problem. We are still in the exploratory stage. We have many ideas, many promising thoughts at the moment, but none of them have been made to work yet. Yes, that's the first point.
The second point is that internally, we value this narrative quite a lot: training our next version of the model, we hope it can help with our own development. It can improve DeepSeek's efficiency; our model's very first goal is to improve DeepSeek's own work efficiency, so that when we develop the next version of the model, it can provide more help.
Or to put it more simply, the models we build are not first meant to be useful for everyone, but to be useful for ourselves. First, they need to be useful to us. Once they are useful to us, then when I develop the next version of the model, it will be faster.
A lot of people inside us think this way: first it has to be useful to ourselves, first it has to work for us. And then that is the fastest way to achieve AGI. When it works well for us, that may mean it also works well for others, but first it has to work well for us. This narrative is a bit strange, but many people really do think that way.
It is not about user satisfaction; I hope it helps us more, so that we can achieve AGI, and much faster. So the logic of this narrative is that it helps us achieve AGI. But first it helps us achieve it. We need this help to achieve AGI. Right now, it is very certain that we really need artificial intelligence to help us achieve AGI.
Although it still does not work autonomously, and still only works in combination with humans, it is already very useful.
投资人
The second question is, just now you mentioned solving continual learning first, and then entering general intelligence. Is this your expectation for the future? I want to understand the technical root behind this. Why do we need to solve continual learning first, and then enter general intelligence? Your understanding of this area later on.
梁文锋
Because solving continual learning can greatly accelerate our R&D progress. If I solve the continual learning problem first, then the problem of general intelligence is no longer a big deal. With AI assistance, if AI can continually learn, its capabilities should be very strong. The current capabilities of agents are limited because they cannot continually learn; they cannot effectively continue learning.
If we can first finish continual learning, then AI's capabilities will be very strong, and it can greatly improve the efficiency of our own research. If continual learning is done first, general intelligence may become very easy, and using it to do that would be easy. So I say this is a result we very much hope to see; it saves us effort, and we are more relaxed.
Otherwise, if you go and manually build general intelligence now, it is relatively tiring and difficult, it is a data-intensive, labor-intensive task, and the cost-effectiveness is not high either.
投资人
Thank you for sharing.
主持人
A small question: please take a look at the chat group for the online questions. How long do you think until AGI? Can domestic hardware catch up by then? It's in the Zoom meeting chat window.
梁文锋
Okay, I saw that. Huawei 950—right now Huawei is giving us 16,000 cards, which should be something we can say publicly. It should be about one order of magnitude less than the big internet companies. Huawei can only give us that many, because the price is not cheap either.
The big internet companies' demand will be larger; for them, they need it more. For us, we can buy some non-compliant cards. So our purpose in buying Huawei 950 is still to help Huawei build out its ecosystem. 16,000 Huawei 950 cards are only equivalent to 4,000 B-series cards. So it is not a very large amount, and the significance is not very big.
It is not enough to train a next-generation model; it is only enough to train our current generation model, not enough for the next generation. But it can help Huawei get this right first, that is, regarding Huawei 950. Then, how long until AGI? Can domestic hardware catch up by then?
I think in the AI endeavor, in this AI matter, domestically it should be possible to get to roughly the same level as abroad in one or two years, or maybe even this year to achieve a substitute for foreign models. In AI, with the current approach and current paradigm, it is not very difficult, so it should be possible this year. But that is still not AGI.
At the very least, I think it has to be able to continually learn. Can domestic hardware catch up by then? I think domestically it may take a few years. First, China needs to solve the ecosystem problem, because the ecosystem is a matter of confidence. Solve the ecosystem problem, then solve the capacity problem, and I think it should be gradually solvable. I don't really believe that five years from now we will still be stuck on capacity.
Right now we are definitely stuck on capacity. This year, next year, and the year after, I think we may still be stuck on capacity, but five years from now, I think that may not necessarily be the case; I remain relatively optimistic. Then the second question is about thinking through the future organizational structure and the scale of personnel planning. First, our previous organizational structure was very decentralized, because there was no organizational structure. But as our people continue to expand, these things definitely need some changes.
At the moment I can only say that many changes will be needed here, but it is hard to express them all at once right now. In the end, we will probably still need different departments, and some departments will need us to establish a relatively rigorous hierarchical structure. Other departments may still maintain a relatively loose and relatively flat structure. As personnel increases, we will make these adjustments. We should be making this adjustment very soon, because I am already making this adjustment.
If we don’t make this adjustment anymore, many things simply can’t be pushed forward. Indeed, there are many departments that should have an organizational structure. Another question everyone has is, CV is before which version, right? I think right now, the online release of this GCV4 version is still relatively rough, and there are still many capabilities that need time.
For us, generally speaking, a comfortable release cadence is about one version every two to three months. The last release may have been at the end of April, so the next release may be at the end of June, roughly like that. If nothing unexpected happens, each version should be better than the last. At the scale of 50B active parameters, I feel that in the end we won’t be too different from the current wave of open source.
In terms of inference speed and performance, I feel there may not be much difference. But compared with that large model, that unpublished model of theirs, the gap should still be quite big. That gap, I think, is something our active-parameter scale probably cannot achieve; it would definitely require a much larger model, perhaps something like 150B.
As for 150B, based on our current training progress, optimistically we could start training by the end of this year; at the very least, it would be next year’s Q1... the gap is still quite large. Yes, that’s the gap with OCE.
Investor
Hello, I actually have a question. You often say that AGI’s realization process is a gradual one, not something that happens in a sudden jump. So can I understand this as a process without a critical point?
梁文锋
It doesn’t have a critical point, but it is nonlinear. We currently believe in a narrative that AI can accelerate AI research, AI can accelerate AI research. That is to say, it’s not linear, because you can use AI to accelerate your own research, so later on it may become nonlinear.
Investor
Understood. So at this stage, my understanding is that the conclusion in front may be that continuing to scale language models is enough, it’s sufficient to reach this state.
梁文锋
I can only say that, for language-model scaling, I currently do not see an upper limit. We have not seen an upper limit to our current intelligence level, or even to the level of intelligence in the United States.
Investor
Understood. Because I’m very curious—what you said is that for the United States, a model with 800B active parameters can be trained, but it can’t be used, so it can only be trained; it’s also hard to put it out for everyone to use because it’s really expensive. One thing I was actually very curious about before is that humans have only had language ability for nearly 100,000 years, while before that evolution took maybe 3.7 billion years.
But in training AI, maybe it is possible—you said that sequence can be reversed—but in the end it may still have to move into what is called the world model, or maybe it’s not necessarily called the world model, but it still has to enter the physical model or the embodied part, right? That is, after this upper limit.
梁文锋
Yes, I think embodiment definitely still has to be entered, ultimately embodiment. So for our company, naturally, the end point may be embodiment. Because for a normal person, their needs are not a computer, right? Because a normal person eats, drinks, plays, dresses, lives, and travels—they don’t need a computer.
What they need is embodied intelligence to solve specific human labor needs. If the goal is to relieve labor demand, then embodiment is something that can’t be avoided.
Investor
Understood. So at a certain stage, if we reach something similar, maybe not even a critical point—something that can self-evolve, can self-evolve relatively well, or close to that point of ascent with SV—I’m very curious what the first landing point of AI in that state would be... maybe it will be different from now.
梁文锋
What we hope is that it can, if there is no embodiment, then our definition of AGI, or what we hope AGI can do, is what? It can help me iterate the next version of the model; it can help me iterate the next version of the model in the same way. And then if embodiment is added, what we hope it can do is also let it iterate the next version of embodiment, let it make the next version of the robot.
Investor
I’m still very curious about one thing, because I saw previous DeepSeek interviews, etc. It seems that in the selection of important directions and research directions, taste and intuition are very important, not just simple engineering optimization. If AI can self-evolve in the future, will taste, taste, intuition still matter, or what will matter?
梁文锋
AI does not currently lack taste or intuition; what it lacks is the ability to continue learning. AI’s taste and intuition are fine. If you ask it to write an article, its taste and intuition, I think, are not a problem. There were a few questions earlier; let me look at them. I saw a few questions on the screen, but I can’t see them here.
Investor
Mr. Liang, I left a question on the screen; let me read it to you again. Actually, I want to ask about continuous learning—you also mentioned it, and many researchers have mentioned that it is still an unsolved research problem. Then the coding agent, especially catching up with MILES, is
reaching the Office level and MIS level a relatively certain target. For an unsolved research Scaling target and a relatively certain Scaling target, how do you think research resources should be allocated, especially the talent resources on the research side, so as to achieve the best balance and effect?
梁文锋
Model Office is a relatively certain target, but model MIS, I think, is still hard to say it’s certain; it can only be said to be a target. Model Office, I think, should be relatively certain.
CoT does not consume resources. Doing any research does not consume cards, it only needs very few cards; what it needs is ideas. It also does not consume talent resources, because you don’t need someone to stay there doing it all the time. It is not a project, but rather something that requires many people to think about the problem. So there is no need to allocate resources to it, because it doesn’t need resources. You need resources to train models, to make models, to release models, to do efficiency experiments for models.
What I just mentioned doesn’t consume much in terms of people or cards. So we call this “mò jiǎng” [trying your luck]. The threshold is very low, anyone can try, but who can get anything out of it—that, I also don’t know whether it depends on talent or what. So there is no need for us to allocate resources here. It’s just that the difference between us and other companies is that we will spend time discussing this problem, thinking about this problem, and treating it as an important matter.
Inside the company, it is an important problem, something we will spend time thinking about, but it does not require a lot of resources to do. Then I also see another question below: the hallucination problem of large models affects the user experience quite a lot. The hallucination problem also has a method that can solve it, but this is a long-term topic.
The hallucination problem can be regarded as something that can be solved, or improved, through better Post-training.
It’s just that everyone hasn’t put in much effort to do it. Or, for me, hallucination is a problem, but we may classify it as a product issue. We will solve it, but it is not a key issue. There was also a question earlier about data labeling. In terms of data annotation, this has to do with our capital investment. Given the structure of our capital investment, we cannot support the cost of so much high-quality data annotation, because the cost is very high.
The cost of data annotation in the United States is not materially different from the cost of data annotation in China. China does not have a cost advantage in data labeling; especially in high-end data labeling, there is no cost advantage, which makes it very difficult for us to invest in data labeling the way the United States does. This path is very difficult in China, because data labeling is simply too expensive—whether we outsource it or do it ourselves, it’s all very painful. So now it’s basically a two-pronged approach.
It’s not that we absolutely cannot label data; it’s just that some kinds of data labeling have low cost, and some have high cost. We start by labeling the low-cost ones. So you could also say that now half the people in our company are labeling data. Half of the core researchers, the most important people, half are labeling data. We are concentrating on data labeling. Solving the AI problem at this stage relies on data labeling. You just need to look at it as a data issue.
Investor
Mr. Liang, thank you. What you said was especially relatable for us, very good. Thank you. I’d like to ask a few questions.
The first question: you just mentioned that Chinese models are definitely stronger than American models in terms of efficiency. You also mentioned some other aspects, and in the future we may also be stronger than the United States in those. What do you think those aspects are, where we may be stronger than the United States in intelligence or in other areas?
梁文锋
I think in many experience-related aspects, it’s possible that we can do better than the United States. Not to mention our own experience, our own product experience feels pretty good. I think in user experience we may not necessarily be worse than the United States. In terms of product, product capability may not necessarily be worse than the United States. Costs should also be lower than in the United States, so China may still have competitiveness. In other areas, if you ask whether there are structural advantages, I think maybe not.
But in cost and product, I think there are indeed certain structural advantages. The cost part is easy to understand, because they don’t have to do it, so they don’t develop this capability. They definitely do not value this as much as we do. We can treat it as a very important thing, but for them, this is not important. As for product, naturally through many companies, product capability is still okay.
So I think these two aspects may have structural advantages.
Investor
Okay. Second question, let me ask you: you just mentioned post-training. Our investment costs are relatively high, and companies like Anthropic and OpenAI are investing huge amounts. After this financing, do you think we will increase investment in post-training?
梁文锋
The gap is mainly in high-quality data labeling, and it’s mainly within AI research. We will definitely increase investment, but high-quality data labeling is typically not a matter of capital investment. The bottleneck in high-quality data labeling, I think, is time, because it takes time. For OpenAI, for overseas companies, for Anthropic, they all started earlier, and they have more capital and more cards.
Under these circumstances, for us domestically, you can think of it as only in the past six months that we really started doing this, so in terms of time, I think more time is needed. This has less to do with capital investment, because even without more capital investment, the existing capital is enough for it to expand as fast as possible. But this speed has an upper bound; the bottleneck is not that I can immediately get more people, and it is not bottlenecked by money or by cards. But it is indeed in a process of rapid expansion.
So we think that within a year, if the high-quality data issue can be done well, I think that should be something that can be expected domestically. I think the outlook may not be that cold, but it does require time.
Investor
Thank you. Then the third question is that we see Anthropic using its own models to build its own products, launching many verticals in finance, law...
and even in the future moving into healthcare. Do you think in such vertical applications, at some stage we will also consider this?
梁文锋
I still haven’t thought it through very clearly—what exactly our domestic business model will be in the future, or what the smoothest path will be. We haven’t reached that stage yet. The domestic situation and the overseas situation may not be the same; what it will look like domestically at that time is still hard to judge now.
Given the current domestic situation, I think the most reasonable approach should be to go all-in on general-purpose agents; other agents should have lower priority, including finance and doctor agents. We should first do Coding, because Coding Agent can do a lot, and there are many more vertical agents.
At this stage, we think the most important thing is still Coding Agent.
Investor
Very clear, thank you. There is another question we’d like to ask you about. Actually, we do DeepSeek, and we also admire you very much for always doing DeepSeek in a very pure research manner. But this industry has indeed moved into the capital markets and onto the path of capitalization. And you are also a very responsible person—whether toward teammates or investors, you are very responsible.
So, how do you balance pure research, the pure AGI direction, and the capital markets? You will definitely still have to go into the capital markets in the future, and there will definitely be public shareholders. How do you think about balancing this in the future?
梁文锋
I now feel that it should be possible to do both. Suppose this year I can have a few hundred million dollars in B-side revenue, and at the same time we have C-side users, then that already gives us a certain business foundation. If next year we have B-side revenue, and if this demand can grow further, the company may not be far from net profit; it may already be net profit.
It may no longer be a pure money-burning stage, so I feel that the actions we can take and operate later should still be relatively large. Or, in the worst case, selling APIs may even be enough to support a listed company. If, say, there is no new technological progress later on and our technology freezes here, then in the end we can just go all-in on selling APIs and doing these services well, and I think that would be enough. So I still have confidence; it really isn’t that hard.
Because we are indeed at a highly leveraged point, and in a field developing very fast, it may just not be that hard. We can only say that we hope to have bigger dreams, but we also have fallback performance that we can take out.
Investor
Thank you, that was very well said. That’s all my questions. Thank you.
Thank you for sharing. You mentioned some questions earlier; I’d like to ask you again. Because I think the biggest difference between DeepSeek and other companies is actually that our organization is different from others, or rather, the organizational form is different. But the organizational form is related to our goals, and we may also need to consider the organization’s own boundaries and efficiency. I don’t know, from a macro perspective, whether our organizational form has a good learning model?
Historically, maybe Bell Labs, or what kind of structure would be ideal? Or do we think there isn’t really an ideal one, and we need to explore it gradually ourselves? Maybe in a new era, we can only rely on ourselves to do the exploration. Because our organizational form is actually definitely different from the big companies in the United States, right? The three companies themselves are also different, but at least they all start from the perspective of a commercial company to explore.
I mainly want to ask about the organization issue again.
梁文锋
First of all, we do not have an object to imitate. Every step is based on our actual situation, seeking truth from facts, making decisions according to the actual situation, and finding out how we should do it. So it is a product of the times, or a reflection of the real situation; it is not the result of imitation. In other words, under this situation, I really think the optimal solution may be like this, or rather, in my own view, the path we chose is like this.
Every step, of course, we thought it through, of course we chose it, and in any case the result of the choice is what it is. It is not because we saw someone else choose this way and then we chose this way; it is because we analyzed the pros and cons and chose this way. In the future it may be the same; we are not imitating anyone.
I think we are still different from Bell Labs, because it clearly did not need commercialization, because it was... but we clearly need commercialization. In the end, we still have to survive; after all, we are a company, and the government will not give me a cent. So we can have very grand missions, but in the final analysis we are a company, and we have to think about how to survive.
So B-side is definitely important to us, because maybe in the future we may have to rely on it to survive. It’s just that it’s not important right now, because right now it is a cost line. So I think that is still different; it is still different from Bell Labs. We are essentially still a company. Historically, there have also been many companies that pursued things beyond profit, but you can’t say that because they pursue things beyond profit, they are not companies.
Many companies are great because they have a pursuit beyond profit. That pursuit not only did not affect its commercialization, but instead allowed it to commercialize better. We are essentially still a company; it’s just that we are considering which money to make, when to make money, how much to make, and what to rely on to make money—we simply have trade-offs.
Investor
Mr. Liang, I have two small questions; let me ask quickly. One is that you just mentioned MILOS—maybe it’s not a very certain target, but we will definitely move in that direction. And you also mentioned that the active parameters may, for example...
The next generation may be in the range of 150 to 250 B. In that case, do you think 150 to 250 B is benchmarking O4.7, or maybe benchmarking something else? That’s the first question. The second question is that you also mentioned earlier that in our entire inference stack, we used some programming languages other than CUDA.
I understand that originally, based on NVIDIA's ecosystem, we might have used a lot of PTX and the like. Now that we're using more things like TileLang you just mentioned, will that significantly reduce some of our efficiency in inference, or reduce efficiency in the short term? I don't know how you would view the efficiency loss brought about by this kind of change in programming languages, or whether in the long run it is actually a complementary state, meaning improved efficiency?
梁文锋
It is improved efficiency, a big improvement in efficiency.
Investor
OK, so there isn't really any negative impact instead?
梁文锋
Yes, it is a big improvement in efficiency. So this is an opportunity. It is like in the past you couldn't do without the CUDA ecosystem; now we can discard that ecosystem and use a simpler method, TileLang. It's a high-level language, and writing programs with it is very fast too; the amount of code needed is very small, and I can rewrite it from scratch.
Investor
I understand. So both of these are actually, as you mentioned, big opportunities brought by AI, not shortcomings that may need to be made up for in the short term.
梁文锋
Yes, it is a major opportunity in technological development. It is not an AI opportunity, because we also have a project in which we are using AI to write TileLang.
Investor
So is it even faster?
梁文锋
Right now all TileLang is written by humans, but it is already much faster than writing CUDA before.
Investor
I understand. So even for the execution efficiency at the hardware level, you think there is no impact either?
梁文锋
A loss of 1% to 2% is acceptable, I think.
Investor
OK, got it, understood. Good, thank you, that was very clear.
Host
Does anyone else have any questions? If not, let's stop here for today.
Investor
Okay, thank you. Thank you all very much for your time, thank you, Mr. Liang, thank you.
Host
Bye-bye.
梁文锋
Bye-bye.
內容僅供參考,不構成投資、法律、稅務或財務建議。