Conversation with Kimi founder Yang Zhilin: Moving toward the long, unknown snow mountain ahead
Moonshot AI founder Yang Zhilin discusses AGI, Scaling Law, Kimi, commercialization, Sora, and why artificial general intelligence needs a different…

Editor’s note: This article is the English translation of an interview with Yang Zhilin published by Zhang Xiaojun on March 1, 2024; the translation was created by Kimi K3, and the author reposted it on X in July 2026. The main content remains in English.
Source: Zhang Xiaojun’s X account

This was my first interview with Yang Zhilin, conducted in early 2024 and published on March 1, 2024—exactly the first anniversary of Kimi’s founding. At the time, Kimi had only 80 people and was working out of its first, rather shabby office. There was no logo at the entrance. Only a white piano stood guard by the door.
This article caused quite a stir in China’s tech community at the time.
Back then, I was still a print journalist, so this interview existed only as text and an audio podcast.
As we can see, many of the views expressed in this article were later vindicated.
Just rereading these lines from two years ago, you really can’t help but marvel at how drastically the world has changed!
(This translation was created by Kimi K3.)
Yang Zhilin: “If everyone thinks you are ordinary—if your dream is a dream that anyone could have—then it adds nothing to the sum total of humanity’s dreams.”
By Zhang Xiaojun
Just one year earlier, AI scientist Yang Zhilin had done some exact arithmetic in Silicon Valley. He realized that if he decided to start an AGI-oriented foundation-model startup, he would need to raise more than US$100 million within the next few months.
But that was only the ticket to enter the game. One year later, the number had increased thirteenfold.
For foundation-model companies, competition is less a scientific contest than, first and foremost, a brutal money war. When investors keep a tight grip on their wallets, you have to race ahead of competitors in raising more money, buying more GPUs, and recruiting more talent.
“It requires the concentration of talent and the concentration of capital,” Yang Zhilin, founder and CEO of Moonshot AI, a foundation-model company founded on March 1, 2023, said.
Over the past year, China’s foundation-model companies have seemed to live on a knife-edge of survival and tightening pressure. On the surface, each company holds considerable cash. But on the one hand, they must immediately pour newly raised money into extremely expensive research to chase OpenAI—first catching up to GPT-3.5, and before they could even reach GPT-4, Sora had already arrived. On the other hand, they must keep racing to find viable real-world use cases, to prove they are companies rather than research institutes that only burn capital. And even that is not enough: for each such project, whether the exit is an IPO or M&A, the path out remains completely unclear.
Among the founders of China’s foundation-model companies, Yang Zhilin is the youngest, born in 1992. The industry describes him as a steadfast believer in AGI and a founder with rare technical charisma. Most of his academic and professional background is tied to general-purpose AI, and his papers have been cited more than 22,000 times.
In mid-2023, China’s tech community abruptly shifted from enthusiasm to chilliness toward foundation models, and accelerating real-world applications became the prevailing, pragmatic tune. This inevitably left foundation-model CEOs torn fiercely between ideals and reality.
In China’s AI ecosystem, where everyone shouts PMF (product/market fit) and everyone shouts commercialization, this founder—a formally trained AI researcher—was not in a rush at all.
With 80 people, Moonshot AI had the smallest headcount among China’s leading foundation-model companies. Unlike its rivals, Yang did not choose the safer B2B route or try to deploy in verticals such as healthcare or gaming. He built one—and only one—consumer product: the AI assistant Kimi, which allows up to 200,000 Chinese characters to be input. Kimi is also Yang Zhilin’s English name.
Yang likes to see his company as a system that combines science, engineering, and business. You can picture it this way: above the human world, he is building an AI laboratory—one hand running experiments, the other bringing advanced technology down into the real world, exploring applications through interaction with people and putting those applications into consumers’ hands. Ideally, the former burns billions to tens of billions of dollars in capital; the latter earns that money back hundreds or thousands of times over. However you look at it, it is as thrilling as it is dangerous, like walking a tightrope.
“AI is not about what PMF I can find in the next one or two years; it is about how to change the world in the next ten to twenty years,” he said.
Such an abstract, idealistic way of thinking makes one break into a sweat worrying for him: can a young AI scientist carve out a space for survival in pragmatic China?
In February 2024, Moonshot AI closed a major funding round against market headwinds. It is understood that the company raised a Series B round of more than US$1 billion at a pre-money valuation of more than US$1.5 billion, led by Alibaba with continued participation from Monolith Management, Xiaohongshu, and others. When the deal was completed, Moonshot AI’s post-money valuation was around US$2.5 billion—making it, at that stage, the highest-valued unicorn in China’s foundation-model race. (The company declined to respond or comment on the matter.)
In this third funding round, we sat down with Yang Zhilin to talk about his first year of entrepreneurship—a condensed slice of a year in which China’s foundation-model companies sprinted forward from the starting line.
His company does not have an office at Sohu Network Plaza in Beijing, where foundation-model companies are concentrated. For a company that has raised a total of around RMB 9 billion, this office in Liangzi Xinzuo Building looks rather spartan and run-down. There isn’t even a company logo at the entrance—just a white piano standing guard by the door.
The conference room is tucked in a corner; with its small windows, it is quite dark inside, and the heater hums as it blows warm air against the winter cold.
In the dim light, Yang described what the past year has felt like to him: “It’s a bit like driving on a road with a stretch of snow-capped mountains ahead. You don’t know what’s inside them. You can only keep going forward, step by step.”
Below is the full interview with Yang Zhilin. (For readability, the author has edited some parts of the text.)
This photo was taken in early 2024 at Kimi’s first office in Beijing. They have now moved out of this location. There was no logo along the entrance—just a white piano quietly standing by the door.
Part 1
Standing at the starting point
“You have to ride the wave”
Zhang Xiaojun:
How have you been lately?
Yang Zhilin:
Busy—there’s a lot going on. But I’m still very excited. We’re standing right at the starting point of an industry, and there is still an enormous amount of room to imagine.
Zhang Xiaojun:
Just now when I walked in, I saw a pristine white piano at the entrance of your company.
Yang Zhilin:
There’s also a Pink Floyd album on it. I don’t even know who put them there—I only happened to notice a few days ago and haven’t had time to ask. (Pink Floyd is the British rock band that released the album
The Dark Side of the Moon
.)
Zhang Xiaojun:
On the day ChatGPT was released in November 2022, what were you doing?
Yang Zhilin:
I had already been preparing for this—hiring people, building the team, discussing new ideas. When I saw ChatGPT, I was incredibly excited. Three to five years earlier—even in 2021—it would have been almost unimaginable. That kind of high-level reasoning was very difficult to achieve before.
I felt that many market variables were about to change: capital on one side, talent on the other—the core factors of production for AI.
If those variables fell into place, then it would be possible to build a proper company to do this—an organization built for AGI that could go from 0 to 1.
That was a huge realization. A standalone company sounded more reasonable, but it wasn’t something you could do the moment you wanted to; ChatGPT shook up the variables and brought the factors of production closer together. You have to ride the wave.
Zhang Xiaojun:
After you decided to start an AGI company, what did you prepare? How did you gather the two factors of production—capital and talent?
Yang Zhilin:
It was a winding process. ChatGPT needed time to spread. Some people learned about it early, some late; some were skeptical at first, then shocked, then became believers. Finding people and finding money were both tightly tied to timing.
We started focusing on the first fundraising round in February 2023. If we had delayed until April, we basically would have had no chance. But doing it in December 2022 or January 2023 also wasn’t right—the pandemic was still lingering then, and people hadn’t yet digested it.
So the real window was only one month.
One night in the U.S., I did a very precise calculation. When I finished, I concluded that we needed to raise at least $100 million within a few months. Many people in the market at the time had not even started fundraising, and many didn’t believe you could raise that amount. But in the end, we did—more than that, even.
The talent market also started moving. Inspired by ChatGPT, many people in March or April 2023 came to this realization: this is the only thing worth doing for the next ten years. You have to reach the right people at the right time. One or two years earlier, talent would not have gathered to this extent. Back then, more people were working on traditional AI or AI-related areas—not general-purpose AI.
Zhang Xiaojun:
So in summary: February was the fundraising window, and March and April were the hiring window?
Yang Zhilin:
Broadly speaking, yes.
Zhang Xiaojun:
That night in the U.S.—where were you when you did this calculation? How exactly did you calculate it?
Yang Zhilin:
From late 2022 to early 2023, I was in the U.S. for about one or two months, talking with people. I did it where I was staying. You calculate how many FLOPs are needed, training costs, inference costs, and the number of users.
Zhang Xiaojun:
What was the prevailing atmosphere in Silicon Valley at that time?
Yang Zhilin:
The product started attracting many early users, concentrated within the tech circle. We were also in that circle, so we could feel it very clearly. At big companies in Silicon Valley, people had to write performance reviews every six months, and many had already started writing them with ChatGPT. Some people who had never been particularly polished in their writing were submitting reviews written by ChatGPT, and everyone looked extremely serious.
The undercurrents were surfacing. Many people were thinking about their next job or about starting a company. Quite a few people we talked with later went on to found startups. And there was a very strong sense of FOMO—fear of missing out. No one could sleep. No matter whether it was midnight, 1 a.m., or 2 a.m., if you reached out, they were always there. A little anxious, a little FOMO-driven, and very excited.
Zhang Xiaojun:
The night you calculated that you needed to raise $100 million—how late did you stay up?
Yang Zhilin:
It was fine—the calculation itself didn’t take much time. But after that, I couldn’t tell too many people. If I said it out loud, no one would believe it was possible.
Part 2
Technical Lineage
“Freeing Yourself from Endless Carving”
Zhang Xiaojun:
When venture capitalists talk about you, they say: “The founder is exceptional, has technical gravitas, and the whole team are technical stars.” So before talking about your foundation-model company, I want to start with your academic background. You studied computer science at Tsinghua as an undergraduate and earned a PhD at Carnegie Mellon’s School of Computer Science. Has AI always been your focus?
Yang Zhilin:
I was born in 1992 and started university in 2011. From my second year until now—more than a decade—I’ve been in this field. At first I explored in a fairly scattered way, looking around everywhere; I did some work related to graphs and multimodality. By 2017, I converged on language models. At that time I felt language models were a relatively important problem; later I gradually came to believe it was the only important problem.
Zhang Xiaojun:
In 2017, how did the AI field generally view language models, and how has that perception changed?
Yang Zhilin:
Back then it was just a model used to rank speech recognition results. (Laughs.) After an audio segment was recognized, you would get multiple candidate results, and then use a language model to see which result had the highest probability and output the most likely one. Its applications were very limited.
But then you realize this is a fundamental problem, because you are modeling the probability of the world. Language is finite, but it is a projection of the world; in theory, if you make the token space—the space of all possible tokens—large enough, you can build a general world model. Everything in this world comes into being and develops in some way, and all of it can be assigned a probability. Every problem can be reduced to estimating probabilities.
Zhang Xiaojun:
Your academic mentors are all very famous: your PhD advisors were Ruslan Salakhutdinov, Apple’s head of AI, and William W. Cohen, Google AI’s chief scientist. Both sit between industry and academia.
Yang Zhilin:
In the years before, industry and academia were more closely linked, but the current trend is shifting:
more valuable breakthroughs will happen in industry.
That is an inevitable law of development. It begins with exploratory research and gradually moves into a more mature industrialization process. That doesn’t mean research is unnecessary during industrialization—only that pure research will be harder to produce valuable breakthroughs.
Zhang Xiaojun:
What did you learn from these famous mentors?
Yang Zhilin:
I learned the most at Google, where I interned for a long time. I started working on Transformer-based language models in late 2018.
The biggest thing I learned was to free myself from endless “carving”—the obsession with over-refining surface details. That is very important.
You should look at what the big direction is, at what the big gradient is. When there are ten paths in front of you, ordinary people will worry about how to brake for the pedestrian ahead on this path—short-term details. But choosing which path among the ten is the right one matters most.
This field really did have that problem before. For example, on a dataset with only one or two million tokens, you would look at how to push perplexity down, how to lower loss, how to improve accuracy—and you’d fall into endless carving. People invented all sorts of strange architectures; those were all carving tricks. After carving, you could do better on that kind of dataset, but you would miss the essence of the problem.
The essence is to analyze what the field is lacking. What are first principles? Why can scaling law serve as a first principle? You only need to find a structure that satisfies two conditions: first, it is general enough; second, it can scale. General means you can model every problem within this framework; scalable means that as long as you pour enough compute into it, it will keep getting better.
This is the mindset I learned at Google: if something can be explained by something more fundamental, you should not over-carve the upper layers. There is a very important sentence I especially agree with: if a problem can be solved by scale, then don’t solve it with a new algorithm. The greatest value of a new algorithm lies in how much better it helps you scale. Once you break free from carving, you can see far more.
Zhang Xiaojun:
Was Google also a scaling-law follower at the time? How did they implement first-principles thinking?
Yang Zhilin:
Many such ideas already existed there, but Google did not execute them especially well. They had that kind of thinking, but could not organize it into a real moonshot. It was more like: here are five people pursuing my first principles, and over there are five people pursuing theirs. There was no top-down command.
Zhang Xiaojun:
During your PhD, you published collaborative papers with Turing Award winners Yann LeCun and Yoshua Bengio—and you were the lead author on those papers. How did those collaborations come together? I mean: they are Turing laureates and they were not your advisors—what did you rely on to attract them?
Yang Zhilin:
Academia is very open. As long as you have a good idea and a meaningful problem, that is enough. Two brains—or n brains—are usually more productive than a single brain. This is also true when developing AGI. An important strategy in AI is called “ensembling”—using multiple models or methods and combining their predictions to achieve better performance. Essentially it is doing the same thing: when you have many diverse perspectives, you can spark many new things. Collaboration brings enormous benefits.
Zhang Xiaojun:
Do you come up with the idea first and then ask them whether they are interested?
Yang Zhilin:
Broadly speaking, yes.
Zhang Xiaojun:
Part 4
What is Moonshot's first step — “long context” — so what is the second step?
“The next two big milestones are coming soon”
Zhang Xiaojun:
Back to when you decided to start the company—did you kick off the first fundraising round right after returning to China?
Yang Zhilin:
It started in the U.S. in February (last year), and was partly done remotely. In the end, domestic investors made up the majority.
Zhang Xiaojun:
The first round raised US$100 million, right?
Yang Zhilin:
The first round wasn't that much; the later rounds exceeded that amount. We completed two rounds in 2023, for a total of nearly 2 billion RMB.
This is currently the third round. We haven't officially announced the financing yet, so I can't comment at this time.
Zhang Xiaojun:
Some people say that from the second half of 2023 onward, no one wanted to invest in foundation model companies anymore. Were they wrong?
Yang Zhilin:
Yes, there is. You can indeed see a shift in sentiment, but it’s not that nobody is investing—at least for now, the market still has quite a lot of investment interest.
Zhang Xiaojun:
Besides capital and personnel, what other key decisions did you make in 2023?
Yang Zhilin:
What to do. That’s the advantage of companies like us—we have the technical vision to make decisions at the highest level.
We do long context. That requires judgment about the future: you need to know what is foundational and where everything will go next. Again, it’s first principles—the “de-carving” process. If you focus on carving, you can only look at what OpenAI has done and find a way to recreate it.
You’ll see that lossless long-text compression in Kimi gives the product a unique experience. When you read English papers, it helps you understand them amazingly well. Using Claude or GPT-4 today, you may not do as well; this requires laying the groundwork in advance. We’ve been working on it for more than half a year. That’s very different from seeing the long-context trend today, quickly pulling two teams together, and developing at maximum speed.
Of course, the marathon has only just begun; there will be more differentiation, and that requires you to anticipate in advance what counts as “a judgment that differs from the majority but is still right.”
Zhang Xiaojun:
In which month was this decision made?
Yang Zhilin:
February or March—the decision was finalized as soon as the company was founded.
Zhang Xiaojun:
Why is long context the first step of moonshot?
Yang Zhilin:
Because it is foundational. It is the memory of the new computer.
The memory of the old computer has increased by several orders of magnitude over the past few decades, and the same will happen with the new computer. It can solve many of today’s problems. For example, current multimodal architectures still need a tokenizer, but with lossless long-context compression, you don’t need it anymore—you can feed raw input directly. Going further, it is the foundation for making the new computation model more general.
The old computer could represent everything with 0 and 1; everything could be digitized. But the new computer of today still can’t do that—it doesn’t yet have enough context, so it isn’t general enough. To become a general world model, you need long context.
Second, it enables personalization. The core value of AI is personalized interaction; the ultimate value lies in personalization, and AGI will be more personalized than previous generations of recommendation systems.
But personalization is not achieved through fine-tuning—it is achieved by supporting very long context. Your entire history with the machine is context, and that context determines the personalization process. It cannot be copied, and it creates a more direct dialogue—a dialogue that generates information.
Zhang Xiaojun:
How much room is there to expand this further?
Yang Zhilin:
Very large. On the one hand, there is still a lot of room to expand the window itself—many orders of magnitude.
On the other hand, you can’t just keep expanding the window, and you can’t just look at the number; whether the window is currently a few million tokens or a few billion tokens doesn’t by itself mean much. You have to look at the reasoning ability it brings within that window, fidelity—meaning how closely it adheres to the original information—and instruction-following ability. You shouldn’t chase a single metric; you have to combine metrics with capabilities.
If those two dimensions continue to improve, you can do a lot. It can follow instructions tens of thousands of words long, and that instruction itself can define many agents—very personalized.
Zhang Xiaojun:
Can the technology behind long context and catching up with GPT-4 be reused for each other? Are they the same thing?
Yang Zhilin:
I don’t think so. It’s more like adding a new dimension—a dimension that GPT-4 doesn’t have.
Zhang Xiaojun:
Many people say that China’s leading foundation model companies are basically all doing the same thing—catching up with GPT-3.5 in 2023 and catching up with GPT-4 in 2024. Do you agree?
Yang Zhilin:
Improving general capability definitely has key milestones, so that statement is true to some extent—as a latecomer, you inevitably have to go through the process of catching up. But it is also one-sided. Beyond general capability, there is still a lot of room to develop specialized capabilities and achieve state-of-the-art in certain directions. Long context is one example. DALL-E 3’s image generation ability is completely outperformed by Midjourney V6. So you have to work on both fronts.
Zhang Xiaojun:
What is the ratio of time and resources devoted to general capability versus new dimensions?
Yang Zhilin:
They have to be combined. A new dimension cannot exist independently of general capabilities, so it is very difficult to give a direct ratio. But you still need to invest enough to do the new dimension well.
Zhang Xiaojun:
Will all these new dimensions be handled by Kimi?
Yang Zhilin:
Kimi is definitely a very important product for us, and we will also have some other experiments.
Zhang Xiaojun:
What do you think of Li Guangmi, founder of Shixiang, saying that the technological distinctiveness of Chinese foundation model companies is still not very high at the moment?
Yang Zhilin:
I think it’s fine—right now we’ve already created quite a lot of differentiation. It’s mainly a matter of time; this year you’ll see many more dimensions. Last year, everyone was just putting up scaffolding and first making everything run.
Zhang Xiaojun:
If the first step of moonshot is long context, what is the second step?
Yang Zhilin:
There will be two major milestones ahead. First, a truly unified world model—a model that unifies all different modalities, a truly scalable and general architecture.
Second, enabling AI to continue evolving without human input data.
Zhang Xiaojun:
How long will it take to achieve these two milestones?
Yang Zhilin:
Two to three years—maybe even faster.
Zhang Xiaojun:
Then in three years, we’ll be looking at a world that’s completely different from the present.
Yang Zhilin:
At the current pace of development, yes. Technology is in the germination stage right now, growing very fast.
Zhang Xiaojun:
Can you imagine what there will be in three years?
Yang Zhilin:
There will be a certain level of AGI. A lot of what we do today, AI will also be able to do—even better than us. But the key is how we use it.
Zhang Xiaojun:
And for you—for Moonshot AI as a company—what is the second step?
Yang Zhilin:
We will do those two things. Everything else stems from those two factors. The reasoning and agents people are talking about today are byproducts of solving those two problems. It needs a bit more refinement, but there are no fundamental barriers.
Zhang Xiaojun:
Will you go all out to catch up with GPT-4?
Yang Zhilin:
GPT-4 is a necessary stop on the road to AGI. The key is not to be satisfied with merely matching GPT-4.
First, you have to ask what the real non-consensus view is right now: beyond GPT-4, what comes next? What should GPT-5 and GPT-6 look like? Second, you have to look at what unique capabilities you have in that—and that matters more.
Zhang Xiaojun:
Other foundation model companies announce their capabilities and model rankings. It seems you haven’t done that?
Yang Zhilin:
Chasing leaderboard rankings doesn’t mean much. The best leaderboard is the users—you should let users vote. A lot of leaderboards have problems.
Zhang Xiaojun:
Is becoming the first foundation model company in China to reach GPT-4 your goal? Does being fast or slow make a difference?
Yang Zhilin:
Definitely. Over a long enough timeframe, everyone will get there. But the question is whether you’re ahead or behind, and by how much. A six-month gap or more is meaningful—and it also depends on what you can do with that time.
Zhang Xiaojun:
When do you expect to reach GPT-4?
Yang Zhilin:
Probably pretty soon, but I can’t announce a specific time to the public.
Zhang Xiaojun:
Will you be the fastest?
Yang Zhilin:
That has to be judged flexibly—but we have a real chance.
Zhang Xiaojun:
After launching Kimi, what is your North Star metric?
Yang Zhilin:
Right now it’s making the product better and adding more dimensions. For example, we shouldn’t just fight to the death over a search scenario—in the future, search will only be a very small part of this product’s value; the product has to have much greater growth potential.
Being 10% or 20% better than a traditional search tool isn’t much—only something truly disruptive deserves the three letters "AGI".
Your unique value is your incremental intelligence. You have to hold onto this point: intelligence is always the core added value. If only 10%–20% of the product’s core value comes from AI, then it won’t hold up.
Part 5
I Am Completely Unconcerned About Commercialization
“User scale and model scale need to happen at the same time”
Zhang Xiaojun:
Mid-2023 was a huge turning point—the market shifted from fever pitch to a rapid cooldown. How do you view that?
Yang Zhilin:
I don’t fully agree with that description—we still completed a funding round in the second half of the year. And new things kept emerging. The model capabilities today are unimaginable compared with the end of last year. The user numbers and revenue of more and more AI companies continue to grow. It keeps proving its value.
Zhang Xiaojun:
For you, what was different between the first half and the second half of the year?
Yang Zhilin:
Not much changed. Of course there are variables, but in the end you come back to first principles—how to deliver a good product to users. After all, we have to meet user needs, not win a race.
We are not a company built to compete.
Zhang Xiaojun:
The industry believes that the notable difference between the first half and the second half of 2023 is a shift in focus: the first half was more AGI-oriented; the second half shifted toward how to bring applications into practice and commercialize them. Did you make that shift?
Yang Zhilin:
Of course I’ll do AGI—that’s the only thing that matters over the next ten years. But that doesn’t mean we won’t build applications. Or rather, it shouldn’t be defined as an “application.”
“Application” sounds like you have a technology and then go looking for a place to use it, with a closed commercial loop. But “application” isn’t the right word. It and AGI complement each other. It is both a means to reach AGI and the purpose of reaching it.
“Application” sounds more like a goal: I want to make it useful. You have to combine Eastern and Western philosophy—you have to make money, and you have to have ideals.
Today, users help us discover scenarios we never thought of. Some people use it to screen résumés—a thing we never thought about when designing the product, but it works naturally. Conversely, user input makes the model better. Why is Midjourney so good? It has scaled on the user side—user scale and model scale need to happen at the same time. Conversely, if you only focus on applications and ignore the iterative improvement of model capabilities—ignore AGI—then your contribution will be very limited.
Zhang Xiaojun:
The management partner at GSR Ventures, Zhu Xiaohu, only invests in foundational model applications. One of his views is: the hardest core problem is AIGC PMF—if ten people can’t find PMF, then a hundred people won’t either; it has nothing to do with headcount or cost, so don’t burn money on it. He said, “Training on LLaMA for two or three months and you can at least reach the level of the top 30 people—it can replace people immediately.” What do you think of this view?
Yang Zhilin:
AI is not about finding some PMF in one or two years; it is about how to change the world in ten to twenty years—these are two different ways of thinking.
We are steadfast long-termists. When AGI or something even more powerful is achieved, everything today will be rewritten. PMF is certainly important, but if you rush to chase PMF, you are very likely to be hit by another “dimensionality-reduction attack”—crushed by a technology of higher dimensionality. That has happened too many times. In the past, many people built customer service and dialogue systems, slot-filling systems—some of those companies were fairly sizable. But all of them were swept away by a strike at a higher dimension. Very painful.
That is not to say that approach has never worked. Suppose today you find a scenario where current technology is enough, where the 0-to-1 value add is huge and the 1-to-n space is not too large—that scenario is fine. Midjourney is like that, or ad content generation—relatively simple tasks with a very clear 0-to-1 effect. That is an opportunity for the application-only camp. But the biggest opportunity is not there. If your premise is commercialization, you cannot think about it separately from AGI. If I only build applications right now—fine, but in one year you may be crushed.
Zhang Xiaojun:
You can quietly upgrade the foundational model underneath, right?
Yang Zhilin:
But that approach can never become bigger than the model itself. Technology is the only new variable of this era; the other variables do not change. Going back to first principles, AGI is the core of everything. From that, we infer: a super app will definitely require the strongest technical capability.
Zhang Xiaojun:
Can open-source models be used? (The latest news is that Google has announced the open-source model Gemma.)
Yang Zhilin:
Open source is lagging behind closed source—that is also a fact.
Zhang Xiaojun:
Could that lag only be temporary?
Yang Zhilin:
At least up to now, it doesn’t seem so.
Zhang Xiaojun:
Why can open source not catch up with closed source?
Yang Zhilin:
Because the way open source develops today is different. In the past, everyone could contribute to open source; today, open source itself is still centralized.
A lot of open-source contributions probably have not been validated by compute. Closed source benefits from the concentration of talent and capital, so in the end closed source will definitely be better—that is a process of centralization.
If I have a leading model today, opening its source is very likely irrational. Maybe only those who are behind would do that, or they open-source a small model—to stir up the market; after all, if you don’t open-source it, then it has no value anyway.
Zhang Xiaojun:
How should one deal with anxiety in China? People say that a foundational model company, if it cannot quickly create commercial scenarios and products that meet investor expectations, will find it hard to raise the next round.
Yang Zhilin:
There needs to be a balance between long term and short term.
Absolutely no users and no revenue definitely won’t do.
As we’ve seen, going from GPT-3.5 to GPT-4 unlocked a lot of applications; from GPT-4 to GPT-4.5 and then GPT-5, it may very likely continue to unlock more—and even more exponentially. The so-called “Moore’s law of scenarios” means the number of usable scenarios will increase exponentially over time. We need to improve model capabilities while also finding more scenarios—a balance like that.
That is a whirlpool. It depends on how much of your investment is allocated to the short term and how much to the long term. You pursue the long term on the condition that you can survive. The long term absolutely cannot be ignored, otherwise you will miss an entire era. Drawing conclusions at this point is truly still too early.
Zhang Xiaojun:
Do you agree with the “two-wheel drive” idea proposed by Wang Huiwen, Meituan co-founder and founder of Light Year?
Yang Zhilin:
That is a good question. To some extent, that logic makes sense. But the way you execute it is what creates a very big difference. Can you really make those “non-consensus bets with favorable odds” or not?
Zhang Xiaojun:
As I understand it, their two-wheel drive also requires quickly finding new application scenarios; otherwise, the technology has no way to land.
Yang Zhilin:
In the end, it is still the difference between scaling models and scaling users.
Zhang Xiaojun:
In China, besides you, who else is taking the model-scaling path?
Yang Zhilin:
That is not something I can assess.
Zhang Xiaojun:
Maybe most people are taking the user-scaling path. Or put it another way: is this the difference between the academic camp and the commercialization camp?
Yang Zhilin:
We are not the academic camp. The academic camp is definitely not efficient.
Zhang Xiaojun:
Many foundational model companies monetize through B2B—after all, B2B offers higher certainty. What about you?
Yang Zhilin:
We do not. From day one, we decided to go B2C.
It depends on what you want. If you know something is not what you want, you won’t have FOMO—because even if you get it, it doesn’t mean much.
Zhang Xiaojun:
Over the past year, have you felt anxious?
Yang Zhilin:
More excited and exhilarated than anxious. Because I’ve been thinking about this for a very long time. Maybe we were one of the earliest people who wanted to explore the dark side of the Moon. Today, you realize you’re actually building a rocket, and every day you discuss what kind of fuel to add to make it go faster—and how to keep it from exploding.
Zhang Xiaojun:
To sum up the “non-consensus favorable-odds bets” you have made—besides B2C and long context, what else is there?
Yang Zhilin:
There are many more being implemented; I hope to share them with everyone soon.
Zhang Xiaojun:
The previous generation of entrepreneurs in China tasted success with applications and scenarios, so they focused more on products, users, and data flywheels. Can the new generation of AI entrepreneurs that you represent create a new future?
Yang Zhilin:
We also care deeply about users. Users are our ultimate goal, but it is also a process of co-creation.
The biggest difference is that this time it will be more technology-led.
It is still the same horse-and-carriage versus automobile question: we are in the leap from horse-drawn carriages to cars, and we should focus as much as possible on how to bring users a car.
Zhang Xiaojun:
Do you feel lonely?
Yang Zhilin:
Ha ha ha… That’s quite an interesting question. I think I’m fine, because I still have dozens—almost a hundred—people fighting alongside me.
Part 6
Before we could catch up with GPT-4, Sora appeared
“Right now is like the GPT-3.5 moment for video generation”
Zhang Xiaojun:
Sora suddenly appeared this year—how much of that was within your expectations, and how much was not?
Yang Zhilin:
That generative AI could achieve this effect was within expectations; what was surprising was the timing—it came earlier than we estimated.
That also reflects how quickly AI is developing now: a lot of the returns from scaling have yet to be fully harvested.
Zhang Xiaojun:
Last year, the industry believed that foundation models in 2024 would definitely compete fiercely around the multimodal narrative, and that video generation quality would improve as quickly as text-to-image did in 2023. Did Sora’s technical capability exceed, meet, or fall short of your expectations?
Yang Zhilin:
It solved many problems that were very difficult before. For example, maintaining consistency in the generation process over a relatively long period—that is the key point, and it is a very big improvement.
Zhang Xiaojun:
What does that mean for the global industry landscape? In 2024, what new narratives will foundation models bring?
Yang Zhilin:
First, short-term application value: it can continue improving efficiency in production workflows, and of course we also expect more extensions to be built on existing capabilities. Second, combining with other modalities. It itself is a world model; with that knowledge, it is a great complement to existing text. On that basis, there is still a lot of room and opportunity, whether in agent or in connecting with the physical world.
Zhang Xiaojun:
Overall, how do you assess Sora?
Yang Zhilin:
We had also planned a similar direction and had been working on it for some time. In terms of direction, it was not particularly surprising—mainly the technical details.
Zhang Xiaojun:
Which technical details are worth learning from?
Yang Zhilin:
OpenAI has also not fully explained many of those points. They only describe the broad strokes; the key details still have to be inferred from its output or from the information available, combined with our previous experiments. At least for us, it provides more data points for the development process—more data input.
Zhang Xiaojun:
Compared with text generation, what are the main bottlenecks in video generation? What solution do you think OpenAI found this time?
Yang Zhilin:
The biggest bottleneck—at its core—is still data: how do you match data at scale? That had not been verified before, especially when the motion is complex and the generated results are as realistic as photographs.
Under those conditions, scaling is possible—that’s what it solved this time.
What still hasn’t been solved includes, for example, the need for a unified architecture. DiT is still not a sufficiently universal architecture.
Modeling the marginal probabilities of purely visual signals can do very well, but how do you generalize that into a new, universal computer? A more unified architecture is still needed—there is still room there.
Zhang Xiaojun:
Have you read OpenAI’s Sora report, “Video generation models as world simulators”? Which points in it are worth emphasizing?
Yang Zhilin:
I have read it. Given the current competitive situation, they definitely would not write down the most important points. But it is still very worth learning from. Basically, this is paid content—things that otherwise you might have to spend a lot of experiments to learn. Now you can know part of it without paying for those experiments, and form a preliminary understanding.
Zhang Xiaojun:
What key signals did you take from it?
Yang Zhilin:
That this thing can scale to a certain extent. In addition, it also gives a fairly specific description of how the architecture is built. But it is also possible that different architectures do not make such an essential difference for this problem.
Zhang Xiaojun:
Do you agree with their statement—“Scaling video generation models is a promising path towards building general-purpose simulators of the physical world”?
Yang Zhilin:
I completely agree. These two things optimize the same objective function—that’s hardly in doubt.
Zhang Xiaojun:
What do you think about Yann LeCun speaking out again against generative AI? His view: “Modeling the world by generating pixels is wasteful and will definitely fail. Generation has always worked with text because text is discrete, with a finite number of symbols. In that case, handling uncertainty in prediction is easy; handling predictive uncertainty in continuous, high-dimensional sensory inputs is intractable.”
Yang Zhilin: Now I think that when you model the marginal probability of video, it is essentially lossless compression — there is no essential difference from next-token prediction in language models. As long as you compress well enough, you can explain anything in this world that can be explained.
But there are still important jobs unfinished: how does it combine with the capabilities that have already been compressed? You can think of that as two different kinds of compression. One is compressing the raw world — that is what video models do. The other is compressing the behaviors created by humans, because human behavior has gone through the human brain — the only thing in the world that creates intelligence.
You can see video models as doing the first kind and text models as doing the second, even though video models also contain part of the second kind: some videos created by humans contain the intelligence of their creators. In the end, it will probably be a mixture — you need to learn from multiple angles through both approaches, and both contribute to the growth of intelligence.
So generation is probably not the goal; it is merely a compression function. If you compress well enough, generation will eventually be very good. Conversely, if a model itself cannot generate, can it still compress extremely well? That is doubtful. Perhaps being able to generate very well is a necessary condition for compressing very well.
Zhang Xiaojun:
Sora and ChatGPT from last year were two different milestones. Which one is bigger?
Yang Zhilin:
Both are very important.
Right now it is quite like the time of GPT-3.5 for video generation — a step change.
The model is still relatively small, and it is foreseeable that there will be larger models, which means capability improvements are definitely on the way.
Zhang Xiaojun:
Some people also believe that, in terms of multimodality, Google Gemini's breakthrough is even more important.
Yang Zhilin:
Gemini follows the GPT-4V direction and also includes that understanding. Both are important; the final step is to put all of these things into one model, and that still has not been solved.
Zhang Xiaojun:
Why is it so hard to put them into one model?
Yang Zhilin:
No one knows how yet. There still is no architecture that has been validated.
Zhang Xiaojun:
What will Sora + GPT create?
Yang Zhilin:
Sora can be applied directly to video production, but if combined with language models, it can connect the digital world and the physical world. You can also complete tasks in a more end-to-end way, because your world modeling is now better than before — it can even be used to improve multimodal input understanding. So, in the end, you can transition quite flexibly between modalities.
In short: you understand the world better; you can do more end-to-end tasks in the digital world; and you can even build a bridge to the physical world to complete tasks there. This is the starting point.
For example, autonomous driving, or some household chores — in theory, all of these are cases that connect to the physical world. So the breakthrough in the digital world is certain, but it also contains the potential to become a path toward the physical world.
Zhang Xiaojun:
What does Sora mean for China's foundation model companies? What is the right way to respond?
Yang Zhilin:
It does not change much. This has always been a certain direction.
Zhang Xiaojun:
China's foundation models still have not caught up with GPT-4, and now Sora has appeared. What do you think? The two worlds seem to be drifting further apart — are you worried?
Yang Zhilin:
That is simply an objective fact. But the actual gap is probably still narrowing — that is the law of technological development.
Zhang Xiaojun:
What do you mean? That the technology curve is steep at first and then gradually flattens out?
Yang Zhilin:
Exactly. I am not too surprised either—OpenAI has always been developing next-generation models. Objectively, that gap will continue to exist for a while, and the gap among different Chinese companies will also continue to exist for a while—this is a period of technological boom.
But in another two or three years, it is very likely that leading Chinese companies will do better in the foundational work here — including technical infrastructure, talent reserves, and the accumulation of organizational culture. With that process of refinement, they will have more chances to lead in certain aspects — but a little patience is needed.
Zhang Xiaojun:
Could China and the U.S. eventually form completely different AI technology ecosystems?
Yang Zhilin:
The ecosystems may be different if viewed from the perspective of product and commercialization. But from a technology perspective, the general capabilities will not follow completely different technical paths — the basic general capabilities will definitely be similar. Because the AGI space is so broad, differentiation built on top of general capabilities is more likely to happen.
Zhang Xiaojun:
In Silicon Valley there has long been a debate: “one model rules them all” or “many specialized models” — one general model for all kinds of tasks, or many small specialized models for specific tasks. What is your view?
Yang Zhilin:
I stand with the first view.
Zhang Xiaojun:
On this point, will China and the U.S. diverge significantly?
Yang Zhilin:
I don't think so, in the long run.
Part 7
I Accept That Failure Is a Possibility
“It Changed My Life”
Zhang Xiaojun:
A startup built on a platform model in China is a rather unusual kind of beast: you’ve raised a lot of money, but it seems most of it has gone into scientific experiments. How did you convince investors to open their wallets under those circumstances?
Yang Zhilin:
Not much different from in the US.
The amount of money we’ve raised so far still isn’t that much.
So there’s still a lot we need to learn from OpenAI.
Zhang Xiaojun:
I want to know: how much more money does it take to get to GPT-4? How much to get to Sora?
Yang Zhilin:
Neither GPT-4 nor Sora requires that much. The money now is mainly reserved for the next generation—or the generation after that—of models, for frontier exploration.
Zhang Xiaojun:
Chinese platform-model startups have taken money from the big tech giants, but the giants are also training their own models. How do you see the relationship between platform-model startups and the giants?
Yang Zhilin:
It’s both competition and cooperation. Giants and startups have different top priorities. Look at any major tech company today: its number-one priority is different from an AGI company’s number-one priority. Your number-one priority shapes your actions and outcomes, and ultimately determines the different relationships in the ecosystem.
Zhang Xiaojun:
Why do the giants make small investments across many platform-model companies instead of placing a big bet on one company?
Yang Zhilin:
That’s a matter of stage.
Later on, there will be more consolidation—and fewer companies.
Zhang Xiaojun:
Some people say the endgame for platform-model companies is to be acquired by a giant. Do you agree?
Yang Zhilin:
I don’t think that’s necessarily the case. But they’re very likely to have very deep partnerships.
Zhang Xiaojun:
For example, how could they work together?
Yang Zhilin:
OpenAI and Microsoft are the classic partnership model. Much of it can be referenced, and some of it can be improved.
Zhang Xiaojun:
Over the past year, where have the twists and turns of entrepreneurship shown up for you?
Yang Zhilin:
There are many external variables—capital, talent, GPU, product, R&D, technology. There are standout moments, and there are also difficulties to overcome. Take GPU as an example.
There have been lots of ups and downs: supply was tight for a while, then it improved again.
The most extreme phase was when prices changed every day—a machine that cost 260 today might be 340 tomorrow, then drop again two days later. It was a constantly changing situation. You had to keep up closely.
Prices kept changing, so strategies kept changing too: which channel to use, whether to buy or rent—there were many different choices.
Zhang Xiaojun:
What drove these fluctuations?
Yang Zhilin:
Geopolitical reasons; the production process itself happens in batches; and market sentiment also played a role. We observed many companies starting to return GPUs, realizing they didn’t necessarily need to use them to train this model. As market sentiment and people’s decisions changed, supply and demand changed too. The good news is that overall supply has improved a lot recently.
My personal assessment is that, at least over the next one to two years, GPUs won’t be a major bottleneck.
Zhang Xiaojun:
You always seem to be thinking about organization. How did you build the team?
Yang Zhilin:
Our hiring approach has changed over time. AGI talent worldwide is extremely scarce, and people with relevant experience are very few. Our initial hiring profile focused on finding geniuses with directly relevant skills. That turned out to be very successful. People who had “operated on” models and had hands-on experience training super-large models could do everything very quickly. Including the launch of Kimi—capital efficiency and organizational efficiency were both really high.
Zhang Xiaojun:
How much did it cost?
Yang Zhilin:
A rather small amount—relative to many other expenditures, it was doing big things with little money. For a long time we only had 30–40 people. Now it’s 80. We pursue talent density.
Later, the talent configuration changed. In the beginning, we hired geniuses, believing their ceiling was very high—the ceiling of a company is determined by the ceiling of the people in it. After that, we expanded the team across more dimensions—people on the product-operations side, leader-type people, people who can push things to the limit. Now it’s a more complete team, with more endurance and the ability to fight.
Zhang Xiaojun:
After a year of platform-model entrepreneurship in China, how do you assess the current milestone result?
Yang Zhilin:
We’ve built a rocket prototype and are now test-firing it. We’ve assembled the team, found some fuel formulas, and basically can see an embryonic PMF. You could say we’ve taken the first step in a Moon mission.
Zhang Xiaojun:
What do you think of Yann LeCun’s view? He’s not optimistic about the current technical path—he believes self-supervised language models cannot acquire real world knowledge, and as models scale, the probability of errors—machine hallucinations—will only increase. He proposed the idea of a “world model”.
Yang Zhilin:
There’s no fundamental bottleneck at all. When the token space is large enough, it becomes a new kind of computer for solving general problems—and that is exactly a general world model.
An important point behind what he said: everyone can see the current limitations. But the solution does not necessarily require a completely new framework. The only thing that works in AI is next-token prediction plus the scaling law. As long as the tokens are sufficiently complete, everything can be done. Of course, the problems he points out do exist right now—but you solve them by making the token space truly general. That’s all.
Zhang Xiaojun:
So he’s exaggerating the limitations.
Yang Zhilin:
I think so. The underlying principle is correct—it's just that there are still some small technical issues that haven't been resolved.
Zhang Xiaojun:
What do you think of Geoffrey Hinton, the godfather of deep learning, constantly calling attention to AI safety?
Yang Zhilin:
The fact that he focuses on safety actually shows that he has great confidence in the upcoming improvement of technical capabilities. The two are opposites.
Zhang Xiaojun:
How do you solve the hallucination problem?
Yang Zhilin:
It's still scaling law—it's just that what gets scaled up is something different.
Zhang Xiaojun:
What is the probability that scaling law ultimately turns out to be a dead end?
Yang Zhilin:
Approximately zero.
Zhang Xiaojun:
What do you think of former CMU alumnus Qi Lu's view: OpenAI will definitely be bigger than Google—it's just a matter of one, five, or ten times bigger?
Yang Zhilin:
The most successful AGI company in the future will definitely be larger than any company today—that is beyond doubt.
In the end, it could be a matter of two or three times GDP. It may not be OpenAI; it could be another company. But there will definitely be such a company.
Zhang Xiaojun:
If you happen to become the CEO of this AI empire, what would you do to protect humanity?
Yang Zhilin:
Thinking about that question right now still lacks some prerequisites. But we are certainly ready to cooperate with and learn from various actors in society, including putting more safety measures into the model.
Zhang Xiaojun:
What are your goals for 2024?
Yang Zhilin:
First, technical breakthroughs—right now we should already be able to build a much better model than in 2023. Second, users and product—I hope for more users at scale and stronger retention.
Zhang Xiaojun:
What is your prediction for the global foundation model industry in 2024?
Yang Zhilin:
This year, more capabilities will emerge, but the landscape will not be very different from the current one—the few leaders will still lead. On the capability side, the second half of the year may see some fairly big breakthroughs, many of them from OpenAI; it definitely has a next-generation model—maybe 4.5, maybe 5. That feels like a high-probability event. Video generation models will certainly still be able to continue scaling.
Zhang Xiaojun:
What about your prediction for China's foundation model industry in 2024?
Yang Zhilin:
First, you will see new, differentiated capabilities emerge. Chinese models—thanks to earlier investment and the right teams—will achieve world-leading capabilities along certain dimensions. Second, products with much larger user bases will emerge—the likelihood of this is very high. Third, there will be more consolidation and differentiation in route choices.
Zhang Xiaojun:
When starting this company, what were you most afraid of?
Yang Zhilin:
Actually, not much—just charge forward fearlessly.
Zhang Xiaojun:
Would you like to say anything to your colleagues?
Yang Zhilin:
Let's keep working hard together.
Zhang Xiaojun:
State one question about the foundation model industry that you هنوز don't know the answer to but most want to know.
Yang Zhilin:
I don't know what the limits of AGI will look like — what kind of company it will create, and what products that company will make. That's what I most want to know right now.
Zhang Xiaojun:
As AGI continues to develop like this, what is the one thing you least want to see?
Yang Zhilin:
I'm quite optimistic about that. It may bring human civilization to the next stage.
Zhang Xiaojun:
Has anyone ever said you're too idealistic?
Yang Zhilin:
We are also very pragmatic. In fact, we have already made some real things—we're not just talking.
Zhang Xiaojun:
If the money you have raised today were the last money you ever had, how would you spend it?
Yang Zhilin:
I hope that will never happen, because in the future we will need even more money.
Zhang Xiaojun:
If you don't succeed, would you consider yourself a failure?
Yang Zhilin:
That's not too important either—I accept that failure is a possibility.
This journey has completely changed my life, and I am immensely grateful.
(End)
Information only. No investment, legal, tax, or financial advice.