梁文峰長談:為什麼 DeepSeek 堅持開源與克制?(投資人交流會全文)
在一場投資人交流會上,梁文峰談到了公司的願景、開源與克制、AGI 路線圖、商業化節奏、c…

梁文鋒
歡迎,各位投資者。當我們最初開始這家公司時,我們的初衷並不是去想我最終能賺多少錢,或者我們是否會走向資本市場、上市,或者諸如此類的事情。所以我們一開始並沒有這樣的意圖。最初加入我們的前幾十個人也從來沒這樣想過。要是他們那樣想,他們就不會來了。
所以總體來說,我們做這件事時,對世界抱有很大的善意,我們也覺得這對人類是有用的,是超越金錢的東西。當然,後來我們發現這件事情帶來的好處非常大之後,也出現了其他誘惑;那是另一回事。但我們的初衷、我們的願景,以及我們一直延續到今天的願景,都不是建立在商業利益最大化之上的。
我認為這一點很重要。大約二十年前,我在管理方面最欽佩的人是通用電氣前首席執行官傑克·韋爾奇。回頭看,他說的大多數話也許都是錯的,但他最重要的一點是對的:對於一家公司來說,最重要的是願景。管理一家大公司需要什麼?不是你的規章制度,而是你的願景。什麼是願景?
願景不是牆上的標語;願景是你怎麼做事,而不是你怎麼說,也就是你實際上如何運作。總之,我忘了傑克·韋爾奇的原話,但大致就是這個意思。那麼我們如何管理這麼多人並把他們組織起來?其實我們不是按傳統意義上的組織方式在組織;我們是由願景驅動的,是圍繞著一個願景來組織的。我們沒有傳統意義上的組織。這既有優點,也有缺點。
未來,我們會努力揚長避短,但這就是我們的特點。我們做事不是那種「我需要達成某個 KPI,沒有考核」的方式,而只有一個願景。這個願景甚至都沒有寫下來;它從來沒有被整理成任何正式文件。這個願景存在於我們做事的方式中,也存在於我們對世界的態度中。
也許我們公司裡每個人對這個願景的理解都不一樣,也許每個人的願景也不同,但在大層面上它們是一致的。
我覺得歸根到底還是要對世界抱有很大的善意,想要做一些事情。這就是我們把自己凝聚在一起的方式。接下來我先講,等我講完大家可以提問。我接下來也許會圍繞這個願景繼續談後面的內容。這個願景是真實的,不是編造出來的;我們真的是這樣想,也真的是這樣做。否則你無法解釋我們做的很多事情。為什麼我們如此堅持開源?
因為這個願景本身就要求開源。沒有這個願景,你就無法組織人。例如, 也開源,但 的開源和我們的不一樣。** 的開源感覺有點勉強;他們覺得這不是他們的初衷。但對我們來說,這就是我們的初衷。而且在開源這件事上,我們從一開始就非常清楚。
第一,是願景;第二,我們認為要讓 AI 在商業上成功,開源是有好處的。這聽起來可能有點矛盾,有點反直覺,因為歷史上開源和商業化一直是對立的。但我認為 AI 和以前不一樣。
歷史上,一家軟件公司的市場一年可能只有幾十億美元;如果它開源了,可能就會消失,最後只剩下幾千萬或幾億美元。但 AI 足夠大,最終它可能會占到人類社會 GDP 的百分之十,例如。那其實是一個非常大的數字。
一個人不可能壟斷這件事;你無法壟斷它。你必須和別人分享,否則你肯定活不下來。這和以前開源一個軟件不一樣,因為那個軟件市場根本沒有那麼大。但 AI 就是太大了。如果我們想壟斷收益,那我們一定會被歷史拋下。我認為這主要是一種客觀規律,一種歷史視角。
不是說我不開源就可以壟斷市場;理論上這不符合客觀實際。你一定會遇到很多阻力,而且一定還會有其他方式阻止你達成那個目標。在這種情況下,我認為你不一定要沿用傳統的商業思維。你需要一種機制,來保證你能獲得的利益是有限的;只有這樣才有可能成功。約束是必要的,我認為約束是必要的。
如果我們想讓 AI 在我們手裡成功,首先我認為就需要約束。我們不能想著人類 GDP 的某個百分比是屬於我的,或者中國 GDP 的某個百分比是屬於我的。你越這麼想,就越不可能成功。
所以從一開始我們就覺得,約束是必要的。你越克制,越有可能把事情做成。這當然是一種商業考量,但同時也是一種宏觀層面的考量。我覺得這是直覺,至少對我來說是直覺,或者至少這確實是我真實的想法。我們沒有太多其他優勢;我們並不特別有能力,也不比別人更有錢,我們的人也不比其他公司的更好。其實,完全不是。
想想看:當我們兩年前創辦這家公司的時候,我們沒有多少錢,沒有多少芯片,沒有多少知名度,也沒有多少影響力。我們只是一群非常普通的人。
我們真的只是一群普通人。如果比較偏好的敘事是:一群普通人做出了不平凡的事情,而不是一群天才做出了不平凡的事情,那這其實與我們的克制非常相關,也與我們的克制和我們的願景是一脈相承的。所以,開源和商業化會衝突嗎?我認為在 AI 裡,如果你不克制,你就起不來。
開源是克制的一部分,而我們的克制不僅體現在開源上;也體現在很多其他方面。但總的來說,我們沒必要去糾結開源或克制。你越克制,越可能成功,至少到目前為止是這樣;到目前為止,這是說得通的。
否則就沒法解釋我們為什麼能成功:我們沒有什麼特殊武器,我們的起點很低,資源也非常有限,而且我們的人其實就是一群隨機湊在一起的普通人。我自己也只是個大學畢業生,不是那種最頂尖學校出來的。這種克制也是我們願景的一部分。AI 太大了;好處也太大了。
我們非常克制。只要能把事情做成,最後的收益就會非常大。哪怕你只拿一小部分,收益也已經很大了,所以根本沒必要去想到底拿哪一部分的收益,或者怎麼去拿。我覺得完全沒必要去想這個,因為收益已經足夠大了。只要你拿一點點,那都已經綽綽有餘了。所以我們之前說只賺合理利潤,其實只是看你的意願,不是看最大化利潤——這是不一樣的。
這不是我們的 API 定價。我們的 API 定價,是按我們認為的合理利潤來定的:大致足以在市場上買一批設備,並在十個月內收回成本。我認為這就是合理利潤。
在當前情況下,考慮到風險、前期投入等等,如果一臺伺服器在財務上按三年或五年折舊,但在商業上我們按十個月收回成本來看,我們覺得這就夠了。OK,我們覺得這就夠了。這就是我們現在 API 定價的邏輯。我們的 V3.2 Flash 和其他產品,都是按照十個月收回設備成本來定價的。
這就是我們的標準。這其實並不是利潤最大化;如果是利潤最大化,我們應該把價格定得更高。因為在這個價格區間裡,用戶需求是缺乏彈性的:如果我把價格提高 50%,甚至翻一倍,token 消耗量的變化也不會有那麼大。如果我把價格提高到兩倍,我的總收入就接近翻倍了。等一下,我確認一下。啊,這太好了。
我給你講個故事:我們的一個模型。一開始,我們擔心需求太大,所以最開始把價格定得比較高,團隊裡大家都不太高興。後來我又把價格降回來,降到四分之一,大家就非常高興。我覺得這反映了我們真正的想法。
我剛才說的關於願景的意思是,我們還是希望這個東西對人有用,不是為了賺最多的錢,而是讓大家都能用得起,同時還能賺取合理利潤。我覺得,從我們公司裡其他人的想法來看,當時我們降價的時候,公司群裡很多人都在歡呼;大家都很高興。因為這就是我們投入那麼多精力和心血把這個模型做好的目的。
目的就是要讓它非常便宜、非常有效,讓大家都能充分使用。我們對此感到非常高興。這就是我們的動機,這就是我們的願景,也是我們公司內部對於為什麼能聚在一起做這件事的共識。我覺得這是一個比較特別的點,因為對我們的競爭對手來說,降價絕對不是什麼好事;他們肯定不會歡呼。
因為你的收入,你的 ARR,如果你砍一半,那 ARR 就減半。對,這就是我們和別人的一個不同。我們覺得這就夠了。從公司內部看,十個月收回成本,在商業上已經很讓人滿意了。對外看,我們也覺得這個價格會讓大家更開心、更願意使用;這是雙贏,對公司、對社會、對大家都是雙贏。
我覺得,OK,剛剛螢幕上有人留言說,十個月收回成本的利潤太高了。確實還有降價空間;確實還有降價空間。模型本身也還有優化空間,所以整體來說,降成本的空間還是挺大的。但這個成本、十個月收回,是我們自己能做到的;其他公司做不到。比如 Alibaba 或 Tencent 不是我們的最優配置;他們的成本應該會比這個高好幾倍。
這裡還有很多優化工作。你剛才問我們為什麼不繼續降價。那是因為沒有彈性。也就是說,如果我再降一次價格,需求不會進一步增加,或者我再降一次,需求只會增加一點點。因為這個價格大家都買得起,而且大家都覺得價格很滿意;不會因為太貴而不用了。
所以降價也不會給公司帶來更多收入,也不會給社會帶來更多價值,因為大家對這個價格已經很滿意了。如果價格再低一些,社會上大家的幸福感也不會再提升多少。對,OK,不過剛才這個問題,說到定價,我們肯定不是從
最大化公司營收或利潤。那不是出發點。這也是我們克制的一部分,因為短期內,如果你把價格定高一點,可能會多賺一點;但長期來看,就不太好說了。因為我認為克制是一種策略。對我來說,克制就是一種策略。有時候你可以放棄一些東西,去換取另一些東西更多的收益。
開源其實也是一樣;它也可以看作是對我們的一種壓力,或者說是我們做出的一種讓步。首先,這種讓步,從公司內部來看,會讓我們非常開心;大家都很開心,員工會有很強的成就感,我們也會因此更加有凝聚力。並且,這種讓步對社會是有好處的;社會也會很開心,其他同行,或者普通人,大家都會很開心。
所以我把這種克制理解成一種東西,從長期來看,它可以提高我們在 AGI 上成功的概率。當我考慮一件事情時,我毫不懷疑 AGI 會有非常大的商業價值。在這個基礎上,我優先考慮的不是我如何能拿到更大的一塊,如何能分到更多的蛋糕;我優先考慮的是如何提高我真正把它做成的概率。這種克制也可能體現在很多其他方面。
比如說,去年春節的時候,我們突然有了很多用戶,但我們並沒有去追求留住這些用戶,也沒有去變現它們,或者試圖從中攫取商業利益。我們沒有去搶用戶或者賺錢,而是很努力地把這些用戶服務好。
我們不會有那種想法,像是“我要做下一個超級應用”,或者“我要去和誰競爭”,或者“我要做下一個 ByteDance 或者下一個 Tencent”。絕對不會。我們可以那麼做,但我們沒有那麼做。我理解這也是克制的一部分。不要覺得你一定要把所有東西都變現;一旦有了用戶,似乎就能做出下一個 ByteDance,然後你就一口把它吃掉。
我覺得這在商業上是可行的;是有可能的。如果去年我們花很多錢去和 ByteDance 搶用戶,這也會是一個可能的策略。但我們選擇了一種非常克制的做法:我不跟你搶那個,因為後面可能是西瓜,前面可能全都是芝麻。
我不應該去抓每一顆芝麻。當然,有些芝麻可能也相當大,但我認為後面的 AI 機會,相比前面的機會並不小。現在回頭看,去年沒有在消費端猛推,也許是對的選擇。因為很明顯,後面確實還有更大的西瓜,前面真的就只是小芝麻而已。如果我去年有很多錢,把這件事做得很大,那又有什麼好處呢?你也不會多得到什麼。
這些都是我真實的想法,因為我認為後面的 AGI 機會應該會非常大;後面的 AGI 機會一定會非常大。
我甚至都不需要去想,我會不會在裡面佔據一個位置,或者我在裡面的商業模式會是什麼;我們真的不需要想這些。只要有這樣巨大的商業機會,你一定會找到辦法。至於前面的芝麻,我們也會撿,但我們只是順手撿一點;我們不會停下來把它當成一件重要的事。所以去年那些消費端日活之類的,我覺得可能就是一件小事。
但我們也撿了,並且以相對較低的成本維持了用戶使用,因為它們之後也許有用。雖然目前我們還不知道這些用戶有什麼用,現在看起來就是純粹的成本支出,但之後也許會有用。既然是順手可以帶上的東西,那我們就順手帶上。包括今年,很有可能我們在 API 或 AI 上也有 ARR 收入的機會。
也就是說,如果這個需求能繼續擴大,並且持續擴大,然後 GPU 能買得更多,那麼做到幾億美元的 ARR 是非常有可能的。如果 AI 能做到 10 億美元,那基本上我公司的現金流可能就能打平了,足夠覆蓋我的 R&D 支出和所有開支。
所以這也是可能的,但我們沒有把它當成優先事項。我們會做,但我認為這是一件重要的事。它不是我們的第一優先級,也不是我們今天真正聚焦的事情。更大的機會還是應該在後面;更早的機會,包括去年的消費端和今年的 B端,我覺得我們需要做,而且要做好,但那不是我們的目標。
或者說,我們公司裡大多數人也不認為這是一件很重要的事,也不認為它和 AGI 一樣重要。再多說一點開源,因為前面很多問題都是關於開源的。首先,我認為我們會開源,我們最強的模型也可能會開源。因為我看不出閉源有任何好的理由,我看不出有任何必然的好處。
ByteDance 的模型是閉源的;它有什麼好處?我看不出有什麼好處。即使模型開源了,你把所有東西都告訴大家,門檻還是很高。別人要真正用起來還是很難。對他們來說,用起來很難;其次,在使用的同時還要把成本壓得很低,也非常困難。這不容易。
這不是說我一開源,他們就能以和我一樣的成本很容易地部署。這裡還有很多工作要做。雖然原理大家都懂,但不是每家公司都願意,或者具備意願和能力,去組織人力和資源達到那個目標。我對這種情況也很習慣了。他們可能就是不擅長做這件事,因為阻力太大了。他們很難控制這些成本;還有很多管理和物理上的約束。
這也是創業公司的優勢,因為如果一家創業公司太小,你沒有這個能力去做這件事;如果你是一家大公司,又很難組織起來。
每種情況都有它的困難,所以這屬於我們這個規模公司的甜蜜點。如果我們再大一些,可能也未必沒有別的問題;如果我們再小一些,實力又不夠。所以對於開源,我覺得應該定價。現在,我們應該承認,我們不會去強迫別人。
因為採用定價模式,我也不會收很高的費用;我大概還是會按十個月回本來收。按照十個月回本來算,這已經能讓我們打敗競爭對手了……按十個月回本來收,我們就已經能讓第三方獨立部署不划算。第三方做不了,他們做不到這個成本,肯定做不到。所以開源不會影響我的收入。當然,如果我想賺一百倍的利潤,那開源就會……
主持人
我聽到了你,但好像視頻掉了,老闆。
可能是剛好有電話進來了。
梁文鋒
所以對於開源,我覺得它對我們的商業模式沒有影響。前提是我們只賺六倍利潤;十個月回本對應大概就是六倍利潤。如果我們只賺六倍利潤,那開源不會有太大影響。但如果你要賺一百倍利潤,那開源確實會影響你賺一百倍利潤的能力,因為第三方會去部署,它們可能比你的成本低二十倍,所以比你便宜。
這種模式能不能長期持續?我覺得是可以的。在我們的願景下,我覺得開源是長期可持續的,或者至少這是我們打算做的事情。你可以說,克制一點,也有長期收益。這個策略讓我們在技術前面有更多機會,做成 AGI 的概率也更大。我們更從容。
你想想,我們甚至都不需要加班,因為這事本來就沒有那麼難。可對別的公司來說,可能就很難,因為他們想的事情太多了。其實一點都不難,真的不難。一開始是一個……從外面看,可能覺得我們選了一個很難的模式:我們做研究,我們去追最難的事情,好像是在困難模式。
但實際上,我們在別的地方放棄了很多,這反而讓我們仍然很強,讓我們可以很輕鬆地做事情。所以我對開源的判斷是,它是可持續的。開源和商業付費並不衝突,前提是這種六倍利潤的情況下沒有衝突。六倍利潤聽起來很高,但其實並不高。因為以現在 AI 的效率來看,今天合理的利潤大概就是這樣。
未來也許會降到,比如四倍或者三倍。我覺得那已經……真的不能再低了。但即便如此,還是會有很大的利潤。單看賣 API,我覺得賣 API 沒那麼有吸引力。但在這點上,你可以說沒有衝突;我沒有看到任何衝突。
而且我一點也不擔心別人部署我們的模型來和我們競爭。我們其實希望他們能部署。我們會盡量給開源社區提供幫助,幫大家部署我們的模型。我不擔心他會把我這門生意搶走,因為市場足夠大。我只擔心他部署不出來,某些細節沒做好,結果變差了,或者他的成本相對較高。
對,這裡沒有衝突。然後去年,我在被問到 To B 業務的時候,一個比較常見的問題就是:如果我把我們的 C端開源了,那 C端會不會又跟我的 C端衝突?
因為我沒有流量優勢,或者說,騰訊本身就有很多流量。如果它部署我們的開源模型,它就會拿走所有的 C端用戶,然後把我所有的 C端用戶都搶走。但其實不會;這裡面有很多原因。而且問題是:我們提供的開源模型,跟我們自己部署的模型是一樣的嗎?是的,是一樣的。
我們不會開源一個更差的模型,然後在自己部署的時候用更好的模型。我們不會這麼做;是一樣的。這也說明其實沒有衝突。去年全年,在 C端我基本上就是一直在開源,也沒有看到 C端服務有什麼衝突。我真的沒有看到衝突。所以開源這部分就是這樣。
然後下面,也有公司的長期願景。我覺得我們的目標應該是 AGI。大家對 AI 的定義可能不太一樣,但這並不妨礙我們把 AGI 作為目標。從技術路線上看,到 AGI 的路線圖其實已經比較清晰了。
以現在這一代 AI 技術來說,如果你能把一個問題描述得非常清楚,並且給它完整的上下文和指令,它已經超過人類了。但這裡有一個定義,一個前提:你給它完整的上下文,你給它完整的指令。而這個前提是很難做到的。
比如今天在會議上,我們其實有一個很長、很強的上下文在前面——也許大家都有幾十年的上下文——而這是 AI 沒有的。AI 現在能做到的是,在有限的上下文裡,它可以比人類做得更好。但它還不能取代人類。現在還缺少的是持續學習。因為人類也可以一直學習。
你招一個員工,他可能要花兩個月熟悉公司環境和工作。兩個月之後,他就能上手了。他能做很多事情。他能聽懂你說的話;比如你說,叫小王過來,他知道小王是誰。但對 AI 來說,因為它沒有那個上下文,它沒有經歷那兩個月的學習。
如果你跟它說,把小王帶過來,你就得告訴它小王是誰、他是什麼職位、他在哪裡、怎麼找到他、找到他的時候要注意什麼。你得把所有上下文都給 AI。這樣的話,AI 是可以做到,但你不可能把所有上下文都給它,這也不現實。所以 AI 不能取代你的員工。
但如果 AI 具備持續學習能力,像你的員工一樣,在公司裡學兩個月,那麼
它就可以替代天下所有人,所以我們離下一個階段還差一步:學會學習。我們可以把 AI 的發展理解成一段階梯。去年我們爬上的那段階梯是 CoT,也就是 chain-of-thought。因為我們發現,通過 chain-of-thought,智能可以達到更高的水平。讓它自己思考,我們就能把天花板抬高,讓 AI 做更多事情。所以我們又跨上了一級階梯。
今年的階梯是 Agent,因為我們發現,使用 Agent 之後,哪怕有很多事情要做,它也還是能做;它的能力範圍變大了,智能天花板也更高了。為什麼說它是一級階梯?因為後面的每一步都是建立在前面的基礎上。Agent 需要 CoT,CoT 也需要更早的那級階梯,也就是語言模型,所以前面的每一步都沒有白費。
所以 AI 的發展、智能的方向,是可追溯的。今年這級階梯是 Agent,但 Agent 這級階梯最終也會完成。等它把所有能解的問題都解完之後,它仍然不能取代你的員工,但它已經達到了能力的上限。
這就像 CoT 一樣,當 CoT 達到天花板時,它已經超過了最頂尖的人類,在做奧數題和寫程式方面都已經超過了。但它還是停在那裡,這個技術還沒有到 AGI。所以你看,AI 智能的方向是可追溯的。
那麼在 Agent 之後,我們認為應該要解決的問題就是持續學習——怎麼讓模型能夠持續不斷地學習,而不是一定要一個很強的訓練環境。它應該能像人一樣,進行相對長期的持續學習。這個問題,跟任務完成等等,其實是一回事;它們彼此相關,解的是同一個問題。
現在站在 Agent 這個位置,我們能看到下一個瓶頸:持續學習。下一個要解的問題,就是怎麼實現持續學習。這個是看得到的,而且相對清晰。它是擋在我們面前的一道坎;你必須跨過去,而且一定有辦法跨過去,只是需要時間。等到持續學習之後,我們可能就會到達奇點。
那個奇點就是,當模型已經能夠持續學習之後,它就已經可以做所有人類能做的事了。它可以自己發展自己的版本,自己做研究,然後再發展出下一個版本,去發展更先進的人工智慧模型。於是它就會到達一個奇點,能夠自我迭代。但這個奇點其實也不算真正的奇點;它同樣是一個漸進的過程。
這個過程可能也是一個相對長期的漸變,它不是突然發生的。但我們習慣上都會覺得它可能是一個奇點。因為很早以前,先知們就認為這裡會有一個奇點,但實際上它不是奇點,它是一個持續的過程。而且當這一步完成之後,我認為那就是具身智能出現的時候。
這是我們的推測。我們認為時間表應該是:先解決學會學習,再到那個自我迭代的奇點,然後才是具身智能。具身智能之後,它進入真實世界,就可以幫你做家務、照顧你的晚年。我們認為這是一條相對理想的路線圖,但每個人的看法不一樣,也沒有對錯。我們只是覺得,這條路線圖是最容易的。
這條路線圖之所以容易,是因為每一步你都只需要做很少的新工作。沿著這條路線圖,我們不需要加班。但如果路線圖是反過來的,比如必須先把具身智能做出來,那麼自己去做就會很累;那會是很辛苦的工作。我們不想要這樣的路線圖,我們想做得輕鬆一點。
如果我們先解決持續學習,再解決那個自我迭代的奇點,然後再解決具身智能,這條路就會很容易。因為到了後面,你可以用前面的技術去幫助發展後面的技術。到了奇點之後,具身智能就不用人去做了,不需要我們去做,它自己就能長出來。所以這就是我們長期目標的答案。
我跟他說,這就是我們的長期目標,也就是我們所說的 AGI。我們回到現實。去年的現實重點是,大家都要做 Chatbot,去爭 C端流量。今年的現實是,大家都要去爭 To B 收入,都要進來參與,因為你不進來,連牌桌都上不了,對吧?
但我們不認為這是一件重要的事情,或者說,在我們公司內部,我們真正關心的是我剛才提到的 AGI 路線圖,以及下一個技術突破會怎麼發生。但有一件很奇怪的事情是,你最想得到的那個東西,往往恰恰是你得不到的。反而是那些你沒那麼在意的東西,通常都比較容易拿到。
這裡面有一個戰略優勢:在我們腦子裡想的是 AGI,我們做的也是 AGI。那麼當我們去做應用,或者做 C端和 B端的時候,其實根本不需要在那些事情上花太多心思;實際上,花很少的力氣就夠了。我認為,當你站在更高的技術層面上,去做相對更底層的技術時,會有一種降維打擊。
至少在去年的 C端,我們看到確實是這樣。我們在 C端 上並沒有花太多精力,甚至一度都不太想維護那些用戶,但用戶就是趕不走。因為真的趕不走,最後都留下來了。不過有一點 ... 而且今年,B端 的收入從增長上看現在也相對樂觀。我覺得這個數字相比同行可能也還不錯吧,我猜。
但我們在這上面其實並沒有花很多力氣。我們基本上沒有
刻意去把這件事做出來;這是我們順手做的事情。在做互聯網智能、走向線上發布的過程中,AGI 這一步是我必須要走的一步。在通往 AGI 的路上,我必須經過這一步。然後我把這些技術都通過 API 提供給大家;我並沒有額外去做什麼。我們還是在做 AI;這只是一個副產品。
我只需要幾個人維護這個 API,甚至連客服都沒有;不需要銷售;什麼都不需要,用户自己就會來。
或者說,被認為是 C端 的用戶——無論 C端 還是 B端——都是我們走向 AGI 過程中的副產品。它們都是中間產物,並不和我的 AGI 工作衝突。我不是為了做 C端,或者做 B端,而去做它們。相反,我們是為了 AGI 而做 AGI,而它恰好產出了這些東西,所以我把它們拿來做商業化。這和其他公司不一樣。
其他公司是為了這個而做這個——要麼是為了服務 C端 用戶,要麼是為了服務 B端 用戶,然後再去做模型。但對我們來說,這並不是最初的意圖;我們最初的意圖還是追求 AGI。我認為,在某種程度上,這是一種降維打擊。AGI 是更大的願景,而這個願景能夠把更多優秀的人聚到一起;它的凝聚力更強。
所以我在組織上是有優勢的,而我用這個優勢去 ... 這是一種降維打擊。但是如果你是一家商業公司,你的願景是把 C端 用戶服務好,那又是另一回事了。它有別的優勢,在產品、用戶服務、流量方面更強,但在技術上不佔優。現在有利的情況是,模型技術是最重要的。
你要把模型做好,剩下的 ... 對,這也能解釋我們之前的發展路徑。我們確實選擇了 AGI,而且我從來沒想過要做很多用戶。去年春節我們突然火了,那完全不在我們的劇本裡;我們從來沒想過會這樣。我們只是想把技術做好。
但當時我發現,相比那些完全商業化、產品導向的組織,我們的組織在人才和組織上其實有額外的優勢。這也非常神奇。那時候,C端 的所有人都在拼命廝殺,最後卻被一個沒去競爭的人給帶走了。那也確實說明我前面說的是有道理的:人才和組織上確實有優勢。
這種人才優勢並不是說我的人比他們更聰明,而是我如何組織這些人才,如何激勵他們,以及他們如何協作。這才是優勢。因為把聰明人聚在一起,並不意味著他們就會自然而然地協作,或者自然而然地擁有朝著某個目標奔跑並完成它的熱情,所以你需要願景。我們之前的經歷教會我的是,AGI 的願景非常強大。
好,這就是這個問題。那下一個問題是,核心利益的重要性是什麼?
我剛才說了,在很多方面我們需要非常克制。但我們的核心利益是什麼?其實我們只有一個核心利益:我們最大的核心利益就是保持團隊穩定。這是我們最大的核心利益,甚至你可以說這是唯一的核心利益。
只要我能保持團隊穩定,我就一定會成功,一定會實現 AGI,就這麼簡單。只要大家都還在,而且我們還能繼續做下去,那我就一定 ... 基本上沒有太大風險。只是遲早會有挫折;如果有挫折,大家還能繼續留下來,那我就還能繼續做。錢肯定不是問題,資源也不是問題,其他所有因素都很容易獲得。
對我們來說,只有一個核心利益,只有一件事不能妥協:我們必須保持團隊穩定。這也是我們面臨的非常大的挑戰,或者說,我認為這是最大的風險。當然,隨著我們最近的融資,這個風險已經大大降低了。因為大家的選擇還是相對較多的,而且金額也還是相當可觀的。
從團隊穩定的角度來看,只要一些最重要的員工和最早的一批員工保持穩定,其他人就不太可能離開。即使其他人的股權少一點,或者收入少一點,他們也不會走。因為他們來不只是為了錢;大家都希望在一個真正能夠實現 AGI 的環境裡工作。所以對人才來說,還是有吸引力的。
從歷史上看,我們的人才流失率一直比較低。和同行相比,我們的人才流失率一直相對較低。但這仍然是我們最大的挑戰,可以說也是唯一的挑戰。其他的都是時間問題;最多只是讓我們晚六個月或者一年,但不會讓我們做不成。我們肯定不缺錢,也肯定不缺資源。實際上,這些都不缺。
So many of the things we are doing now are to maintain team stability. Aside from that, I think we can give up on everything else; we can restrain ourselves on everything else. We have always been very restrained, and we do not want to become adversaries of any large or small internet company. I hope I can empower them, or help everyone do this, and hope I can assist everyone in doing this.
This is also part of the commercial meaning we talked about earlier. The premise is that everyone should not ...
Under that premise, we are very willing to assist and help anyone, even our competitors, including Alibaba, Zhipu, and Moonshot AI, to do better. Because we do not lose anything; we were open source to begin with, and open source also does not try to make the boundaries as clear as possible. As for how to do it, we hope you can reproduce it; if you can’t reproduce it, say so, and I’ll tell you how to reproduce it. That is originally part of open source, and it won’t change just because you are a competitor.
Of course, if it is a partner, I will do more. But at the level of big interests, there is no conflict.
When dealing with the outside world, our attitude is: we only do the main line of AGI. That is what I just said about GPT, CoT, Agent, and so on—we only do the main line. The AI field is very broad, and there are many things that we think are not on this main line, such as 3D and video generation. I think they may not have much to do with the main line of intelligence, so we will not do them.
There are also some things, such as world models. I think they currently do not have much to do with the upper limit of intelligence either, so we will not do them. But if others do them, we are also very happy to help. Whether we have time is one thing, but there is no conflict in interests. We also hope that these AI technologies can be used in all kinds of production environments, can improve social productivity, and can help all industries improve productivity.
We are very motivated to do this. Whether I have time, whether I have enough people, or whether our teammates themselves are interested—that is another matter. But there is no conflict in interests; we hope to achieve this goal. And we believe this has no conflict with business at all; the benefits I should get have not been reduced at all.
I think that, in the attitude we held before, we actually didn’t lose anything because of it. We didn’t lose anything because I open-sourced things, or because of our goodwill, or because I helped other people. For example, last year’s C-end users — we still have quite a lot of C-end users now, and they are still relatively stable. This year’s B-end, I think, is also relatively optimistic.
I have not hurt my business interests because of our goodwill at all; absolutely not. On the contrary, it may even have added points. This seems counterintuitive, but it really is like that. Or, if we think about it the other way around, even if we violated it, would it necessarily mean I could get more? No. There is a question: how should we understand that world models have nothing to do with raising the upper limit of AI intelligence? What I’m saying is that, at this stage, this is our judgment.
We do have a roadmap; this is our own AI roadmap, not the only roadmap. From our understanding and judgment, what matters most right now is doing AI training well. Doing AI training well does not require world models, and does not even require multimodality. Because if you narrow the scope of AI training a bit, without multimodality, you just won’t be able to do some tasks, but that does not affect the validity of the algorithm. Multimodality will ultimately still need to be done.
What matters now is training, and then the next step is to solve the problem of continual learning, and then the next step after that is getting it to ask questions on its own. But this roadmap does not include world models, and does not include video generation. When video generation first came out, it was very hot, as if it was something that had to be done; if you didn’t do it, you didn’t seem like an AI company.
So I found that very strange. Actually, if you think about it carefully, it has nothing to do with the roadmap to intelligence.
In fact, you also saw that after Sora came out at the beginning of video generation, everyone did it — big companies and small companies all did it. But small companies later cut it all off. It has nothing to do with the upper limit of intelligence. But commercially, it is a good business; commercially, it is a good business. But it has nothing to do with intelligence. We would not do it just because it is a good business; we would only do it if it is on the intelligence roadmap.
Video generation is relatively clear, so I’ll use it as an example. As for world models, the meaning of world models is not so clear, because many things can be said to be world models. From our judgment, world models and intelligence are not yet the most important things at this stage. What is most important is AI training, and then after AI training, how to solve continual learning.
This is our company’s judgment. Of course, every company has a different judgment. What I just said is that, for our company, the most important issue is staff stability. From another dimension, what do we lack? What is the gap between us and the United States? Actually, there is only one gap: resources. We don’t have that many cards; the number of our cards is still relatively small.
We currently have about 20,000 H-equivalent compute. Most of this only just arrived; it came in the last one or two months, and there may still be many machines that have not arrived yet. Our total compute was relatively small last year; this year we are expanding compute very aggressively. We are now at about 20,000 H-equivalent, and over the next few months we will also buy a large batch of machines, basically all NVIDIA.
我們需要多少卡?現在當然是越多越好。在我們負擔得起的範圍內,卡越多越好,這毫無疑問。所以我們現在的策略就是:在合理的價格內,能買多少卡就買多少卡。這次融資如果花完了,我能買多少卡,就買多少卡。花錢的速度不在計畫之內;反而是只要價格合理,有多少我就買多少。
所以如果我能在半年內把錢花完,我覺得那是一件好事。如果我能在半年內把錢花完,那就非常開心,非常理想。其實要花掉這麼多錢是很難的;你買不到那麼多卡。很難買,而且價格也高。你也不能用極高的價格去買;還是得確保價格合理。
如果我能在半年內把錢花完,那也許就是最理想的結果。因為把我的錢換成 NVIDIA 卡,肯定比放在銀行裡好。如果放在銀行裡,我大概只能拿到 2%,好像是這樣。但買 NVIDIA 卡是一個十個月的社會成本。
所以當然是能買多少就買多少。先把卡買下來之後,後面我就有很多空間了。通過提供服務或者做其他任何事情,我總能有現金流。只要有現金流,我就能活下來;我不需要在資產負債表上保留很多現金。所以我們擔心的,只是買不到那麼多卡。
如果我們能把所有的錢都換成卡,那我們就會毫不猶豫地把所有的錢都換成卡,而且我們也願意為此支付一定的溢價。我們願意支付一定的溢價把它換成卡,因為這太划算了。即使支付了溢價,其實我們還是很難達到這個目標。所以客觀來說,如果我今年能花掉 20 billion,那我們的採購團隊就已經表現得超級好了。
我們和美國的差距主要在資源上,而在人員上的差距並不大。人員上幾乎沒有差距,因為基本上就是同一批人,大概就是中國人。中國人出國之後,有些留在中國,有些留在海外,也有些出去;不是說聰明人都出國了。不是。其實是很隨機的。
最聰明的人也未必有超過一半出國,稍微不到一半的人留在中國。中國不缺人才,而且我們基數大;每年都有那麼多新的人進來。人才不是瓶頸,資源才是最大的瓶頸。資源首先影響人才培養,因為算力少了,我們做實驗的機會就少,所以我們整體的人才是落後於美國的。人才差距,從本質上講,也是由算力差距造成的。
就目前最大的模型來說,我們其實根本訓練不起。即使我們把 50 billion 全部花掉,還是訓練不起。即使能把它湊起來,我們也用不起。今天最大的模型大概有 800B activations;國內我們還停留在 tens of B 的規模。國內最大的模型,activations 可能也只有 tens of B,所以我們差了一個數量級。
如果我要訓練一個和 AI 一樣大的模型,我大概需要 50,000 GB300s,或者 Huawei 950,200,000 張卡。而這還只是訓練,還不包括研究。所以我們和美國最大的差距就在資源上。我們現在的資源,加上今年之內以及未來幾個月內會有的資源,包括很快就會到來的那些大資源,也只夠我們在 10B activation scale 上做更多實驗。
因為在 tens of B activation scale 這個層面上,還有很多實驗是我需要去做的,還有很多事情需要去弄清楚。我們距離訓練一個 800B 模型還是很遠;還有很多時間,而且卡也沒那麼多。所以我認為,我們和美國的差距,就是資源的差距。
我們可能會覺得,我們看到的所有差異——包括人才差異、模型能力差異,以及應用差異——都可以歸結為算力資源的差異。
算力資源,一方面是因為卡在國內就是買不到;另一方面是因為我們的資本投入比美國少得多。我們投入的資本少很多。在資本上,這裡人才薪資所佔的比例很低。你看到像 100 million US dollars 這樣的薪資,但算下來,人才薪資仍然只佔很小的一部分;主要成本是算力。這個問題目前基本上無解,因為華為的產出也是有限的。
因為如果我要訓練 800B,我需要 200,000 張華為最新的卡;而這還只是訓練,還沒把研究算進去。所以現在我們根本不會考慮和美國在這麼大的規模上競爭。我們現在失去的,還是在我們能夠負擔得起訓練和使用的規模上——也就是 tens of B activation scale——我們需要先把這個做好。
等到下一步我們有更多資源之後,就可以把它提升到 150B、156B 或者 250B activation scale。所以我們和美國之間目前存在一個看起來非常難以彌補的差距。如果你堅持要訓練這麼大的模型,它是可以訓練的,但你做不到充分研究。也就是說,在訓練之前,你做不到充分研究。
另一個大家都很關心的問題是:在所有大模型競爭中,最後的差距會體現在什麼地方?也就是說,當大模型最終分化的時候,差異會在哪裡顯現出來?我認為這個差異最終可能不會很大。最後的差異應該會體現在三個方面:成本、時間,以及用戶體驗。除此之外,可能就沒什麼太大的差別了。
成本很容易理解:如果你提供同樣的服務,在同樣的質量下,你能以什麼成本提供?同樣的服務,以比亞迪電池為例,在同樣的技術水平下,其他公司能不能以那個價格提供?我認為那是一件很難的事情。這不是那麼容易做到的;它肯定是一道護城河。所以成本絕對是一個差異,我認為成本可能是第一大差異。
然後第二是時間:你什麼時候能做到?如果你早幾個月或者晚幾個月,情況就不同。第三可能是經驗;用戶體驗還是有些不同的。中間還是會有一些用戶黏性和用戶壁壘,但那可能不是根本。從根本上說,還是第一:成本。第二是時間。你是最先做出來,還是後來做出來;你是更快還是更慢地實現同樣的事情。
在長期商業化路徑和產品線進一步豐富之後,定價應該怎麼設置?我們認為,現在最值得做、最值得花精力的,還是去做 AGI。現在需要做的是把 AGI 再往前推,也就是把智能的下限往上抬、往前走。在這個階段,這應該比去做更多產品線、考慮更多商業化路徑更具成本效益,或者說,這是一條回報更大的路徑。
我認為未來一段時間,以及過去任何一段時間,可能都是一樣的。也就是說,如果我們
我們花了很多時間思考,豐富產品到底有沒有用?半年前的商業化討論是什麼?一定是:我要做廣告,然後做電商,然後把電商嵌到產品裡,再和本地生活服務之類的深度結合。這肯定不行,因為變化太快了。
如果你的產品已經領先,或者在我們現在這個階段,你花很多時間去思考商業化、商業化路徑,產品線的生命周期非常短。我不認為我們已經到了那一步。這是我們的判斷,或者至少以往所有經驗都支持這個判斷。過去三年任何時候,如果你來跟我談商業化路徑和產品線,那都是在浪費時間,因為你無法預測,也無法預見未來。
你能預見的東西非常少。因為我在看時間,不知道大家想不想中場休息一下?要不要先吃點東西?還是我們繼續聊?如果大家沒有意見,我就繼續。要在全球領先的 AGI 研究以及合適的時機進行商業化。我認為我們一直都在做商業化;我覺得我們一直都在做商業化,只是不是以商業化為目標。
我們以 AGI 為目標,但我們一直都在做商業化,所以才會有 C 端用戶和 B 端收入。從歷史經驗看,這個策略是成功的。而且我認為,完全轉向商業化的那個節點還非常遠。所以最大的事情還是技術延伸,然後去做下一代技術,去解決更多當前問題。
這些回報比,在我個人看來,在可預見的未來裡,現在任何時候把重心放在產品上都還太早。所以這也是我們克制的一部分。我希望這些商業機會可以由別人來做。我們希望這些商業機會,以及如何使用這個 AI,可以由整個社會和我們所有的合作夥伴一起來做,共同分享回報,而不是我想去壟斷它們;那是不可能的。
而且我們也沒有那麼多精力,我們的組織也沒有那麼多人去做這些。對於合作夥伴,其實我們的融資是經過仔細挑選的。具體怎麼合作是一回事,但首先我認為,利益還是比較一致的——和我們利益最一致、對我們最不敵視,或者最希望我們成功的那些。
不是每個人都希望我們成功,因為我們確實會損害很多其他人的利益。公司重大戰略、技術和業務決策的過程與決策機制。我們公司整體是建立在共識上的。我不是說我自己一個人決定一切;相反,我需要去尋求共識。我在公司內的權威和影響力,是建立在共識之上的。
比如說,如果我想做一件事,我一定會先看我們的共識是什麼,看看大家想不想做。然後我可能會有一些引導或者傾向,但這種引導的作用是有限的;非常有限。它還是必須建立在共識之上。
這個決策機制,其實就是一個尋求共識的機制。不是我可以推動任何一件事,必須先有共識,然後我才能推動,接著我會把它推動下去。目前主要精力基本上都在 DeepSeek 上。
Moderator
我們休息五分鐘。
Investor
我快點記一下。
Moderator
大家可以打開麥克風,給我一些回饋。
Investor
我們這邊沒問題,也許先休息五分鐘。
Liang Wenfeng
不同公司之間模型性能的最終差距,應該是一個綜合性的差距。比較模型性能時,肯定要在同樣成本下比較;那才有意義。因為你比較兩輛車的時候,都是在同一價格區間內比較。做得好和做得不好的模型之間的差異,不應該體現在某一個具體環節上;應該是整體上的。現在領先 OpenAI 的 Anthropic,這是一種長期性的東西嗎?
我不這麼認為。這肯定是週期性的。OpenAI 和 Google 以後很可能還會繼續交替;它們應該在上升中交替。事實上,Anthropic 在 Code Agent 上的優勢現在並沒有那麼大;也不是說它正在碾壓 OpenAI。
也許我們公司有一半的人,在日常工作中都覺得 OpenAI 更好。事實上,Anthropic 其實有先發優勢,但這種先發優勢應該很快就會消失;它不是一個能長期保持的優勢。這三家都很強;在這三家裡,效率最高的那家,就是花費成本最低、花錢最少、燒錢最少的那家。
當全球 AI 處於一種主導分工時,中國公司很可能扮演的角色仍然是最大的生產者。按常識來說,我們的產能是最大的,包括晶片——我們的晶片產能可能是最大的——而且我們的電力最多,所以我們的 AI 很可能會是三體之一。
中國人會把這個產品做得最便宜,然後在性能上,畢竟對外國貨來說,現在很多中製和美製產品差距也沒有那麼大。未來 AI 也可能是這樣,但中國做的 AI 可能會更便宜。這種便宜可能是系統性的更低,就像中國在其他產業提供的服務可能更便宜一樣。
我做事情的時候,通常的思路是:我現在應該做什麼,才能拿到最高回報?如果我認為現在做產品的回報最高,那我就做產品;如果我認為現在先把 AGI 做出來的回報最高,那我就先去做 AGI。顯然,我認為現在做產品不是回報最高的事情。如果是做國產晶片適配,那國產晶片其實現在有一個歷史性的機會。
因為以前國產晶片適配有一個很難的問題,叫做生態不好。你把卡買來了,但你用不了;它沒有 NVIDIA 生態。所以 NVIDIA 的護城河很強。
但這件事正在改變。NVIDIA CUDA 的護城河正在被快速瓦解,而且它之所以被瓦解得這麼快,可能有三個原因。一個是現在有了 AI,有了 AI 之後,搭建這種生態比以前容易太多了,因為 AI 會寫程式。
我可以用 AI 來搭建這個生態,然後搭出一個和 NVIDIA 完全一樣的生態。第一,是因為 AI;第二,是因為一些新技術。比如我們公司出了一個叫 TileLang 的技術,它是一種高級語言。
用這種高級語言去寫 CUDA operator,可以很快把 NVIDIA 的整個生態寫出來,然後再結合 AI,似乎就沒有障礙了。但它還沒有完全完成;還沒做完。不過,這條技術路線看起來沒有障礙。
另外,因為 CUDA、NVIDIA 是從遊戲卡演化出來的,所以很多地方,遊戲卡的設計和遊戲卡的架構是一脈相承的。CUDA 和遊戲卡是相容的。過去因為 AI 計算是一個很小的領域,比遊戲卡市場還小,所以這是合理的。但現在算力卡市場已經比遊戲卡更大了,所以兩者沒有理由還要綁在一起。
現在的趨勢是,未來它們不會再綁在一起了。那麼專用晶片——不管是華為的,還是 NVIDIA 自己的——未來都會是專用晶片,沒有一個還是老東西了。在這種情況下,原來 NVIDIA 生態的作用就大大降低了。因為對於專用晶片來說,它和 CUDA 沒有關係;它不是綁定在 CUDA 上的。
或者說,當這個晶片在設計的時候,其實已經把怎麼搭建這個生態考慮進去了。這個有點複雜,但總之,國產 AI 晶片替代現在是有歷史機遇的。我們認為,在未來一年內,我們會看到一件事被驗證:國產晶片的生態是完全沒問題的。
以前大家認為有問題的地方——就是不能用、不好用——我覺得一年之內我們會把這個認知扭轉過來,或者用事實把它扭轉過來。國產 AI 晶片的硬體和生態都沒有問題;唯一的問題是產能不足。國產卡適配沒有門檻,Nvidia 也攔不住。
如果我們處在一個正常的商業環境裡,我能買到 Nvidia 卡,那國產替代會相對比較難;但當 Nvidia 卡買不到的時候,大家就被迫去做國產晶片了。在這種情況下,適配國產卡完全沒有門檻。把國產卡的生態做成和 Nvidia 一樣,甚至比 Nvidia 更好,我覺得沒有門檻,但它還是需要時間。
我們現在主要和華為合作。華為那邊是自己做適配,但我們會親自參與這個生態,也會深度介入華為那邊。華為的問題還是產能不夠。華為給我們大概 16,000 張卡的產能,而網際網路大廠可能有 100,000-plus;我們是 10,000 多張,我覺得這個比例也還比較……但這可能已經是華為全部的產能了。
所以我們也不能真正指望華為去訓練更大的模型,或者去訓練一個有幾百 B activated parameters 的模型;有時候有人說那今年就可能發生。但明年或者後年,也許會有機會。對華為卡適配來說,我們的主要工作是把它的高級語言編譯器做好,把 TileLang 做好。等 TileLang 做好了,這個問題可能就自然解決了。
這個解釋起來有點複雜,但我們正在做,等做完之後,解釋起來就會清楚很多。你可以這樣理解:V3 訓練的時候,還是用了 Nvidia 卡,但它已經不再用 Nvidia 的生態了。
V3 uses Nvidia cards, but not Nvidia’s ecosystem. Instead, we first write a high-level compiler called TileLang, and then based on the TileLang ecosystem we complete everything else, and we are almost no longer dependent on Nvidia’s ecosystem. As long as I take this whole set of things and redo the process on Huawei cards, then it’s done.
I think this may be a historic mission, namely, it can completely reverse everyone’s previous perception that domestic card ecosystems are bad. Right now, the timing, the geography, and the people are basically all in place; the only thing missing is time. I think within a year, many people should have it, or rather, everyone will understand that this problem has been solved, and what remains is a capacity issue. I am relatively optimistic about domestic computing power. On this point, I think Nvidia is digging its own grave.
Huawei’s supernode, Huawei’s 950 supernode, can completely replace Nvidia’s GB200 and GB300 in performance and price. The price will definitely be higher, but the increase is limited. If the price is 50% higher, 100% higher, 100% higher is fine, 200% higher is fine. For example, if it’s 100% higher, I think it can already be considered a replacement in terms of price.
It can also be a replacement in tasks; all the tasks GB300 can do, Huawei supernodes can do too, latency and everything are the same. The only cost is that four Huawei cards equal one Nvidia card, while being two years behind. Four to one is understandable. Being two years behind means that four Huawei 950s can match one GB300. Two years behind means two years behind in time.
Huawei 950 supernodes are shipping in this year’s Q3 or Q4, while Nvidia GB200 is from Q3 two years ago — a two-year gap. Nvidia may already have a new generation this Q3. So the gap between us and the U.S. in chips, I think, will no longer be an ecosystem gap in the future, but on chips it is four times plus two years.
There is also the question of whether we will vertically integrate upstream. I hope not. So there is a question: will we vertically integrate upstream into applications? We hope not to do that; I hope other people will do that. I don’t want to eat everything. I just want to eat one piece, the piece I’m best at, or the piece we ourselves think is the most core, and then the piece that is directly related to users
I think for many of our industry partners, they care about similar things more than we do…… this should be someone else’s significance; it shouldn’t be interpreted by me. Will we build large-scale clusters ourselves in the future? I think building large-scale clusters ourselves is definitely necessary; we have always been doing this ourselves, and all our clusters are self-built.
But whether we will develop chips ourselves in the future, I think, depends on how large the returns are, depends on how large the returns are. Tesla, nurturing…… this sort of thing. If you are operating a power plant, you don’t necessarily need to make the generators yourself, right? The power-generation equipment can be made by others; as long as the price is reasonable, why make it yourself? So I hope we won’t need to do chips. I hope we can buy chips at a reasonable price, so I don’t have to make chips myself.
I think this is very likely how it will be: recently, the profits on Nvidia chips may not necessarily be…… although…… we hope to only do one piece. I think AI is a very big thing; it doesn’t require me to…… I just do one piece. If we focus, and I believe the business value here is already big enough — if the AI era will produce many trillion-dollar companies, I think we are one of them.
We’ve always done one small piece of it, and being one of them is already no problem; there’s no need for me to…… I’m not brimming with confidence that I have to do everything. I think the most likely thing is that it’s hard to change, and if you really want to do something else, you’ll still have more ideas, and that will…… that’s our attitude.
For example, at least in the To B and To C businesses, from what we can see now, the ones that really want to build a closed loop in To C, in fact we do better; the ones that really want to build a closed loop in To B, maybe they still don’t do better than us. The more you want…… and also, To B, I can say a bit more.
The upper limit of the To B business should still be demand. Under the current generation of AGI and AI technology, To B demand should be limited. It will grow rapidly, but it is not something infinite; in the end, it is still constrained by demand, not computing power. Under the current technical conditions, how much revenue there can be ultimately still depends on demand.
Because think about it this way: if I can recoup my costs in ten months, then if I have a destination, I would definitely buy one piece; it’s because there aren’t that many things. Right, demand should keep getting bigger, and if technology keeps making breakthroughs, that demand will keep getting bigger.
Multimodal layout, we’ve always been doing it. For products, it is very important; for C-end user products, it is very important. But for the upper limit of intelligence, it is a component; it is not the main line itself. But as a component, we will definitely do multimodal, and we are doing it. We should release related models — our V4 and the subsequent versions of V4 will support native multimodality.
But for multimodality, for intelligence, it is a component; we don’t treat it as intelligence itself, and its main line is like search. Search is also a component; multimodality can be too, and in our understanding it is also a component. Scaling — we believe in Scaling. The larger the scale, the better the effect, and the more capabilities it can unlock.
真正阻止我們擴展的是算力;不是我們不想擴展,而是我們沒有那麼多算力去做這件擴展。我們還沒有觸及上限。我們訓練這麼大的模型,不是因為我認為這個尺寸已經足夠,而是因為我剛好只有這麼多資源。
我根據我的資源來算:我能負擔並訓練多大的模型?就是這樣算出來的,不是因為這個模型已經足夠。到目前為止,模型回報還是非常明顯的。我們仍然沒有機會碰到 Scaling 的牆;我們離那個還很遠。
當 Silicon Valley 說 Scaling 已經到極限時,那是對 Silicon Valley 而言;對中國人來說,我們還離那個非常遠。我們根本還沒有把 Scaling 擴展到那個程度。這個 Scaling 包括資料 Scaling、模型規模 Scaling,然後還有訓練成本。我們離探索這個 Scaling 的上限還相對很遠;只是沒有那麼多算力。
但我們也會不遺餘力地去推動這個 Scaling 的上限。
包括在我們融資完成之後,我們會有更多算力,也可能訓練更大的模型。能在那段時間內越早買到就越早買;如果能在半年內把所有精力都投入進去,那當然最好,但現實中做不到。我認為下一代模型的核心能力,一定要有持續學習能力;只有這樣才能稱得上是下一代模型。在那之前,我們能做的是降低成本,讓效果更好,並且讓速度更快。
但要有大的突破,應該要有持續學習。還有一個問題:為什麼我們似乎這麼在意模型的計算效率?因為我真的發現,不是每個人都那麼在意這個模型的效率。對一家商業公司來說,沒有動機去追求模型效率,因為如果模型效率高一點……所以並沒有那麼大的動機去把模型效率追求得很高。
所以幾家創業公司不會自己……你也沒聽他們說低成本是他們追求的目標,因為那不符合他們的利益。成本低了,你還怎麼賺錢?成本低了,你就收不到太多錢。
相反,這是我們願景的一部分。或者說,我們的願景——我們的夥伴很在意這個成本,因為我們的夥伴都是普通人,他們知道使用這個都要花錢。他們能感同身受:別人使用也要付費,所以如果便宜一點,別人就會更願意接受。因此我們很多夥伴還是希望我們能把成本再往下壓。
但如果你從商業角度來看,你就不會這樣想。從商業角度來看……
從成本或商業角度來看,這不是優先級最高的事情。不論是對創業公司還是對大公司而言,服務成本其實都不高。但我認為我們希望它相對輕量;我希望它價格親民,特別是在中國算力短缺的背景下,能夠在國產卡上價格親民且可用。我認為低成本首先是一個結果。
我們的模型在架構上確實一直都在朝更低成本的方向發展,這也和我們的願景有關。我們也有很多算法方法,成本還可以繼續下降。成本下降的另一個原因是,成本越低,我能訓練的模型就越大,我能負擔的模型也越大。在相同的算力下,如果我的計算效率更高,我就能負擔更大的模型。
對大公司來說,他們可能不會這樣想。對大公司來說,可以往裡加資源;問題可以靠增加資源來解決。但我們會優先考慮成本效率。數據類模型的價值——數據是一個相對寬泛的類別;數據幾乎應該等於模型的一半。
為什麼我認為如果我想要那樣,或者說如果 AI 能佔到 GDP 的 20%,然後如果我想從中拿走 5%,那絕對不行?因為我一定會被別人打敗;如果別人說他們只想要 1%,那他們一定會打敗我。如果我的目標是拿走 5% 的 AI、拿走全人類 GDP 的 5%,那從理論上講數學還是算得通的。
你可以看 OpenAI,似乎它的數學是算得通的;理論上沒有問題。但它有一個問題:它會被另一個只願意拿 1% 的人打敗。因為那個人會說,我把這件事做得這麼好,但我只需要全球 GDP 的 1%,那他就會打敗他。
到了那個時候,如果又有另一個人出來說,我只需要 0.1%,那這個人又會把前面的人打敗。
從宏觀角度來看,無論那個百分比來自哪裡,彼此之間都沒有差別;反正都一樣。拿得更多的人會被拿得更少的人打敗。你甚至不需要真的拿得更多——只要你的願景是拿得更多,你就會被那個願景是拿得更少的人打敗。事實上,沒有人真的拿到任何錢;這只是一個願景。如果你的願景是拿得更多,那你就先輸了,你會面臨更大的困難。世界就是這樣。
OpenAI 原本以為自己真的可以壟斷世界,但實際上它會遇到很多很多挑戰者。它會面臨挑戰,事情不會那麼容易。美國會面臨挑戰,未來也可能會面臨來自中國的挑戰,因為中國人願意拿得更少,但仍然可以為你提供這個服務。在中國,也會有人願意少拿一點。
But in the end, there will be a balance here, because if you take too little, the company’s business logic won’t hold, and it won’t survive. So if you take too little, you can’t survive; if you take too much, you will be beaten by those who take less. So for us, it’s not about taking the most profit, or maximizing return in pricing, but only earning a reasonable return. That’s an explanation.
I believe in this matter. I’m not finding reasons for it, because there’s no need to find reasons. This is just how I do it; if I do it this way, there must be reasons. Those reasons may not be very conventional, but I think the company itself is not conventional. Our company’s management actually has two lines: one from top to bottom, and one from bottom to top. Bottom-up means each person decides what they want to do and does it themselves; nobody manages them, and there is no KPI.
Top-down means doing the important work; if we need to collectively do something, everyone in the company needs to coordinate. For example, if we want to release V4, then we need division of labor, and everyone needs to do a part. That is top-down, and what we call top-down here is “doing the important work.” Generally, we hope that the “important work” does not take up half of an employee’s time — it should not exceed half.
They still have half of their time unassigned; they can do whatever they want. This is a research scope, allowing them to explore on their own, according to what they think is important, with no prior requirements. As long as the company can support it, and the company’s computing power can support them doing it, or if they don’t need computing power and only need to do very little, then they don’t even need to coordinate. So this is the organizational approach we have now.
Some people think we are top-down, some people think we are bottom-up; I think both are true. My criterion is that “important work” should preferably not exceed half.
We usually don’t work overtime much either. There are two reasons for overtime. First, research needs a relatively relaxed environment. If you pressure people too tightly, they can’t do research. Since you need your own interest and you need to think about these problems normally, it has to be in a relatively relaxed environment for exploration to be possible. That’s a need from research culture. Second, we are very focused.
Being very focused means that the things we need to do are very few. Then I don’t have that many things to do, so I don’t need to work overtime. This is consistent with the restraint I mentioned earlier. Because I restrain myself, many times I just don’t do things. Then with fewer things to do, the work assigned to each person is less. You can see that many of our products are not very complete, and we haven’t gone to fix them. That is also part of our culture.
OK, because there are many questions, I’ve mostly skimmed through them. If everyone still has questions, please ask.
Host
Please feel free to speak up and exchange ideas, investors. Let me remind everyone that Liang Wenfeng mentioned quite a few sensitive pieces of information just now, so please do not share any numbers or situations externally, including card counts and such. Also please do not screen-record or share externally. Thank you very much, and if you have questions, please feel free to speak up.
Investor
Yang-ge, could you share more about the timeline for when continuous learning might bring breakthroughs? Also, if continuous learning is achieved, what other architectural or algorithmic innovations, and other key factors, are still needed?
梁文鋒
Fewer people; research is needed. Right now the whole world is studying this problem. Or rather, for investors, what investors see most now is AGENT; but for us researchers, what we see more now is learning, and how to solve the problem of learning. In fact, learning may not be a technology; it is a problem. How to solve this problem may involve many kinds of technologies. It is not a single technology; it is not one thing, it will be many things.
Or rather, AGI is composed of many things. It needs the model, and it also needs many other things. In fact, it is also an engineering and algorithm problem. This problem is quite specialized, but there are many methods and also a lot of research.
Investor
Mr. Liang, thank you. Thank you very much for today’s opportunity. First, I really want to respond with gratitude. What you said at the beginning moved me deeply and also gave us a lot of inspiration. You mentioned that this team is carrying the greatest goodwill, hoping to make some contribution to the development of this industry and of human intelligence in this society.
And within that, it carries this sense of mission and vision. I think that is very similar to the corporate culture of the company we serve, which is “to cultivate oneself and help others.” I deeply understand why you lead the team to do open source. I can vividly imagine it like creating a little bird paradise ecosystem with a banyan tree — benefiting all things without contention, and in that way it will be accepted by everyone, all things living together in symbiosis and coexistence, and ultimately it will be everywhere.
So, through this investment, we also want to express our recognition of, support for, and respect toward this mission and vision. At the same time, we also hope that in the future of this industry, in some of the areas we are good at, we can contribute some strength. On this point, I also want to continue asking for your guidance and discussion. For example, in the co-building of the future ecosystem, now that it has been open-sourced, how many partners, talents, and teams in this industry can fairly well reproduce some of our currently open-sourced models and results?
下一步,當我們希望這個生態進一步發展時,您覺得我們在哪些方面需要更多高質量的人才來接入我們的模型並復現它們?還是說,現在 GPU 算力相對比較稀缺?未來大家會不會都走向模型矩陣的方式?比如說,隨著我們把大模型的基座模型做得越來越好,產業鏈上的合作方和團隊再去做模型矩陣裡的垂直行業模型,或者一些應用模型。
這方面現在發展得怎麼樣了?兩年、三年之後,你覺得它會長成一個什麼樣的生態?這是我的第一個問題。第二個問題,剛才您也跟很多合作夥伴分享了很多對 AI 硬件的觀察。比如全球 AI 巨頭們可能每一家都在做百億美金級別的投入,而中國目前似乎在硬件和算力上有一些短板。
您覺得這個需要多長時間才能解決,並支撐我們 AI 的發展,讓算力和硬件上的短板不至於拖 AGI 的後腿?您覺得這是不是中國人最終一定能做到的事情,肯定能把它做成,只是時間和資本投入的問題?但同時,它可能也是一個雙向的過程。
一方面,模型進步會提升模型的智能度,導致單個任務或者某些智能單體對硬件和算力的消耗逐漸下降,不再需要那麼大的計算能力,因為模型進步會讓它思考得更高效。我不知道我理解得對不對。另一方面,硬件技術的進步會讓算力的計算能效更強。這會不會是一條雙方相向而行的路徑?
在這個時間點,如果我們用現在的 960 或 H200,做百億美金級別的算力投入,您剛才也提到是按三年攤銷。您覺得它的實際生命週期是多久?技術迭代是四年還是五年?或者說得更直白一點,會不會出現今天用現在的卡和 1 萬卡集群建一個算力中心,但三年後它其實相對來說已經不算很先進的算力了?
會不會出現這樣一種情況:現在正在建設、當下還不夠用,但三年後它相對變成了高質量算力不足、甚至有過剩產能?我不知道這種現象會不會存在。請您對這兩個問題給點建議。謝謝。謝謝。
梁文锋
第一個問題是關於生態。現在我們的感受是,也許每個公司都面臨人才不夠的問題。但我覺得這個人才短缺是暫時的。任何一個行業在發展早期,人才都不夠。包括最早做網站的時候,做網站的人非常少,人才也很稀缺。後來互聯網需要做服務器端開發,人才也很稀缺。
但是這種人才短缺很快就能解決,兩三年就能解決,因為會有大量的人被培養出來。AI 人才的短缺也是暫時的,而且我們已經看到它大大緩解了。因為確實不缺 AI 人。每家公司都會很快培養出人來,培養人是很快的。所以整體來看,在 AI 產業裡,不管是生態,還是模型公司,還是其他什麼,人都不稀缺。人才稀缺肯定是一個短期現象。
歷史上,從來沒有哪一類人會長期短缺。我還記得十多年前大家說飛行員短缺。飛行員的培養週期很長,但這個問題也很快就解決了。所以大家不用擔心人才短缺。還有一點,中國現在做模型的公司也有點太多了,還是太多。美國可能就三家公司。在中國,做基座模型的努力太多了。最後肯定不需要那麼多公司去做基座模型,肯定會收斂。
所以資源也相對比較分散,在某種程度上相對浪費。就在前不久
大家都得做同樣的事情,但在美國只需要三家公司去做,資源就集中在那三家。中國這邊資源分得很散,每家分到的就少很多。我覺得這個一定會收斂,一定會發生,只是需要時間,最後一定會收斂。沒有必要那麼多家公司,因為現在大家可能覺得做這個的利潤空間很高,所以都必須自己做。
但當他們發現這可能不是一個那麼高毛利的生意時,就可能不做了。最近肯定沒有那麼高毛利了。我不相信有那麼高的毛利,因為這不符合客觀規律。那就說明我們現在處在一個如果有非常高的利潤空間,那肯定不符合客觀規律的階段。我們應該有合理的利潤。所以這就是行業當前的狀態,我覺得它肯定會收斂。
大家都在做大模型這一塊,更不用說一家獨大、說「我要把所有利潤都拿走」,那肯定不行。如果每家公司只拿合理的利潤,那其實也沒必要那麼多公司去做大模型。最後中國有三四家公司競爭就已經夠了,競爭已經非常充分了,而且定價絕對足夠打價格戰。
For large models, it may not even be two big companies and two small companies; maybe that is already quite enough. As for the ecosystem, I don't have too many ideas. We hope to support more people, but we don't have that much energy. We do have this intention, and there will be no conflict of interest, but whether we actually do it is another matter. But at least there is no conflict of interest here; we hope for cooperation and a win-win outcome.
First, I absolutely do not believe that large model companies can take away most of the profit; that is impossible, because there are so many large model companies. The gap does not need to be that big now. There are only two things that create a gap: one is time, the other is cost. So it is not like any one company can have excess profits; I don't think there can be excess profits. Those who control costs well earn a bit more, and those who control costs poorly earn a bit less, and that is all. Did that answer it? Was the first question answered?
投資人
Everyone can... you believe that in the future there will definitely be many people who can... in the future it will actually... everyone’s data applications and such will iterate in a cycle... can you hear me now? Thank you. Right, thank you for your answer, and I also deeply understand and respect your ecosystem strategic positioning within the industry. For example, on the data side, publicly available data—I'm sure model companies already have channels to obtain it, so that method should not be a problem; it's just a matter of time and cost.
Then later, for example, when we really get to AGI, one possible imagined or ideal state is that the model can self-iterate and self-learn, meaning it trains itself. On this point, regarding the current data, do you think simulated data can be used, or is real data still of the highest quality?
If real data is still needed, would that limit AI intelligence to human history... at that level, because it depends on real data that humans truly had in the past? Or can this upper bound be broken through using simulated data, synthetic data, generated data, and other methods, allowing model capabilities to surpass all of humanity's past real...
梁文鋒
I think it can surpass it. I think there are two kinds of surpassing, for example Go, AlphaGo made a move that humans had never seen before. That is to say, it definitely surpassed humans within a certain range. But it may also have an upper bound, and it may also have limitations. But we cannot see those limitations now.
We believe, so broadly speaking, that it can surpass what humans already have, and the knowledge we can already express.
投資人
Then later on, will this rely on real data or simulated data? Will it work?
梁文鋒
There are many methods; it is not impossible.
投資人
Okay, thank you. I’ve also taken up your time, and I’d like to continue asking about the AI Infra issue just now.
梁文鋒
What was the second question? A bit...
投資人
Okay, I’ll repeat it briefly and quickly. I wanted to ask about AI Infra. In the future, we believe computing power is now being invested at the hundred-billion-dollar scale by everyone. On this point, we may believe that Chinese people will eventually achieve their mission on hardware, and one day there may be highly efficient compute, but in practice, it is still currently a bottleneck. So in the future, will the two sides move downward toward each other?
On one hand, after model capabilities iterate, it actually shifts from brute-force compute to clever compute, so the requirements and consumption of compute per unit model or task will gradually decrease at the margin. On the other hand, if hardware such as card capabilities improve, iteration speeds will become faster and single-card efficiency will improve. What kind of phenomenon is this in the process?
Will it be that now you build a 10,000-card cluster and buy H200 or 960, but after two or three years it becomes relatively less high-quality compute and relatively becomes an obsolete device?
梁文鋒
NVIDIA cards can basically be depreciated over five years. Huawei cards should be depreciated over at most three years. Huawei 950 is still pretty good to use this year; I think it is still okay to use it next year, but using it after that I think it may really be too power-hungry. The lifecycle of Huawei cards is definitely shorter, because they are already two years behind NVIDIA to begin with. But I think the gap is not that big. If B200 can be bought now, I think it is all worthwhile.
If it were Tencent, and Alibaba could buy it, then it depends on the scale; if it can be bought at a reasonable price, it is definitely worthwhile. But when calculating costs, believe me, you can't buy it.
投資人
Understood.
梁文鋒
Our compute is lagging behind; that is a fact. This fact is alleviated through three aspects. The first aspect is that we accept model lag; we can only use smaller models than theirs, and how much smaller is a training issue. We need to accept a certain degree of model lag, as well as smaller model sizes. This lag has one advantage: lag means you have more time. You have a certain amount of technology, and then you can use clever methods
投資人
Thank you.
梁文鋒
So our gap with the US may be that we are 12 months behind the US, perhaps 12 to 18 months behind, or 6 to 12 months behind. Simply put, we are two years behind the US, and then we do this with only one-twentieth of the compute. That narrative is: one to two years behind, but using only one-twentieth of their compute.
然後在未來,我們需要重寫那個敘事:我們只用了他們一小部分的算力,但把時間進一步縮短,縮到 6 個月、3 個月。我覺得這是一個目標。而且我們甚至可能在某些領域超過他們。但是,當整體算力還存在一個數量級的差距時,全面超越是不現實的;不過,在一些我們做出取捨的關鍵領域裡,超過某些地方是有可能的。
投資人
明白了,謝謝梁總。很有信心,我們一起加油。時間就留給其他夥伴,謝謝。感謝剛才的分享。我有兩個很快的技術問題。在剛才討論的技術路線圖裡,提到目前這個階段我們需要解決的核心問題是持續學習,這個目前也是海外研究的熱點,叫做 RecursiveImprovement。我想請問,從技術角度來看,現在最大的難點是什麼?
在您看來,這個問題什麼時候能解決?這是第一個問題。第二個問題是,您剛才提到先解決持續學習,再去做智能。我也想了解一下您這麼說背後的技術根源。這是否意味著,解決了持續學習之後,DeepSeek 也會進一步去做通用智能?請幫忙解釋這兩個問題。
梁文鋒
技術問題其實有點難講……難點在於,我們現在還沒有找到一個非常可行的方法。全世界現在都還沒找到一個好的方法,大家都還在探索。所以我們還是在探索階段,也就是說,我們不知道下一個解決這個問題的方法會被誰找到。我們還是在探索階段。我們現在有很多想法,很多看起來很有希望的想法,但目前都還沒有真正跑通。對,這是第一點。
第二點是,內部其實我們很看重這個敘事:訓練我們下一版的模型,我們希望它能幫助我們自己的發展。它可以提升 DeepSeek 的效率;我們模型的第一個目標就是提升 DeepSeek 自己的工作效率,這樣我們在開發下一版模型的時候,它就能提供更多幫助。
或者更直白地說,我們做出來的模型,首先不是為了對所有人有用,而是為了對我們自己有用。先要對我們有用。等它對我們有用了之後,那我再去開發下一版模型,它就會更快。
我們裡面很多人都是這樣想的:先要對自己有用,先要對我們能起作用。然後這就是實現 AGI 的最快方式。當它對我們很好用時,可能也意味著它對別人也很好用,但首先它得先對我們好用。這個敘事有點奇怪,但很多人確實是這樣想的。
這不是為了用戶滿意度;我希望它更多地幫助我們,這樣我們才能更快地實現 AGI。所以這個敘事的邏輯是,它幫助我們實現 AGI。但首先它要幫助我們實現。我們需要這種幫助來實現 AGI。現在非常確定的是,我們真的需要人工智能來幫助我們實現 AGI。
雖然它目前還不能自主工作,仍然只能和人類結合起來使用,但它已經非常有用了。
投資人
第二個問題是,剛才您提到先解決持續學習,再進入通用智能。這是您對未來的預期嗎?我想了解一下這背後的技術根源。為什麼我們需要先解決持續學習,再進入通用智能?您後面對這個領域的理解。
梁文鋒
因為解決持續學習可以極大加速我們的研發進度。如果我先把持續學習問題解決了,那麼通用智能的問題就不再是什麼大問題了。有了 AI 的幫助,如果 AI 能夠持續學習,它的能力應該會非常強。現在 agent 的能力有限,就是因為它不能持續學習;它不能有效地持續學下去。
如果我們能先把持續學習做完,那麼 AI 的能力會非常強,而且它可以大大提升我們自己研究的效率。如果先把持續學習做好,通用智能可能就變得很容易,用它來做這件事也會很容易。所以我說這是一個我們非常希望看到的結果;它能幫我們省力,我們也更輕鬆。
否則,如果你現在去手工做通用智能,會比較累、比較難,這是一個數據密集、勞動密集的工作,而且性價比也不高。
投資人
謝謝分享。
主持人
一個小問題:請看一下線上問題的聊天群。您覺得還要多久才到 AGI?到時候國產硬體能追上嗎?在 Zoom 會議的聊天窗口裡。
梁文鋒
好的,我看到了。華為 950——現在華為給我們 16,000 張卡,這個應該是可以公開說的。這大概比大網際網路公司少一個數量級左右。華為只能給我們這麼多,因為價格也不便宜。
大網際網路公司的需求會更大;對他們來說,他們更需要這個。對我們來說,我們可以買一些非合規的卡。所以我們買華為 950 的目的,還是希望幫助華為把生態做起來。16,000 張華為 950 卡只相當於 4,000 張 B 系列卡。所以數量不算很大,意義也不是特別大。
這還不足以訓練下一代模型;只夠訓練我們當前這一代模型,不夠下一代。但它可以幫華為先把這件事做對,也就是關於華為 950。那麼,還要多久才到 AGI?到時候國產硬體能追上嗎?
我認為在 AI 這件事上,國內大概一兩年內就能達到和海外差不多的水平,甚至也許今年就能實現對海外模型的替代。在 AI 上,按照現在的路線和現在的範式,這並不算很難,所以今年應該就有可能。但那還不是 AGI。
至少我覺得,它得能夠持續學習。那國內硬件到那時候能追上嗎?我覺得國內可能還要幾年。首先中國要先把生態問題解決掉,因為生態是信心問題。把生態問題解決了,再解決算力問題,我覺得應該是可以逐步解決的。我不是很相信五年後我們還會卡在算力上。
現在我們肯定是卡在算力上。今年、明年、後年,我覺得可能還是會卡在算力上,但五年後,我覺得未必還是這樣;我還是相對樂觀的。然後第二個問題是要想清楚未來的組織架構,以及人員規劃的規模。首先,我們之前的組織架構是非常鬆散的,因為根本就沒有組織架構。但隨著我們的人繼續擴大,這些東西肯定都需要一些變化。
目前我只能說,這裡面會有很多調整,但現在一時半會兒也很難全部表達清楚。最終,我們大概還是會需要不同的部門,其中有些部門需要我們建立相對嚴謹的層級結構;而有些部門可能還是會保持相對寬鬆、相對扁平的結構。隨著人員增加,我們會做這些調整。這個調整應該很快就會做,因為我已經在做這個調整了。
如果再不做這個調整,很多事情根本推不下去。確實有很多部門應該有組織架構。還有一個大家都有的問題是,CV 是在前哪個版本,對吧?我覺得現在這個 GCV4 版本的線上發布還是比較粗糙,還有很多能力需要時間。
對我們來說,一般比較舒服的發版節奏大概是兩到三個月一版。上一次發布可能是在四月底,那麼下一次發布可能在六月底,差不多這樣。如果沒有什麼意外,每個版本都應該比上一個版本更好。到 50B 活躍參數這個規模,我感覺最後和現在這波開源不會差太多。
在推理速度和性能上,我感覺可能差別也不會太大。但和那個大模型、他們那個未公開的模型相比,差距應該還是會挺大的。這個差距,我覺得以我們的活躍參數規模可能做不到;肯定還是需要更大的模型,可能要到 150B 這種級別。
至於 150B,按照我們目前的訓練進度,樂觀的話,年底可以開始訓練;至少也要到明年 Q1... 差距還是很大。對,這就是和 OCE 的差距。
投資人
你好,我其實有一個問題。你經常說 AGI 的實現過程是一個漸進的過程,不是突然一下子跳躍過去的。那我能不能理解為,這是一個沒有臨界點的過程?
梁文锋
它沒有臨界點,但是它是非線性的。我們現在相信的一個敘事是,AI 可以加速 AI 研究,AI 可以加速 AI 研究。也就是說,它不是線性的,因為你可以用 AI 來加速你自己的研究,所以後面它可能就會變成非線性的。
投資人
明白。那這個階段,我的理解是前面的結論可能就是繼續把語言模型做大就夠了,就足以達到這個狀態。
梁文锋
我只能說,對於語言模型的擴大,我目前沒有看到上限。我們還沒有看到我們現在智能水平的上限,甚至沒有看到美國智能水平的上限。
投資人
明白。因為我很好奇——你說的是,對美國來說,800B 活躍參數的模型是能訓練出來的,但不能用,所以只能訓練;也很難把它拿出來給大家用,因為真的很貴。我之前其實很有好奇的一點是,人類有語言能力也就近 10 萬年,而在那之前演化可能用了 37 億年。
但在訓練 AI 的時候,可能是可以的——你說那個順序可以反過來——但最後它還是要進入所謂的 world model,或者說不一定叫 world model,但最後還是要進入物理模型或者具身部分,對吧?也就是說,在這個上限之後。
梁文锋
對,我覺得具身肯定還是要進去的,最終還是具身。所以對我們公司來說,自然最終點可能就是具身。因為對一個正常人來說,他的需求不是一個電腦,對吧?因為正常人要吃、要喝、要玩、要穿、要住、要行,他不需要一個電腦。
他需要的是具身智能來解決具體的人類勞動需求。如果目標是解決勞動力需求,那具身就是繞不開的。
投資人
明白。那在某一個階段,如果我們達到類似的程度,也許甚至不是臨界點——就是那種能夠自我演化、能夠比較好地自我演化,或者接近 SV 那個上升點的狀態——我很好奇 AI 在那個狀態下第一個落地的地方會是什麼……也許會和現在不一樣。
梁文锋
What we hope is that it can, if there is no embodiment, then our definition of AGI, or what we hope AGI can do, is what? It can help me iterate the next version of the model; it can help me iterate the next version of the model in the same way. And then if embodiment is added, what we hope it can do is also let it iterate the next version of embodiment, let it make the next version of the robot.
Investor
I’m still very curious about one thing, because I saw previous DeepSeek interviews, etc. It seems that in the selection of important directions and research directions, taste and intuition are very important, not just simple engineering optimization. If AI can self-evolve in the future, will taste, taste, intuition still matter, or what will matter?
梁文锋
AI does not currently lack taste or intuition; what it lacks is the ability to continue learning. AI’s taste and intuition are fine. If you ask it to write an article, its taste and intuition, I think, are not a problem. There were a few questions earlier; let me look at them. I saw a few questions on the screen, but I can’t see them here.
Investor
Mr. Liang, I left a question on the screen; let me read it to you again. Actually, I want to ask about continuous learning—you also mentioned it, and many researchers have mentioned that it is still an unsolved research problem. Then the coding agent, especially catching up with MILES, is
reaching the Office level and MIS level a relatively certain target. For an unsolved research Scaling target and a relatively certain Scaling target, how do you think research resources should be allocated, especially the talent resources on the research side, so as to achieve the best balance and effect?
梁文锋
Model Office is a relatively certain target, but model MIS, I think, is still hard to say it’s certain; it can only be said to be a target. Model Office, I think, should be relatively certain.
CoT does not consume resources. Doing any research does not consume cards, it only needs very few cards; what it needs is ideas. It also does not consume talent resources, because you don’t need someone to stay there doing it all the time. It is not a project, but rather something that requires many people to think about the problem. So there is no need to allocate resources to it, because it doesn’t need resources. You need resources to train models, to make models, to release models, to do efficiency experiments for models.
What I just mentioned doesn’t consume much in terms of people or cards. So we call this “mò jiǎng” [trying your luck]. The threshold is very low, anyone can try, but who can get anything out of it—that, I also don’t know whether it depends on talent or what. So there is no need for us to allocate resources here. It’s just that the difference between us and other companies is that we will spend time discussing this problem, thinking about this problem, and treating it as an important matter.
Inside the company, it is an important problem, something we will spend time thinking about, but it does not require a lot of resources to do. Then I also see another question below: the hallucination problem of large models affects the user experience quite a lot. The hallucination problem also has a method that can solve it, but this is a long-term topic.
The hallucination problem can be regarded as something that can be solved, or improved, through better Post-training.
It’s just that everyone hasn’t put in much effort to do it. Or, for me, hallucination is a problem, but we may classify it as a product issue. We will solve it, but it is not a key issue. There was also a question earlier about data labeling. In terms of data annotation, this has to do with our capital investment. Given the structure of our capital investment, we cannot support the cost of so much high-quality data annotation, because the cost is very high.
The cost of data annotation in the United States is not materially different from the cost of data annotation in China. China does not have a cost advantage in data labeling; especially in high-end data labeling, there is no cost advantage, which makes it very difficult for us to invest in data labeling the way the United States does. This path is very difficult in China, because data labeling is simply too expensive—whether we outsource it or do it ourselves, it’s all very painful. So now it’s basically a two-pronged approach.
It’s not that we absolutely cannot label data; it’s just that some kinds of data labeling have low cost, and some have high cost. We start by labeling the low-cost ones. So you could also say that now half the people in our company are labeling data. Half of the core researchers, the most important people, half are labeling data. We are concentrating on data labeling. Solving the AI problem at this stage relies on data labeling. You just need to look at it as a data issue.
Investor
Mr. Liang, thank you. What you said was especially relatable for us, very good. Thank you. I’d like to ask a few questions.
The first question: you just mentioned that Chinese models are definitely stronger than American models in terms of efficiency. You also mentioned some other aspects, and in the future we may also be stronger than the United States in those. What do you think those aspects are, where we may be stronger than the United States in intelligence or in other areas?
梁文锋
I think in many experience-related aspects, it’s possible that we can do better than the United States. Not to mention our own experience, our own product experience feels pretty good. I think in user experience we may not necessarily be worse than the United States. In terms of product, product capability may not necessarily be worse than the United States. Costs should also be lower than in the United States, so China may still have competitiveness. In other areas, if you ask whether there are structural advantages, I think maybe not.
但在成本和產品上,我認為確實存在某些結構性優勢。成本這部分很容易理解,因為他們不需要做,所以也不會去發展這個能力。他們肯定沒有我們這麼重視這件事。我們可以把它看得很重要,但對他們來說,這並不重要。至於產品方面,畢竟經過很多家公司,產品能力還是可以的。
所以我認為這兩方面可能有結構性優勢。
Investor
好的。第二個問題,我想問你:你剛才提到了後訓練。我們的投資成本相對比較高,像 Anthropic 和 OpenAI 這樣的公司都在投入巨額資金。這次融資之後,你覺得我們會增加對後訓練的投入嗎?
梁文锋
差距主要還是在高品質資料標註上,而且主要是在 AI research 內部。我們肯定會加大投入,但高品質資料標註通常不是一個資本投入的問題。我認為高品質資料標註的瓶頸,其實在於時間,因為它需要時間。對於 OpenAI,對於海外公司,對於 Anthropic,他們都起步更早,而且他們的資本更多、卡也更多。
在這種情況下,對於我們國內來說,你可以理解為其實是最近六個月我們才真正開始做這件事,所以從時間上看,我覺得還需要更多時間。這跟資本投入關係沒那麼大,因為即使不再增加資本投入,現有的資本也足夠讓它盡可能快地擴張。但是這個速度是有上限的;瓶頸不在於我能不能立刻招到更多人,也不受錢或卡的限制。不過它確實是在一個快速擴張的過程中。
所以我們認為,未來一年之內,如果高品質資料這個問題能做好,我覺得那應該是國內可以期待的事情。我覺得前景可能沒那麼冷,但它確實需要時間。
Investor
謝謝。那第三個問題是,我們看到 Anthropic 用自己的模型去做自己的產品,在金融、法律上推出了很多垂直場景……
甚至未來進入醫療保健。你覺得在這種垂直應用上,未來某個階段我們也會考慮這樣做嗎?
梁文锋
我到現在其實還沒有想得很清楚——我們未來國內到底會是什麼商業模式,或者最順的路徑會是什麼。我們還沒到那個階段。國內情況和海外情況可能不一樣;到時候在國內會是什麼樣子,現在還很難判斷。
按照目前國內的情況,我覺得比較合理的方式應該是先全力做通用 Agent;其他 Agent 的優先級應該低一些,包括金融和醫生 Agent。我們應該先做 Coding,因為 Coding Agent 能做很多事,而且還有更多垂直 Agent 可以做。
在這個階段,我們認為最重要的還是 Coding Agent。
Investor
非常清楚,謝謝。還有一個問題我們也想問你。其實我們做 DeepSeek,也非常敬佩你一直以一種非常純粹的研究方式在做 DeepSeek。但這個行業確實已經進入資本市場,走上資本化的道路。而且你也是一個非常負責的人——無論對隊友還是投資者,你都非常負責。
那麼,你怎麼平衡純研究、純 AGI 方向,以及資本市場?你未來肯定還是要進入資本市場,肯定也會有公眾股東。你未來怎麼看待這種平衡?
梁文锋
我現在感覺兩者應該都能做。假設今年我能有幾億美元的 B 端收入,同時我們還有 C 端使用者,那其實就已經有了一定的商業基礎。如果明年我們還有 B 端收入,而且這個需求還能繼續增長,公司可能離淨利潤就不遠了;甚至可能已經是淨利潤。
這時候可能就不再是純燒錢階段了,所以我覺得我們後面能做的事情、能運營的空間還是比較大的。或者最差的情況下,賣 API 甚至足以支撐一家上市公司。如果說後面沒有新的技術進展,我們的技術就停在這裡,那最後我們就專心把 API 賣好,把這些服務做好,我覺得也就夠了。所以我還是有信心的;真的沒有那麼難。
因為我們確實處在一個槓桿很高的位置,而且又是在一個發展非常快的領域,可能真的沒有那麼難。我們只能說希望有更大的夢想,但我們也有可以拿得出手的保底表現。
Investor
謝謝,說得非常好。我的問題就這些。謝謝。
感謝分享。你前面提到了一些問題;我還想再問你一下。因為我覺得 DeepSeek 和其他公司最大的不同,其實是我們的組織和別人不一樣,或者說組織形態不一樣。但組織形態又和我們的目標相關,我們可能也需要考慮組織自身的邊界和效率。我不知道,從宏觀來看,我們的組織形態是否有一個好的學習模型?
歷史上,也許是 Bell Labs,或者什麼樣的結構會比較理想?還是說我們覺得其實並沒有真正理想的形式,而是需要自己慢慢去探索?也許在新時代,我們只能靠自己來做這個探索。因為我們的組織形態其實肯定和美國的大公司不同,對吧?這三家公司本身也都不同,但至少它們都是從商業公司的角度出發去探索的。
我主要還是想再問一下組織這個問題。
梁文锋
First of all, we do not have an object to imitate. Every step is based on our actual situation, seeking truth from facts, making decisions according to the actual situation, and finding out how we should do it. So it is a product of the times, or a reflection of the real situation; it is not the result of imitation. In other words, under this situation, I really think the optimal solution may be like this, or rather, in my own view, the path we chose is like this.
Every step, of course, we thought it through, of course we chose it, and in any case the result of the choice is what it is. It is not because we saw someone else choose this way and then we chose this way; it is because we analyzed the pros and cons and chose this way. In the future it may be the same; we are not imitating anyone.
I think we are still different from Bell Labs, because it clearly did not need commercialization, because it was... but we clearly need commercialization. In the end, we still have to survive; after all, we are a company, and the government will not give me a cent. So we can have very grand missions, but in the final analysis we are a company, and we have to think about how to survive.
So B-side is definitely important to us, because maybe in the future we may have to rely on it to survive. It’s just that it’s not important right now, because right now it is a cost line. So I think that is still different; it is still different from Bell Labs. We are essentially still a company. Historically, there have also been many companies that pursued things beyond profit, but you can’t say that because they pursue things beyond profit, they are not companies.
Many companies are great because they have a pursuit beyond profit. That pursuit not only did not affect its commercialization, but instead allowed it to commercialize better. We are essentially still a company; it’s just that we are considering which money to make, when to make money, how much to make, and what to rely on to make money—we simply have trade-offs.
投資者
Mr. Liang, I have two small questions; let me ask quickly. One is that you just mentioned MILOS—maybe it’s not a very certain target, but we will definitely move in that direction. And you also mentioned that the active parameters may, for example...
The next generation may be in the range of 150 to 250 B. In that case, do you think 150 to 250 B is benchmarking O4.7, or maybe benchmarking something else? That’s the first question. The second question is that you also mentioned earlier that in our entire inference stack, we used some programming languages other than CUDA.
I understand that originally, based on NVIDIA's ecosystem, we might have used a lot of PTX and the like. Now that we're using more things like TileLang you just mentioned, will that significantly reduce some of our efficiency in inference, or reduce efficiency in the short term? I don't know how you would view the efficiency loss brought about by this kind of change in programming languages, or whether in the long run it is actually a complementary state, meaning improved efficiency?
梁文鋒
It is improved efficiency, a big improvement in efficiency.
投資者
OK, so there isn't really any negative impact instead?
梁文鋒
Yes, it is a big improvement in efficiency. So this is an opportunity. It is like in the past you couldn't do without the CUDA ecosystem; now we can discard that ecosystem and use a simpler method, TileLang. It's a high-level language, and writing programs with it is very fast too; the amount of code needed is very small, and I can rewrite it from scratch.
投資者
I understand. So both of these are actually, as you mentioned, big opportunities brought by AI, not shortcomings that may need to be made up for in the short term.
梁文鋒
Yes, it is a major opportunity in technological development. It is not an AI opportunity, because we also have a project in which we are using AI to write TileLang.
投資者
So is it even faster?
梁文鋒
Right now all TileLang is written by humans, but it is already much faster than writing CUDA before.
投資者
I understand. So even for the execution efficiency at the hardware level, you think there is no impact either?
梁文鋒
A loss of 1% to 2% is acceptable, I think.
投資者
OK, got it, understood. Good, thank you, that was very clear.
主持人
Does anyone else have any questions? If not, let's stop here for today.
投資者
Okay, thank you. Thank you all very much for your time, thank you, Mr. Liang, thank you.
主持人
Bye-bye.
梁文鋒
Bye-bye.
內容僅供參考,不構成投資、法律、稅務或財務建議。