카테고리 없음

240112 월지암면 양지린 AGI로 가는 길은 새로운 조직 패러다임이다

클로vㅏ 컴퓨터 2026. 8. 2. 08:12

Hello everyone. Hello everyone.
President Zhang, thank you for the introduction.
How about it? Don’t you all think what I said is right?
At first glance he’s young and handsome, and also an extremely capable AI expert.
It’s a great honor to sit down and chat with Zhilin this evening.
안녕하세요 여러분. 안녕하세요 여러분.
장 총장님, 소개해 주셔서 감사합니다.
어떠세요? 제가 말한 게 맞지 않나요?
한눈에 봐도 젊고 잘생겼고, 게다가 매우 뛰어난 AI 전문가입니다.
오늘 저녁 지린과 함께 이야기 나눌 수 있어 매우 영광입니다.
Now, your company is said to be a bit mysterious, especially the name.
“Moonshot AI” — Dark Side of the Moon — is a very interesting name.
Looking at this wave of AI startups, the way they name themselves is quite different from the past.
In the early internet era it was basically two-character names: Alibaba, Tencent, Baidu, right?
But when we get to you it’s “Dark Side of the Moon,” and my good friend Li Zhifei’s company is called “Sequence Monkey” — those big model names.
Everyone is coming up with very interesting names.
So let’s first uncover the secret behind the name “Moonshot / Dark Side of the Moon.” Is there any particular meaning?
이제 당신 회사는 좀 신비롭다고들 하죠. 특히 그 이름이요.
‘문샷 AI’ — 월지암면(달의 어두운 면) — 매우 흥미로운 이름입니다.
이번 AI 창업 물결을 보면 이름을 짓는 방식도 예전과 많이 다르네요.
초기 인터넷 시대에는 기본적으로 두 글자 이름이었어요. 알리바바, 텐센트, 바이두 같은 거 말이죠.
그런데 당신네에 와서는 ‘월지암면’이 되고, 제 좋은 친구 리쯔페이의 회사는 ‘시퀀스 몽키’라고 불리죠. 그런 대형 모델 이름들.
다들 이름을 아주 재미있게 짓습니다.
그럼 먼저 ‘문샷 / 월지암면’이라는 이름 뒤에 어떤 의미가 있는지 비밀을 밝혀볼까요?
I think people might be quite interested in this.
Our name actually comes from a rock album, because we all like rock music.
We used to be in a band as well.
And this year happened to be a rather special opportunity — fifty years ago Pink Floyd released the album Dark Side of the Moon.
This year is exactly the 50th anniversary.
I feel that for AGI or large models, it itself is a very special symbol.
Normally when we look at the moon we can only see the illuminated side; the other side is invisible and extremely mysterious.
Yet there is a strong impulse to explore the unknown, to find out what the dark side of the moon really is.
I think large models are quite similar: you want to explore something mysterious and unknown, and at the same time it is a very difficult and challenging task.
It can also combine with that underlying rock spirit — constantly innovating, constantly challenging the existing shape of things, and imagining what might come next.
So we thought of all kinds of different names at the time, but in the end I felt that the combination of “Moonshot” and “Dark Side of the Moon” in Chinese and English could better reflect our determination to do AGI and define what kind of people we are.
사람들이 여기에 꽤 관심이 있을 것 같습니다.
우리 이름은 사실 록 앨범에서 유래했습니다. 우리 모두 록을 좋아하거든요.
예전에는 밴드도 했습니다.
그리고 올해가 마침 특별한 계기가 됐어요. 50년 전 핑크 플로이드가 Dark Side of the Moon 앨범을 발표했거든요.
올해가 정확히 50주년입니다.
AGI나 대형 모델에 있어서 이것 자체가 매우 특별한 상징이라고 느낍니다.
평소에 달을 보면 빛나는 면만 보이죠. 뒷면은 보이지 않고 매우 신비롭습니다.
그런데도 미지를 탐험하고, 달의 어두운 면이 도대체 무엇인지 알고 싶은 강한 충동이 있습니다.
대형 모델도 비슷하다고 생각합니다. 신비롭고 未知한 것을 탐험하고 싶고, 동시에 매우 어렵고 도전적인 일입니다.
그 기저의 록 정신과도 결합할 수 있어요. 끊임없이 혁신하고, 기존의 형태에 도전하며, 다음이 어떤 모습일지 상상하는 것.
그래서 당시 여러 이름을 생각했지만, 결국 ‘Moonshot’과 ‘월지암면’의 중영문 조합이 AGI를 하겠다는 결심과 우리가 어떤 사람들인지를 더 잘 반영할 수 있다고 느꼈습니다.
Because I’m also a Pink Floyd fan, when you mention The Dark Side of the Moon, from my perspective that album, although it uses an astronomical name, is essentially talking about many things closer to the human subconscious.
I think that album is quite profound.
The main track is called “Brain Damage,” right?
It talks about what happens after a person develops hallucination.
Now fifty years later we are trying to save this “hallucination,” especially the hallucination problem of large models.
I feel that… many people call you Professor Yang, right?
But from the process of choosing this name, I can see that you are not only the scientist side.
Although Moonshot itself is about doing something challenging — the meaning already contains difficulty and exploration — I feel that “Dark Side of the Moon” also expresses something deeper, and somehow connects with large models on an artistic level.
You can tell you have a band background.
Let me gossip a bit: what position did you play in the band?
저도 핑크 플로이드 팬이라서, Dark Side of the Moon을 말씀하시니까 제 관점에서는 그 앨범이 천문학적인 이름을 썼지만 본질적으로는 인간의 잠재의식에 더 가까운 많은 것들을 다루고 있다고 느껴집니다.
그 앨범이 꽤 심오하다고 생각해요.
메인 트랙이 “Brain Damage”죠?
사람이 환각이 생긴 후의 이야기를 하고 있습니다.
이제 50년이 지나 우리는 이 ‘환각’, 특히 대형 모델의 환각 문제를 구하려고 하고 있죠.
많은 사람들이 당신을 양 교수라고 부르잖아요.
그런데 이 이름을 짓는 과정에서 보면 과학자로서의 면만 있는 게 아닌 것 같습니다.
문샷 자체가 도전적인 일을 하는 것이지만 — 의미에 이미 어려움과 탐험이 들어 있죠 — ‘월지암면’은 더 깊은 무언가를 말하면서 대형 모델과 어떤 경지에서 연결되는 것 같아요.
밴드 배경이 있다는 게 느껴집니다.
조금 가십을 해볼게요. 밴드에서 어떤 포지션이었나요?
I was the drummer.
Drummer, right?
So from the perspective of a band, what kind of role does the drummer roughly play?
I’m digressing a bit here.
Basically, controlling the rhythm — deciding how fast the tempo should be.
I think that provides a very important framework for the whole band when performing a song.
It might feel something like that.
드러머였습니다.
드러머군요.
밴드 관점에서 드러머는 대략 어떤 포지션인가요?
좀 옆길로 새는 거지만요.
기본적으로 리듬을 잡는 역할이에요. 템포를 얼마나 빠르게 할지 결정하는 거죠.
전체 밴드가 곡을 연주할 때 매우 중요한 프레임워크를 제공하는 것 같습니다.
그런 느낌일 수 있어요.
Interesting.
I feel that this perspective might also be transferable horizontally.
Indeed, the development of large models or doing something like Moonshot — rhythm is probably very important too.
흥미롭네요.
이 관점이 가로로도 전이될 수 있을 것 같아요.
실제로 대형 모델의 발전이나 문샷 같은 일을 할 때도 리듬이 매우 중요할 수 있겠죠.
Let’s talk about some more concrete things.
Actually, looking at you, because we have some mutual friends, even before the news of your entrepreneurship came out I roughly knew you were going to throw yourself into this large-model wave.
But I still want to ask: around the turn of last year and this year, when many people were deciding whether to jump into this wave,
going back to that time, how did you make that decision and choice?
Obviously it was an exciting thing, but how did you choose to build an organization and throw yourself into it?
Can you share the decision logic of that year?
좀 더 구체적인 이야기를 해보죠.
사실 당신을 보면, 공통 친구가 몇 명 있어서 창업 소식이 나오기 전에도 대략 당신이 이 대형 모델 물결에 뛰어들 것을 알고 있었습니다.
그래도 묻고 싶어요. 작년과 올해 경계 무렵, 많은 사람들이 이 물결에 뛰어들기로 결심하던 시기에
그때로 돌아가서, 당신은 어떻게 그 결정과 선택을 했나요?
분명히 흥분되는 일이었지만, 어떻게 조직을 만들어서 뛰어들기로 선택했나요?
당시의 결정 논리를 공유해 주실 수 있을까요?
This is a particularly interesting question.
I think for large models, last year people started hearing about GPT and related trends, right?
But we have actually been working in this field for many years — perhaps starting from 2018–2019, really training language models based on Transformers.
From the very beginning we felt that language models could be tools that improve performance in many different scenarios.
Later the second stage was thinking that language models might be useful for many many tasks.
Until later it became: language models might become the only problem of AI — all other problems don’t matter, or all problems can be solved by making the language model better, or more generally by making next-token prediction better.
So AI became a single problem.
My cognition actually underwent a very big change over the past few years.
The most important thing was probably starting from 2019 at Google, using thousands of cards to train Transformer models.
In that process we observed a great many phenomena, and those phenomena further provided more evidence that this might be the correct path.
You just need to keep walking along this path, continuously scaling, and continuously searching for more efficient ways to scale, and you can get very good results and solve many previously difficult problems — whether memory problems, reasoning problems, common-sense problems, or even more complex multi-hop problems.
So these things in previous experience gave me a very deep impact or foundation.
Then starting from 2020, we began in China — at that time I went to find many people to cooperate, found many institutions, and we trained large models together. We were also among the earliest in China to train models like PanGu and WuDao.
So in that process we were always brewing for a real opportunity.
But at the same time we also saw many challenges of large models — one side technical, the other side organizational.
We discovered that if you still use traditional methods and traditional organizational structures to train large models, it is probably very hard to succeed.
Looking at OpenAI’s success today, essentially it is also because their organization made extremely large innovations that led to today’s success.
So I feel that before that it could be understood as constantly searching for such an opportunity — how to build a new organization from zero.
Because I believe this is the key to AI, even more important than the various detailed technical points we touch today.
Organization is a more fundamental thing. Only when you get the organization right do you have the possibility to truly walk far on the path of AGI.
So when we saw that last year both the capital market and the talent market underwent extremely large changes, under those circumstances I felt the timing was more mature. We could have the chance to build an organization from scratch to do this.
이 질문은 특히 흥미롭습니다.
대형 모델에 있어서 작년에 GPT나 관련 흐름을 듣기 시작했겠죠.
하지만 우리는 사실 이 분야에서 매우 오랫동안 일해 왔습니다. 아마 2018–2019년부터 Transformer 기반 언어 모델을 실제로 훈련하기 시작했죠.
처음부터 언어 모델이 여러 다른 시나리오의 성능을 향상시킬 수 있는 도구가 될 수 있다고 느꼈습니다.
나중에는 두 번째 단계로, 언어 모델이 많은 많은 작업에 유용할 수 있다고 생각하게 됐죠.
결국에는 언어 모델이 AI의 유일한 문제가 될 수 있다 — 다른 모든 문제는 중요하지 않거나, 언어 모델을 더 좋게, 더 일반적으로 next-token prediction을 더 좋게 만들면 모든 문제를 해결할 수 있다 — 는 생각까지 이르렀습니다.
그래서 AI가 단일 문제가 된 거죠.
제 인식은 지난 몇 년 동안 매우 큰 변화를 겪었습니다.
가장 중요했던 것은 아마 2019년부터 구글에서 수천 장의 카드로 Transformer 모델을 훈련하던 과정이었습니다.
그 과정에서 매우 많은 현상을 관찰했고, 그 현상들이 이것이 올바른 길일 수 있다는 더 많은 증거를 제공해 주었습니다.
이 길을 계속 걸어가기만 하면, 계속 스케일하고, 더 효율적인 스케일 방식을 계속 찾아가면 매우 좋은 결과를 얻고 이전에 어려웠던 많은 문제들 — 기억 문제, 추론 문제, 상식 문제, 심지어 더 복잡한 다중 홉 문제까지 — 을 해결할 수 있습니다.
그래서 이전 경험 속의 이런 것들이 저에게 매우 깊은 충격이나 기초를 주었습니다.
그 후 2020년부터 국내에서 시작했습니다. 당시 많은 사람을 찾아 협력하고, 여러 기관을 찾아 함께 대형 모델을 훈련했죠. 중국에서 가장 먼저 판구(PanGu)나 우다오(WuDao) 같은 모델을 훈련한 팀 중 하나였습니다.
그 과정에서 계속 진짜 기회를 준비하고 있었던 거죠.
동시에 대형 모델의 많은 도전도 보았습니다. 한쪽은 기술, 다른 한쪽은 조직이었습니다.
전통적인 방식과 전통적인 조직 구조로 대형 모델을 훈련하면 성공하기 매우 어렵다는 것을 발견했습니다.
오늘날 OpenAI의 성공을 보면, 본질적으로도 그들의 조직이 극단적으로 큰 혁신을 했기 때문에 오늘의 성공이 있는 것입니다.
그래서 그 이전에는 그런 기회를 계속 찾고 있었던 것으로 이해할 수 있어요 — 어떻게 제로에서 새로운 조직을 만들 것인가.
저는 이것이 AI로 가는 열쇠라고 믿습니다. 오늘날 우리가 접하는 각종 세밀한 기술보다도 더 중요할 수 있어요.
조직이 더 근본적인 것입니다. 조직을 제대로 해야만 AGI 길에서 진정으로 멀리 갈 가능성이 생깁니다.
그래서 작년에 자본 시장과 인재 시장이 모두 극단적으로 큰 변화를 겪는 것을 보았을 때, 그 상황에서 타이밍이 더 성숙했다고 느꼈습니다. 제로에서 조직을 만들어 이 일을 할 기회가 생긴 거죠.
I feel the perspective you just shared gave me quite a few insights.
To a certain extent, you decided to throw yourself into this not because the market needed a certain type of large model in China or the world, but because you believe AGI needs innovative organizations to drive it forward.
In other words, develop technology by a new kind of organization, right?
Driving technological development from the organizational angle.
You felt that when the market and capital environment were ready, that gave you the opportunity, so you chose to start a company.
This perspective is something I rarely heard before.
What triggered such a strong idea in you?
I believe you are definitely a deep believer in AGI, otherwise you wouldn’t have named it Dark Side of the Moon and wouldn’t have this logic.
But what made you decide that developing the organization is a core problem, and that you should seize this opportunity to build such an organization?
당신이 말씀하신 관점이 저에게 꽤 많은 영감을 주었습니다.
어느 정도, 당신이 이 일에 뛰어들기로 한 것은 시장이 중국이나 세계에 어떤 유형의 대형 모델을 필요로 해서가 아니라, AGI가 혁신적인 조직에 의해 추진되어야 한다고 믿기 때문인 것 같습니다.
즉, 새로운 종류의 조직으로 기술을 발전시키는 것이죠.
조직의 각도에서 기술 발전을 추진하는 것.
시장과 자본 환경이 준비되었을 때 그 기회가 주어졌다고 느껴서 창업을 선택한 거군요.
이 관점은 이전에 거의 들어본 적이 없습니다.
무엇이 당신에게 이렇게 강한 생각을 촉발했나요?
당신은 분명히 AGI의 깊이 있는 신봉자일 거라고 믿습니다. 그렇지 않으면 월지암면이라고 이름 짓지도, 이런 논리를 가지지도 않았을 테니까요.
그런데 무엇이 조직 발전이 핵심 문제이며, 이 기회를 이용해 그런 조직을 구축해야 한다고 결정하게 만들었나요?
I think it is more from practice.
Because before this year we also tried many different ways, different organizations, to research and develop large models.
Not only large models, but also a whole series of technologies and products centered around large models.
In the past few years we did many similar things, with different attempts.
There were several different possibilities: one is doing it inside traditional enterprises; another is independent research institutions — I also participated in some before; even some university-style modes.
You discover that these different modes are all hard to make fundamental innovations at the organizational level.
They all have their problems — I won’t expand on exactly what the problems are today.
But I can give a very simple example: it is very hard to innovate on large models in a planned way.
For example, in the mobile internet era I can plan what demands I will do next, and once those demands are defined they can be produced deterministically.
There is almost no case where an app suddenly cannot be developed today, or a defined demand ends up not being realized, because it is a deterministic event — you just need people to express the meaning through coding on a computer and it can be realized.
But AGI is different.
It is hard to plan that today I want to complete a certain demand to a certain degree.
It cannot be hardcoded, cannot be expressed by rules.
So the way of doing things needed is not this front-loaded planning type of innovation, but more of a posterior type — I may need to try it to know, or rather, I am not pointing to something right now and saying “do this,” but instead you may need to first distinguish several different possibilities, and then solve them through a general method.
That method is not solving case by case one by one, because that is the speed of traditional mobile internet product development — I define one demand, then define another; my AI engineers label one batch of data, then label another batch.
It is case by case, so it is not AGI.
But AGI needs an underlying machine — a more general way that can do many things at once.
So I feel this is a very very mental difference.
And because your organization must match the way you do things, when the underlying logic of how you do things changes, you need a new organization to be able to do it.
I think in the internet era many excellent organizations emerged that might be extremely good at certain things, for example anything related to recommendation systems.
But in the new era there will probably be some organizations that are extremely good at AGI.
I feel this is something that will most likely happen.
더 많은 부분은 실천에서 온 것 같습니다.
올해 이전에도 우리는 여러 다른 방식, 다른 조직으로 대형 모델을 연구개발하려고 시도했습니다.
대형 모델뿐만 아니라 대형 모델을 중심으로 한 일련의 기술과 제품들도요.
지난 몇 년 동안 비슷한 일을 많이 했고, 여러 시도가 있었습니다.
여러 가능성이 있었죠. 하나는 전통 기업 내부에서 하는 방식, 또 하나는 독립 연구기관 — 이전에 저도 일부 참여했습니다 — 심지어 대학 스타일 모드도 있었습니다.
이런 다른 모드들이 조직 차원에서 근본적인 혁신을 하기 어렵다는 것을 발견했습니다.
그들 모두 문제가 있어요. 오늘은 구체적으로 어떤 문제인지 펼쳐 말하지 않겠습니다.
하지만 아주 간단한 예를 들 수 있습니다. 대형 모델을 계획적인 방식으로 혁신하기가 매우 어렵다는 점입니다.
예를 들어 모바일 인터넷 시대에는 다음에 어떤 수요를 할지 계획할 수 있고, 그 수요가 정의되면 확정적으로 생산할 수 있습니다.
앱이 갑자기 오늘 개발되지 않는다거나, 정의한 수요가 실현되지 않는 경우는 거의 없죠. 확정적 사건이기 때문입니다. 사람이 컴퓨터에서 코딩으로 그 의미를 표현하기만 하면 실현됩니다.
하지만 AGI는 다릅니다.
오늘 어떤 수요를 어느 정도까지 완성하겠다고 계획하기 어렵습니다.
하드코딩할 수 없고, 규칙으로 표현할 수 없습니다.
그래서 필요한 일하는 방식은 이런 앞선 계획형 혁신이 아니라, 더 사후적인 방식입니다. 해봐야 알 수 있거나, 지금 무언가를 가리키며 “이것을 하라”고 하는 게 아니라, 먼저 여러 다른 가능성을 구분한 다음 일반적인 방법으로 해결하는 것이죠.
그 방법은 case by case로 하나씩 해결하는 게 아닙니다. 그건 전통 모바일 인터넷 제품 개발 속도예요. 수요를 하나 정의하고 또 다른 수요를 정의하고, AI 엔지니어가 데이터 한 배치를 라벨링한 뒤 또 다른 배치를 라벨링하는 것.
case by case이므로 AGI가 아닙니다.
하지만 AGI는 기저의 기계가 필요합니다. 한 번에 많은 일을 할 수 있는 더 일반적인 방식 말입니다.
그래서 이것은 매우 매우 멘탈적인 차이라고 느낍니다.
그리고 조직은 일하는 방식과 맞춰야 하기 때문에, 일하는 기저 논리가 변하면 새로운 조직이 필요해집니다.
인터넷 시대에는 특정 일에 매우 능한 훌륭한 조직들이 많이 등장했습니다. 예를 들어 추천 시스템과 관련된 모든 것에 능한 조직들.
하지만 새 시대에는 AGI에 매우 능한 조직들이 등장할 가능성이 높습니다.
이것이 아마 일어날 일이라고 느낍니다.
I feel the perspective you just shared makes me think of something: when we look at the dark side of the moon, we can’t see it.
If we don’t have that kind of data, it’s hard to say we will do R&D based on something already determined, right?
Indeed it may need an organization that can match more general, more uncertain problems.
I can understand what you mean.
But let me ask one more: for example, in your eyes, is OpenAI a relatively good model of an organization oriented toward AGI?
What do you think they got right, and what might not be completely the best in your view?
당신이 말씀하신 관점이 저를 생각하게 만드네요. 달의 어두운 면을 볼 때 우리는 그것을 볼 수 없습니다.
그런 데이터가 없다면, 이미 확정된 것에 기반해서 R&D를 하겠다고 말하기 어렵겠죠?
실제로 더 일반적이고 더 불확실한 문제에 맞출 수 있는 조직이 필요할 수 있습니다.
말씀하신 의미를 이해할 수 있습니다.
하지만 하나 더 묻겠습니다. 예를 들어 당신 눈에 OpenAI는 AGI를 향한 조직 형태의 비교적 좋은 모델인가요?
그들이 잘한 점은 무엇이고, 당신이 보기에 완전히 최선이 아닐 수 있는 점은 무엇인가요?
I think, first of all from the results, they certainly made extremely large breakthroughs, right?
Without this company, I feel the progress of humanity might have been different.
So from the results they definitely did many correct things.
If we go deeper, of course I have not specifically participated in their organizational workflow, but there is some related information.
First, an excellent organization needs very high talent density, and then an extremely strong shared vision — everyone should have a common vision and be able to focus efficiently around one goal.
I feel these points are things they did extremely well.
But the most core point, and one that outsiders may not so easily see, and also the point we are now most focused on wanting to do better, is: after you have these premises, how do you find a systematic way of doing things.
I feel this point transcends all technology; it should be the prerequisite condition for all technology.
When you say “systematic,” do you mean it can be replicated and scaled?
It can, but the “replication” here means you can apply it to different things.
It may not necessarily be replicable to other places — forming it inside one company and replicating it elsewhere might be hard — but one company can repeatedly use this system to do different things.
That is my definition of system.
For example, today I can use it to overcome the long-context challenge; tomorrow I can use this thing to overcome the ability of autonomous AI; later I can still use it for multimodal; today I can use it to make a productivity product, tomorrow a generative product.
It should be a usable system.
Essentially what precipitates down becomes your core asset.
I feel this is something every AGI company should spend the most time polishing.
우선 결과적으로 그들은 분명히 매우 큰 돌파를 이뤘습니다.
이 회사가 없었다면 인류의 진보가 달라졌을 수도 있다고 느낍니다.
그래서 결과적으로 그들은 분명히 많은 올바른 일을 했습니다.
더 깊이 들어가면, 물론 저는 그들의 조직 워크플로우에 구체적으로 참여하지는 않았지만 관련 정보가 있습니다.
먼저, 훌륭한 조직은 매우 높은 인재 밀도가 필요하고, 그다음 극도로 강한 공유된 비전 — 모든 사람이 공통의 비전을 가지고 하나의 목표를 중심으로 효율적으로 집중할 수 있어야 합니다.
이 점들이 그들이 매우 잘한 부분이라고 느낍니다.
하지만 가장 핵심적인 점, 그리고 외부에서는 쉽게 보기 어려울 수 있는 점, 또한 우리가 지금 가장 잘하고 싶은 점은: 이런 전제들을 갖춘 후에 어떻게 체계적인 일하는 방식을 찾느냐입니다.
이 점이 모든 기술을 초월한다고 느낍니다. 모든 기술의 전제 조건이어야 합니다.
‘체계적’이라고 하신 것은 복제되고 확대될 수 있다는 의미인가요?
가능합니다. 하지만 여기서의 ‘복제’는 다른 일에 적용할 수 있다는 뜻입니다.
다른 곳으로 복제하기는 어려울 수 있어요 — 한 회사 안에서 형성한 것을 다른 곳에 복제하기는 힘들 수 있지만 — 한 회사가 이 시스템을 반복적으로 사용해 다른 일들을 할 수 있어야 합니다.
그게 제가 정의하는 시스템입니다.
예를 들어 오늘은 long-context 도전을 극복하는 데 사용할 수 있고, 내일은 자율 AI 능력을 극복하는 데 사용할 수 있으며, 나중에는 멀티모달에, 오늘은 생산성 제품을, 내일은 생성형 제품을 만드는 데 사용할 수 있어야 합니다.
사용 가능한 시스템이어야 합니다.
본질적으로 침전되는 것이 당신의 핵심 자산이 됩니다.
모든 AGI 회사가 가장 많은 시간을 들여 연마해야 할 것이라고 느낍니다.
I feel what you said provides a pretty good perspective.
Because previously when we talked we might all be talking about the large-model technology itself, but whether the technology behind large models can continue to develop better may also be related to the organization at its root.
On this point, previously when people talked about OpenAI or organizations, especially facing uncertain innovation, we often had two either-or methods: one called bottom-up, one called top-down, right?
I wonder, when you conceive the innovative system you mentioned — that is, inside this organization — can it be defined with the original things, or will it have some new definitions, not just bottom-up and top-down?
당신이 말씀하신 것이 꽤 좋은 관점을 제공합니다.
이전에는 대형 모델 기술 자체에 대해 이야기했을 수 있지만, 대형 모델 기술 뒤에 있는 것이 계속 더 잘 발전할 수 있는지는 근본적으로 조직과도 관련이 있을 수 있습니다.
이 점에서, 이전에 OpenAI나 조직에 대해 이야기할 때, 특히 불확실한 혁신에 직면했을 때 우리는 자주 두 가지 양자택일 방법을 가졌습니다. 하나는 bottom-up, 하나는 top-down이죠.
궁금한 게, 당신이 구상하는 혁신 시스템 — 즉 이 조직 안에서 — 원래의 것들로 정의될 수 있을까요, 아니면 bottom-up과 top-down만이 아닌 새로운 정의들이 있을까요?
I think the big framework is still applicable.
Especially for large models, having a top-down framework is very important.
It’s like moon-landing engineering — going to the moon is a very complex project. It’s not one move or one breath, nor one soldier or one clan.
It is actually a huge system that requires a relatively long time and many people engaged in extremely complex and mutually coupled things.
So some top-level settings are definitely needed.
On one hand it strongly emphasizes top-down; top-down is about the leadership’s vision — whether you can judge what is the right thing to do and what you should not do right now.
This kind of top-down design is definitely necessary.
Then under this framework, whether the organization can get the thing done well is the system I just talked about.
Inside it there may be many small units, each doing different things, but if you have a good system, each unit should be able to efficiently produce something.
It might be more like a factory — an AI factory.
Each unit produces different things, but there is a top-down framework that can integrate them.
For example, today you are exploring three product lines at the same time; those three product lines should all be able to effectively utilize your system.
For instance my system can quickly produce data so that in a new scenario it can quickly know whether there is possible PMF.
But everything happens under a top-down framework: first you decide which three directions, which three product lines to explore — that should have a top-level design in advance.
But once you have an excellent system you can quickly make these attempts.
And the process of making these attempts is an AI-native process.
Any traditional organization cannot do it, because the characteristic of AI-native is different from others: essentially you are using data to define a new product, and this point is different from any product that has existed in history.
That is why this system may also need a new system.
큰 프레임워크는 여전히 적용 가능하다고 생각합니다.
특히 대형 모델에 있어서 top-down 프레임워크가 매우 중요합니다.
달 착륙 공학과 같아요. 달에 가는 것은 매우 복잡한 프로젝트입니다. 한 동작이나 한 숨이 아니고, 한 병사나 한 부족도 아닙니다.
비교적 긴 시간과 많은 사람들이 극도로 복잡하고 서로 결합된 일에 종사하는 거대한 시스템입니다.
그래서 일부 최상위 설정이 분명히 필요합니다.
한쪽에서는 top-down을 강하게 강조합니다. top-down은 리더십의 비전 — 무엇이 해야 할 올바른 일이고 지금은 하지 말아야 할 일인지를 판단할 수 있는가 — 에 관한 것입니다.
이런 top-down 설계는 분명히 필요합니다.
그 프레임워크 아래에서 조직이 일을 잘할 수 있는지가 제가 방금 말한 시스템입니다.
그 안에는 많은 작은 유닛이 있을 수 있고, 각각 다른 일을 하지만, 좋은 시스템이 있으면 각 유닛이 효율적으로 무언가를 생산할 수 있어야 합니다.
더 공장 같을 수 있어요 — AI 공장.
각 유닛이 다른 것을 생산하지만, 그것들을 통합할 수 있는 top-down 프레임워크가 있습니다.
예를 들어 오늘 세 개의 제품 라인을 동시에 탐색한다면, 그 세 제품 라인 모두 당신의 시스템을 효과적으로 활용할 수 있어야 합니다.
제 시스템이 빠르게 데이터를 생산해서 새로운 시나리오에서 가능한 PMF가 있는지 빠르게 알 수 있게 하는 식이죠.
하지만 모든 것은 top-down 프레임워크 아래에서 일어납니다. 먼저 어떤 세 방향, 어떤 세 제품 라인을 탐색할지 — 미리 최상위 설계가 있어야 합니다.
하지만 훌륭한 시스템이 있으면 이런 시도를 빠르게 할 수 있습니다.
그리고 이런 시도를 하는 과정은 AI-native 과정입니다.
어떤 전통 조직도 할 수 없습니다. AI-native의 특징이 다른 것과 다르기 때문입니다. 본질적으로 데이터로 새로운 제품을 정의하는 것이고, 이 점은 역사상 존재했던 어떤 제품과도 다릅니다.
그래서 이 시스템도 새로운 시스템이 필요할 수 있는 이유입니다.
I see that in our live-stream room many people are still very impressed by this perspective.
Someone wrote a long comment saying technology is lower priority than economy, lower than institutions, lower than culture.
This is an important perspective — you cannot only follow technology, but must drive things at the organizational and institutional level, because culture is hard to change, so organizational institutions are a very important point.
I feel some people resonate with you.
Actually on the matter of innovation, if we look at the past — I have been watching technological revolutions for more than twenty years, watched several waves — I indeed had a perception at the time: sometimes a technological breakthrough is the intersection of several technologies, emerging from different fields, becoming the correct recipe at the correct moment, and then it explodes.
To a certain extent I feel you are actually, under a big framework inside an organization, simulating the ability to have suitable innovations emerge in different directions, and finally they together may lead to AGI.
Because today we cannot precisely define which path, by what timeline, completing it means AGI.
There are still large amounts of uncertain things inside, so it needs some emergence.
So the organization needs to support this kind of emergence, right?
But it also needs a framework, and also needs to support emergence.
I don’t know if I understand it correctly this way.
우리 라이브 방송실에서 많은 사람들이 여전히 이 관점에 매우 인상 깊어하는 것을 봅니다.
누군가 긴 댓글을 썼어요. 기술은 경제보다, 제도보다, 문화보다 우선순위가 낮다고.
이건 중요한 관점입니다. 기술만 따라갈 게 아니라 조직과 제도 차원에서 추진해야 한다. 문화는 바꾸기 어렵기 때문에 조직 제도가 매우 중요한 점이라고.
일부 사람들이 당신과 공명하는 것 같습니다.
사실 혁신에 있어서, 과거를 보면 — 저는 20년 넘게 기술 분야의 변혁을 지켜봤고 여러 물결을 봤습니다 — 당시 한 가지 인식이 있었습니다. 어떤 기술 돌파는 여러 기술의 교차이고, 다른 분야에서 출현해서 올바른 순간에 올바른 레시피가 되어 폭발하는 것.
어느 정도 당신은 사실 큰 프레임워크 아래 조직 안에서, 다른 방향에서 적합한 혁신이 출현할 수 있게 시뮬레이션하고 있고, 결국 그것들이 함께 AGI로 이어질 수 있다고 느낍니다.
오늘날 우리는 어떤 길을, 어떤 일정으로 완성하면 AGI인지 정밀하게 정의할 수 없습니다.
안에는 여전히 많은 불확실한 것들이 있어서 어떤 출현이 필요합니다.
그래서 조직이 이런 출현을 지원해야 하겠죠?
하지만 프레임워크도 필요하고, 출현도 지원해야 합니다.
제가 이렇게 이해한 게 맞는지 모르겠네요.
Yes, I feel this is a very good understanding.
For example we can look at how Transformer was produced.
Essentially Google provided an environment for emergence for that group of people.
Because looking at the background when Transformer was proposed, several technologies already existed in the world: attention mechanism, residual connections / residual networks, layer normalization, basic training setups like SGD, learning-rate schedules — all these things were prepared in advance.
Then Google provided such an environment that allowed these people to freely combine inside it: today I take these two things and combine them, tomorrow I take those things and combine them.
Then when they combined to that place, suddenly this architecture emerged.
But different environments can emerge different things.
Google’s environment could only emerge a scientific result — it could emerge a Transformer architecture, but it could not emerge a great systems engineering work, that is, it could not emerge something like GPT that takes one thing to the extreme while also precisely capturing demand on the product side — a cross-era work.
Google’s environment could not emerge that because its organization could not.
But OpenAI emerged something else.
They did not emerge the greatest scientific discovery in history; they did not invent any new thing.
But they emerged an industrialized operation.
What they combined was not the dimensions just mentioned; what they combined was: they could see what the existing stock in the world is — one is Transformer architecture, two is the computing centers that can support 10^25-scale computation, three is the data accumulated by the internet over twenty years.
That may be the greatest value of the internet.
So they saw these three factors, then provided an environment that allowed these three factors to be combined, and thus emerged an AGI milestone.
So this is what I just said: different organizations allow you to emerge different things.
But whatever you want it to emerge, you should adjust the organization in that direction.
네, 매우 좋은 이해라고 느낍니다.
예를 들어 Transformer가 어떻게 나왔는지 볼 수 있습니다.
본질적으로 구글이 그 사람들에게 출현을 위한 환경을 제공한 것입니다.
Transformer가 제안된 배경을 보면, 이미 세상에 여러 기술이 존재했습니다. 어텐션 메커니즘, residual connection / residual network, layer normalization, SGD 같은 기본 훈련 설정, learning-rate schedule — 이 모든 것들이 미리 준비되어 있었죠.
그 다음 구글이 그런 환경을 제공해서 그 사람들이 안에서 자유롭게 조합할 수 있게 했습니다. 오늘은 이 두 가지를 가져와서 조합하고, 내일은 저 것들을 가져와서 조합하고.
그러다 그 지점에 조합되었을 때 갑자기 이 아키텍처가 출현했습니다.
하지만 다른 환경은 다른 것을 출현시킬 수 있습니다.
구글의 환경은 과학적 결과만 출현시킬 수 있었습니다. Transformer 아키텍처를 출현시킬 수 있었지만, 위대한 시스템 엔지니어링 작품 — 즉 GPT처럼 하나를 극한까지 밀면서 제품 쪽에서 수요를 정밀하게 잡는 시대를 가로지르는 작품 — 을 출현시킬 수는 없었습니다.
구글의 환경은 그것을 출현시킬 수 없었습니다. 조직이 맞지 않았기 때문입니다.
하지만 OpenAI는 다른 것을 출현시켰습니다.
그들은 역사상 가장 위대한 과학적 발견을 출현시키지 않았습니다. 어떤 새로운 것도 발명하지 않았죠.
하지만 산업화 운영을 출현시켰습니다.
그들이 결합한 것은 방금 말한 차원이 아니었습니다. 그들이 결합한 것은 세상의 기존 재고가 무엇인지 볼 수 있었다는 점입니다. 하나는 Transformer 아키텍처, 둘은 10^25 규모의 연산을 지원할 수 있는 컴퓨팅 센터, 셋은 인터넷이 20년 동안 축적한 데이터.
그것이 인터넷의 가장 큰 가치일 수 있습니다.
그래서 그들은 이 세 요인을 보고, 이 세 요인이 결합될 수 있는 환경을 제공해서 AGI의 이정표를 출현시켰습니다.
그래서 이것이 제가 방금 말한 것입니다. 다른 조직은 다른 것을 출현시킬 수 있게 합니다.
하지만 무엇을 출현시키고 싶은지에 따라 조직을 그 방향으로 조정해야 합니다.
Ah, I feel first adjust the genes of the organization, then a new species may appear.
This is actually very Moonshot.
Looking at large models, in the process of emergence they emerged chain-of-thought, right?
Looking ahead, this organization may also need an emergence culture or architecture to do this.
I see someone in the live-stream room again wrote a long comment: in the past defining a demand and landing it was measured in months, but with the scalable characteristic of AI, new scenario capabilities can land in days, even many applications can land in one day.
So I feel he understood what you meant.
Actually in the future this kind of emergence may become the norm.
The organizations of the next era may need this emergence even more.
You provided a very good perspective today.
아, 먼저 조직의 유전자를 조정하면 새로운 종이 나타날 수 있다고 느낍니다.
이건 실제로 매우 문샷스럽네요.
대형 모델을 보면 출현 과정에서 chain-of-thought가 출현했잖아요.
앞을 보면 이 조직도 출현 문화나 아키텍처가 필요할 수 있습니다.
라이브 방송실에서 또 누군가 긴 댓글을 썼어요. 과거에는 수요를 정의하고 착지하는 게 달 단위였지만, AI의 scalable 특성 때문에 새로운 시나리오 능력이 일 단위로 착지할 수 있고, 심지어 하루에 많은 응용을 착지할 수 있다고.
그래서 그가 당신 말씀을 이해한 것 같습니다.
사실 미래에는 이런 출현이 상시가 될 수 있습니다.
다음 시대의 조직은 이런 출현을 더 필요로 할 수 있어요.
오늘 매우 좋은 관점을 제공해 주셨습니다.
Let me jump out and ask one more.
There might be another side.
A while ago many people were saying, at least in the first half of the year, there was a popular saying called “name-brand heavy industry” — that this large-model matter we can all see clearly: just follow the Transformer route, right?
The architecture has already been run to this point by someone, so we just follow it.
And inside it there doesn’t seem to be too much room for innovation left — compute power, engineering capability, and application capability.
The game looks very clear.
Even people said that for future models to become stronger just add parameters — at that time some people even proposed things like hundreds of trillions of parameters, which made us all a bit stunned. Our brains couldn’t count the zeros anymore, how many powers we couldn’t count.
So how do you look at it?
One side thinks the technology is determined, and next is just name-brand heavy industry, right?
Do you accept this view? Why don’t you accept it?
Looking at it you seem to go toward the organizational layer, and you think there is still a lot of uncertainty, because from what I hear your viewpoint is…
하나 더 뛰어나와 물어보겠습니다.
다른 쪽도 있을 수 있어요.
얼마 전 많은 사람들이, 적어도 상반기에는 ‘명품 중공업’이라는 유행 말이 있었죠. 이 대형 모델 일은 우리가 다 봤다고. Transformer 노선을 따라가면 된다고.
아키텍처는 이미 누군가 여기까지 달렸으니 그냥 따라가면 된다.
그리고 안에는 혁신 여지가 많지 않은 것 같다 — 연산력, 엔지니어링 능력, 응용 능력.
게임이 매우 명확해 보인다.
심지어 미래 모델이 더 강해지려면 파라미터만 추가하면 된다고 말한 사람들도 있었어요. 당시 어떤 사람들은 백조 파라미터 같은 말도 해서 우리 모두 깜짝 놀랐죠. 머릿속에서 0을 셀 수 없었고, 몇 승인지 셀 수 없었습니다.
당신은 어떻게 보시나요?
한쪽은 기술이 확정되었고 다음은 명품 중공업이라고 생각합니다.
이 관점을 받아들이시나요? 왜 받아들이지 않으시나요?
보기에는 조직 층으로 가시는 것 같고, 여전히 많은 불확실성이 있다고 생각하시는 것 같습니다. 제가 듣기로는 당신의 관점이…
I think this question is particularly interesting.
This is how I think: the first principles are clear.
There is no fundamental blocker in principle.
The first principle is: as long as you can keep making lossless compression better, you can produce higher degrees of intelligence, even intelligence that surpasses humans.
I feel this first principle itself already has a large amount of evidence proving it.
And now more and more people support and endorse this basic assumption.
This is a basic assumption of the AGI field — if you believe it holds, you can oppose it, you can disagree, that is all fine.
But those who believe it holds are actually based on such a basic judgment.
I myself of course also believe this.
So under this first-principle situation it is actually already determined.
From this point, according to the “name-brand” view you just mentioned, I think it has a certain degree of reasonableness, because the big things are already roughly set — a bit like the three laws of mechanics, they are already almost done, and what remains is to deduce some second-layer things.
Although your first principles are determined, under this big principle there are still some concrete things that need to be figured out how to do.
For example, how to do a truly lossless long context — this problem may not be that simple.
You may do many length extrapolations or some methods, but the effect may not necessarily be good.
Although you can quickly raise that number, the actual effect may be relatively poor.
Or how to do a truly cross-modal model.
Today even OpenAI’s current state may only have taken the first step.
What exactly each subsequent step should be, I feel there is still some uncertainty.
Then setting aside the technical layer, looking at things at the overall agent level — today we are still quite far from the many super-intelligent agents shown in science-fiction movies, useful agents, agents that can establish long-term emotional connections.
And current products are not necessarily developing toward the correct solutions.
You don’t know.
So I feel it is like this: every era will have the greatest people and some next-greatest teams.
The greatest people discover a correct first principle — next-token prediction is like someone discovering electromagnetic induction that can generate electricity; suddenly everything is different.
Your first principle is discovered.
But there will also be a batch of people who are also very great, though they may not reach that level, who solve many technical challenges, many product challenges, many even commercial challenges inside.
So I feel at the second layer there is still a lot of space, a lot of questions, and also a lot of opportunities for both old companies and new companies to capture value and create more new value.
So I feel imagining the future it is certainly a huge space, a blue ocean, but how exactly to play inside still has quite a few positions and challenges.
이 질문이 특히 흥미롭다고 생각합니다.
저는 이렇게 생각합니다. 제1원리는 명확합니다.
원리적으로 fundamental한 blocker는 없습니다.
제1원리는: 무손실 압축을 계속 더 잘하기만 하면 더 높은 정도의 지능, 심지어 인간을 초월하는 지능까지 만들어낼 수 있다는 것입니다.
이 제1원리 자체는 이미 많은 증거로 증명되었다고 느낍니다.
그리고 지금 더 많은 사람들이 이런 기본 가정을 지지하고 옹호합니다.
이것이 AGI 분야의 기본 가정입니다. 성립한다고 믿으면 반대할 수도, 동의하지 않을 수도 있습니다. 다 괜찮습니다.
하지만 성립한다고 믿는 사람들은 사실 그런 기본 판단에 기반합니다.
저 자신은 물론 그렇게 믿습니다.
그래서 이 제1원리 상황에서는 사실 이미 확정된 것입니다.
이 점에서, 방금 말씀하신 ‘명품’ 관점에 따르면 어느 정도 합리성이 있다고 생각합니다. 큰 것들은 이미 대략 정해졌기 때문입니다. 역학의 세 법칙과 비슷해요. 거의 다 됐고, 남은 것은 두 번째 층의 것들을 연역하는 것입니다.
제1원리가 확정되었더라도, 이 큰 원리 아래에는 구체적으로 어떻게 해야 할지 알아내야 할 것들이 있습니다.
예를 들어 진정으로 무손실인 long context를 어떻게 할 것인가 — 이 문제는 그렇게 간단하지 않을 수 있습니다.
길이 외삽이나 여러 방법을 많이 할 수 있지만 효과가 반드시 좋지는 않을 수 있습니다.
숫자를 빠르게 올릴 수 있어도 실제 효과는 비교적 나쁠 수 있죠.
또는 진정으로 크로스모달 모델을 어떻게 할 것인가.
오늘날 OpenAI의 현재 상태도 첫 단계만 밟은 것일 수 있습니다.
이후 각 단계가 구체적으로 어떻게 되어야 하는지는 여전히 불확실성이 있다고 느낍니다.
기술 층을 제쳐두고 전체 에이전트 층에서 보면 — 오늘날 우리는 SF 영화에 나오는 많은 초지능 에이전트, 유용한 에이전트, 장기적인 감정 연결을 맺을 수 있는 에이전트와는 아직 꽤 거리가 있습니다.
그리고 현재 제품이 반드시 올바른 해결책으로 발전하고 있는 것도 아닙니다.
모릅니다.
그래서 저는 이렇게 느낍니다. 모든 시대에는 가장 위대한 사람들과 그다음으로 위대한 팀들이 있습니다.
가장 위대한 사람들은 올바른 제1원리를 발견합니다. next-token prediction은 전자기 유도를 발견해 전기를 만들 수 있게 된 것과 같습니다. 갑자기 모든 게 달라지죠.
제1원리가 발견된 것입니다.
하지만 그 정도에는 미치지 못하더라도 매우 위대한 사람들 한 무리가 있어서 그 안의 많은 기술적 도전, 많은 제품적 도전, 많은 심지어 상업적 도전을 해결합니다.
그래서 두 번째 층에서는 여전히 많은 공간, 많은 의, 그리고 오래된 회사이든 새로운 회사이든 가치를 포착하고 더 많은 새로운 가치를 창조할 기회가 많다고 느낍니다.
그래서 미래를 상상하면 분명히 거대한 공간, 블루오션이지만, 그 안에서 구체적으로 어떻게 플레이할지는 여전히 꽤 많은 포지션과 도전이 있다고 느낍니다.

So every era will have the greatest people and some next-greatest teams.
The greatest people discover a correct first principle — next-token prediction is like someone discovering electromagnetic induction that can generate electricity; suddenly everything is different.
Your first principle is discovered.
But there will also be a batch of people who are also very great, though they may not reach that level, who solve many technical challenges, many product challenges, many even commercial challenges inside.
So I feel at the second layer there is still a lot of space, a lot of questions, and also a lot of opportunities for both old companies and new companies to capture value and create more new value.
So I feel imagining the future it is certainly a huge space, a blue ocean, but how exactly to play inside still has quite a few positions and challenges.
모든 시대에는 가장 위대한 사람들과 그다음으로 위대한 팀들이 있습니다.
가장 위대한 사람들은 올바른 제1원리를 발견합니다. next-token prediction은 전자기 유도를 발견해 전기를 만들 수 있게 된 것과 같습니다. 갑자기 모든 게 달라지죠.
제1원리가 발견된 것입니다.
하지만 그 정도에는 미치지 못하더라도 매우 위대한 사람들 한 무리가 있어서 그 안의 많은 기술적 도전, 많은 제품적 도전, 많은 심지어 상업적 도전을 해결합니다.
그래서 두 번째 층에서는 여전히 많은 공간, 많은 질문, 그리고 오래된 회사이든 새로운 회사이든 가치를 포착하고 더 많은 새로운 가치를 창조할 기회가 많다고 느낍니다.
그래서 미래를 상상하면 분명히 거대한 공간, 블루오션이지만, 그 안에서 구체적으로 어떻게 플레이할지는 여전히 꽤 많은 포지션과 도전이 있다고 느낍니다.

@@@@@@@

전체 인터뷰 내용 생략 없는 상세 요약
인터뷰는 진행자(장 총, 장펑으로 추정)와 문샷 AI(월지암면) 창업자 겸 CEO 양지린의 대화로 시작된다. 진행자는 양지린을 “한눈에 봐도 젊고 잘생겼으며 매우 뛰어난 AI 전문가”라고 소개하며, 오늘 저녁 함께 이야기할 수 있어 영광이라고 인사한다. 이어 문샷 AI 회사가 다소 신비롭다는 평과 함께, 특히 ‘월지암면(Dark Side of the Moon)’이라는 이름이 흥미롭다고 지적한다. 과거 인터넷 시대 창업 회사들이 알리바바·텐센트·바이두처럼 간결한 두 글자 이름을 선호했던 것과 달리, 이번 AI 물결에서는 이름 짓는 방식이 크게 달라졌음을 언급하며, 양지린의 친구 리쯔페이 회사 이름 ‘시퀀스 몽키’ 등을 예로 든다. 그리고 월지암면이라는 이름 뒤에 어떤 의미가 있는지 먼저 밝혀달라고 요청한다.
양지린은 회사 이름이 록 앨범에서 유래했다고 밝힌다. 팀원들 모두 록을 좋아했고 과거 밴드 활동도 했으며, 올해가 핑크 플로이드의 《The Dark Side of the Moon》 발매 50주년이라는 특별한 계기가 있었다고 설명한다. AGI나 대형 모델 자체도 매우 특별한 상징이라고 느낀다. 평소 달을 보면 빛나는 면만 보이지만 뒷면은 보이지 않고 매우 신비롭기 때문에, 미지를 탐험하고 달의 어두운 면이 무엇인지 알고 싶은 강한 충동이 생긴다. 대형 모델도 이와 유사하다. 신비롭고 未知한 것을 탐험하면서 동시에 매우 어렵고 도전적인 일이며, 끊임없이 혁신하고 기존 형태에 도전하며 다음을 상상하는 록의 기저 정신과 결합할 수 있다. 여러 이름을 고민했지만 결국 ‘Moonshot’과 ‘월지암면’의 중영문 조합이 AGI를 하겠다는 결심과 자신들이 어떤 사람들인지를 가장 잘 반영한다고 판단했다.
진행자는 자신도 핑크 플로이드 팬이라며, 앨범이 천문학적 이름을 썼지만 본질적으로 인간 잠재의식에 가까운 내용을 다루고 있다고 느낀다고 말한다. 메인 트랙 ‘Brain Damage’가 환각이 생긴 후의 이야기를 하고 있으며, 50년이 지난 지금 우리는 특히 대형 모델의 환각 문제를 구하려고 한다고 연결 짓는다. 많은 사람들이 양지린을 ‘양 교수’라고 부르지만, 이름을 짓는 과정에서 과학자로서의 면만 있는 것이 아니라 더 깊은 차원과 대형 모델의 경지가 연결되는 느낌이 든다고 평가한다. 밴드 배경이 느껴진다며, 밴드에서 어떤 포지션이었는지 묻는다.
양지린은 드러머였다고 답한다. 드러머는 기본적으로 리듬을 잡고 템포를 결정하는 역할이며, 전체 밴드가 곡을 연주할 때 매우 중요한 프레임워크를 제공한다고 설명한다. 진행자는 이 관점이 대형 모델 발전이나 문샷 같은 일에도 전이될 수 있다며, 리듬이 매우 중요할 수 있다고 공감한다.
대화는 더 구체적인 창업 결심 과정으로 넘어간다. 진행자는 공통 친구가 있어 창업 소식이 나오기 전부터 양지린이 대형 모델 물결에 뛰어들 것을 대략 알고 있었다고 밝히며, 작년과 올해 경계 무렵 많은 사람들이 결심하던 시기로 돌아가 어떻게 결정을 내렸는지, 왜 조직을 만들어서 뛰어들기로 선택했는지 당시 논리를 공유해 달라고 요청한다.
양지린은 이 질문이 특히 흥미롭다고 말한다. 대형 모델에 대해 작년에 GPT 관련 흐름을 듣기 시작했지만, 자신들은 이미 2018~2019년부터 Transformer 기반 언어 모델을 실제로 훈련해 왔다. 초기에는 언어 모델이 여러 시나리오의 성능을 향상시키는 도구가 될 수 있다고 생각했고, 이후 많은 작업에 유용할 수 있다고 여겼으며, 결국 언어 모델이 AI의 유일한 문제가 될 수 있다는 인식에 이르렀다. next-token prediction을 더 잘하면 모든 문제를 해결할 수 있다는 생각으로 AI가 단일 문제로 수렴했다. 지난 몇 년 동안 인식이 크게 바뀌었으며, 가장 중요했던 경험은 2019년부터 구글에서 수천 장의 카드로 Transformer 모델을 훈련하던 과정이었다. 그 과정에서 관찰한 수많은 현상이 이 길이 올바른 길일 수 있다는 증거를 더 많이 제공했다. 이 길을 계속 걸어가며 스케일하고 더 효율적인 스케일 방식을 찾으면 기억·추론·상식·다중 홉 문제 등 이전에 어려웠던 많은 문제를 해결할 수 있다. 2020년부터는 국내에서 여러 사람과 기관을 찾아 협력하며 대형 모델을 훈련했고, 중국에서 가장 먼저 판구(PanGu)·우다오(WuDao) 같은 모델을 훈련한 팀 중 하나였다. 그 과정에서 계속 진짜 기회를 준비하면서 동시에 기술적 도전과 조직적 도전을 모두 보았다. 전통적인 방식과 전통적인 조직 구조로는 대형 모델을 성공시키기 매우 어렵다는 것을 깨달았다. OpenAI의 성공도 본질적으로 조직이 극단적으로 큰 혁신을 했기 때문에 가능했다. 그래서 제로에서 새로운 조직을 만드는 것이 AI로 가는 열쇠이며, 세밀한 기술보다도 더 근본적인 문제라고 믿게 되었다. 조직을 제대로 해야만 AGI 길에서 진정으로 멀리 갈 가능성이 생긴다. 작년에 자본 시장과 인재 시장이 모두 극단적으로 큰 변화를 겪으면서 타이밍이 더 성숙했다고 판단했고, 제로에서 조직을 만들어 이 일을 할 기회가 생겼다고 느꼈다.
진행자는 이 관점이 큰 영감을 준다고 평가한다. 양지린이 이 일에 뛰어든 것은 시장이 특정 유형의 대형 모델을 필요로 해서가 아니라, AGI가 혁신적인 조직에 의해 추진되어야 한다고 믿기 때문이라는 점이 핵심이라는 것이다. “새로운 종류의 조직으로 기술을 발전시킨다”는 관점이며, 시장과 자본 환경이 준비되었을 때 그 기회가 주어져 창업을 선택했다는 해석이다. 이전에 거의 들어본 적 없는 시각이라며, 무엇이 조직 발전을 핵심 문제로 보고 이 기회를 이용해 그런 조직을 구축해야 한다고 결심하게 만들었는지 묻는다. 양지린이 AGI의 깊이 있는 신봉자인 것은 분명하지만(그렇지 않으면 월지암면이라는 이름을 짓지도 않았을 것이다), 그 결심의 촉발점을 알고 싶다는 것이다.
양지린은 더 많은 부분이 실천에서 왔다고 답한다. 올해 이전에도 여러 다른 방식과 조직으로 대형 모델과 이를 중심으로 한 기술·제품을 연구개발하려고 시도했다. 전통 기업 내부, 독립 연구기관(일부 참여 경험 있음), 대학 스타일 모드 등 여러 가능성을 실험했으나, 이 모드들은 조직 차원에서 근본적인 혁신을 하기 어렵다는 것을 발견했다. 구체적인 문제는 오늘 자세히 펼치지 않겠지만, 간단한 예를 든다. 모바일 인터넷 시대에는 수요를 정의하면 확정적으로 생산할 수 있었다. 앱이 갑자기 개발되지 않거나 정의한 수요가 실현되지 않는 경우는 거의 없었다. 사람이 코딩으로 의미를 표현하면 실현되는 확정적 사건이었기 때문이다. 그러나 AGI는 다르다. 오늘 어떤 수요를 어느 정도까지 완성하겠다고 계획하기 어렵고, 하드코딩하거나 규칙으로 표현할 수 없다. 그래서 앞선 계획형 혁신이 아니라 사후적인 방식, 즉 해봐야 알 수 있는 방식, 여러 가능성을 먼저 구분한 뒤 일반적인 방법으로 해결하는 방식이 필요하다. case-by-case로 하나씩 해결하는 것은 전통 모바일 인터넷 제품 개발 속도이며, 이는 AGI가 아니다. AGI는 한 번에 많은 일을 할 수 있는 더 일반적인 기저 기계가 필요하다. 이것이 매우 멘탈적인 차이다. 조직은 일하는 방식과 맞춰야 하므로, 일하는 기저 논리가 변하면 새로운 조직이 필요해진다. 인터넷 시대에는 추천 시스템 등에 매우 능한 훌륭한 조직들이 등장했지만, 새 시대에는 AGI에 매우 능한 조직들이 등장할 가능성이 높다.
진행자는 달의 어두운 면을 볼 수 없는 것처럼, 확정된 데이터가 없으면 확정된 것에 기반해 R&D를 하기 어렵다는 점과 연결하며, 더 일반적이고 불확실한 문제에 맞출 수 있는 조직이 필요하다는 의미를 이해한다고 말한다. 이어 OpenAI가 양지린 눈에 AGI를 향한 조직 형태의 비교적 좋은 모델인지, 잘한 점과 완전히 최선이 아닐 수 있는 점은 무엇인지 묻는다.
양지린은 결과적으로 OpenAI가 매우 큰 돌파를 이뤘으며, 이 회사가 없었다면 인류의 진보가 달라졌을 수도 있다고 평가한다. 결과적으로 많은 올바른 일을 했다. 더 깊이 들어가면(직접 조직 워크플로우에 참여하지는 않았지만 관련 정보가 있다), 훌륭한 조직은 매우 높은 인재 밀도와 극도로 강한 공유된 비전이 필요하다. 모든 사람이 공통의 비전을 가지고 하나의 목표를 중심으로 효율적으로 집중할 수 있어야 한다. 이 점들은 OpenAI가 매우 잘한 부분이다. 그러나 가장 핵심적이며 외부에서 쉽게 보기 어렵고, 자신들이 지금 가장 잘하고 싶은 점은 “이런 전제들을 갖춘 후에 어떻게 체계적인 일하는 방식을 찾느냐”이다. 이 점이 모든 기술을 초월하며 모든 기술의 전제 조건이어야 한다. ‘체계적’이라는 것은 복제되고 확대될 수 있다는 의미인가라는 질문에, 가능하지만 여기서의 복제는 다른 일에 적용할 수 있다는 뜻이라고 답한다. 다른 곳으로 복제하기는 어려울 수 있으나, 한 회사가 이 시스템을 반복적으로 사용해 다른 일들을 할 수 있어야 한다. 예를 들어 오늘은 long-context 도전을, 내일은 자율 AI 능력을, 나중에는 멀티모달을, 오늘은 생산성 제품을, 내일은 생성형 제품을 만드는 데 사용할 수 있는 시스템이어야 한다. 본질적으로 침전되는 것이 핵심 자산이 되며, 모든 AGI 회사가 가장 많은 시간을 들여 연마해야 할 것이다.
진행자는 이전에는 대형 모델 기술 자체에 대해 이야기했지만, 기술이 계속 더 잘 발전할 수 있는지는 근본적으로 조직과도 관련이 있을 수 있다는 관점을 높이 평가한다. 불확실한 혁신에 직면했을 때 자주 등장하는 bottom-up과 top-down이라는 양자택일 프레임을 언급하며, 양지린이 구상하는 혁신 시스템이 기존 정의로 충분한지, 아니면 새로운 정의가 필요한지 묻는다.
양지린은 큰 프레임워크는 여전히 적용 가능하다고 본다. 특히 대형 모델에 있어서 top-down 프레임워크가 매우 중요하다. 달 착륙 공학과 같다. 달에 가는 것은 한 동작이나 한 병사가 아니라, 긴 시간과 많은 사람들이 극도로 복잡하고 서로 결합된 일에 종사하는 거대한 시스템이다. 그래서 최상위 설정이 분명히 필요하다. top-down은 리더십의 비전, 즉 무엇이 해야 할 올바른 일이고 지금은 하지 말아야 할 일인지를 판단하는 능력에 관한 것이며, 이런 설계는 필수적이다. 그 프레임워크 아래에서 조직이 일을 잘할 수 있는지가 바로 시스템이다. 그 안에는 많은 작은 유닛이 각각 다른 일을 하지만, 좋은 시스템이 있으면 각 유닛이 효율적으로 무언가를 생산할 수 있어야 한다. 더 공장 같은, AI 공장 같은 모습이다. 각 유닛이 다른 것을 생산하지만 top-down 프레임워크가 그것들을 통합한다. 예를 들어 세 개의 제품 라인을 동시에 탐색할 때 모든 라인이 시스템을 효과적으로 활용할 수 있어야 하며, 시스템이 빠르게 데이터를 생산해 새로운 시나리오에서 PMF 가능성을 빠르게 알 수 있게 해야 한다. 모든 것은 top-down 프레임워크 아래에서 일어난다. 먼저 어떤 방향과 제품 라인을 탐색할지 최상위 설계가 있어야 하고, 훌륭한 시스템이 있으면 시도를 빠르게 할 수 있다. 이 시도 과정은 AI-native 과정이다. 전통 조직으로는 불가능하다. AI-native의 특징은 본질적으로 데이터로 새로운 제품을 정의하는 것이며, 이는 역사상 존재했던 어떤 제품과도 다르다. 그래서 새로운 시스템이 필요하다.
진행자는 라이브 방송실 반응을 전하며, 기술은 경제·제도·문화보다 우선순위가 낮다는 댓글이 나왔고, 문화는 바꾸기 어렵기 때문에 조직·제도 차원의 추진이 중요하다는 공명이 있다고 말한다. 자신이 20년 넘게 기술 변혁을 지켜본 경험상, 어떤 기술 돌파는 여러 기술의 교차가 올바른 순간에 올바른 레시피가 되어 폭발하는 경우가 많았다고 회고한다. 양지린이 큰 프레임워크 아래 조직 안에서 다른 방향에서 적합한 혁신이 출현할 수 있게 시뮬레이션하고 있으며, 결국 그것들이 함께 AGI로 이어질 수 있다는 느낌이 든다고 해석한다. 오늘날 AGI로 가는 길을 정밀하게 정의할 수 없고 많은 불확실성이 남아 있기 때문에 출현이 필요하며, 조직이 출현을 지원하면서도 프레임워크를 유지해야 한다는 이해가 맞는지 확인한다.
양지린은 매우 좋은 이해라고 동의한다. Transformer가 나온 과정을 예로 든다. 본질적으로 구글이 그 사람들에게 출현을 위한 환경을 제공한 것이다. Transformer 제안 당시 이미 어텐션 메커니즘, residual connection, layer normalization, SGD, learning-rate schedule 등이 준비되어 있었고, 구글이 자유롭게 조합할 수 있는 환경을 제공해 그 조합 과정에서 아키텍처가 출현했다. 그러나 다른 환경은 다른 것을 출현시킨다. 구글의 환경은 과학적 결과(Transformer 아키텍처)를 출현시킬 수 있었지만, GPT처럼 하나를 극한까지 밀면서 제품 쪽에서 수요를 정밀하게 잡는 시대를 가로지르는 시스템 엔지니어링 작품을 출현시킬 수는 없었다. 조직이 맞지 않았기 때문이다. OpenAI는 역사상 가장 위대한 과학적 발견을 출현시키지 않았고 새로운 것을 발명하지 않았지만, 산업화 운영을 출현시켰다. 그들이 결합한 것은 세상의 기존 재고—Transformer 아키텍처, 10^25 규모 연산을 지원하는 컴퓨팅 센터, 인터넷이 20년 동안 축적한 데이터—였다. 이것이 인터넷의 가장 큰 가치일 수 있다. 이 세 요인을 보고 결합될 수 있는 환경을 제공해 AGI의 이정표를 출현시켰다. 다른 조직은 다른 것을 출현시킬 수 있게 하며, 무엇을 출현시키고 싶은지에 따라 조직을 그 방향으로 조정해야 한다. 조직의 유전자를 먼저 조정하면 새로운 종이 나타날 수 있다는 것이다.
진행자는 대형 모델이 출현 과정에서 chain-of-thought를 출현시킨 것처럼, 미래 조직도 출현 문화나 아키텍처가 필요할 수 있다고 연결한다. 라이브 댓글 중 “과거에는 수요 정의와 착지가 달 단위였지만 AI의 scalable 특성 때문에 새로운 시나리오 능력이 일 단위, 심지어 하루에 많은 응용이 착지할 수 있다”는 내용이 양지린의 의도를 잘 이해한 것 같다고 전한다. 미래에는 출현이 상시가 될 수 있으며, 다음 시대의 조직은 출현을 더 필요로 할 수 있다는 관점을 높이 평가한다.
마지막으로 진행자는 또 다른 측면을 묻는다. 상반기 유행했던 ‘명품 중공업’ 관점—대형 모델은 Transformer 노선을 따라가면 되고, 아키텍처는 이미 여기까지 달렸으니 혁신 여지는 많지 않으며, 연산력·엔지니어링·응용 능력의 게임이 명확해 보이고, 더 강해지려면 파라미터만 추가하면 된다는 식의 주장—을 어떻게 보는지, 이 관점을 받아들이는지, 왜 받아들이지 않는지, 조직 층으로 가면서 여전히 많은 불확실성이 있다고 생각하는 이유는 무엇인지 질문한다.
양지린은 이 질문이 특히 흥미롭다고 말한다. 제1원리는 명확하다. 원리적으로 fundamental한 blocker는 없다. 제1원리는 “무손실 압축을 계속 더 잘하기만 하면 더 높은 정도의 지능, 심지어 인간을 초월하는 지능까지 만들어낼 수 있다”는 것이다. 이 원리 자체는 이미 많은 증거로 증명되었고, 더 많은 사람들이 지지하고 있다. 이것이 AGI 분야의 기본 가정이며, 성립한다고 믿으면 반대할 수도 있다. 자신은 성립한다고 믿는다. 따라서 제1원리 차원에서는 이미 확정된 것이다. 이 점에서 ‘명품 중공업’ 관점에는 어느 정도 합리성이 있다. 큰 것들은 역학의 세 법칙처럼 이미 대략 정해졌고, 남은 것은 두 번째 층의 것들을 연역하는 일이다. 그러나 제1원리가 확정되었더라도 그 아래에는 구체적으로 어떻게 해야 할지 알아내야 할 것들이 많다. 진정으로 무손실인 long context를 어떻게 구현할 것인가(길이 외삽 등으로 숫자는 올릴 수 있어도 실제 효과가 나쁠 수 있다), 진정으로 크로스모달 모델을 어떻게 만들 것인가 등이 그 예다. OpenAI의 현재 상태도 첫 단계만 밟은 것일 수 있으며, 이후 각 단계가 어떻게 되어야 하는지는 여전히 불확실하다. 기술 층을 넘어 전체 에이전트 층으로 보면, SF 영화에 나오는 초지능 에이전트·유용한 에이전트·장기적 감정 연결이 가능한 에이전트와는 아직 꽤 거리가 있으며, 현재 제품이 반드시 올바른 해결책으로 발전하고 있는지도 알 수 없다. 모든 시대에는 가장 위대한 사람들(올바른 제1원리를 발견하는 사람들)과 그다음으로 위대한 팀들(기술적·제품적·상업적 도전을 해결하는 사람들)이 있다. next-token prediction은 전자기 유도를 발견한 것과 같다. 제1원리가 발견된 것이다. 그러나 두 번째 층에서는 여전히 많은 공간과 질문, 그리고 오래된 회사와 새로운 회사 모두가 가치를 포착하고 새로운 가치를 창조할 기회가 많다. 미래는 분명히 거대한 공간, 블루오션이지만, 그 안에서 구체적으로 어떻게 플레이할지는 여전히 꽤 많은 포지션과 도전이 남아 있다.


이 인터뷰는 극객공원(Geek Park) 장펑(张鹏)과 양지린의 대담으로, 팟캐스트 「AI局内人」 Vol.13 「张鹏对谈月之暗面杨植麟:大模型创业需要新的组织范式」