https://youtu.be/91fmhAnECVc?si=mMl6YFoJr13mwD3F
Kimi Founder Yang Zhilin: K2, Agentic LLMs, Brains in Vats, and the Beginning of Infinity
This is an interview with Kimi founder Yang Zhilin, recorded shortl...
www.youtube.com
At the end of your first year of entrepreneurship, our 2024 interview was titled “Marching Toward an Endless, Unknown Snow Mountain.”
창업 첫해 말, 2024년 우리 인터뷰 제목은 《끝없이 이어지는 미지의 설산을 향해 전진》이었다.
Another year has passed.
또 한 해가 지났다.
Standing here now, in July 2025, how do you feel?
지금 2025년 7월, 이 자리에 서서 당신의 가장 최근 느낌은 어떤가?
The phrase you just mentioned… it feels like ages ago.
당신이 방금 말한 그 표현… 꽤 오래전 일처럼 느껴진다.
One day in AI is a year in the human world.
AI에서의 하루는 인간 세상의 일 년이다.
I don’t even know how many human days make up one AI year.
AI의 일 년이 인간 세상의 며칠에 해당하는지도 모르겠다.
Many things have indeed changed, but the “snow mountain feeling” you describe is pretty much the same.
많은 것들이 실제로 변했지만, 당신이 말한 ‘설산의 느낌’은 거의 비슷하다.
Toward the summit, we have walked another stretch of the way.
정상으로 향하는 길에서 우리는 또 한 구간을 걸어왔다.
Where have you gotten to now?
지금은 어디까지 와 있는가?
From where we stand today, the model’s progress has been substantial — two years ago it couldn’t even write a coherent article; now it can not only write very good articles, but also work continuously for hours to help you complete a highly complex coding task.
오늘날 우리가 서 있는 지점에서 보면 모델의 진전은 상당하다 — 2년 전만 해도 일관된 글조차 제대로 쓰지 못했는데, 지금은 아주 좋은 글을 쓸 수 있을 뿐만 아니라 몇 시간 동안 연속으로 작업하며 매우 복잡한 코딩 작업을 도와줄 수 있다.
Two years ago it was hard to imagine.
2년 전에는 상상하기 어려웠다.
In the process of climbing the snow mountain, we unlocked some new scenarios and roughly understood what the path in the middle looks like; but at the same time, while going upward, you still observe roughly the same scenery — there will still be many unknown technical problems to solve next.
설산을 오르는 과정에서 새로운 시나리오들을 열어냈고 중간 길이 어떤 모습인지 대략 알게 됐지만, 동시에 위로 올라가는 과정에서 여전히 비슷한 풍경을 관찰한다 — 앞으로 여전히 해결해야 할 많은 미지의 기술적 문제들이 있다.
On a mountain peak surrounded by heavy snow on all sides, are you clearer now, or more confused?
사방이 큰 눈으로 둘러싸인 산봉우리에서, 지금은 더 명확해졌는가, 아니면 더 혼란스러워졌는가?
There are definitely many things that have become clearer.
분명히 더 명확해진 것들이 많다.
Two years ago, how various reinforcement learning paradigms should be done, how to give the model stronger reasoning ability or stronger agentic ability, was not that clear.
2년 전에는 다양한 강화학습 패러다임을 어떻게 해야 하는지, 모델에 더 강한 추론 능력이나 더 강한 에이전트적 능력을 어떻게 부여해야 하는지가 그렇게 명확하지 않았다.
At that time we focused more on how to do pre-training of the model better, and how to use RLHF to improve the conversation experience.
당시에는 모델의 사전훈련을 어떻게 더 잘 할지, RLHF를 사용해 대화 경험을 어떻게 향상시킬지에 더 집중했다.
But now some questions have gotten answers, and at the same time those answers unfold and bring new other questions.
하지만 지금은 어떤 문제들이 답을 얻었고, 동시에 그 답들이 펼쳐지면서 새로운 다른 문제들을 가져온다.
Although we can now do reinforcement learning, it ultimately still relies on a good evaluation or verification.
지금은 강화학습을 할 수 있지만, 결국 좋은 평가나 검증에 여전히 의존한다.
If you let the model do a math problem, or do some programming tasks with test cases, it may do relatively well.
모델에게 수학 문제를 풀게 하거나, 테스트 케이스가 있는 프로그래밍 작업을 하게 하면 비교적 잘할 수 있다.
But if you let it do a more complex end-to-end task, sometimes it is hard to find an evaluation or measurement method.
하지만 더 복잡한 엔드투엔드 작업을 하게 하면, 때로는 평가나 측정 방법을 찾기 어렵다.
So this system will produce new problems.
그래서 이 시스템은 새로운 문제들을 만들어낸다.
This is a bit like a book I’ve been reading recently, which I’ve read several times, called The Beginning of Infinity.
이건 내가 최근에 여러 번 읽은 책, 《The Beginning of Infinity》(무한의 시작)과 좀 비슷하다.
He says there are two sentences that can be carved in stone: one is “Problems are inevitable,” and the other is “Problems are solvable.”
그는 돌에 새길 수 있는 두 문장이 있다고 말한다. 하나는 “문제는 불가피하다”, 다른 하나는 “문제는 해결될 수 있다”.
You can think that before the Enlightenment, this society was static.
계몽운동 이전의 이 사회는 정적이었다고 생각할 수 있다.
People did not pursue innovation; you would use a lot of mysticism to explain the phenomena you saw, but these explanations were not good explanations.
사람들은 혁신을 추구하지 않았고, 본 현상을 설명하기 위해 많은 신비주의를 사용했지만, 그런 설명들은 좋은 설명이 아니었다.
For example, when you saw thunder in the sky, you would think the Thunder God was angry; when you saw snow in winter, you would say some god was in a bad mood — the entire social structure was static, and only a very small number of people were truly doing scientific research or knowledge creation.
예를 들어 하늘에서 천둥을 보면 뇌신이 화가 났다고 생각하고, 겨울에 눈이 내리면 어떤 신이 기분이 나쁘다고 말했다 — 전체 사회 구조가 정적이었고, 극소수만이 진짜로 과학 연구나 지식 창조를 하고 있었다.
But after the Enlightenment, society became dynamic, and new knowledge was continuously created.
하지만 계몽운동 이후 사회는 동적으로 변했고, 새로운 지식이 계속 창조되었다.
Whenever you solve a problem, it brings new problems.
문제를 하나 해결할 때마다 새로운 문제들이 생긴다.
Problems keep coming because your knowledge boundary is expanding, so you encounter new problems.
문제가 끊임없이 생기는 이유는 지식의 경계가 확장되기 때문이고, 그래서 새로운 문제들을 만나게 된다.
AI research and development right now happens to be in exactly this state.
지금 AI 연구개발도 정확히 이런 상태에 있다.
You solved some problems of reinforcement learning, then next you encounter problems of evaluation, measurement, and verification, and again need to find new answers.
강화학습의 일부 문제를 해결하면, 다음으로 평가·측정·검증 문제들을 만나게 되고, 다시 새로운 답을 찾아야 한다.
But that is exactly what makes it interesting — you always have new problems to solve, and every time you solve one, technology can climb another few hundred meters higher.
하지만 바로 그게 흥미로운 점이다 — 항상 해결할 새로운 문제들이 있고, 하나를 해결할 때마다 기술이 다시 수백 미터를 더 올라갈 수 있다.
Maybe one day we’ll discover that this snow mountain has no summit — I don’t know — I hope it never does.
언젠가 이 설산에는 정상이 없다는 것을 발견할지도 모른다 — 모르겠다 — 나는 그것이 영원히 없기를 바란다.
That is what The Beginning of Infinity means: it is an infinite mountain.
그것이 《The Beginning of Infinity》의 의미다. 그것은 무한한 산이다.
Looking back, what were the most important things in global foundation models in your mind over the past year? Which of them were paradigm-level changes in artificial intelligence?
되돌아보면, 지난 1년 동안 글로벌 기초 모델에서 당신의 머릿속에 가장 중요했던 몇 가지는 무엇인가? 그중 인공지능의 패러다임 수준의 변화는 어떤 것들이었나?
One is the long-thinking reasoning model, represented by o1 as the first of its kind.
하나는 긴 사고 추론 모델로, o1이 그 첫 번째 대표작이다.
Essentially, it lets the model make many attempts and reflections in the process; reflection is the key among them.
본질적으로 모델이 과정에서 많은 시도와 반성을 하게 한다. 그중 반성이 핵심이다.
Reflection is two abilities: one is proposing new conjectures, the other is verifying conjectures.
반성은 두 가지 능력이다. 하나는 새로운 가설을 제안하는 것, 다른 하나는 가설을 검증하는 것이다.
You can understand it this way: in the process of solving a problem, the model continuously proposes new conjectures, and this conjecture gets self-verified.
이렇게 이해할 수 있다. 문제를 푸는 과정에서 모델이 계속 새로운 가설을 제안하고, 그 가설은 자기 검증을 받는다.
For example, after it proposes this conjecture, to judge whether it is right or wrong, it needs a certain verification ability.
예를 들어 이 가설을 제안한 후 맞는지 틀린지 판단하려면 일정한 검증 능력이 필요하다.
Although you do not explicitly train a verification model, it implicitly performs verification during the reasoning process.
명시적으로 검증 모델을 훈련시키지는 않지만, 추론 과정에서 암묵적으로 검증을 수행한다.
It makes multiple conjectures and verifications on this problem, and finally gets an answer.
이 문제에 대해 여러 번의 가설과 검증을 하고 최종적으로 답을 얻는다.
This greatly improves the model’s capability.
이것이 모델의 능력을 크게 향상시킨다.
Originally you could only do it once and directly give an answer, which might be right or wrong.
원래는 한 번만 하고 바로 답을 냈는데, 그 답이 맞을 수도 틀릴 수도 있었다.
This greatly improves the model’s capability. You originally could only do it once and directly give an answer, which might be right or wrong; you didn’t have this process. But now you can continuously propose conjectures to verify, which is equivalent to trying several times.
이것이 모델의 능력을 크게 향상시킨다. 원래는 한 번만 하고 바로 답을 냈는데, 그 답이 맞을 수도 틀릴 수도 있었다. 이런 과정이 없었다. 하지만 지금은 계속 가설을 제안하고 검증할 수 있어서, 여러 번 시도한 것과 같다.
You can turn Pass@k into Pass@1. That’s the essential idea.
Pass@k를 Pass@1로 바꿀 수 있다. 본질적으로 그런 이치다.
This is very similar to the process of humans doing scientific research or solving problems: continuously proposing new conjectures and then verifying them.
이건 사람이 과학 연구를 하거나 문제를 푸는 과정과 매우 비슷하다. 계속 새로운 가설을 제안하고 검증하는 것이다.
It is a free exploration process, not a linear process.
자유로운 탐색 과정이지, 선형적인 과정이 아니다.
Its effective working method is still quite linear in many cases. If we don’t consider parallel sampling and assume serial sampling, it is a linear process.
효과적인 작업 방식은 많은 경우 여전히 비교적 선형적이다. 병렬 샘플링을 고려하지 않고 직렬 샘플링을 가정하면 선형 과정이다.
Each time you propose a new conjecture, this conjecture may be based on previous conjectures, or even on conjectures you have already negated, and then propose a new one. It approaches a more linear process.
매번 새로운 가설을 제안할 때, 이 가설은 이전 가설에 기반할 수도 있고, 이미 부정한 가설에 기반할 수도 있으며, 그다음 새로운 가설을 제안한다. 더 선형적인 과정에 가까워진다.
Now you can also combine the linear process with parallel strategies, for example sampling many at the same time, combining both parallel and serial methods. Recently there have also been some papers saying that the upper limit of serial sampling is higher, which is related to our experimental conclusions.
지금은 선형 과정과 병렬 전략을 결합할 수도 있다. 예를 들어 동시에 여러 개를 샘플링해서 병렬과 직렬 두 방식을 결합한다. 최근 일부 논문에서도 직렬 샘플링의 상한이 더 높다고 하는데, 이는 우리 실험 결론과 관련이 있다.
What was described above is one paradigm, but it is still a “brain in a vat”. It does not need to interact with the outside world.
위에서 말한 것은 하나의 패러다임이지만, 여전히 “통 속의 뇌”다. 외부 세계와 상호작용할 필요가 없다.
Brain in a vat?
통 속의 뇌?
Imagine a fish tank. You put a brain inside it, with no connection to the outside world. It just thinks inside its own brain, keeps thinking, and can solve a problem without any interaction with the outside.
어항을 상상해 보라. 그 안에 뇌를 넣고 외부 세계와 아무런 연결이 없다. 그냥 자기 뇌 안에서만 생각하고, 계속 생각해서 외부와 아무런 상호작용 없이 문제를 풀 수 있다.
But there is another very important paradigm: the multi-turn Agent reinforcement learning paradigm, or Agentic models trained through reinforcement learning techniques. Its characteristic is that it interacts a lot with the outside world.
하지만 또 하나의 매우 중요한 패러다임이 있다. 바로 다회차 Agent 강화학습 패러다임, 또는 강화학습 기술로 훈련된 Agentic 모델이다. 특징은 외부 세계와 많은 상호작용을 한다는 것이다.
For example, while thinking I also perform some operations, possibly many rounds of operations: calling a search engine one moment, using a browser the next, writing a few lines of code after that, solving a problem through multiple turns.
예를 들어 생각하면서 동시에 어떤 조작을 한다. 많은 라운드의 조작을 할 수 있다. 한순간 검색을 호출하고, 다음엔 브라우저를 사용하고, 그다음엔 코드 몇 줄을 쓰면서, 다회차로 문제를 해결한다.
It is no longer a “brain in a vat”; it interacts with the outside world — my next action is related to the feedback obtained from the interaction and the update of the new state given by the outside world.
더 이상 “통 속의 뇌”가 아니다. 외부 세계와 상호작용한다 — 내 다음 행동은 상호작용에서 얻은 피드백과 외부 세계가 준 새로운 상태 업데이트와 관련이 있다.
But both of these point to the same thing: test-time scaling. Meaning that at test time, or during inference, better scaling can be achieved.
하지만 이 두 가지는 모두 같은 것을 가리킨다. 바로 test-time scaling이다. 테스트 시, 또는 추론 시에 더 나은 규모 확장을 할 수 있다는 뜻이다.
For example, previously when doing Chat, it was more about outputting a result in a single turn: I ask you to write an article, you write an article; I ask you to polish it again, you output a few hundred more tokens, and the number of tokens is small. But whether it is reinforcement learning based on long thinking or Agent reinforcement learning, essentially both are ways of scaling tokens at prediction time.
예를 들어 이전의 Chat은 대부분 단회차로 결과를 출력했다. 글을 써 달라고 하면 글을 쓰고, 다시 다듬어 달라고 하면 또 수백 개의 토큰을 출력하는데, 토큰 수가 적다. 하지만 긴 사고 기반 강화학습이든 Agent 강화학습이든, 본질적으로 둘 다 예측 시 토큰을 규모 확장하는 방식이다.
Whether you increase the number of turns or have more thinking tokens in each turn, both are methods of scaling tokens, allowing you to complete more complex tasks.
회차 수를 더 늘리든, 각 회차에서 더 많은 사고 토큰을 가지든, 모두 토큰을 규모 확장하는 방법으로, 더 복잡한 작업을 완성할 수 있게 한다.
This is also accompanied by longer completion times. Now you can spend several hours doing a complex thing without human intervention in the process. For example, cloning a code repository, translating it into another new language, debugging, testing, fixing all the bugs, and making it run normally. Such work can be completed end-to-end, thanks to the scaling of test-time computation.
이는 완료 시간이 더 길어지는 것과도 함께한다. 지금은 복잡한 일을 몇 시간 동안 할 수 있고, 과정에서 사람이 개입할 필요가 없다. 예를 들어 코드 저장소를 클론하고, 다른 새로운 언어로 번역하고, 디버깅하고, 테스트하고, 모든 버그를 고쳐서 정상적으로 돌아가게 하는 것. 이런 작업을 엔드투엔드로 완성할 수 있는 것은 테스트 시 계산의 규모 확장 덕분이다.
There is also a very interesting trend: now more model companies are making “first-party Agent products.”
또 하나의 매우 흥미로운 추세가 있다. 지금 더 많은 모델 회사들이 “1차 Agent 제품”을 만들고 있다는 것이다.
“First-party Agent products” means: the model company itself goes into the product field, controlling the context environment, tool interfaces, prompt structure, etc., that is, becoming the “user” itself, rather than only providing model APIs to others.
“1차 Agent 제품”이란 모델 회사가 스스로 제품에 뛰어들어 컨텍스트 환경, 도구 인터페이스, 프롬프트 구조 등을 통제하는 것, 즉 스스로 “사용 측”이 되는 것이지, 남에게 모델 API만 제공하는 것이 아니다.
Initially, half a year ago or last year, many products were based on foundation models. You build some scaffolding on top, or design some tools to better let the model use them, thereby building a product.
처음에는 반년 전이나 작년에 많은 제품이 기초 모델 위에 기반했다. 그 위에 어떤 스캐폴딩을 쌓거나, 모델이 더 잘 사용할 수 있도록 도구를 설계해서 제품을 구축했다.
Enjoying the overflow capability of the model.
모델의 넘쳐흐르는 능력을 즐기는 것이다.
Yes, what it essentially does is reverse-engineer the model’s training process. Because the model training process is also through various means — you can think that Anthropic, in its internal environment, tools, and scaffolding, may have trained such a model, but it did not directly open it to you.
맞다. 본질적으로 하는 일은 모델의 훈련 과정을 역공학하는 것이다. 모델 훈련 과정도 다양한 수단을 통해서인데 — Anthropic이 내부 환경, 도구, 스캐폴딩으로 그런 모델을 훈련했을 수 있지만, 당신에게 직접 개방하지는 않았다.
Through reverse engineering, you get closer to fitting its distribution — what tools work better? What kind of System Prompt works better? What kind of Context Engineering works better? This is a reverse process.
역공학을 통해 그 분포에 더 가깝게 맞춘다 — 어떤 도구가 효과가 좋은가? 어떤 System Prompt가 효과가 좋은가? 어떤 Context Engineering이 효과가 좋은가? 이것은 역방향 과정이다.
But you will find that if the model company does “first-party products,” the logic is completely different.
하지만 모델 회사가 “1차 제품”을 하면 논리가 완전히 다르다.
You no longer need this reverse process; it is more of a forward approach. I first design these tools well, design the Context Engineering methods well, and then train the model in this environment, so the model naturally performs better in your environment.
더 이상 이 역방향 과정이 필요 없다. 더 순방향 접근이다. 먼저 이 도구들을 잘 설계하고, Context Engineering 방법을 잘 설계한 다음, 이 환경에서 모델을 훈련시킨다. 그래서 모델이 자연스럽게 당신의 환경에서 더 잘 작동한다.
These are two different approaches, but the second approach may have a higher ceiling.
이것은 두 가지 다른 사고방식인데, 두 번째 사고방식의 상한이 더 높을 수 있다.
You can better integrate tools and the model. If the model has some places it cannot solve well, you can adjust the tool design to make it better, and at the same time train end-to-end. This is also a relatively large variable in the development method.
도구와 모델을 더 잘 통합할 수 있다. 모델이 잘 해결하지 못하는 부분이 있으면 도구 설계를 조정해서 더 좋게 만들고, 동시에 엔드투엔드로 훈련할 수 있다. 이것 또한 개발 방식에서 비교적 큰 변수다.
To make it easier for everyone to understand. The first scaffolding you mentioned is like the development method of Manus; the second is the end-to-end model training method like yours.
모두가 이해하기 쉽게. 당신이 말한 첫 번째 스캐폴딩은 Manus 같은 개발 방식이고, 두 번째는 여러분처럼 엔드투엔드로 모델을 훈련하는 방식이다.
Yes. Of course, we are not investing particularly much in “first-party products” yet; the main line is still the model. But Claude Code or ChatGPT Agent are “first-party products,” and this should also be a big trend.
맞다. 물론 우리는 아직 “1차 제품”에 특별히 많이 투자하고 있지는 않다. 주선은 여전히 모델이다. 하지만 Claude Code나 ChatGPT Agent는 “1차 제품”이고, 이것 또한 큰 추세가 될 것이다.
Later it will depend on how “first-party” and “third-party” products cooperate and what state they will be in the ecosystem.
나중에는 “1차”와 “3차” 제품이 어떻게 협력하고 생태계에서 어떤 상태가 될지가 관건이다.
Speaking of the main line, OpenAI has set levels from L1 to L5. I have always been curious about the internal logic behind it.
주선 얘기가 나와서, OpenAI가 L1부터 L5까지의 등급을 설정했다. 그 뒤의 내재 논리가 늘 궁금했다.
L1 is Chatbot, L2 is Reasoner, L3 is Agent, L4 is Innovator, L5 is Organizer.
L1은 Chatbot, L2는 Reasoner, L3는 Agent, L4는 Innovator, L5는 Organizer다.
Why is it that only after having Chatbot and Reasoner do we have Agent? Why next are Innovator and Organizer? How does this order progress in terms of capability structure?
왜 Chatbot과 Reasoner가 있은 후에야 Agent가 나오는가? 왜 그다음은 Innovator와 Organizer인가? 이 순서는 능력 구조상 어떻게 점진하는가?
It is a step-by-step dependency of capabilities. The upper limit of Agent (L3) depends on having strong Reasoning (L2) ability, but it is not necessary to have Reasoning first.
능력의 한 단계씩 의존 관계다. Agent(L3)의 상한은 강한 Reasoning(L2) 능력에 달려 있지만, Reasoning이 반드시 먼저 있어야 하는 것은 아니다.
Suppose we slightly change the order of technological development: you first develop Agent ability, then do narrow Reasoning, that is, long CoT Reasoning, it might also be valid.
기술 발전 순서를 조금 바꾼다고 가정하자. 먼저 Agent 능력을 만들고, 그다음 좁은 의미의 Reasoning, 즉 long CoT Reasoning을 한다면, 그것도 성립할 수 있다.
You can consider Claude’s path as betting on this point: it did not do particularly much on Reasoning, but did very well on Agent. Behind this are bets on different technical paths. But ultimately you cannot get around it: if you want to climb a few more steps toward the summit, you need both abilities; it is only a matter of time.
Claude의 경로를 이것에 베팅한 것으로 볼 수 있다. Reasoning에서는 특별히 많이 하지 않았지만 Agent에서는 매우 잘했다. 그 뒤에는 다른 기술 경로에 대한 베팅이 있다. 하지만 결국 피할 수 없다. 정상으로 몇 걸음 더 오르고 싶다면 이 두 능력이 모두 필요하다. 단지 시간 문제일 뿐이다.
So they are not necessarily interdependent. It is not that you must first do Reasoning and then do Agent. But to achieve the best Agent, you must also make Reasoning the best. And with Agent ability, you can do some subsequent things.
그래서 그것들은 반드시 상호 의존적이지는 않다. Reasoning을 먼저 하고 나서 Agent를 해야 하는 것은 아니다. 하지만 최고의 Agent를 하려면 Reasoning도 최고로 만들어야 한다. 그리고 Agent 능력이 있으면 이후의 어떤 일들을 할 수 있다.
Why is the next stage (L4) Innovation? The most critical point here is: when can the model participate in the development of the model itself? Only when the model participates in the development process can the true Innovator stage be unlocked.
왜 다음 단계(L4)가 Innovation인가? 여기서 가장 핵심적인 점은: 모델이 언제 모델 자체의 개발에 참여할 수 있는가? 모델이 개발 과정에 참여해야만 진정한 Innovator 단계가 열릴 수 있다.
We hope K2 can participate in the development of K3. Without Agentic ability, it is hard to do this. But when you have Agentic ability, it can propose some new ideas, conduct corresponding experiments, analyze experimental results, draw conclusions, iterate the next version of ideas, or optimize the performance of some Infra. All of these rely on strong Agentic ability to do.
우리는 K2가 K3 개발에 참여하기를 바란다. Agentic 능력이 없으면 이런 일을 하기 어렵다. 하지만 Agentic 능력이 있으면 새로운 아이디어를 제안하고, 해당 실험을 하고, 실험 결과를 분석하고, 결론을 도출하고, 다음 버전 아이디어를 반복하거나, 어떤 Infra 성능을 최적화할 수 있다. 이 모든 것은 강한 Agentic 능력에 의존한다.
Innovation (L4) and Organization (L5) are also not necessarily completely linear relationships; some are parallel. We have already seen such trends —
Innovation(L4)과 Organization(L5)도 반드시 완전히 선형적인 관계는 아니다. 일부는 병렬적이다. 우리는 이미 그런 추세를 보았다 —
For example, when you have an Agent, you can expand it into a Multi-Agent System. You can fork many different Agents from one Agent and let them do different things. Some are serial, some are parallel, then merge, then split into different tasks. Some write tests, some write documentation, some design software frameworks, each with their own division of labor.
예를 들어 Agent가 하나 있으면 Multi-Agent System으로 확장할 수 있다. 하나의 Agent에서 여러 다른 Agent를 포크해서 각각 다른 일을 하게 할 수 있다. 일부는 직렬, 일부는 병렬로 한 다음 병합하고, 다시 다른 작업으로 나눈다. 어떤 것은 테스트를 쓰고, 어떤 것은 문서를 쓰고, 어떤 것은 소프트웨어 프레임워크를 설계하는 식으로 각자의 분업이 있다.
So it is not necessarily a linear relationship; the two (L4 and L5) may occur simultaneously.
그래서 반드시 선형 관계는 아니다. 둘(L4와 L5)이 동시에 일어날 수 있다.
But Reasoning (L2) and Agent (L3) may be prerequisites for Innovation (L4) and Organization (L5).
하지만 Reasoning(L2)과 Agent(L3)는 Innovation(L4)과 Organization(L5)의 전제일 수 있다.
The hallmark of Innovation (L4) is the model’s self-iteration. What about Organization (L5)?
Innovation(L4)의 표지는 모델의 자기 반복이다. 그렇다면 Organization(L5)은?
A relatively simple way of thinking is that it would be a Multi-Agent system.
비교적 간단한 생각은 그것이 Multi-Agent 시스템일 것이라는 점이다.
Of course, how to well train a Multi-Agent system end-to-end without overfitting to certain types of Agents, so that it has better generalization, is quite challenging.
물론 Multi-Agent 시스템을 엔드투엔드로 잘 훈련하면서 특정 Agent 유형에 과적합되지 않게 하여 더 나은 일반화를 갖게 하는 것은 꽤 도전적이다.
Is Organization the “summit of the snow mountain” for the model?
Organization이 모델의 “설산 정상”인가?
Not really. Maybe it truly has no summit.
그렇지도 않다. 정말로 정상이 없을 수도 있다.
So what kind of scale do you regard the L1 to L5 levels as?
그렇다면 L1부터 L5까지의 등급을 어떤 척도로 보는가?
They are several important technical milestones, but they are not necessarily serial relationships. It is not that we expect one ability to be solved before solving the next problem.
그것들은 몇 가지 중요한 기술적 마일스톤이지만, 반드시 직렬 관계는 아니다. 어떤 능력이 해결된 후에야 다음 문제를 해결한다고 예상하는 것이 아니다.
For example, Reasoning. If you really want to solve open-ended reasoning problems, you need very strong Innovation, proposing new model architectures, and this in turn places higher requirements on reasoning ability. Essentially, it is using L4 methods to solve L2 problems and make L2 better.
예를 들어 Reasoning. 정말로 개방형 추론 문제를 해결하려면 매우 강한 Innovation이 필요하고, 새로운 모델 아키텍처를 제안해야 하며, 이는 다시 추론 능력에 더 높은 요구를 한다. 본질적으로 L4의 방법으로 L2 문제를 해결하고 L2를 더 좋게 만드는 것이다.
These several abilities will continuously become better over time.
이 몇 가지 능력은 시간이 지나면서 지속적으로 더 좋아질 것이다.
However, different technical bets inside will cause some differences in short-term paths. These differences in short-term paths will also have an impact — because you are facing a dynamic market.
다만 그 안의 다른 기술적 베팅이 단기 경로에 약간의 차이를 만들 것이다. 이 단기 경로의 차이도 영향을 미친다 — 왜냐하면 당신은 동적 시장을 마주하고 있기 때문이다.
Previously we solidified the idea that the endpoint is AGI. Today the endpoint is no longer AGI. Then what is AGI?
이전에는 종점이 AGI라고 고착화해서 생각했다. 오늘날의 종점은 더 이상 AGI가 아니다. 그렇다면 AGI는 무엇인가?
AGI is not a certain level of step. You climb to this level of step and suddenly overnight reach AGI. Rather, it is a direction.
AGI는 어떤 단계의 계단이 아니다. 이 단계의 계단에 올라서서 갑자기 하룻밤 사이에 AGI에 도달하는 것이 아니다. 그것은 하나의 방향이다.
Today in many fields, you can consider it already AGI, performing better than 99% of humans. In many math or programming competitions, at the current rate of improvement, it is expected that many problems will be fully solved soon.
오늘날 많은 분야에서 이미 AGI라고 볼 수 있다. 99%의 인간보다 더 잘한다. 많은 수학이나 프로그래밍 대회에서 현재의 향상 속도라면 곧 많은 문제가 충분히 해결될 것으로 예상된다.
AGI has two levels: on one hand technology is continuously improving; on the other hand is the impact of technology on human society. The latter is a longer-cycle matter, and is also part of AGI.
AGI에는 두 가지 차원이 있다. 한편으로는 기술이 계속 향상되고 있고, 다른 한편으로는 기술이 인간 사회에 미치는 영향이다. 후자는 더 긴 주기의 일이며, 이것 또한 AGI의 일부다.
This is a bit like after the steam engine was produced, social change needed decades or hundreds of years to digest. Some jobs become no longer necessary, but new jobs are created. Everyone becomes a “superhuman” and can do more things. The way society works and its operating efficiency will undergo huge changes.
이것은 증기기관이 생긴 후 사회 변화가 수십, 수백 년이 걸려 소화된 것과 비슷하다. 어떤 일은 더 이상 필요하지 않게 되지만 새로운 일이 생긴다. 모든 사람이 “초인”이 되어 더 많은 일을 할 수 있다. 사회의 일하는 방식과 운영 효율이 거대한 변화를 겪을 것이다.
Although we named the company after “Moonshot,” it is different from landing on the moon. Landing on the moon is the moment you stand on the moon and can claim “I have achieved it.” But with AI it is hard to suddenly shout a slogan at a certain time saying “we have achieved AGI at this very moment.”
우리가 회사 이름을 “Moonshot”으로 지었지만, 달 착륙과는 다르다. 달 착륙은 달에 선 그 순간 “내가 달성했다”고 말할 수 있다. 하지만 AI는 어떤 시점에 갑자기 “우리는 바로 이 순간 AGI를 실현했다”고 구호를 외치기 어렵다.
You keep climbing upward.
당신은 계속 위로 오른다.
Even after a period of time, it may not be you yourself climbing, but you using AI to climb. Now we let K2 do data processing, model analysis, model training — things that previously all required human work — and gradually hand them over to the model. It is like an amplifier, helping you better climb this mountain.
심지어 시간이 좀 지나면, 당신이 직접 오르는 것이 아니라 AI를 사용해서 오르게 될 수도 있다. 지금은 K2에게 데이터 처리, 모델 분석, 모델 훈련을 시키고 있는데, 이전에는 모두 사람이 해야 했던 일이다. 점차 모델에게 넘긴다. 그것은 증폭기와 같아서, 당신이 이 산을 더 잘 오르게 도와준다.
If the snow mountain is endless, what are you pursuing?
설산이 끝이 없다면, 당신은 무엇을 추구하는가?
It is the process of climbing.
오르는 과정 그 자체다.
You were originally at the foot of the mountain; now you have risen a bit higher, and the scenery you can see is different.
원래는 산 아래에 있었는데, 지금은 조금 더 위로 올라와서 볼 수 있는 풍경이 다르다.
It is a dynamically evolving process.
그것은 동적으로 진화하는 과정이다.
Let’s review the key decisions in your two years of entrepreneurship. From 2023 to 2024, your key decisions were — deciding to start the company in February 2023, beginning fundraising, building the team; in the second half of the year, Kimi launched and you bet on long context.
창업 이 년간의 핵심 결정을 복기해 보자. 2023년에서 2024년까지 당신의 핵심 결정은 — 2023년 2월에 창업을 결정하고, 자금 조달을 시작하고, 팀을 구성한 것; 하반기에 Kimi가 출시되고 긴 컨텍스트에 베팅한 것이었다.
From 2024 to 2025, what were several key decisions this year?
2024년에서 2025년까지, 이 한 해 당신의 몇 가지 핵심 결정은 무엇이었나?
A very important point is that technically we shifted from a research and development paradigm focused on pre-training and SFT to one focused on pre-training and reinforcement learning. This requires doing many things — whether talent reserves or changes in R&D methods.
매우 중요한 점은 기술적으로 우리가 사전훈련과 SFT를 중점으로 한 연구개발 패러다임에서 사전훈련과 강화학습을 중점으로 한 방식으로 전환했다는 것이다. 이것에는 많은 일이 필요하다 — 인재 확보든 연구개발 방식의 변화든.
Another point is that the shift from conversation to Agent is an important paradigm change that greatly affects our actual working methods.
또 다른 점은 대화에서 Agent로의 전환이 중요한 패러다임 변화이며, 우리의 실제 작업 방식에 큰 영향을 미친다는 것이다.
In the past half year you launched K1.5 and K2. What do they respectively mean for Kimi?
지난 반년 동안 여러분이 K1.5와 K2를 출시했다. 그것들은 각각 Kimi에게 무엇을 의미하는가?
K1.5 is more of a verification of reinforcement learning technology.
K1.5는 강화학습 기술의 검증에 더 가깝다.
Catching up with o1?
o1을 따라잡는 것인가?
Yes, we invested relatively early in this technical route, got some results, and saw how the underlying technology actually works.
맞다. 우리는 이 기술 경로에 비교적 일찍 투자했고, 일부 결과를 얻었으며, 그 뒤의 기술이 실제로 어떻게 작동하는지 보았다.
At that time we discovered that we didn’t need too much process reward or value function; they even had some side effects during training.
당시 우리는 process reward나 value function이 너무 많이 필요하지 않다는 것을 발견했다. 심지어 그것들은 훈련 과정에서 일부 부작용도 있었다.
We found that you might be able to train very well using end-to-end reward directly. This was not very clear in the early stages. In this process, we accumulated some reinforcement learning infrastructure and some algorithmic know-how.
우리는 엔드투엔드 reward를 직접 사용해도 훈련을 매우 잘할 수 있다는 것을 발견했다. 초기에는 이것이 매우 명확하지 않았다. 이 과정에서 우리는 일부 강화학습 인프라와 일부 알고리즘 know-how를 축적했다.
K2 has several priorities: First, we hope it is a very good Base Model. If you want a better Base Model, you have to look at where the current bottleneck of pre-training in the entire field lies.
K2의 중점은 몇 가지다. 첫째, 우리는 그것이 매우 좋은 Base Model이 되기를 바란다. 더 나은 Base Model을 원한다면, 현재 전체 분야의 사전훈련 병목이 어디에 있는지 봐야 한다.
We found that the growth of high-quality data is indeed very slow, and multimodal data cannot well improve the “intelligence” of text itself. You can consider high-quality data as approaching a constant. In this situation, we hope to maximize the use of every piece of data, which is the so-called token efficiency.
우리는 고품질 데이터의 성장이 실제로 매우 느리다는 것을 발견했고, 멀티모달 데이터는 텍스트 자체의 “지능”을 잘 향상시키지 못한다. 고품질 데이터를 상수에 가깝다고 생각할 수 있다. 이런 상황에서 우리는 모든 데이터를 최대한 사용하기를 바라며, 이것이 바로 token efficiency다.
You hope that with the same amount of data fed in, the “brain” can grow more, and you can obtain more intelligence.
같은 양의 데이터를 먹여도 “뇌”가 더 많이 자라고, 더 많은 지능을 얻기를 바란다.
This is different from previous thinking. Suppose you do a lot of performance optimization on the training system to make training faster; this of course has value. But training faster itself cannot raise the upper limit of intelligence, because the number of tokens is still the same. Training faster only means completing training in a shorter time, but the model effect may not become better. This is optimization on training efficiency or compute efficiency.
이것은 이전 사고와 다르다. 훈련 시스템에서 많은 성능 최적화를 해서 훈련을 더 빠르게 한다고 가정하자. 물론 가치가 있다. 하지만 훈련을 더 빠르게 하는 것 자체는 지능의 상한을 높이지 못한다. 토큰 수는 여전히 같기 때문이다. 훈련을 더 빠르게 하면 더 짧은 시간에 훈련을 끝내는 것일 뿐, 모델 효과가 반드시 더 좋아지지는 않는다. 이것은 훈련 효율 또는 compute efficiency상의 최적화다.
Previously some people did work in this area. Now we hope more to improve token efficiency, treating one piece of data as several pieces.
이전에 어떤 사람들이 이 방면에서 작업을 했다. 지금은 우리는 token efficiency를 더 향상시키기를 바라며, 한 조각의 데이터를 여러 조각처럼 취급한다.
We pay a lot of attention to things like the Muon optimizer. It is very interesting and greatly improves token efficiency. Optimizers like Adam have been used for 10 years; most model training uses Adam, but its token efficiency is not good enough.
우리는 Muon 최적화기 같은 것에 많은 관심을 기울인다. 그것은 매우 흥미롭고 token efficiency를 크게 향상시킨다. Adam 같은 최적화기는 10년 동안 사용되었고, 대부분의 모델 훈련이 Adam을 사용하지만, 그 token efficiency는 충분히 좋지 않다.
The Muon optimizer does not consider each element independently, but considers the parameters of a matrix as a whole and their dependencies. Through this method, it obtains better learning efficiency — learning the same piece of data, you can learn more intelligence.
Muon 최적화기는 각 요소를 독립적으로 고려하지 않고, 행렬의 매개변수를 전체적으로 고려하며 그들 사이의 의존 관계를 본다. 이 방식을 통해 더 나은 학습 효율을 얻는다 — 같은 데이터를 학습해도 더 많은 지능을 배울 수 있다.
In our early experiments, under compute optimal conditions, there is basically a twofold improvement. That is, learning one piece of data is equivalent to learning two pieces of data with Adam.
초기 실험에서 compute optimal 조건하에서 기본적으로 두 배의 향상이 있다. 즉, 한 조각의 데이터를 학습하는 것이 Adam으로 두 조각의 데이터를 학습하는 것과 같다.
Assuming you have 30T of high-quality tokens, it is equivalent to now having 60T of high-quality tokens.
30T의 고품질 토큰이 있다고 가정하면, 지금은 60T의 고품질 토큰이 있는 것과 같다.
But it is still the same amount of data in reality.
하지만 실제로는 여전히 같은 양의 데이터다.
After it learns, the brain grows faster. Because the learning efficiency is higher, the optimizer is better, and absorption is faster.
학습한 후 뇌가 더 빨리 자란다. 학습 효율이 더 높고, 최적화기가 더 좋으며, 흡수가 더 빠르기 때문이다.
You feed it the same amount of data, it absorbs better, the compression rate rises faster, and the loss decreases faster.
같은 양의 데이터를 먹여도 더 잘 흡수하고, 압축률이 더 빨리 올라가고, loss가 더 빨리 내려간다.
Is the Muon optimizer original to you?
Muon 최적화기는 여러분의 독창적인 것인가?
The Muon optimizer was proposed by Keller Jordan (a computer scientist and machine learning engineer who joined OpenAI in December 2024). We did a lot of optimization on top of it to adapt it and train very large-scale language models.
Muon 최적화기는 Keller Jordan(컴퓨터 과학자이자 머신러닝 엔지니어로, 2024년 12월 OpenAI에 합류)이 제안한 것이다. 우리는 그 위에서 많은 최적화를 해서 매우 대규모 언어 모델을 훈련할 수 있도록 적응시켰다.
Previously we had a Moonlight work that enabled it to be trained for the first time on language models of a certain scale. Later in the process of further scaling, we discovered many new pitfalls, such as the max logit possibly exploding.
이전에 우리는 Moonlight 작업을 해서 일정 규모의 언어 모델에서 처음으로 훈련할 수 있게 했다. 나중에 더 규모를 키우는 과정에서 많은 새로운 함정을 발견했다. 예를 들어 max logit이 폭발할 수 있는 문제.
This problem is hard to discover in small-scale experiments but will be encountered in large-scale training. So we proposed some new methods, such as clipping technology, so that it can still train well in very large-scale situations.
이 문제는 소규모 실험에서는 발견하기 어렵지만 대규모 훈련에서 만나게 된다. 그래서 우리는 clipping 기술 같은 새로운 방법을 제안해서 매우 대규모 상황에서도 잘 훈련할 수 있게 했다.
This is very important — because the number of tokens is limited, you hope every token can produce greater value.
이것은 매우 중요하다 — 토큰 수가 제한되어 있기 때문에, 모든 토큰이 더 큰 가치를 생산하기를 바란다.
I read your technical report. You tried rewriting existing data with existing models to generate new corpora. What is the specific rewriting strategy? — This is not mentioned in the report.
여러분의 기술 보고서를 읽었다. 기존 모델로 기존 데이터를 재작성해서 새로운 코퍼스를 생성하는 시도를 했다. 구체적인 재작성 전략은 어떤가? — 보고서에는 언급되지 않았다.
We do a lot of Rephrase operations on the data. For example, you have 30T tokens, but high-quality data is less, perhaps only at the level of tens of B or hundreds of B. You hope these high-quality data can be well utilized. We did some rewriting on these data so that they can be better absorbed by the model and have better generalization ability.
우리는 데이터에 많은 Rephrase 작업을 한다. 예를 들어 30T 토큰이 있지만 고품질 데이터는 더 적어서 수십 B 또는 수백 B 수준일 수 있다. 이 고품질 데이터가 잘 활용되기를 바란다. 우리는 이 데이터에 일부 재작성을 해서 모델이 더 잘 흡수하고 더 나은 일반화 능력을 갖게 했다.
The main idea is that if you learn the same piece of data many times, generalization may not be that good; there may be some overfitting problems. We hope that through rewriting, it has a certain degree of generalization.
주요 생각은 같은 데이터를 여러 번 학습하면 일반화가 그렇게 좋지 않을 수 있고, 일부 과적합 문제가 있을 수 있다는 것이다. 우리는 재작성을 통해 어느 정도의 일반화를 갖기를 바란다.
There are very many specific rewriting methods. We found one that works relatively well in experiments.
구체적인 재작성 방식은 매우 많다. 우리는 실험에서 효과가 비교적 좋은 하나를 찾았다.
Which one?
어떤 것인가?
This space is also large; there are very many research opportunities.
이 공간도 크다. 매우 많은 연구 기회가 있다.
How do you view a certain viewpoint — “Rewriting and expansion are actually useless. Being able to write out the knowledge means the knowledge is already inside; there is no new knowledge, unless other methods are used during rewriting.”
어떤 관점을 어떻게 보는가 — “재작성과 확장은 사실 쓸모없다. 지식을 쓸 수 있다는 것은 지식이 이미 안에 있다는 뜻이며, 새로운 지식은 없다. 재작성 때 다른 방법을 사용하지 않는 한.”
This is a very good question, and it is indeed related to the rewriting method. Theoretically, it still depends on whether you have new entropy input. It has some requirements on the rewriting method. But we may not have used the best rewriting method yet; there is a lot of room for exploration.
이것은 매우 좋은 질문이며, 실제로 재작성 방식과 관련이 있다. 이론적으로는 여전히 새로운 엔트로피 입력이 있는지에 달려 있다. 재작성 방식에 일부 요구사항이 있다. 하지만 우리는 아직 최고의 재작성 방식을 사용하지 않았을 수 있으며, 많은 탐색 공간이 있다.
Returning to the point just mentioned, for the K2 model, on one hand we hope it becomes a good Base Model. We very much hope to improve its token efficiency. These are the corresponding designs we made, including adding more parameters through greater sparsity.
방금 말한 점으로 돌아가서, K2 모델에 대해 한편으로는 좋은 Base Model이 되기를 바란다. 우리는 그 token efficiency를 매우 향상시키기를 바란다. 이것들은 우리가 한 해당 설계이며, 더 큰 희소성을 통해 더 많은 매개변수를 추가하는 것을 포함한다.
Then its token efficiency will also be higher, because after having more parameters, although learning the same amount of data, you will absorb better. Anyway, experiments can verify that there is indeed better token efficiency.
그러면 그 token efficiency도 더 높아질 것이다. 매개변수가 더 많아진 후, 같은 양의 데이터를 학습해도 더 잘 흡수하기 때문이다. 어쨌든 실험을 통해 실제로 더 나은 token efficiency가 있음을 검증할 수 있다.
Second, we hope it has good Agentic ability. Through various reinforcement learning or simulation of tools and environments, let it have relatively good generalization.
둘째, 우리는 그것이 좋은 Agentic 능력을 갖기를 바란다. 다양한 강화학습 또는 도구와 환경의 시뮬레이션을 통해 비교적 좋은 일반화를 갖게 한다.
For an Agentic model, the biggest challenge now is the generalization of the model.
Agentic 모델에 대해 지금 가장 큰 도전은 모델의 일반화다.
Because the limitation of current RL technology is that whether training tasks or evaluation metrics, many times they are single-point. For example, if you only train data of the same distribution as SWE-bench, it improves SWE-bench; it is a very certain thing. But after your metrics improve, it does not mean the model’s generalization becomes better.
현재 RL 기술의 한계는 훈련 과제든 평가 지표든 많은 경우 단일 지점이라는 것이다. 예를 들어 SWE-bench와 같은 분포의 데이터만 훈련하면 SWE-bench가 향상된다. 매우 확실한 것이다. 하지만 지표가 향상된 후 모델의 일반화가 더 좋아진다는 뜻은 아니다.
We are also trying to solve part of the generalization problem. We do not hope to overfit to certain tools, or overfit to certain environments, or overfit to certain specific tasks. These tasks may be good observations, but we do not hope to overfit them.
우리는 또한 일반화 문제의 일부를 해결하려고 시도하고 있다. 특정 도구에 과적합되거나, 특정 환경에 과적합되거나, 특정 구체적 과제에 과적합되기를 바라지 않는다. 이 과제들은 좋은 관찰일 수 있지만, 우리는 그것에 과적합되기를 바라지 않는다.
This problem is more serious in Agent training. Compared to conversation models, the generalization of Agents is a bigger challenge.
이 문제는 Agent 훈련에서 더 심각하다. 대화 모델에 비해 Agent의 일반화는 더 큰 도전이다.
Agentic ability is now mostly trained in the Post-Train stage. Why not train it in the Pre-Train stage?
Agentic 능력은 지금 대부분 Post-Train 단계에서 훈련된다. 왜 Pre-Train 단계에서 훈련하지 않는가?
This is also something we want to explore next.
이것 또한 다음에 우리가 탐색하고 싶은 것이다.
Is it possible that this can improve generalization?
이것이 일반화를 향상시킬 가능성이 있는가?
It depends on your approach, for example whether the data distribution is wide enough, and whether there are good evaluation methods.
당신의 접근 방식에 달려 있다. 예를 들어 데이터 분포가 충분히 넓은지, 좋은 평가 방법이 있는지.
Currently overall evaluation is an important bottleneck hindering Agent models from becoming more generalized. You will gradually observe that there are not very many Benchmarks available for Agents now. The score you observe on those Benchmarks is often not a reflection of this ability; it is rather one-sided. This is a problem everyone needs to find ways to solve.
현재 전체적인 평가는 Agent 모델이 더 일반화되는 것을 방해하는 중요한 병목이다. 점차 관찰하게 될 것이다. 지금 Agent에 사용할 수 있는 Benchmark가 매우 많지 않다. 그 Benchmark에서 관찰하는 점수는 많은 경우 이 능력의 반영이 아니며, 비교적 일면적이다. 이것은 모두가 해결 방법을 찾아야 할 문제다.
One potential idea is that we need to train AI in a more AI-native way. We hope the model participates in more of the training process. For example, if your AI can do good alignment research, theoretically it will have better generalization, not just optimizing some single-point tasks.
하나의 잠재적 아이디어는 우리가 더 AI-native한 방식으로 AI를 훈련해야 한다는 것이다. 우리는 모델이 더 많은 훈련 과정에 참여하기를 바란다. 예를 들어 당신의 AI가 좋은 alignment research를 할 수 있다면, 이론적으로 더 나은 일반화를 가질 것이며, 단지 일부 단일 지점 과제를 최적화하는 것만이 아니다.
Today Agents do not yet have as good generalization as conversation — the next few hundred steps on the snow mountain may be this.
오늘날 Agent는 아직 대화만큼 좋은 일반화를 가지고 있지 않다 — 설산의 다음 수백 계단이 이것일 수 있다.
It sounds like K1.5 was running after OpenAI, while K2 is taking the lead.
K1.5는 OpenAI를 따라 달리고, K2는 선두를 잡는 것처럼 들린다.
We borrowed many technical directions, but also hope to have some of our own innovations.
우리는 많은 기술적 방향을 차용했지만, 또한 우리만의 일부 혁신을 갖기를 바란다.
At least in publicly available materials, we are the first to use a non-Adam, or matrix orthogonalization-based method, to train with a new optimizer on such a large-scale model. This is an innovation point.
적어도 공개된 자료에서 우리는 비 Adam, 또는 행렬 직교화 기반 방식을 사용해 새로운 최적화기로 이렇게 대규모 모델을 훈련한 첫 번째다. 이것은 혁신 포인트다.
Our practices on some Agent data are also relatively early at least in publicly searchable materials.
일부 Agent 데이터에 대한 우리의 실천도 적어도 공개적으로 검색 가능한 자료에서 비교적 이르다.
What is interesting is that the higher you climb on the snow mountain, the larger the space becomes. Because the number of tokens used to complete the same task is increasing, and the complexity of problems is becoming more complex.
흥미로운 점은 설산에서 더 높이 오를수록 공간이 커진다는 것이다. 같은 과제를 완성하는 데 사용되는 토큰 수가 늘어나고, 문제의 복잡도가 더 복잡해지기 때문이다.
Just like what was said earlier: problems are inevitable, but problems can always be solved.
앞서 말한 것처럼: 문제는 불가피하지만, 문제는 항상 해결될 수 있다.
These inevitable problems will seem more than before, but your research space will also become broader accordingly.
이 불가피한 문제들은 이전보다 더 많아 보이겠지만, 당신의 연구 공간도 그에 따라 더 넓어질 것이다.
Let’s talk specifically about the K2 project. How was it initiated? How long was the preparation in between?
K2 프로젝트에 대해 구체적으로 이야기해 보자. 어떻게 입안되었나? 중간 준비 기간은 얼마나 되었나?
The preparation took a relatively long time; many of the technologies involved have been under research since last year.
준비는 비교적 긴 시간이 걸렸다. 관련된 많은 기술이 작년부터 연구되어 왔다.
For technologies like Muon, research requires a relatively long cycle. You first do early experiments and discover that this idea has potential. We will have some small experiments to verify how much potential this idea has.
Muon 같은 기술은 연구에 비교적 긴 주기가 필요하다. 먼저 초기 실험을 해서 이 아이디어에 잠재력이 있다는 것을 발견한다. 우리는 이 아이디어의 잠재력이 얼마나 되는지 검증하는 작은 실험들을 한다.
After having the idea, until finally you can put it into training a trillion-parameter model, you need to verify its effectiveness through different scaling experiments. Some problems will only be discovered after you scale to a certain scale — so the cycle is relatively long.
아이디어를 가진 후, 최종적으로 그것을 조 단위 매개변수 모델 훈련에 넣을 수 있을 때까지, 다양한 scaling 실험을 통해 그 유효성을 검증해야 한다. 어떤 문제는 일정 규모로 scale한 후에야 발견된다 — 그래서 주기가 비교적 길다.
Of course, if you only look at the model training itself, from pressing the training button to the end of training, the time is not that long. But R&D needs to do many things in advance to finally ensure that the training goes relatively smoothly.
물론 모델 훈련 자체만 보면, 훈련 버튼을 누른 시점부터 훈련이 끝날 때까지 시간이 그렇게 길지는 않다. 하지만 연구개발은 사전에 많은 일을 해서 최종적으로 훈련이 비교적 순조롭게 진행되도록 보장해야 한다.
When was the bet on doing Agentic LLM made?
Agentic LLM을 하는 베팅은 언제 했나?
It also requires a lot of accumulation; it’s just that the approach differs at different points in time.
이것 또한 많은 축적이 필요하다. 단지 다른 시점에서의 접근 방식이 다를 뿐이다.
At the beginning you don’t necessarily do it end-to-end, but you accumulate some environments and data; later you do reinforcement learning more end-to-end. In between you need a lot of infrastructure and data accumulation. It’s hard to say that it can be done very well in just one or two months.
처음에는 반드시 엔드투엔드로 하지 않지만, 일부 환경과 데이터를 축적한다. 나중에 더 엔드투엔드로 강화학습을 한다. 중간에 많은 인프라와 데이터 축적이 필요하다. 한두 달 만에 매우 잘 할 수 있다고 말하기 어렵다.
Overall I feel that large models and related technologies really need time accumulation. Still, “be a friend of time.” The technology curve is still a bit steep; it’s not something you can produce just because you want to do it today.
전체적으로 나는 대규모 모델과 관련 기술이 정말 시간 축적이 필요하다고 느낀다. 여전히 “시간의 친구가 되라.” 기술 곡선은 여전히 약간 가파르다. 오늘 하고 싶다고 해서 바로 만들어낼 수 있는 것이 아니다.
So when was the project initiated? Why initiate this project?
그래서 언제 입안했나? 왜 이 프로젝트를 입안했나?
We accumulated various technologies a year ago, but K2 was definitely a decision made in the last few months — we decided to train such a model and choose which technologies to put into it. Roughly such a decision. But it is not that I want to train this model today and start from zero.
일 년 전에 다양한 기술을 축적했지만, K2는 확실히 최근 몇 달 동안의 결정이다 — 우리는 이런 모델을 훈련하기로 결정하고, 어떤 기술을 넣을지 선택했다. 대략 그런 결정이다. 하지만 오늘 이 모델을 훈련하고 싶어서 0에서부터 시작하는 것은 아니다.
We are always training the next-generation model. It’s just a decision: which technologies should I add to the next-generation model? What kind of model do you expect it to be? Just like now we are also considering what the next-generation model after K2 should look like — this is something that needs continuous thinking and decision-making.
우리는 항상 다음 세대 모델을 훈련하고 있다. 단지 하나의 결정일 뿐이다: 다음 세대 모델에 어떤 기술을 추가해야 하는가? 어떤 종류의 모델을 기대하는가? 지금처럼 K2 이후의 다음 세대 모델이 어떤 모습이어야 하는지 고려하는 것 — 이것은 지속적으로 생각하고 결정해야 할 일이다.
Each time you look at it, the toolbox now has many more new things; which ones to take out and use? It is such a process.
매번 보면, 지금 도구 상자에 많은 새로운 것들이 더 있다. 어떤 것을 꺼내 사용할 것인가? 그런 과정이다.
Are the research and training teams separate? If research on these technologies started a year ago, when doing formal training, is it one team doing the whole thing?
연구 팀과 훈련 팀은 분리되어 있나? 일 년 전부터 이 기술들을 연구하기 시작했다면, 정식 훈련을 할 때 한 팀이 전체 일을 하는가?
It is one team; these things are hard to separate. You will encounter problems in actual training; if you didn’t understand them before, you cannot solve them.
한 팀이다. 이 것들은 분리하기 어렵다. 실제 훈련에서 문제를 만나게 된다. 이전에 이해하지 못했다면 해결할 수 없다.
What challenges did you encounter during the R&D process of K2?
K2 연구개발 과정에서 어떤 도전을 만났나?
When you train with Muon, it will explode.
Muon으로 훈련할 때 폭발한다.
We have drawn some figures in the paper; your max logit will rise very high, rising to several hundred or even higher. We believe this thing has an impact on training stability; if you train for a long time, many so-called internal metrics become abnormal, which is harmful to the upper limit of the model.
우리는 논문에 일부 그래프를 그렸다. max logit이 매우 높게 올라가며, 수백 또는 그 이상으로 올라간다. 우리는 이것이 훈련 안정성에 영향을 미친다고 믿는다. 오랫동안 훈련하면 많은 이른바 내부 지표가 비정상적이 되어 모델의 상한에 해롭다.
We essentially went back to revisit and fix it. Because this thing cannot be predicted in small-scale experiments; there is no explosion problem on small scale.
우리는 본질적으로 돌아가서 재검토하고 고쳤다. 이 것은 소규모 실험에서 예측할 수 없기 때문이다. 소규모에서는 폭발 문제가 없다.
Others are basically fine; many experiments were done on small scale and are transferable; the problems are not big. The only problem is this one that cannot be verified on small scale and needs to be temporarily solved during the scaling process.
다른 것들은 기본적으로 괜찮다. 소규모에서 많은 실험을 했고 이전 가능하다. 문제는 크지 않다. 유일한 문제는 소규모에서 검증할 수 없고 scaling 과정에서 임시로 해결해야 하는 이것이다.
Recently K2 became popular. Has your mood had ups and downs?
최근에 K2가 인기를 끌었다. 당신의 기분에 기복이 있었나?
Not really, it’s okay — this is a long journey.
별로 없다. 괜찮다 — 이것은 긴 여정이다.
You need to continuously do the next-generation model, still returning to those two sentences carved on the stone: new problems will arise, and then go solve them. This is also the most interesting part.
다음 세대 모델을 지속적으로 해야 한다. 여전히 돌에 새긴 그 두 문장으로 돌아간다: 새로운 문제가 생길 것이고, 그러면 그것을 해결하러 간다. 이것 또한 가장 흥미로운 부분이다.
I heard that in the internal group you described K2 as meaning K2 mountain (乔戈里峰).
내부 그룹에서 K2를 乔戈里峰(K2 산)을 의미한다고 묘사했다고 들었다.
K2 was originally one of the hardest mountains in the world to climb; the names happen to coincide.
K2는 원래 세계에서 가장 어려운 산봉우리 중 하나다. 이름이 우연히 겹친다.
It is not the endpoint, because it is not the highest mountain, but it may be the hardest. This is because there are many paradigm shifts now, from conversation to Agent, and the Base Model scale further increases; there is inherent difficulty.
그것은 종점이 아니다. 가장 높은 산이 아니기 때문이다. 하지만 가장 어려울 수 있다. 지금은 대화에서 Agent로의 많은 패러다임 전환이 있고, Base Model 규모가 더 커지기 때문에 본질적인 어려움이 있다.
Did the release results of K2 exceed your expectations?
K2의 출시 결과가 당신의 기대를 초과했나?
About as expected. You already know how well the model trained while it’s in progress. There were no surprises, pleasant or otherwise.
대략 기대했던 대로다. 진행 중에 모델이 얼마나 잘 훈련되는지 이미 알고 있다. 유쾌하거나 그렇지 않거나 놀랄 일은 없었다.
What were the most important pieces of know-how you gained from training K2?
K2를 훈련하면서 얻은 가장 중요한 know-how는 무엇이었나?
One is the optimization of token efficiency. How to make every token produce more value. This includes the optimizer, architecture design, and data processing methods.
하나는 token efficiency의 최적화다. 모든 토큰이 더 많은 가치를 생산하게 하는 방법. 여기에는 최적화기, 아키텍처 설계, 데이터 처리 방법이 포함된다.
Another is the training of Agentic ability. How to make the model have better generalization in multi-turn interactions with the environment.
또 하나는 Agentic 능력의 훈련이다. 모델이 환경과의 다회차 상호작용에서 더 나은 일반화를 갖게 하는 방법.
These are things that need continuous accumulation and iteration.
이것들은 지속적인 축적과 반복이 필요한 것들이다.
Looking back at the entire process of K2, what was the biggest risk?
K2의 전체 과정을 되돌아보면, 가장 큰 위험은 무엇이었나?
The biggest risk is still technical uncertainty. Many things you cannot fully predict until you actually scale up.
가장 큰 위험은 여전히 기술적 불확실성이다. 실제로 규모를 키우기 전까지는 많은 것을 완전히 예측할 수 없다.
Especially the stability of new optimizers and the generalization of Agent training. These require a lot of trial and error.
특히 새로운 최적화기의 안정성과 Agent 훈련의 일반화. 이것들은 많은 시행착오를 필요로 한다.
But this is also the charm of doing frontier research.
하지만 이것 또한 최전선 연구를 하는 매력이다.
You mentioned earlier that K2 escaped the “brain in a vat” through coding ability and grew “hands.” Can you elaborate more on this?
앞서 K2가 코딩 능력을 통해 “통 속의 뇌”에서 탈출하고 “손”을 키웠다고 말했다. 이것에 대해 더 자세히 설명할 수 있나?
Previous models were more like thinking inside a closed space. They could reason, write, answer questions, but they could not actively manipulate the external digital world.
이전 모델들은 더 닫힌 공간 안에서 생각하는 것과 같았다. 추론하고, 쓰고, 질문에 답할 수 있었지만, 외부 디지털 세계를 능동적으로 조작할 수 없었다.
Through strong coding ability and tool-calling ability, the model can now write code, execute code, call APIs, browse the web, operate files, and so on.
강한 코딩 능력과 도구 호출 능력을 통해 모델은 이제 코드를 쓰고, 코드를 실행하고, API를 호출하고, 웹을 브라우징하고, 파일을 조작하는 등을 할 수 있다.
It is like growing hands and being able to interact with the environment, get feedback, and continuously adjust its behavior.
손을 키운 것과 같아서 환경과 상호작용하고, 피드백을 얻고, 행동을 지속적으로 조정할 수 있다.
This is a key step from pure language models to Agentic models.
이것은 순수 언어 모델에서 Agentic 모델로의 핵심 단계다.
How do you view the relationship between open source and closed source now?
지금 오픈소스와 클로즈드소스의 관계를 어떻게 보는가?
Open source has very important value. It can accelerate the progress of the entire community and let more people participate.
오픈소스는 매우 중요한 가치가 있다. 전체 커뮤니티의 진전을 가속화하고 더 많은 사람들이 참여하게 할 수 있다.
At the same time, first-party products and closed-source models also have their advantages, especially in vertical integration and experience optimization.
동시에 1차 제품과 클로즈드소스 모델도 장점이 있다. 특히 수직 통합과 경험 최적화에서.
The two can coexist and promote each other.
둘은 공존하고 서로를 촉진할 수 있다.
K2 is open source. What considerations did you have when deciding to open source it?
K2는 오픈소스다. 오픈소스로 결정했을 때 어떤 고려가 있었나?
We hope that through open source, more people can use it, improve it, and build on top of it.
우리는 오픈소스를 통해 더 많은 사람들이 그것을 사용하고, 개선하고, 그 위에 구축하기를 바란다.
This is also consistent with the idea of climbing the infinite mountain. The more people climb together, the faster the overall progress may be.
이것 또한 무한한 산을 오르는 생각과 일치한다. 더 많은 사람들이 함께 오를수록 전체 진전이 더 빠를 수 있다.
Of course, we will also continue to do our own product and model iterations.
물론 우리도 계속해서 우리만의 제품과 모델 반복을 할 것이다.
K2 was originally one of the hardest mountains in the world to climb; the names happen to coincide.
K2는 원래 세계에서 가장 어려운 산봉우리 중 하나다. 이름이 우연히 겹친다.
It is not the endpoint, because it is not the highest mountain, but it may be the hardest. This is because there are many paradigm shifts now, from conversation to Agent, and your Base Model scale further becomes larger; there is inherent difficulty itself.
그것은 종점이 아니다. 가장 높은 산이 아니기 때문이다. 하지만 가장 어려울 수 있다. 지금은 대화에서 Agent로의 많은 패러다임 전환이 있고, Base Model 규모가 더 커지기 때문에 본질적인 어려움이 있다.
Did the release results of K2 exceed your expectations?
K2 출시 결과가 당신의 기대를 초과했나?
About the same. How the model trained was already known during the process; there were no surprises or accidents.
거의 같다. 모델이 어떻게 훈련되는지는 과정 중에 이미 알고 있었다. 놀랄 일이나 우연은 없었다.
What were the most important several pieces of know-how you harvested from training K2?
K2 훈련에서 수확한 가장 중요한 몇 가지 know-how는 무엇인가?
We wrote them all in the paper.
우리는 그것들을 모두 논문에 썼다.
We are all very open; we still want to share more with the community.
우리는 모두 매우 개방적이다. 여전히 커뮤니티와 더 많이 공유하고 싶다.
Because K2 is an Agentic large language model, how would you define Agent and classify Agents?
K2가 Agentic 대규모 언어 모델이기 때문에, 당신은 Agent를 어떻게 정의하고 Agent를 분류하겠는가?
It may be turning from a “brain in a vat” into a system that can interact with the world, because the most important feature of a so-called Agent is that it can use tools in multiple turns.
그것은 “통 속의 뇌”에서 세계와 상호작용할 수 있는 시스템으로 변하는 것일 수 있다. 이른바 Agent의 가장 중요한 특징은 다회차로 도구를 사용할 수 있다는 것이기 때문이다.
There are two key points: one is multi-turn, and the other is tools.
두 가지 핵심 포인트가 있다. 하나는 다회차, 다른 하나는 도구다.
Multi-turn means you can do it many times; it is a way of test-time scaling. Tools are the way to connect this “brain” with the external world.
다회차는 여러 번 할 수 있다는 뜻이며, test-time scaling의 한 방식이다. 도구는 이 “뇌”를 외부 세계와 연결하는 방식이다.
For example, if you use a search engine, you can connect the model with the entire internet; if you can write code, you can let the “brain” connect with the digital world, because almost all automation in the digital world can be described with code, so it can possess this automation ability.
예를 들어 검색 엔진을 사용하면 모델을 전체 인터넷과 연결할 수 있다. 코드를 쓸 수 있으면 “뇌”가 디지털 세계와 연결될 수 있다. 디지털 세계의 거의 모든 자동화가 코드로 기술될 수 있기 때문에, 이러한 자동화 능력을 가질 수 있다.
These two are the features of Agent in my imagination. Next there will be more and more tools. Of course, tools will present a long-tail distribution. If the model generalizes well, it will not only use common tools but also be able to use highly personalized tools.
이 두 가지가 내가 상상하는 Agent의 특징이다. 다음에는 점점 더 많은 도구가 생길 것이다. 물론 도구는 긴 꼬리 분포를 보일 것이다. 모델이 잘 일반화되면 흔한 도구만 사용하는 것이 아니라 매우 개인화된 도구도 사용할 수 있다.
For example, the model can access the company’s internal database, personal documents, or even access customized APIs to complete business operations such as ticket refunds or placing orders. It should be able to generalize to tools it has never seen. I have always felt that what Agents lack the most is generalization ability.
예를 들어 모델이 회사 내부 데이터베이스, 개인 문서에 접근하거나, 심지어 맞춤형 API에 접근해서 환불이나 주문 같은 비즈니스 작업을 완성할 수 있다. 본 적 없는 도구에도 일반화할 수 있어야 한다. 나는 항상 Agent가 가장 부족한 것이 일반화 능력이라고 느꼈다.
If the generalization ability is strong, the various vertical Agents that everyone discusses become less necessary. Because a general Agent generalizing to long-tail tools can solve many domain-specific problems by connecting different tools. As long as you add a customized database, customized API, or customized document interface to it, you can make a very vertical Agent. Its universality will be much stronger.
일반화 능력이 강하면 모두가 논의하는 다양한 수직 Agent가 덜 필요해진다. 일반 Agent가 긴 꼬리 도구에 일반화되면 다양한 도구를 연결해서 많은 영역 특화 문제를 해결할 수 있기 때문이다. 맞춤형 데이터베이스, 맞춤형 API, 맞춤형 문서 인터페이스만 추가하면 매우 수직적인 Agent를 만들 수 있다. 그 보편성이 훨씬 강해질 것이다.
Multi-turn is mainly to realize test-time scaling and can do complex tasks. Unlike conversation models that output one round at a time, this can do different things — just like humans — a person’s daily work can be considered a sequence of multi-turn tool usage. You hope to fit the human sequence into it, but you cannot collect such digitized data, so you can construct it with reinforcement learning.
다회차는 주로 test-time scaling을 실현하고 복잡한 작업을 할 수 있게 한다. 한 번에 한 라운드를 출력하는 대화 모델과 달리, 이것은 다른 일들을 할 수 있다 — 사람처럼 — 사람의 일상 작업을 다회차 도구 사용의 시퀀스로 볼 수 있다. 그 사람의 시퀀스를 맞추고 싶지만, 그런 디지털화된 데이터를 수집할 수 없으니 강화학습으로 구성할 수 있다.
Its essence is simulating human behavior — however, you cannot simply say it is simulating humans; calling it “simulating human behavior” is not very accurate; it is actually general.
본질은 인간 행동을 시뮬레이션하는 것이다 — 그러나 단순히 인간을 시뮬레이션한다고 말할 수 없다. “인간 행동을 시뮬레이션한다”고 부르는 것은 그다지 정확하지 않다. 그것은 실제로 일반적이다.
What does it mean that you cannot simply say it is simulating humans? Humans are also very general.
단순히 인간을 시뮬레이션한다고 말할 수 없다는 것은 무슨 뜻인가? 인간도 매우 일반적이다.
Yes, humans are general; humans are the so-called universal constructor.
맞다. 인간은 일반적이다. 인간은 이른바 universal constructor(만능 구성자)다.
But its main purpose is not to simulate humans; the main purpose is universality; that is the purpose of the design.
하지만 주요 목적은 인간을 시뮬레이션하는 것이 아니다. 주요 목적은 보편성이다. 그것이 설계의 목적이다.
It is similar to how humans do things; it is just a coincidental result, not the purpose of designing the system.
인간이 하는 방식과 유사하다. 단지 우연한 결과일 뿐, 시스템을 설계하는 목적이 아니다.
This is like designing an airplane: the purpose is to make it a means of transportation, not to fly like a bird.
이것은 비행기를 설계하는 것과 같다. 목적은 그것을 교통수단으로 만드는 것이지, 새처럼 날게 하는 것이 아니다.
We make Agent systems more to create a general intelligence, aligned with this goal; but it just happens to be similar to humans.
우리는 Agent 시스템을 만드는 것이 더 일반 지능을 만들기 위함이며, 이 목표에 맞춰져 있다. 하지만 우연히 인간과 유사할 뿐이다.
How to improve the universality of Agents? Have you explored any methods?
Agent의 보편성을 어떻게 향상시키는가? 어떤 방법을 탐색해 보았나?
This is a very difficult problem. Today the generalization of Agents has a risk of falling into overfitting on certain Benchmarks, but now there is a lack of good Benchmarks. This is the next challenge.
이것은 매우 어려운 문제다. 오늘날 Agent의 일반화에는 특정 Benchmark에 과적합될 위험이 있지만, 지금은 좋은 Benchmark가 부족하다. 이것이 다음 도전이다.
However, there may be some solutions. I still feel that using more AI to train AI can alleviate this problem to a certain extent.
그러나 일부 해결책이 있을 수 있다. 나는 여전히 더 많은 AI를 사용해 AI를 훈련하면 이 문제를 어느 정도 완화할 수 있다고 느낀다.
When can we achieve using AI to train AI? What is the current bottleneck?
언제 AI로 AI를 훈련할 수 있게 되는가? 현재 병목은 무엇인가?
Part of it has already been achieved now, but you hope it does more. Much of it still relies on human design now.
지금은 일부가 이미 달성되었지만, 더 많이 하기를 바란다. 지금은 많은 부분이 여전히 인간 설계에 의존한다.
This then reaches the next stage of Innovator (L4).
그러면 다음 단계인 Innovator(L4)에 도달한다.
This is very interesting. You need to use some Innovation methods to solve Agent problems. Because Agent generalization is not enough, you have to use innovation to solve it — using L4 technology to solve L3 problems.
이것은 매우 흥미롭다. Agent 문제를 해결하기 위해 일부 Innovation 방법을 사용해야 한다. Agent 일반화가 충분하지 않기 때문에 혁신으로 해결해야 한다 — L4 기술을 사용해 L3 문제를 해결하는 것이다.
So the definition from L1 to L5 may really not be linear. Without good Innovation, without using AI to train or using AI to align AI, it is hard for Agents to achieve good generalization.
그래서 L1에서 L5까지의 정의는 정말로 선형적이지 않을 수 있다. 좋은 Innovation 없이, AI로 훈련하거나 AI로 AI를 정렬하는 방식 없이 Agent가 좋은 일반화를 달성하기 어렵다.
You artificially define some tasks and only fit those tasks, but performance on other unseen tasks is not good. You only brush up the scores of a few tasks, but users do not feel that good in more OOD scenarios.
일부 작업을 인위적으로 정의하고 그 작업에만 맞추면, 다른 보이지 않는 작업에서의 성능이 좋지 않다. 몇 가지 작업 점수만 올리고, 더 많은 OOD 시나리오에서 사용자의 체감이 그렇게 좋지 않다.
The field is now facing a stage where Benchmarks are insufficient or ineffective, and Agent generalization has problems.
이 분야는 지금 Benchmark가 부족하거나 무효하고, Agent 일반화에 문제가 있는 단계에 직면해 있다.
Why are math and code relatively easy domains for generalization?
왜 수학과 코드가 상대적으로 일반화하기 쉬운 영역인가?
Actually not. If you do reinforcement learning, there are similar problems now.
사실 그렇지 않다. 강화학습을 하면 지금도 비슷한 문제가 있다.
The generalization of reinforcement learning itself is better than doing SFT, because there are more on-policy samples in the process; the model learns from its own samples, generalization looks better, and there is negative gradient. These two factors lead to better generalization performance from the evidence.
강화학습 자체의 일반화는 SFT를 하는 것보다 낫다. 과정에서 더 많은 on-policy 샘플이 있기 때문이다. 모델이 자체 샘플에서 학습하고, 일반화가 더 나아 보이며, 음의 그래디언트가 있다. 이 두 요인이 증거상 더 나은 일반화 성능을 가져온다.
But generalization is limited. For example, if you achieve 99 points on a certain type of math competition, other math problems may improve by 5 points, but it is hard to directly achieve 99 points. Without doing the corresponding RL tasks, it is hard to directly achieve such generalization.
하지만 일반화는 제한적이다. 예를 들어 특정 유형의 수학 대회에서 99점을 달성하면, 다른 수학 문제는 5점 향상될 수 있지만, 직접 99점을 달성하기는 어렵다. 해당 RL 작업을 하지 않으면 그런 일반화를 직접 달성하기 어렵다.
So doing math problems also has similar problems; it is constrained by the distribution — still “you reap what you sow.”
그래서 수학 문제를 푸는 것도 비슷한 문제가 있다. 분포에 의해 제약받는다 — 여전히 “심은 대로 거둔다.”
But we hope that reinforcement learning or post-training uses more AI, so that the model can escape the situation of “you reap what you sow.”
하지만 우리는 강화학습이나 후훈련이 더 많은 AI를 사용해서 모델이 “심은 대로 거두는” 상황을 벗어나기를 바란다.
Is it possible that ultimately it cannot escape, and generalization cannot be greatly improved?
결국 벗어날 수 없고, 일반화를 크게 향상시킬 수 없을 가능성이 있는가?
Still returning to what was just said: problems are inevitable, but problems can be solved. You advance forward every time — generalization will become better; it does not necessarily have an end; there will always be better generalization.
여전히 방금 말한 것으로 돌아간다: 문제는 불가피하지만, 문제는 해결될 수 있다. 매번 앞으로 전진한다 — 일반화는 더 좋아질 것이다. 끝이 있을 필요는 없다. 항상 더 나은 일반화가 있을 것이다.
For Agents, tasks and environments are very important. How to define good tasks? How to define good environments? Do you have some thoughts during the exploration process?
Agent에게 작업과 환경은 매우 중요하다. 좋은 작업을 어떻게 정의하는가? 좋은 환경을 어떻게 정의하는가? 탐색 과정에서 어떤 생각이 있었나?
One way is: given a model, then design some environments to reverse-fit this model. Of course you can also design forward: assuming you are a first-party developer, design tools and environments forward, so that the model improves ability in these environments.
한 가지 방식은: 모델을 주고, 그다음 일부 환경을 설계해서 이 모델을 역으로 맞추는 것이다. 물론 순방향으로 설계할 수도 있다: 당신이 1차 개발자라고 가정하고, 도구와 환경을 순방향으로 설계해서 모델이 이 환경들에서 능력을 향상시키게 한다.
The key is to make this design have better universality. It can do many tasks and should not specially design tools and environments for certain specific tasks. When the design is general enough, the model can learn in it, rather than reverse-fitting the model. This may be a better approach.
핵심은 이 설계가 더 나은 보편성을 갖게 하는 것이다. 많은 작업을 할 수 있어야 하며, 특정 작업을 위해 특별히 도구와 환경을 설계해서는 안 된다. 설계가 충분히 보편적일 때 모델이 그 안에서 학습할 수 있으며, 모델을 역으로 맞추는 것이 아니다. 이것이 더 나은 접근일 수 있다.
I noticed one point: generally people tend to design a sufficiently challenging task in task design, which will give birth to some more essential new methods; but K2 designed some medium-difficulty tasks. What was the consideration for this? Will this affect universality?
한 가지 점을 주목했다: 일반적으로 사람들은 작업 설계에서 충분히 도전적인 작업을 설계하는 경향이 있으며, 이는 더 본질적인 새로운 방법을 낳는다. 하지만 K2는 일부 중간 난이도 작업을 설계했다. 이에 대한 고려는 무엇인가? 이것이 보편성에 영향을 미치는가?
It is also a mountain-climbing process. You cannot let the model prove a mathematical problem that no one has proved yet right from the start; the sample efficiency will be very low.
이것 또한 산을 오르는 과정이다. 처음부터 모델에게 아직 아무도 증명하지 못한 수학 문제를 증명하게 할 수 없다. 샘플 효율이 매우 낮을 것이다.
A relatively good method now is that if reinforcement learning is paired with a good sampling strategy, it is essentially an implicit curriculum learning mechanism, hoping the model starts learning from an appropriate difficulty and gradually increases difficulty, rather than learning very difficult tasks from the beginning. Otherwise sampling efficiency is low, basically learning nothing, and computing power may all be wasted.
지금 비교적 좋은 방법은 강화학습이 좋은 샘플링 전략과 결합되면 본질적으로 암묵적인 커리큘럼 학습 메커니즘이 되어, 모델이 적절한 난이도부터 학습을 시작하고 점차 난이도를 높이는 것을 바라는 것이다. 처음부터 매우 어려운 작업을 학습하는 것이 아니다. 그렇지 않으면 샘플 효율이 낮아 기본적으로 아무것도 배우지 못하고, 컴퓨팅 파워가 모두 낭비될 수 있다.
But the challenge is that many of today’s tasks are still based on human stock data or artificially designed tasks; the AI-native part is still relatively small, which will bring generalization problems.
하지만 도전은 오늘날의 많은 작업이 여전히 인간 축적 데이터나 인위적으로 설계된 작업에 기반한다는 것이다. AI-native 부분이 여전히 비교적 적어서 일반화 문제를 가져올 것이다.
In your eyes, what is the relationship between Coding Agent and general Agent?
당신의 눈에 Coding Agent와 일반 Agent의 관계는 무엇인가?
Coding Agent is a subset of tasks, but it may be a very important subset.
Coding Agent는 작업의 하위 집합이지만, 매우 중요한 하위 집합일 수 있다.
In the end we still hope not only to do Coding. Including the models we train now, we also do not only let it do Coding, because it itself has some limitations.
결국 우리는 Coding만 하는 것을 원하지 않는다. 지금 훈련하는 모델도 Coding만 하게 하지 않는다. 그것 자체에 일부 한계가 있기 때문이다.
Can it be said this way? — Coding is equivalent to the human hand.
이렇게 말할 수 있는가? — Coding은 인간의 손에 해당한다.
Relatively speaking, is Coding a relatively easy task for Agents?
상대적으로 Coding이 Agent에게 비교적 쉬운 작업인가?
It is relatively easy to verify, so relatively easy to learn. It will also face similar challenges — generalization problems; even Coding Agents will encounter the same challenges.
검증하기 비교적 쉽기 때문에 학습하기 비교적 쉽다. 비슷한 도전도 직면할 것이다 — 일반화 문제. Coding Agent조차 같은 도전에 직면할 것이다.
The importance of Coding Agent as a subset lies in that it represents the automation of the digital world. Now many Agent tool collections are fixed; if you want to create a new tool, the essence is to write a piece or a large piece of code to implement it. Or if you want to do better context management (Context Engineering), behind it also corresponds to a tool, and this tool may also be implemented with code. Code has a unique position and role here.
Coding Agent가 중요한 하위 집합인 이유는 그것이 디지털 세계의 자동화를 대표하기 때문이다. 지금 많은 Agent 도구 집합은 고정되어 있다. 새로운 도구를 만들고 싶다면 본질은 코드 한 조각 또는 큰 조각을 써서 구현하는 것이다. 또는 더 나은 컨텍스트 관리(Context Engineering)를 하고 싶다면 그 뒤에도 도구가 대응되며, 이 도구도 코드로 구현될 수 있다. 코드는 여기서 독특한 위치와 역할을 가진다.
But it is not enough just to do a Coding Agent. Because many non-programmers also use Claude Code to complete tasks, such as lawyers, product managers, designers; they use Claude Code because the model has a certain degree of generalization ability, not only writing code.
하지만 Coding Agent만 하는 것으로는 충분하지 않다. 많은 비프로그래머도 Claude Code를 사용해 작업을 완성한다. 변호사, 제품 매니저, 디자이너 등. 그들이 Claude Code를 사용하는 이유는 모델이 어느 정도의 일반화 능력을 가지고 있기 때문이며, 코드만 쓰는 것이 아니다.
What you want to do is a general Agent, not a Coding model?
여러분이 하고 싶은 것은 일반 Agent이지, Coding 모델이 아닌가?
We still hope to do a general model.
우리는 여전히 일반 모델을 하기를 바란다.
From writing code to manipulating the entire digital world, what abilities does Agent currently still lack?
코드를 쓰는 것에서 전체 디지털 세계를 조작하는 것까지, Agent가 현재 여전히 부족한 능력은 무엇인가?
The use of these high-frequency tools is still not good enough now; there is large room in ability. This also shows that better Benchmark observations are currently lacking. SWE-bench may soon become saturated; many Benchmarks are not good enough and do not truly reflect actual user experience.
지금 이러한 고빈도 도구 사용이 아직 충분히 좋지 않다. 능력에 큰 여지가 있다. 이것 또한 현재 더 나은 Benchmark 관찰이 부족하다는 것을 보여준다. SWE-bench는 곧 포화될 수 있다. 많은 Benchmark가 충분히 좋지 않고 실제 사용자 경험을 진정으로 반영하지 않는다.
High-frequency tools themselves will have room. For long-tail tools, in situations you have never seen, completely OOD, how to have better generalization? This is also a very important problem that needs to be solved.
고빈도 도구 자체에 여지가 있을 것이다. 긴 꼬리 도구의 경우, 본 적 없는 완전히 OOD 상황에서 어떻게 더 나은 일반화를 가질 것인가? 이것 또한 해결해야 할 매우 중요한 문제다.
For Agents, are Long Context and Long-Term Memory important?
Agent에게 Long Context와 Long-Term Memory가 중요한가?
Long Context is also very important. Because many tasks now cannot be completely solved by 128K or 256K Context; you need million-level or even more.
Long Context도 매우 중요하다. 지금 많은 작업이 128K나 256K Context로 완전히 해결되지 않기 때문이다. 백만 수준 또는 그 이상이 필요하다.
And the challenge is that you not only need to be able to handle such long Context, but also ensure that the “brain works well”; the intelligence needs to be very high.
그리고 도전은 그렇게 긴 Context를 처리할 수 있을 뿐만 아니라, “뇌가 잘 작동”하도록 보장해야 한다는 것이다. 지능이 매우 높아야 한다.
This is a very big challenge for model training. On one hand you hope the compression rate is high enough and the model is large enough; on the other hand you hope it is relatively long. There is a natural conflict between the two, so better architecture is needed.
이것은 모델 훈련에 매우 큰 도전이다. 한편으로는 압축률이 충분히 높고 모델이 충분히 크기를 바라고, 다른 한편으로는 비교적 길기를 바란다. 둘 사이에 자연스러운 충돌이 있으므로 더 나은 아키텍처가 필요하다.
But for some architectures you will find that the effect improves under longer Context, but under short Context it may not improve or even decline; this involves the balance problem of the architecture.
하지만 일부 아키텍처에서는 더 긴 Context에서 효과가 향상되지만, 짧은 Context에서는 향상되지 않거나 심지어 하락하는 것을 발견할 것이다. 이것은 아키텍처의 균형 문제를 수반한다.
However, these problems can be gradually solved next; I feel there are some solutions.
그러나 이 문제들은 다음에 점차 해결될 수 있다. 일부 해결책이 있다고 느낀다.
In addition, the current RL training methods still have large room for improvement. For example, when training complex multi-agent systems, if you only use end-to-end reward, it is likely not enough. How is the intermediate reward generated? Can it escape some human design?
또한 현재 RL 훈련 방식에는 여전히 큰 향상 여지가 있다. 예를 들어 복잡한 다중 에이전트 시스템을 훈련할 때 엔드투엔드 reward만 사용하면 충분하지 않을 가능성이 높다. 중간 reward는 어떻게 생성되는가? 일부 인간 설계를 벗어날 수 있는가?
This is also a direction very worth exploring.
이것 또한 매우 탐색할 가치가 있는 방향이다.
Looking back at our conversation last year, there is a question I really want to ask you.
작년 우리 대화를 되돌아보니, 정말 묻고 싶은 질문이 하나 있다.
You said last year that open source would lag behind closed source. Because the way of open source is different from before; previously everyone could contribute to open source, but now large model open source is essentially centralized, and community contributions have not been verified by computing power. In contrast, the closed-source camp gathers talent and capital, which is an integration of market resources.
당신은 작년에 오픈소스가 클로즈드소스에 뒤처질 것이라고 말했다. 오픈소스 방식이 이전과 다르기 때문이다. 이전에는 모두가 오픈소스에 기여할 수 있었지만, 지금은 대규모 모델 오픈소스가 본질적으로 중앙집권적이며, 커뮤니티 기여가 컴퓨팅 파워로 검증되지 않았다. 그에 비해 클로즈드소스 진영은 인재와 자본을 모아 시장 자원을 통합한다.
You said at the time: “Leaders will not open source; only laggards will do so.”
당신은 당시 이렇게 말했다: “선도자는 오픈소스하지 않는다. 뒤처진 자만 그렇게 한다.”
But today you open-sourced.
하지만 오늘 여러분은 오픈소스했다.
Because we are not yet completely leading on a global scale (laughs).
왜냐하면 우리는 지금 전 세계 범위에서 아직 완전히 선도하고 있지 않기 때문이다 (웃음).
Some judgments hold true in the big direction: when your model is released, the community can contribute some things. For example, on the inference side you can do many things; you can let more people use the model for free.
일부 판단은 큰 방향에서 성립한다: 모델이 출시되면 커뮤니티가 일부 기여를 할 수 있다. 예를 들어 추론 측면에서 많은 일을 할 수 있고, 더 많은 사람들이 모델을 무료로 사용하게 할 수 있다.
But if it is to contribute to the model itself and make the model stronger, currently only the original factory can do it.
하지만 모델 자체에 기여해서 모델을 더 강하게 만드는 것은, 현재로서는 원제작사만이 할 수 있다.
Of course, if you look at the Base Model, this is indeed the case; but if based on an open-source model you do a large amount of post-training, especially Agentic post-training, it may give birth to new opportunities.
물론 Base Model을 보면 확실히 그렇다. 하지만 오픈소스 모델을 기반으로 대량의 후훈련, 특히 Agentic 후훈련을 하면 새로운 기회가 생길 수 있다.
Suppose you now really want to make a law-related Agent, and you are a startup company, then you can completely based on K2, under your specific tool set, train a Specialized Agent, which can perform very well in the scenarios you care about. This kind of opportunity exists.
지금 법률 관련 Agent를 정말 만들고 싶고, 당신이 스타트업 회사라면, K2를 기반으로 특정 도구 집합 아래에서 Specialized Agent를 훈련할 수 있다. 그것은 당신이 관심 있는 시나리오에서 매우 잘 작동할 수 있다. 이런 기회가 존재한다.
It is more about empowering downstream applications, rather than feeding back to the improvement of the base model. Of course this issue needs to be observed dynamically.
더 많은 것은 다운스트림 애플리케이션을 강화하는 것이지, 기초 모델 향상에 피드백하는 것이 아니다. 물론 이 문제는 동적으로 관찰해야 한다.
Will you choose open source for the long term?
장기적으로 오픈소스를 선택할 것인가?
This is what we hope to do for the long term, but not necessarily only do open source. We hope to share technical know-how with the community; this is an important point for accelerating technological improvement.
이것이 우리가 장기적으로 하기를 바라는 것이지만, 반드시 오픈소스만 하는 것은 아니다. 우리는 커뮤니티와 기술 know-how를 공유하기를 바란다. 이것이 기술 향상을 가속화하는 중요한 점이다.
Everyone does not have to be completely competitive; there can also be cooperation, and even all open-source companies form an ecosystem to better promote technological development — the snow mountain can be climbed better, race to the top.
모두가 완전히 경쟁할 필요는 없다. 협력도 있을 수 있고, 심지어 모든 오픈소스 회사가 생태계를 형성해서 기술 발전을 더 잘 추진할 수 있다 — 설산을 더 잘 오를 수 있다, race to the top.
But not everything has to be open-sourced. For example, when cooperating with certain companies, not everything has to be opened.
하지만 모든 것을 오픈소스할 필요는 없다. 예를 들어 특정 회사와 협력할 때 모든 것을 공개할 필요는 없다.
Overall, is open source a belief in a technical system, or a strategy of market game?
전반적으로 오픈소스는 기술 체계에 대한 신념인가, 아니면 시장 게임의 전략인가?
Objectively speaking both, and both have benefits. But ultimately we hope through this to make technology safer and reach a better level faster.
객관적으로 말하면 둘 다이며, 둘 다 이점이 있다. 하지만 궁극적으로 우리는 이를 통해 기술을 더 안전하게 만들고 더 빠르게 더 나은 수준에 도달하기를 바란다.
How will the open-closed source ecosystem evolve? In your cognition, how many open-source and closed-source will ultimately remain globally?
오픈-클로즈드 소스 생태계는 어떻게 진화할 것인가? 당신의 인지에서 최종적으로 전 세계에 오픈소스와 클로즈드소스가 몇 개 남을 것인가?
Not many, but there will still be several. If you look at the past two years, this trend is relatively clear — the market gradually becomes more concentrated, more convergent, more focused. Perhaps at the beginning there were hundreds, then dozens, then a few.
많지 않지만, 몇 개는 있을 것이다. 지난 2년을 보면 이 추세가 비교적 명확하다 — 시장이 점차 더 집중되고, 더 수렴되고, 더 초점이 맞춰진다. 처음에는 수백 개였을 수 있고, 수십 개, 그다음 몇 개.
Several, perhaps the final stable number; looking at it now, it is a high-probability matter.
몇 개, 아마도 최종 안정적인 수일 것이다. 지금 보면 높은 확률의 일이다.
Do you belong to the open-source side or the closed-source side?
여러분은 오픈소스 쪽에 속하는가, 클로즈드소스 쪽에 속하는가?
This needs to be observed dynamically; we hope to share more technology in the long term.
이것은 동적으로 관찰해야 한다. 우리는 장기적으로 더 많은 기술을 공유하기를 바란다.
Why have most Chinese companies open-sourced?
왜 중국 회사 대부분이 오픈소스했는가?
Objectively speaking, there are market game factors. But this is a good thing for the community.
객관적으로 말하면 시장 게임 요인이 있다. 하지만 이것은 커뮤니티에 좋은 일이다.
How do you view products in the AI era? What is different about making AI products compared to making mobile internet products? — You used to like saying “model is product.”
AI 시대의 제품을 어떻게 보는가? AI 제품을 만드는 것과 모바일 인터넷 제품을 만드는 것은 무엇이 다른가? — 당신은 이전에 “모델이 제품이다”라고 말하기를 좋아했다.
I can only talk about AI products; I have not done mobile internet products.
나는 AI 제품에 대해서만 말할 수 있다. 모바일 인터넷 제품은 해보지 않았다.
(Model is product) There is no change now. When you make an Agent product, you need to combine the model with tools and Context. But you will find that when training the model, you basically have to build this entire system before you can train this model.
(모델이 제품이다) 지금은 변화가 없다. Agent 제품을 만들 때 모델과 도구, Context를 결합해야 한다. 하지만 모델을 훈련할 때 기본적으로 이 전체 시스템을 구축해야 모델을 훈련할 수 있다는 것을 알게 될 것이다.
Once the model training is completed, the product is basically completed. Making some interaction improvements on this basis of course has value, but that is an icing-on-the-cake step.
모델 훈련이 완료되면 제품도 기본적으로 완료된다. 이 기반 위에서 일부 상호작용 개선을 하는 것은 물론 가치가 있지만, 그것은 금상첨화의 단계다.
Your model performance has already been polished during training and has very good adaptation with tools and environments — that is, the product is completed during the training process.
모델 성능은 훈련 중에 이미 다듬어졌고 도구와 환경과 매우 잘 적응되어 있다 — 즉, 제품은 훈련 과정에서 완성된다.
Last year you mentioned that the current development method has evolved into — you need to make a huge system, just like Google making the search engine system at the beginning of the 20th century. Today, do you have more imagination for the huge systems of the AI era?
작년에 당신은 현재의 개발 방식이 — 20세기 초 Google이 검색 엔진 시스템을 만든 것처럼 거대한 시스템을 만들어야 하는 것으로 진화했다고 언급했다. 오늘 AI 시대의 거대한 시스템에 대해 더 많은 상상이 있는가?
The complexity of the current system lies in that you want this model to become general. On one hand it becomes simpler, on the other hand it becomes more complex.
현재 시스템의 복잡성은 이 모델을 일반적으로 만들고 싶다는 데 있다. 한편으로는 더 단순해지고, 다른 한편으로는 더 복잡해진다.
Simpler in that you only need to put everything in the same model; you do not need to maintain so many models, nor do you need to set up a bunch of routing strategies. Conceptually, or from the engineering implementation perspective, it becomes simpler.
더 단순한 점은 모든 것을 같은 모델에 넣으면 된다는 것이다. 그렇게 많은 모델을 유지할 필요도 없고, 많은 라우팅 전략을 설정할 필요도 없다. 개념적으로, 또는 엔지니어링 구현 관점에서 더 단순해진다.
But at the same time, it also becomes complex. If you hope it is general, you hope this model can work in various scenarios. For example, when you make an Agent model, you do not hope it only works in your tool set, but hope that when others use this model, even with other tool sets, or even tools you have never seen, or tools with different definitions and implementations, it can still work. This requirement is very high.
하지만 동시에 복잡해지기도 한다. 일반적이기를 바란다면 이 모델이 다양한 시나리오에서 작동하기를 바란다. 예를 들어 Agent 모델을 만들 때 당신의 도구 집합에서만 작동하기를 바라지 않고, 다른 사람들이 이 모델을 사용할 때 다른 도구 집합, 심지어 본 적 없는 도구, 또는 정의와 구현 방식이 다른 도구에서도 작동하기를 바란다. 이 요구사항은 매우 높다.
Like in current Agents, there may be several different types of tasks, whether Coding Agent, Search Agent, or other Agents; when you put them into the same general model, there may be fighting problems. Perhaps the tool definitions are different, or the data patterns are different.
현재 Agent에서 Coding Agent, Search Agent 또는 다른 Agent처럼 여러 다른 유형의 작업이 있을 수 있다. 그것들을 같은 일반 모델에 넣으면 충돌 문제가 있을 수 있다. 도구 정의가 다르거나 데이터 패턴이 다를 수 있다.
That is, the process of making it into a general model has many technical challenges.
즉, 그것을 일반 모델로 만드는 과정에 많은 기술적 도전이 있다.
But if you do not make it a general model, its generalization is not that good, and it can only do one thing. Especially current Agents need many steps to complete a task. Even programmers are not only writing code; even when writing code, they are not only doing SWE-bench. To make something very general and truly usable, as the number of steps increases, the requirement for generality becomes higher.
하지만 일반 모델로 만들지 않으면 일반화가 그렇게 좋지 않고 한 가지 일만 할 수 있다. 특히 현재 Agent는 작업을 완성하는 데 많은 단계가 필요하다. 프로그래머조차 코드만 쓰는 것이 아니며, 코드를 쓸 때도 SWE-bench만 하는 것이 아니다. 매우 일반적이고 진정으로 사용 가능한 것을 만들려면 단계 수가 늘어날수록 보편성에 대한 요구가 높아진다.
Its system complexity is reflected in the process of training the model: making this model general enough, rather than only fitting to certain single-point abilities. If you only fit single-point abilities, the Benchmark scores may look good, but the generality is not enough — this is a relatively large challenge I can observe in this system now.
그 시스템 복잡성은 모델을 훈련하는 과정에서 반영된다: 이 모델을 충분히 일반적으로 만드는 것이지, 특정 단일 지점 능력에만 맞추는 것이 아니다. 단일 지점 능력에만 맞추면 Benchmark 점수는 좋아 보일 수 있지만 보편성이 부족하다 — 이것이 지금 이 시스템에서 관찰할 수 있는 비교적 큰 도전이다.
One example is: if you want to add multimodal capability into the model, you need to make this multimodal capability not damage its “brain.”
한 가지 예는: 모델에 멀티모달 능력을 추가하고 싶다면, 이 멀티모달 능력이 그 “뇌”를 손상시키지 않게 해야 한다.
Can multimodality only achieve not damaging?
멀티모달은 손상시키지 않는 정도까지만 가능한가?
Yes, being able not to damage is already very good.
맞다. 손상시키지 않을 수 있는 것만으로도 이미 매우 좋다.
You hope that in multimodal mode and in text mode, they share the same “brain”; you hope that in multimodal mode, it can also stimulate the intelligence of the text part, rather than entering another part of parameters, where it may completely lose the part learned from text.
멀티모달 모드와 텍스트 모드에서 같은 “뇌”를 공유하기를 바란다. 멀티모달 모드에서도 텍스트 부분의 지능을 자극할 수 있기를 바라며, 다른 매개변수 부분으로 들어가서 텍스트에서 학습한 부분을 완전히 잃어버리는 것이 아니기를 바란다.
When you make a general model, you will face such challenges. When you have various modalities, various task types, plus Agent, Reasoning, Chat, and so on, to fuse them all together, there are challenges.
일반 모델을 만들 때 이런 도전에 직면한다. 다양한 모달리티, 다양한 작업 유형, 그리고 Agent, Reasoning, Chat 등을 모두 하나로 융합하는 데 도전이 있다.
And now it is not only doing SFT, but also doing RL; the challenge is further aggravated.
그리고 지금은 SFT만 하는 것이 아니라 RL도 한다. 도전이 더욱 가중된다.
General Pre-Training is relatively easy to do; you just put all the text together, and there will not be too many problems. But the later you go into Post-Train, the more into RL, this problem becomes more severe — this is its system complexity.
일반 Pre-Training은 비교적 하기 쉽다. 모든 텍스트를 함께 넣으면 큰 문제가 없다. 하지만 Post-Train 후반으로 갈수록, RL로 갈수록 이 문제가 더 심각해진다 — 이것이 그 시스템 복잡성이다.
This requires creating a new interaction paradigm.
이것은 새로운 상호작용 패러다임을 창조할 필요가 있다.
Yes. But this interaction also needs to adapt to the development of model capabilities. Your interaction cannot exceed the model’s capabilities; it should design a good interaction within the current range of model capabilities.
맞다. 하지만 이 상호작용은 모델 능력의 발전에 맞춰야 한다. 당신의 상호작용이 모델 능력을 초과해서는 안 된다. 현재 모델 능력 범위 내에서 좋은 상호작용을 설계해야 한다.
This is worth trying. Just looking at today, scaling the FLOPs dimension or improving learning efficiency is a method with higher certainty and more effectiveness.
이것은 시도할 가치가 있다. 단지 오늘을 보면 FLOPs 차원을 scale하거나 학습 효율을 향상시키는 것이 더 높은 확실성과 더 효과적인 방법이다.
If according to Yan Junjie (founder and CEO of MiniMax)’s statement, user data cannot improve the model’s intelligence, then is it unnecessary to do To C products today, and just focus wholeheartedly on improving intelligence?
闫俊杰(MiniMax 창업자 겸 CEO)의 말에 따르면 사용자 데이터가 모델의 지능을 향상시킬 수 없다면, 오늘은 To C 제품을 할 필요가 없고 오로지 지능 향상에만 전념하면 되는가?
This depends on how you understand it. You may not be able to directly use user feedback for training; but the benefit of having a certain user volume is that you know what the demand distribution is like, know where users use it well or poorly, and can abstract these things into evaluation, then optimize the model. If the model is completely unused by anyone, you do not know in which direction to optimize.
이것은 어떻게 이해하느냐에 달려 있다. 사용자 피드백을 직접 훈련에 사용할 수 없을 수도 있다. 하지만 일정 사용자 규모가 있는 이점은 수요 분포가 어떤지 알고, 사용자가 잘 사용하는 곳과 그렇지 않은 곳을 알며, 이것들을 evaluation으로 추상화해서 모델을 최적화할 수 있다는 것이다. 모델이 완전히 아무도 사용하지 않으면 어느 방향으로 최적화해야 할지 모른다.
In addition, it also depends on the commercial value of users. Now it has reached a new watershed: users are possible to generate commercial value. Look at OpenAI; C-end users have generated large commercial value, accounting for a relatively large proportion of its revenue.
또한 사용자의 상업적 가치도 봐야 한다. 지금은 새로운 분수령에 도달했다: 사용자가 상업적 가치를 생성할 가능성이 있다. OpenAI를 보라. C단 사용자가 큰 상업적 가치를 생성했으며, 그 수익의 비교적 큰 비율을 차지한다.
Especially now many Agent products can generate value end-to-end, so it also depends on what kind of users you have. If it is only chatting or checking the weather, the commercial value is not that large. But if they are professional Agent users, they themselves have very good productivity value.
특히 지금 많은 Agent 제품이 엔드투엔드로 가치를 생성할 수 있으므로, 어떤 사용자인지도 봐야 한다. 단지 잡담이나 날씨 확인이라면 상업적 가치가 그렇게 크지 않다. 하지만 Agent 전문 사용자라면 그 자체로 매우 좋은 생산성 가치가 있다.
In the recent year, what new thoughts do you have on C-end products?
최근 1년, C단 제품에 대해 어떤 새로운 생각이 있는가?
More still thinking about how to make the model, because once the model is trained well, the product is basically done. We will continue to do it along this way.
더 많은 것은 여전히 모델을 어떻게 만들지에 대해 생각한다. 모델이 잘 훈련되면 제품은 기본적으로 완성되기 때문이다. 우리는 이 방식으로 계속할 것이다.
Some people say that Kimi has shifted from initially wanting to be “China’s OpenAI” — of course you previously did not agree with this statement — to wanting to be “China’s Anthropic.” Has there been such a positioning shift internally?
어떤 사람들은 Kimi가 처음 “중국의 OpenAI”를 하고 싶어했던 것에서 — 물론 당신은 이전에 이 말에 동의하지 않았다 — “중국의 Anthropic”을 하고 싶어하는 것으로 전환했다고 말한다. 내부적으로 그런 포지셔닝 전환이 있었나?
It is hard to define in this way. The contexts and soils of China and the US are different; today we think more from a global perspective. “Being China’s so-and-so” does not really hold.
이런 방식으로 정의하기 어렵다. 중국과 미국의 맥락과 토양이 다르다. 오늘은 더 글로벌 관점에서 문제를 생각한다. “중국의 모모”가 되는 것은 그다지 성립하지 않는다.
Actually simpler: we hope to continue climbing the mountain, be a friend of time, and accelerate the advancement of technology together with the community.
사실 더 간단하게: 우리는 계속 산을 오르고, 시간의 친구가 되며, 커뮤니티와 함께 기술 추진을 가속화하기를 바란다.
As a Founder, what is your current life rhythm like?
Founder로서 지금 생활 리듬은 어떤가?
Perhaps I sleep relatively late, haha, every day is different.
비교적 늦게 자는 것 같다, 하하, 매일 다르다.
But it’s still okay; I spend a lot of time looking at how to train the model better.
하지만 괜찮다. 모델을 어떻게 더 잘 훈련할지에 많은 시간을 쓴다.
Is your time mainly invested in model training?
당신의 시간은 주로 모델 훈련에 투입되는가?
Yes, but model training is an abstract concept; what is important is technical strategy, which is the most critical part of the company strategy — what to do next, what not to do, because the technical space is large, you always have to select some directions for focused investment.
맞다. 하지만 모델 훈련은 추상적인 개념이다. 중요한 것은 기술 전략이며, 이것이 회사 전략에서 가장 핵심적인 부분이다 — 다음에 무엇을 할지, 무엇을 하지 않을지. 기술 공간이 크기 때문에 항상 일부 방향을 선택해서 집중 투자해야 한다.
Our bets in many directions were relatively early and effective. We did long CoT RL very early, reacted relatively quickly; did the optimizer; did larger-scale Pre-Training; did the first Open Agentic model — these are all key technical decisions. These decisions can determine 50-60% of the company’s direction.
많은 방향에서의 우리 베팅은 비교적 일찍이었고 효과적이었다. 우리는 long CoT RL을 매우 일찍 했고, 비교적 빠르게 반응했다; 최적화기를 했고; 더 대규모 Pre-Training을 했고; 첫 번째 Open Agentic 모델을 했다 — 이것들은 모두 핵심 기술 결정이다. 이 결정들이 회사 방향의 50-60%를 결정할 수 있다.
But to make good decisions, you need a lot of evidence, and also need to do many experiments. You need to understand the specific results of experiments very well; you cannot just decide off the top of your head; you need to know more information.
하지만 좋은 결정을 하려면 많은 증거가 필요하고, 또한 많은 실험을 해야 한다. 실험의 구체적 결과를 매우 잘 이해해야 한다. 머리로만 결정할 수 없다. 더 많은 정보를 알아야 한다.
Among these decisions, which one made you the most tangled?
이 결정들 중 당신을 가장 고민하게 만든 것은 무엇인가?
It’s still okay. The key is a process of collecting data. Do experiments, see if the experiments are solid. Plus your understanding of technology to judge. Many times as long as the data is sufficient enough, the judgment is relatively obvious.
그래도 괜찮다. 핵심은 데이터를 수집하는 과정이다. 실험을 하고, 실험이 탄탄한지 본다. 그리고 기술에 대한 이해를 더해 판단한다. 많은 경우 데이터가 충분히 충분하면 판단이 비교적 명백하다.
Next — at least now, the performance potential of K2 has not been fully squeezed out. What we released before is closer to a Base Model. We can add more FLOPs in the Post-Training stage; the upper limit should be much higher than now.
다음으로 — 적어도 지금, K2의 성능 잠재력은 아직 완전히 짜내어지지 않았다. 우리가 이전에 낸 것은 Base Model에 더 가깝다. Post-Training 단계에서 더 많은 FLOPs를 추가할 수 있다. 상한은 지금보다 훨씬 높을 것이다.
We will also do the next-generation model, but specifically how to do it, we decide through experiments.
우리는 다음 세대 모델도 할 것이다. 하지만 구체적으로 어떻게 할지는 실험을 통해 결정한다.
Will multimodality also be added?
멀티모달도 추가할 것인가?
Multimodality is relatively certain.
멀티모달은 비교적 확실하다.
But doing the multimodal capability itself well is not easy. There is a lot of work inside: how to let it borrow the text brain, rather than opening a separate brain by itself. For example, if in MoE there are 20 experts specially doing multimodality, you may not hope this situation appears — this way, the multimodality you learn may be a “stupid multimodality.”
하지만 멀티모달 능력 자체를 잘 하는 것은 쉽지 않다. 안에 많은 작업이 있다: 텍스트 뇌를 차용하게 하는 방법, 스스로 별도의 뇌를 여는 것이 아니라. 예를 들어 MoE에 멀티모달을 전문으로 하는 전문가 20개가 있다면, 이런 상황이 나타나기를 바라지 않을 수 있다 — 이렇게 하면 학습한 멀티모달이 “바보 멀티모달”이 될 수 있다.
We hope it is a “smart multimodality.”
우리는 그것이 “똑똑한 멀티모달”이 되기를 바란다.
What other important technical milestones will there be next?
다음에 어떤 다른 중요한 기술 마일스톤이 있을 것인가?
The generalization of Agent is the most important.
Agent의 일반화가 가장 중요하다.
Support for Long Context, we will continue to research. Especially under the condition of high intelligence, still being able to have longer Context, is also a very important problem.
Long Context 지원, 우리는 계속 연구할 것이다. 특히 지능이 높은 상황에서 여전히 더 긴 Context를 가질 수 있는 것도 매우 중요한 문제다.
Now many Long Context architectures still affect the “intelligence.”
지금 많은 Long Context 아키텍처는 여전히 “지능”에 영향을 미친다.
Why does the Long Context architecture affect “intelligence”?
왜 Long Context 아키텍처가 “지능”에 영향을 미치는가?
Pure Linear Attention may affect intelligence, because this architecture has some biases, and these biases do not perform that well in some scenarios. But to a certain extent it can be solved.
순수한 Linear Attention은 지능에 영향을 미칠 수 있다. 이 아키텍처에 일부 bias가 있기 때문이며, 이 bias들이 일부 시나리오에서 그렇게 잘 작동하지 않는다. 하지만 어느 정도는 해결될 수 있다.
How do you view Zhang Xiangyu (chief scientist of StepFun) saying about the essential defect of next token prediction?
张祥雨(阶跃星辰 수석 과학자)가 말한 next token prediction의 본질적 결함에 대해 어떻게 보는가?
His meaning is that as the model scale expands, dialogue ability, knowledge volume, and emotional intelligence are all becoming stronger, but reasoning ability especially data performance rises first then plateaus, and expanding further instead declines. Using a larger model to do math problems, it is easy to skip steps and not be honest. This is the essential defect of next token prediction.
그의 뜻은 모델 규모가 확대되면서 대화 능력, 지식량, 감성 지능이 모두 강해지지만, 추론 능력 특히 데이터 성능은 먼저 상승한 후 평탄해지고, 더 확대하면 오히려 하락한다는 것이다. 더 큰 모델로 수학 문제를 하면 단계를 건너뛰고 정직하지 않기 쉽다. 이것이 next token prediction의 본질적 결함이다.
So it needs to be paired with reinforcement learning scaling. If today you do not do reinforcement learning, it is hard to say the model is very smart. Like math problems, it may not do them very well.
그래서 강화학습 scaling과 짝을 이뤄야 한다. 오늘 강화학습을 하지 않으면 모델이 매우 똑똑하다고 말하기 어렵다. 수학 문제처럼 매우 잘하지 않을 수 있다.
But a larger base, the upper limit of reinforcement learning will be higher. Because there is more knowledge, the essence is activating a reasoning paradigm, letting it unlock the knowledge. Its upper limit is higher, but it needs to be paired with RL to activate.
하지만 더 큰 base는 강화학습의 상한이 더 높아질 것이다. 지식이 더 많기 때문에, 본질은 추론 패러다임을 활성화해서 지식을 잠금 해제하게 하는 것이다. 상한이 더 높지만, RL과 짝을 이뤄 활성화해야 한다.
How do you view world models? — Some people say making world models is creating the world, making Agents is creating humans.
세계 모델을 어떻게 보는가? — 어떤 사람들은 세계 모델을 만드는 것은 세계를 창조하는 것이고, Agent를 만드는 것은 인간을 창조하는 것이라고 말한다.
Using AI to train AI has a bit of this meaning. If you have a good world model, you can simulate these things; it is a way of using AI to train AI.
AI로 AI를 훈련하는 것은 이런 뜻이 조금 있다. 좋은 세계 모델이 있으면 이런 것들을 시뮬레이션할 수 있다. AI로 AI를 훈련하는 방식이다.
It may be a path leading to better generalization.
그것은 더 나은 일반화로 가는 경로일 수 있다.
Now, let’s discuss some practical issues.
이제 일부 현실적인 문제를 논의해 보자.
Where is the long-term boundary between base model companies and application companies making Agent products?
기초 모델 회사와 Agent 제품을 만드는 애플리케이션 회사의 장기적 경계는 어디에 있는가?
I do not have a definite answer. I can only say that today, “first-party products” have one advantage: vertical integration. You put the model inside and train it; the model and the tools become one, rather than being built separately and then reverse-engineered.
나는 명확한 답이 없다. 오늘 “1차 제품”에 한 가지 이점이 있다고만 말할 수 있다: 수직 통합. 모델을 안에 넣고 훈련한다. 모델과 도구가 하나가 되며, 따로 만들고 역공학하는 것이 아니다.
But because the Agent field is broad, “first-party products” may not be able to cover everything. If some space can be found, for example the implementation of tools requires a lot of domain know-how, or evaluation is something that “first-party products” cannot consider, there is opportunity.
하지만 Agent 분야가 넓기 때문에 “1차 제품”이 모든 것을 커버하지 못할 수 있다. 일부 공간을 찾을 수 있다면, 예를 들어 도구 구현에 많은 영역 know-how가 필요하거나, evaluation이 “1차 제품”이 고려하지 못하는 것이라면 기회가 있다.
Because there are open-source models like K2, everyone can fine-tune on top of them, making it easier to produce Specialized Agents and vertical Agents.
K2 같은 오픈소스 모델이 있기 때문에 모두가 그 위에서 fine-tune할 수 있고, Specialized Agent와 수직 Agent를 더 쉽게 만들 수 있다.
That depends on how general the general model becomes.
그것은 일반 모델이 얼마나 일반화되느냐에 달려 있다.
No matter how general, there will always be some tools you have to build. You may not necessarily make the model, but make the tools very well.
아무리 일반적이라도 항상 만들어야 할 도구가 있다. 모델을 만들지 않고 도구를 매우 잘 만들 수도 있다.
If this tool is made too general, it will have a large overlap with “first-party products” — in this case, the advantage of vertical integration is greater.
이 도구를 너무 일반적으로 만들면 “1차 제품”과 크게 겹친다 — 이런 경우 수직 통합의 이점이 더 크다.
But if the tool is specifically targeted at a certain scenario, even others cannot make it. For example, if you master some offline service entrances, your order-placing or deal-closing tools others cannot make, you may generate unique value.
하지만 도구가 특정 시나리오를 특별히 대상으로 하고, 심지어 다른 사람이 만들 수 없다면. 예를 들어 일부 오프라인 서비스 입구를 장악하고, 주문이나 거래 체결 도구를 다른 사람이 만들 수 없다면, 독특한 가치를 생성할 수 있다.
Of course there is another possibility: when the traffic and business model of “first-party products” or general Agents are mature enough, many proprietary, originally monopolized tools are also willing to connect, because the overall commercialization efficiency will be higher. But the improvement of commercialization efficiency needs time. Within this time window, proprietary Agents will also have space.
물론 또 다른 가능성이 있다: “1차 제품”이나 일반 Agent의 트래픽과 비즈니스 모델이 충분히 성숙하면, 많은 전용, 원래 독점된 도구도 연결을 원할 것이다. 전체 상업화 효율이 더 높아지기 때문이다. 하지만 상업화 효율 향상에는 시간이 필요하다. 이 시간 창 안에서 전용 Agent도 공간을 가질 것이다.
Ultimately, the reason generality is effective is because the overall commercialization efficiency is higher. Today, including many content platforms, ultimately it is possible that connecting content to general Agents will have higher commercialization efficiency than today — but it may take a long time.
궁극적으로 일반성이 효과적인 이유는 전체 상업화 효율이 더 높기 때문이다. 오늘 많은 콘텐츠 플랫폼을 포함해서, 최종적으로 콘텐츠를 일반 Agent에 연결하면 오늘보다 상업화 효율이 더 높아질 가능성이 있다 — 하지만 오랜 시간이 걸릴 수 있다.
Companies like Manus, will they be your potential customers or competitors?
Manus 같은 회사는 당신의 잠재 고객인가, 경쟁자인가?
It is still very early; it is hard to judge exactly what it is; the product itself will also evolve.
아직 매우 초기이다. 정확히 어떤지 판단하기 어렵다. 제품 자체도 진화할 것이다.
In the short term, more cooperation than competition. Today you can also see the figure of K2 in Cursor, Perplexity, Genspark.
단기적으로는 경쟁보다 협력이 더 많다. 오늘 Cursor, Perplexity, Genspark에서도 K2의 모습을 볼 수 있다.
But the future will evolve, a bit like the relationship between Claude and Cursor. Cursor may also need to dynamically adjust its product strategy. On one hand it may need a certain level of model capability, because the technology curve is still steep; on the other hand, can it have tools or environments that others cannot build?
하지만 미래는 진화할 것이다. Claude와 Cursor의 관계와 좀 비슷하다. Cursor도 제품 전략을 동적으로 조정해야 할 수 있다. 한편으로는 일정 수준의 모델 능력이 필요할 수 있다. 기술 곡선이 여전히 가파르기 때문이다. 다른 한편으로는 다른 사람이 만들 수 없는 도구나 환경을 가질 수 있는가?
It cannot be directly answered now. I can only say that currently, the integration advantage of first-party still exists.
지금 직접 답할 수 없다. 현재로서는 1차의 통합 이점이 여전히 존재한다고만 말할 수 있다.
How do you think about the business model today? Is API a good business?
오늘 비즈니스 모델을 어떻게 생각하는가? API는 좋은 사업인가?
Currently clear business models: one is API service, two is “first-party products.” We will all make some attempts. Today the main priority is still to make the model better; this remains the primary goal.
현재 명확한 비즈니스 모델: 하나는 API 서비스, 둘은 “1차 제품”. 우리는 모두 일부 시도를 할 것이다. 오늘 주요 우선순위는 여전히 모델을 더 좋게 만드는 것이다. 이것이 여전히 주요 목표다.
In the process of model improvement, if leading in some aspects, there is indeed commercialization space. Today the market scale is growing very fast; leading companies have tens of billions or even over a hundred billion USD ARR, possibly achieving two or three times growth every one or two quarters. We will observe dynamically and make corresponding attempts.
모델 향상 과정에서 일부 측면에서 선도하면 확실히 상업화 공간이 있다. 오늘 시장 규모가 매우 빠르게 성장하고 있다. 선도 회사들은 수십억 또는 심지어 천억 달러 ARR을 가지고 있으며, 한두 분기마다 두세 배 성장을 달성할 수 있다. 우리는 동적으로 관찰하고 해당 시도를 할 것이다.
One of your users said he likes Kimi very much, but also worries that Kimi cannot make money. Can you make money?
여러분의 사용자 한 명이 Kimi를 매우 좋아하지만, Kimi가 돈을 벌지 못할까 걱정한다고 말했다. 여러분은 돈을 벌 수 있는가?
Still invest first. Whether we can make money depends on how the model effect is.
여전히 먼저 투자한다. 돈을 벌 수 있는지는 모델 효과가 어떤지에 달려 있다.
We are also willing to serve the last-mile experience of users, delivering high-value problem deliverables for users.
우리는 또한 사용자의 마지막 1마일 경험을 서비스하고, 사용자에게 고가치 문제의 deliverable을 제공하기를 원한다.
In the global AI market of tens of billions of USD with high-speed growth, focusing on doing the technology well, other things instead have more certainty.
전 세계 수백억 달러의 고속 성장 AI 시장에서 기술을 잘 하는 데 집중하면, 다른 것들은 오히려 더 확실성이 있다.
In the organization internally, how do you define reward?
조직 내부에서 reward를 어떻게 정의하는가?
You establish more observation metrics, and try as much as possible not to overfit; this is effective to a certain degree. This way you can have better generalization and will not be hacked.
더 많은 관찰 지표를 세우고, 가능한 한 과적합하지 않도록 한다. 이것은 어느 정도 효과적이다. 이렇게 하면 더 나은 일반화를 가질 수 있고, 해킹당하지 않는다.
The biggest problem of managing a team with RL is that you are easily hacked. Everyone looks like various results are very good, but in reality it has not achieved what you ultimately want — this is the risk. The risk of managing a team with SFT is that everyone loses creativity. In the end these several things need a certain degree of balance.
RL로 팀을 관리하는 가장 큰 문제는 쉽게 해킹당한다는 것이다. 모두가 다양한 결과가 매우 좋아 보이지만, 실제로는 당신이 궁극적으로 원하는 것을 달성하지 못했다 — 이것이 위험이다. SFT로 팀을 관리하는 위험은 모두가 창의력을 잃는 것이다. 결국 이 몇 가지 것들은 어느 정도의 균형이 필요하다.
Of course I am also learning; today it is not done very perfectly.
물론 나도 학습 중이다. 오늘은 매우 완벽하게 하지 못했다.
Over the past year Kimi has oscillated back and forth between peaks and valleys. Being in the midst of it, what is your mentality like? Do you need to balance your own mentality?
지난 1년 Kimi가 파고와 파저 사이를 왔다 갔다 했다. 그 가운데에 있으면서 당신의 마음 상태는 어떤가? 자신의 마음 상태를 균형 잡을 필요가 있는가?
The mentality is just to be a friend of time.
마음 상태는 그냥 시간의 친구가 되는 것이다.
A real person’s mentality is not this simple.
진짜 사람의 마음 상태는 이렇게 간단하지 않다.
As you said, there will be high points and low points. What is perhaps very important is still liking to do this thing and wanting to do it well. So you do not need to think about anything else; just think about how to do it well. It seems relatively simple; there is nothing particularly complex.
당신이 말한 것처럼 높은 점과 낮은 점이 있을 것이다. 아마도 매우 중요한 것은 여전히 이 일을 좋아하고 잘하고 싶다는 것이다. 그래서 다른 것을 생각할 필요 없다. 어떻게 잘할지에만 생각하면 된다. 비교적 간단해 보인다. 특별히 복잡한 것은 없다.
A lot of complexity is artificially forced on; in reality it is not that complex.
많은 복잡성은 인위적으로 억지로 더한 것이다. 실제로는 그렇게 복잡하지 않다.
Do you have more understanding of human nature?
인간성에 대해 더 많은 이해가 있는가?
This still needs the polishing of time; I dare not say it is that deep.
이것도 여전히 시간의 연마가 필요하다. 그렇게 깊다고 감히 말할 수 없다.
I can only say that it is inside one’s own story — you continuously feel what kind of person you really are, why you want to do this thing — continuously thinking about these questions.
자신의 이 이야기 안에 있다고만 말할 수 있다 — 당신은 자신이 도대체 어떤 사람인지, 왜 이 일을 하는지 계속 느끼고 — 이 질문들을 계속 생각한다.
What kind of person are you? Why do you want to do this kind of thing?
당신은 어떤 사람인가? 왜 이런 일을 하고 싶은가?
Just because it is interesting.
그냥 재미있기 때문이다.
What is interesting? Is doing experiments interesting, doing scientific research interesting, or doing AI interesting?
무엇이 재미있는가? 실험을 하는 것이 재미있는가, 과학 연구를 하는 것이 재미있는가, 아니면 AI를 하는 것이 재미있는가?
The process of seeking the truth. The process of continuously discovering new problems and solving them.
진실을 찾는 과정. 새로운 문제를 계속 발견하고 그것을 해결하는 과정.
Then you can also solve other problems; why must you solve this problem?
그러면 다른 문제도 해결할 수 있다. 왜 반드시 이 문제를 해결해야 하는가?
Because this thing is very important; AI is very important.
이 것이 매우 중요하기 때문이다. AI가 매우 중요하다.
I also asked Kimi this question. It said this thing is “an amplifier of human civilization,” and I think that makes a lot of sense.
나도 Kimi에게 이 질문을 물어봤다. 그것은 이 것이 “인간 문명의 증폭기”라고 말했고, 나는 그것이 매우 이치에 맞는다고 생각한다.
Back again to The Beginning of Infinity: from the Enlightenment to now, humanity has been searching for new methods to break through the boundaries of knowledge. But perhaps the next breakthrough of the boundary will rely on AI; it is a huge lever.
다시 The Beginning of Infinity로 돌아간다: 계몽운동부터 지금까지 인류는 지식의 경계를 돌파할 새로운 방법을 계속 찾아왔다. 하지만 다음 경계 돌파는 AI에 의존할 수 있다. 그것은 거대한 지렛대다.
Today in any frontier discipline, it takes twenty or thirty years to learn the most frontier knowledge. But AI can learn it overnight and then go on to make new breakthroughs.
오늘 어떤 최전선 학문이든 가장 최전선 지식을 배우는 데 이삼십 년이 걸린다. 하지만 AI는 하룻밤 사이에 배울 수 있고, 그다음 새로운 돌파를 해나간다.
AI will become Meta science.
AI는 Meta science가 될 것이다.
It is an amplifier of human civilization.
그것은 인간 문명의 증폭기다.
Is it possible that it destroys human civilization?
그것이 인간 문명을 파괴할 가능성이 있는가?
This risk cannot be said not to exist, but there are many things we can do. Whether safer alignment or better social mechanisms.
이 위험이 존재하지 않는다고 말할 수 없지만, 우리가 할 수 있는 많은 일이 있다. 더 안전한 정렬이든 더 나은 사회 메커니즘이든.
For example, when AI can do some things, it is very likely to create some new jobs; we need some methods to complete this transition.
예를 들어 AI가 어떤 일을 할 수 있을 때, 그것은 매우 새로운 직업을 창조할 가능성이 높다. 우리는 이 전환을 완성할 일부 방법이 필요하다.
I exactly asked Kimi this question. It said that although there is such a risk, we cannot give up. Because if we give up, it is equivalent to giving up the upper limit of human civilization — you do not know what the upper limit can achieve; there is a bit of a feeling of “giving up eating for fear of choking.”
나는 정확히 Kimi에게 이 질문을 물어봤다. 그것은 비록 그런 위험이 있지만 우리는 포기할 수 없다고 말했다. 포기하면 인간 문명의 상한을 포기하는 것과 같기 때문이다 — 상한이 어디까지 갈 수 있는지 모른다. “질식할까 봐 먹는 것을 포기하는” 느낌이 조금 있다.
But I admit that we have to do a lot to cope. Because the many AI capabilities you see today are a bit shocking; half a year ago you could not even imagine it could do them.
하지만 나는 우리가 대처하기 위해 많은 일을 해야 한다는 것을 인정한다. 오늘 보는 많은 AI 능력이 조금 충격적이기 때문이다. 반년 전에는 그것이 할 수 있을 거라고 상상조차 하지 못했다.
At the same time I believe that the unique value of humans will continue to exist in this process. Human experience and emotion cannot be replaced by AI. So there may be different ways of living; I hope it is being able to live better.
동시에 나는 인간의 독특한 가치가 이 과정에서 계속 존재할 것이라고 믿는다. 인간의 경험과 감정은 AI로 대체될 수 없다. 그래서 다른 삶의 방식이 있을 수 있다. 더 잘 살 수 있기를 바란다.
What kind of different ways of living?
어떤 다른 삶의 방식인가?
I previously felt that a person’s life has several meanings: creation, experience, and love. Of course everyone is different; for me it is like this.
나는 이전에 사람의 삶에는 몇 가지 의미가 있다고 느꼈다: 창조, 경험, 그리고 사랑. 물론 사람마다 다르다. 나에게는 이렇다.
A large part of “creation” perhaps AI can do. I enjoy this process, but I have to admit that one day many creative works will be done by AI. But the latter two, “experience” and “love,” will be human-centered.
“창조”의 큰 부분은 아마도 AI가 할 수 있다. 나는 이 과정을 즐기지만, 언젠가 많은 창조적 작업이 AI에 의해 이루어질 것이라는 것을 인정해야 한다. 하지만 후자 두 가지, “경험”과 “사랑”은 인간을 중심으로 할 것이다.
If AI takes away creation, it also takes away productivity.
AI가 창조를 가져가면 생산성도 가져간다.
It does not matter; people can enjoy the results of production, if we have good mechanisms.
상관없다. 좋은 메커니즘이 있다면 사람들이 생산의 결과를 즐길 수 있다.
But it is a slow process; it will not be finished in one or two years; it needs ten or twenty years of gradual adjustment.
하지만 그것은 느린 과정이다. 일 이년 안에 끝나지 않는다. 십이십 년의 점진적 조정이 필요하다.
Do you frequently chat with Kimi?
Kimi와 자주 채팅하는가?
Of course. I need to test the model.
물론. 모델을 테스트해야 한다.
Do you chat with it about some very profound topics or self-exploration topics?
그것과 매우 심오한 주제나 자기 탐색 주제를 채팅하는가?
Sometimes yes. It’s still okay; there are also some work-related questions.
가끔 그렇다. 그래도 괜찮다. 일부 업무 관련 질문도 있다.
Over the past year, have you experienced moments when your emotions were very low?
지난 1년, 감정이 매우 낮아진 순간을 경험한 적이 있는가?
I feel it’s still okay. More is: some things will work, some things will not work; some problems will be solved, new problems will arise — continuously in this process.
나는 그래도 괜찮다고 느낀다. 더 많은 것은: 어떤 것은 작동하고, 어떤 것은 작동하지 않는다; 일부 문제가 해결되고, 새로운 문제가 생긴다 — 이 과정에서 계속.
As long as you feel this thing is interesting, you always want to continue doing it.
이 것이 재미있다고 느끼는 한, 계속 하고 싶다.
Over the past year, have you taken any detours?
지난 1년, 어떤 우회로를 간 적이 있는가?
Definitely yes. In the process there will be many many decisions, some technical decisions, some business decisions.
확실히 있다. 과정에서 매우 많은 결정이 있을 것이다. 일부 기술 결정, 일부 비즈니스 결정.
What is very important is a company’s ability to gradually adjust in this process. The process of knowledge creation is also such a process — it is impossible that the knowledge created is all correct; you will find that some things are also wrong, but it may be correct within a certain time, and may be wrong again after a certain time. But when it is wrong, you have to make adjustments.
매우 중요한 것은 이 과정에서 회사가 점진적으로 조정하는 능력이다. 지식 창조의 과정도 그런 과정이다 — 창조된 지식이 모두 옳을 수는 없다. 어떤 것도 틀릴 수 있다는 것을 알게 될 것이다. 하지만 일정 시간 내에는 옳을 수 있고, 일정 시간 후에는 다시 틀릴 수 있다. 하지만 틀렸을 때 조정을 해야 한다.
For example, many things Newton did were the best theories at the time, but not perfect; in some scenarios they were completely wrong. Universal gravitation needs some other explanations, needs some explanations of relativity, explained through the distortion of space-time.
예를 들어 Newton이 한 많은 것들은 당시 최고의 이론이었지만 완벽하지 않았다. 일부 시나리오에서는 완전히 틀렸다. 만유인력은 일부 다른 설명이 필요하고, 상대성 이론의 일부 설명이 필요하며, 시공간의 왜곡을 통해 설명된다.
I feel that the evolution of the organization and the development of the company are the same; it is a dynamic process. Any intermediate point that is correct at a certain time point may be wrong at another time point.
조직의 진화와 회사의 발전도 같다고 느낀다. 그것은 동적 과정이다. 어떤 중간 점도 어떤 시점에서는 옳지만 다른 시점에서는 틀릴 수 있다.
This is also what Kimi told me — any intermediate state is possible to become an object of criticism. You will always have the limitations of this era.
이것 또한 Kimi가 나에게 말한 것이다 — 어떤 중간 상태도 비판의 대상이 될 수 있다. 당신은 항상 이 시대의 한계를 가질 것이다.
What is more important is how in this process, on one hand you invest in some “unchanging” things, such as talent and technical accumulation; on the other hand adapt and adjust, making adjustments according to environmental changes and feedback signals. Both are very important.
더 중요한 것은 이 과정에서 한편으로는 인재, 기술 축적 같은 일부 “변하지 않는” 것에 투자하고, 다른 한편으로는 적응하고 조정하며, 환경 변화와 피드백 신호에 따라 조정을 하는 것이다. 둘 다 매우 중요하다.
Internet products expand DAU and expand market scale through market promotion; AI products seem somewhat different — growth and customer acquisition rely more on large leaps in model capability — which is more essential, intelligence or promotion?
인터넷 제품은 시장 홍보를 통해 DAU를 확대하고 시장 규모를 확대한다. AI 제품은 다소 다른 것 같다 — 성장과 고객 확보는 모델 능력의 큰 도약에 더 의존한다 — 지능과 홍보 중 어느 것이 더 본질적인가?
It still depends on which of the two variables is larger. In the stage of rapid technological development, it is hard to win the war through market promotion. It is more of an auxiliary means.
여전히 두 변수 중 어느 것이 더 큰지에 달려 있다. 기술 급속 발전 단계에서는 시장 홍보 방식으로 전쟁을 이기기 어렵다. 그것은 더 보조 수단이다.
Just saying what kind of ratio is between this auxiliary means and your main means? It needs dynamic adjustment, and also depends on how strong your current commercialization progress or PMF is. Different time points have different strategies.
단지 이 보조 수단과 주요 수단 사이의 비율이 어떤 것인가? 동적 조정이 필요하며, 현재 상업화 진전이나 PMF가 얼마나 강한지에 달려 있다. 다른 시점에 다른 전략이 있다.
Even this strategy perhaps after another one or two years, you find it is again a good strategy. I feel it is not necessarily so.
심지어 이 전략이 아마 일 이년 후에 다시 좋은 전략이라고 발견할 수도 있다. 반드시 그렇지는 않다고 느낀다.
We look with a more open mentality, but at each time point the most important is to grasp — which is the largest variable.
우리는 더 열린 마음으로 보지만, 각 시점에서 가장 중요한 것은 파악하는 것이다 — 어느 것이 가장 큰 변수인가.
Another year has passed; do you feel the success probability of Kimi has increased or the failure probability has increased?
또 한 해가 지났다. Kimi의 성공 확률이 커졌다고 느끼는가, 실패 확률이 커졌다고 느끼는가?
I feel (the success probability) has increased. As long as you climb up each time, the success probability becomes larger, because some people will not continue climbing.
나는 (성공 확률이) 커졌다고 느낀다. 매번 위로 오르기만 하면 성공 확률이 커진다. 어떤 사람들은 계속 오르지 않을 것이기 때문이다.
Do you fear falling down?
떨어지는 것을 두려워하는가?
Definitely there is fear.
확실히 두려움이 있다.
More need to focus on what you can do in the current step — thinking about this question is more important.
더 현재 이 단계에서 무엇을 할 수 있는지에 집중해야 한다 — 이 질문을 생각하는 것이 더 중요하다.
I have been asking about your emotions, and you always say: ah, still okay, still okay. What was the situation of your most recent one or two “intracranial self-highs”?
나는 계속 당신의 감정을 물어봤는데, 당신은 항상 말한다: 아, 그래도 괜찮다, 괜찮다. 가장 최근 한두 번의 “두개내 자뻑” 상황은 무엇이었나?
Not pleased by external gains, not saddened by personal losses. Although it is hard to achieve, one must avoid emotional decisions.
외부의 이득에 기뻐하지 않고, 개인적인 손실에 슬퍼하지 않는다. 비록 달성하기 어렵지만, 감정적인 결정을 피해야 한다.
Do you get emotional?
당신은 감정적이 되는가?
To some extent definitely yes — you are a human being. But one must avoid some emotional decisions. Ultimately when it comes to decision and execution, one needs to be a bit more rational.
어느 정도는 확실히 그렇다 — 당신은 인간이니까. 하지만 일부 감정적인 결정을 피해야 한다. 최종적으로 결정과 실행에 있어서는 조금 더 이성적일 필요가 있다.
Over the past year, what is your biggest growth?
지난 1년, 당신의 가장 큰 성장은 무엇인가?
Recognizing this point: problems are inevitable; they will always exist; continuously solving new problems is the most important, and perhaps also the most interesting — this is a change in mentality; it will change many ways of doing things.
이 점을 인식한 것: 문제는 불가피하다; 그것들은 항상 존재할 것이다; 새로운 문제를 지속적으로 해결하는 것이 가장 중요하며, 아마도 가장 흥미롭기도 하다 — 이것은 마음 상태의 변화다; 그것은 많은 일 처리 방식을 바꿀 것이다.
This sounds like a kind of mindfulness.
이것은 일종의 마음챙김처럼 들린다.
I do not know how to understand it, but perhaps about the same. (laughs)
어떻게 이해해야 할지 모르겠지만, 아마도 비슷하다. (웃음)
I ask you the last few rapid Q&A.
마지막 몇 가지 빠른 질문과 답을 묻겠다.
One food you like globally.
전 세계적으로 좋아하는 음식 하나.
Ramen!
라면!
Why?
왜?
Delicious!
맛있다!
A piece of knowledge that few people know but must know.
거의 모르는 사람이 없지만 반드시 알아야 할 지식 한 가지.
I seem not very good at answering this kind of question.
나는 이런 종류의 질문에 답하는 데 별로 능숙하지 않은 것 같다.
Based on all the books you have read, recommend a must-read book.
읽은 모든 책을 기반으로, 필독서를 추천해 달라.
There is one book I have been talking about just now; I recommend this one.
방금 계속 이야기하던 책이 하나 있다. 이 책을 추천한다.
In your mind, what are several papers that influenced the process of AI?
당신의 마음속에 AI 과정에 영향을 준 몇 편의 논문은 무엇인가?
The most important several papers are Backpropagation, Transformer, GPT-3.
가장 중요한 몇 편의 논문은 Backpropagation, Transformer, GPT-3이다.
Of course there are some that are building blocks, also very important, such as ResNet, which may be the foundation of optimization. And Adam. Now perhaps also Muon.
물론 일부는 building block이며, 매우 중요하다. 예를 들어 ResNet, 그것은 최적화의 기초일 수 있다. 그리고 Adam. 지금은 아마도 Muon도.
Based on current cognition, what is the most critical bet?
현재 인식을 기반으로, 가장 핵심적인 베팅은 무엇인가?
Generalized Agent.
일반화된 Agent.
Using Innovation, using L4 to do L3.
Innovation을 사용해서, L4로 L3를 한다.
This year, have you had any moment of sudden realization?
이 해에 어떤 깨달음의 순간이 있었는가?
I do not know; I feel my brain is already muddled.
모르겠다. 내 뇌가 이미 흐릿해진 것 같다.
I have already said a whole year’s worth of words.
나는 이미 일 년치 말을 다 했다.
A third person on site: This may be Long Context affecting intelligence.
현장의 세 번째 사람: 이것이 Long Context가 지능에 영향을 미치는 것일 수 있다.
Yang Zhilin: No way.
Yang Zhilin: 어쩔 수 없다.
The limitations of carbon-based organisms. (laughs)
탄소 기반 생물의 한계. (웃음)
ㅁㅁㅁㅁㅁㅁㅁㅁㅁㅁㅁㅁ
Yang Zhilin (Kimi / Moonshot AI 창업자) 인터뷰 전체 내용 상세 요약
(「站在无限的开端 / The Beginning of Infinity」 – 2025년 7월 말 K2 출시 직후 진행, 약 100분 분량)
1. 서두와 철학적 토대: 무한한 설산과 《The Beginning of Infinity》
Yang Zhilin은 인터뷰 시작부터 David Deutsch의 저서 《The Beginning of Infinity(무한의 시작)》를 반복적으로 인용한다. 책의 핵심 두 문장인 “문제는 불가피하다(Problems are inevitable)”와 “문제는 해결될 수 있다(Problems are solvable)”를 돌에 새길 수 있는 문장으로 제시한다.
계몽운동 이전의 사회는 정적이었고, 현상을 신비주의적으로 설명했다. 계몽운동 이후 사회는 동적으로 변했고, 하나의 문제를 해결하면 새로운 문제가 생기며 지식의 경계가 확장된다. AI 연구개발도 정확히 이 상태에 있다. 강화학습의 일부 문제를 해결하면 평가·검증·일반화라는 새로운 문제가 등장한다.
Yang은 이를 “설산을 오르는 과정”으로 비유한다. 정상은 없을 수도 있으며, 정상이 없기를 바란다고 말한다. 이것이 《The Beginning of Infinity》의 의미이며, AI 발전은 무한한 산을 오르는 과정이다. 2년 전보다 모델은 크게 발전했지만(글을 쓰지 못하던 수준에서 복잡한 코딩 작업을 수 시간 동안 수행할 수 있는 수준으로), 여전히 미지의 기술적 문제가 가득하다. 명확해진 것도 많지만 동시에 새로운 미지의 영역이 열린다.
2. 패러다임 전환: Reasoning Model과 Agentic Model, “통 속의 뇌”
지난 1년 글로벌 기초모델의 가장 중요한 변화로 두 가지를 꼽는다.
• 긴 사고 추론 모델(Reasoning Model): o1이 대표적이다. 모델이 과정에서 여러 번 가설을 세우고 자기 검증을 수행한다. Pass@k를 Pass@1로 바꾸는 효과를 낸다. 이는 인간이 과학 연구나 문제 해결을 하는 과정과 유사하다. 본질적으로 여전히 “통 속의 뇌(brain in a vat)”다. 외부와 상호작용 없이 내부에서만 생각한다.
• 다회차 Agent 강화학습 패러다임: 모델이 외부 세계와 실제로 상호작용한다. 검색, 브라우저, 코드 작성 등을 다회차로 수행하며 피드백을 받아 다음 행동을 결정한다. “통 속의 뇌”에서 벗어나 “손”을 기른 상태다.
둘 다 test-time scaling으로 수렴한다. 추론 시 토큰을 크게 늘려 복잡한 작업을 가능하게 한다. 대화 모델의 단회차 출력과 달리, Agent는 수 시간 동안 인간 개입 없이 코드 저장소를 클론하고 번역·디버깅·테스트·버그 수정까지 엔드투엔드로 수행할 수 있다.
3. OpenAI의 L1~L5 등급과 비선형 관계
OpenAI가 제시한 L1(Chatbot) → L2(Reasoner) → L3(Agent) → L4(Innovator) → L5(Organizer)는 반드시 직렬적이지 않다.
• Agent(L3)의 상한은 Reasoning(L2) 능력에 의존하지만, Reasoning을 먼저 해야만 Agent를 할 수 있는 것은 아니다. Claude의 경로는 Reasoning보다 Agent에 더 무게를 둔 사례다.
• L4(Innovation)의 핵심은 모델이 모델 자체의 개발에 참여하는 것이다(K2가 K3 개발에 참여하는 것).
• L5(Organization)는 Multi-Agent System으로 확장되는 것이다. L4와 L5는 병렬적으로 진행될 수 있다.
• Reasoning과 Agent는 Innovation과 Organization의 전제 조건이 될 수 있으나, 관계는 비선형적이다. L4의 방법으로 L2 문제를 더 잘 풀 수도 있다.
AGI는 특정 계단에 도달하는 순간이 아니라 “방향”이다. 이미 많은 영역에서 99%의 인간보다 뛰어난 성능을 보이며, 기술 발전과 사회 영향이라는 두 층위가 있다. 증기기관이 사회를 변화시키는 데 수십~수백 년이 걸린 것처럼, AI의 사회적 영향도 긴 주기를 필요로 한다.
4. K2 개발의 핵심 기술 결정과 Know-how
20232024년의 핵심 결정이 창업·장문맥 베팅이었다면, 20242025년의 핵심 결정은 사전훈련+SFT 중심에서 사전훈련+강화학습 중심으로의 전환, 그리고 대화에서 Agent로의 패러다임 전환이다.
• K1.5: 강화학습 기술 검증. process reward나 value function이 반드시 필요하지 않으며, 엔드투엔드 reward만으로도 잘 학습된다는 것을 확인. 강화학습 인프라와 알고리즘 know-how 축적.
• K2의 우선순위:
1 뛰어난 Base Model: 고품질 데이터 증가가 정체된 상황에서 token efficiency를 극대화. Muon 최적화기를 대규모에 적용(Keller Jordan 제안 후 Moonshot이 대규모 안정화). Adam 대비 compute optimal 조건에서 약 2배의 효율 향상. 같은 데이터를 먹여도 더 많은 지능을 흡수. max logit 폭발 문제를 clipping 등으로 해결.
2 강력한 Agentic 능력: 일반화가 가장 큰 과제. 현재 RL은 단일 포인트(예: SWE-bench 동분포)에 과적합되기 쉽다. 도구·환경·과제에 과적합되지 않도록 노력. Pre-Train 단계에서 Agentic 능력을 넣는 것도 향후 탐구 대상.
데이터 재작성(rephrase)을 통해 고품질 데이터의 활용도를 높이고 일반화를 개선하려 했다. 오픈소스로 공개한 이유는 커뮤니티와 know-how를 공유하고 기술 발전을 가속하기 위해서다. 다만 Base Model 자체 개선은 여전히 원제작사 중심이며, 다운스트림 Specialized Agent 생성에는 큰 기회가 있다.
5. 오픈소스 vs 클로즈드소스, 제품과 시스템 복잡성
Yang은 이전에 “선도자는 오픈소스하지 않는다”고 말했지만, 현재 글로벌에서 완전히 선도하지 않기 때문에 오픈소스를 선택했다. 오픈소스는 추론 측 기여와 다운스트림 활용에는 유리하나, 모델 자체를 더 강하게 만드는 것은 원제작사 중심이다. 장기적으로 기술 공유를 희망하지만 모든 것을 오픈하는 것은 아니다. 시장은 점차 집중되어 최종적으로 몇 개 정도의 주요 플레이어만 남을 가능성이 높다.
AI 제품은 “모델이 제품”이라는 원칙이 여전히 유효하다. Agent 제품을 만들려면 모델·도구·Context를 통합한 시스템을 먼저 구축해야 하며, 모델 훈련이 끝나면 제품의 핵심도 거의 완성된다. 상호작용 개선은 금상첨화다.
시스템은 한편으로 단순해지고(하나의 일반 모델에 모든 것을 통합), 다른 한편으로 복잡해진다(다양한 모달리티·과제·도구에 대한 일반화 요구). 멀티모달을 추가할 때 텍스트 지능을 손상시키지 않는 것이 이미 매우 어려운 목표다. “똑똑한 멀티모달”을 만들어야 하며, 별도의 전문가 그룹으로 분리되면 “바보 멀티모달”이 될 위험이 있다. Post-Train과 RL로 갈수록 이 융합의 어려움이 커진다.
6. Scaling Law, 데이터 플라이휠, 상호작용 패러다임
Scaling Law는 데이터 벽에 부딪혔다. token efficiency 향상과 RL에 더 많은 컴퓨팅을 투입하는 것이 대응책이다. 그럼에도 모델이 좋아지는 속도는 줄지 않았고 오히려 가속되는 느낌이다.
데이터 플라이휠이 아직 강하게 형성되지 않은 이유는 컴퓨팅 기반 scaling의 효율이 너무 높기 때문이다. RL은 on-policy이고 음의 그래디언트가 있어 Pre-Training보다 scaling 효율이 높다. 대모델은 노이즈에 민감하다. 새로운 상호작용 패러다임을 통해 신호의 노이즈를 줄일 수 있다면 균형이 바뀔 수 있다. 상호작용은 현재 모델 능력 범위 내에서 설계되어야 한다.
사용자 데이터는 직접 훈련에 쓰기 어려울 수 있으나, 수요 분포와 평가 지표를 파악하는 데는 유용하다. C단 사용자는 상업적 가치를 창출할 수 있으며, 특히 Agent 전문 사용자의 생산성 가치는 높다.
7. 조직 관리: RL 방식으로 관리하기
조직 관리도 RL과 유사하다. SFT처럼 “이렇게 하라”고 지시하는 방식이 너무 많으면 창의성이 죽는다. RL처럼 목표(reward)를 주고 자율성을 주되, 적절한 사전(SFT)으로 과도한 이탈을 막는 균형이 필요하다. reward 정의가 핵심이며, 과적합과 reward hacking을 막기 위해 다양한 관찰 지표를 세워야 한다.
8. 개인적 성찰과 AI의 문명적 의미
지난 1년의 기복 속에서도 “시간의 친구가 되자”는 마음가짐을 유지한다. 좋아서 하는 일이기 때문에 복잡하게 생각하지 않는다.
AI는 “인간 문명의 증폭기”다. Meta science가 되어 지식 경계를 돌파하는 거대한 지렛대 역할을 할 것이다. 위험은 존재하지만 포기하면 문명의 상한을 포기하는 것과 같다. 인간의 고유 가치(경험과 사랑)는 대체되지 않을 것이며, 창조의 상당 부분은 AI가 담당하더라도 인간은 결과를 즐기는 다른 삶의 방식을 찾을 수 있다. 이는 수십 년에 걸친 점진적 과정이다.
성장의 핵심은 “문제는 불가피하며 지속적으로 새로운 문제를 해결하는 것이 가장 중요하고 흥미로운 일”이라는 인식이다.
9. 마무리 빠른 Q&A
• 좋아하는 음식: 라면 (맛있으니까)
• 추천 도서: 《The Beginning of Infinity》
• 영향력 있는 논문: Backpropagation, Transformer, GPT-3 (그 외 ResNet, Adam, 최근 Muon)
• 가장 중요한 베팅: 일반화된 Agent (Innovation으로 L3를 푸는 것)
• 최근 깨달음: 특별히 기억나는 순간은 없으며, 이미 1년치 이야기를 다 했다.
인터뷰 전체는 “무한한 산을 오르는 과정”이라는 일관된 메타포로 관통된다. 기술적으로는 token efficiency, Agentic 일반화, test-time scaling, 모델-도구 통합이 핵심 축이며, 조직·제품·문명 차원에서도 같은 철학(문제 해결의 무한성, 시간의 친구, RL적 접근)이 적용된다. K2는 그 여정의 중요한 중간 봉우리(乔戈里峰)일 뿐, 종점이 아니다.