카테고리 없음

KIMI K2.5 we believe in democratizing intelligence. 그리고 저희는 지능을 민주화하는 것을 믿습니다.

클로vㅏ 컴퓨터 2026. 8. 1. 00:13

https://youtu.be/CwePo4847ho?si=YAH-nKHFCygxUpcp

How We Scaled Kimi K2.5 | Zhilin Yang's full GTC 2026 Keynote

If you're curious about the "how" behind scaling Kimi's latest mode...

www.youtube.com


Hi everyone. 안녕하세요 여러분.
Thank you so much for the introduction. 소개해 주셔서 정말 감사합니다.
It’s great to be here. 이곳에 있게 되어 기쁩니다.
It’s great to have this opportunity to share with you guys some of our latest progress and efforts. 여러분께 저희의 최신 진전과 노력들을 공유할 수 있는 이 기회를 갖게 되어 기쁩니다.
So, one of our major pursuits is to build better open models. 그래서 저희의 주요 추구 중 하나는 더 나은 오픈 모델을 구축하는 것입니다.
And we believe in democratizing intelligence. 그리고 저희는 지능을 민주화하는 것을 믿습니다.
With open models, you can deploy anywhere. 오픈 모델로 여러분은 어디에든 배포할 수 있습니다.
It can be on your local servers. 로컬 서버에 있을 수 있습니다.
It can be on the cloud. 클라우드에 있을 수 있습니다.
And you can access every single bit of the weights in the model instead of just using a black box. 그리고 블랙박스만 사용하는 대신 모델의 가중치 모든 비트를 접근할 수 있습니다.
And this is one of the slides that I took from Jensen’s talk earlier this year at CES. 그리고 이것은 올해 초 CES에서 Jensen의 발표에서 가져온 슬라이드 중 하나입니다.
So, as you can see, open models are quickly closing the gap with proprietary models and it’s reaching the frontier. 보시다시피, 오픈 모델이 독점 모델과의 격차를 빠르게 좁히고 있으며 최전선에 도달하고 있습니다.
And we believe that with better and better open models, we’re going to make intelligence more accessible to anybody in the world, in every corner of the world. 그리고 저희는 점점 더 나은 오픈 모델로 전 세계 누구든, 세상의 모든 구석에 있는 사람들에게 지능을 더 접근 가능하게 만들 것이라고 믿습니다.
But open models cannot be just open. 하지만 오픈 모델은 그냥 오픈이기만 해서는 안 됩니다.
They have also to be great. 그것들은 또한 훌륭해야 합니다.
So, in this talk, we’re going to discuss how we make open models great. 그래서 이번 발표에서 저희가 어떻게 오픈 모델을 훌륭하게 만드는지 논의할 것입니다.
So, as we know, scaling is a primary driver for a lot of progress, maybe all of the major AI developments that we have witnessed in the last few years. 아시다시피, 스케일링은 많은 진전의 주요 동인이며, 아마도 지난 몇 년간 저희가 목격한 모든 주요 AI 발전의 동인입니다.
And here we’re going to discuss how we scale our model in different dimensions. 그리고 여기서 저희가 모델을 다른 차원으로 어떻게 스케일하는지 논의할 것입니다.
So, on the left-hand side, the first figure you see here is kind of the standard scaling law. 왼쪽 편에, 여기 보시는 첫 번째 그림은 일종의 표준 스케일링 법칙입니다.
So, on the x-axis you have the log of the number of training tokens. x축에는 훈련 토큰 수의 로그가 있습니다.
And on the y-axis you have the log loss. 그리고 y축에는 로그 손실이 있습니다.
And as you scale the number of training tokens, you get a lower loss. 그리고 훈련 토큰 수를 스케일하면 더 낮은 손실을 얻습니다.
But here the point is we’re not going to just scale the number of training tokens, but we also want to improve the token efficiency. 하지만 여기서 포인트는 저희가 그냥 훈련 토큰 수를 스케일하는 것이 아니라, 토큰 효율성도 개선하고 싶다는 것입니다.
Meaning that we want to move this curve to the left-hand side so that we can achieve a lower loss, a much lower loss using the same number of training tokens. 즉, 이 곡선을 왼쪽으로 이동시켜 같은 수의 훈련 토큰으로 더 낮은 손실, 훨씬 더 낮은 손실을 달성하고 싶다는 의미입니다.
And this can be achieved by having better architectures and optimizers as we’ll discuss in our later slides. 그리고 이것은 나중에 슬라이드에서 논의할 더 나은 아키텍처와 옵티마이저를 가짐으로써 달성될 수 있습니다.
And the second scaling dimension that we’re very interested in is to scale the context length. 그리고 저희가 매우 관심 있는 두 번째 스케일링 차원은 컨텍스트 길이를 스케일하는 것입니다.
So, as you can see in the second figure, if we increase the context length, then we can have a much higher accuracy in terms of predicting the token loss at a given position. 두 번째 그림에서 보시다시피, 컨텍스트 길이를 늘리면 주어진 위치에서 토큰 손실을 예측하는 측면에서 훨씬 더 높은 정확도를 가질 수 있습니다.
And this means that we can increase the capability of the model to achieve more complex tasks by increasing the context length. 그리고 이것은 컨텍스트 길이를 늘림으로써 더 복잡한 작업을 달성하는 모델의 능력을 증가시킬 수 있다는 의미입니다.
So, this is the second scaling dimension that we’re going to talk about. 이것이 저희가 이야기할 두 번째 스케일링 차원입니다.
And the third scaling dimension is the number of agents. 그리고 세 번째 스케일링 차원은 에이전트의 수입니다.
So, we introduce this new learning paradigm of agent swarms where we don’t just rely on a single agent, but we also orchestrate a swarm of agents that can accomplish the subtasks in parallel so that we can increase the task complexity. 그래서 저희가 에이전트 스웜이라는 새로운 학습 패러다임을 도입합니다. 여기서 단일 에이전트에만 의존하지 않고, 하위 작업을 병렬로 수행할 수 있는 에이전트 스웜을 오케스트레이션하여 작업 복잡성을 증가시킵니다.
And we can translate all of this into the language of agents. 그리고 이 모든 것을 에이전트의 언어로 번역할 수 있습니다.
So, if you look at token efficiency, it’s mostly about having a stronger prior so that you can have more efficiency when you do agent RL to search for a better solution. 토큰 효율성을 보면, 주로 더 강한 사전을 가져서 에이전트 RL을 할 때 더 나은 솔루션을 검색하는 데 더 많은 효율성을 갖는 것입니다.
And when you think about long context, it’s mostly about increasing the context length so that you can have a longer running agent. 그리고 긴 컨텍스트를 생각하면, 주로 컨텍스트 길이를 늘려서 더 오래 실행되는 에이전트를 갖는 것입니다.
It can probably run for days or even weeks or months to accomplish more complex tasks. 그것은 아마도 며칠, 심지어 몇 주나 몇 달 동안 실행되어 더 복잡한 작업을 수행할 수 있습니다.
And for agent swarms, it’s another dimension that is added, and at the end of the day we’re going to have a swarm of agents that each of them have a super long context and each of them have a very strong prior for us to search in this entire agent RL system. 그리고 에이전트 스웜의 경우, 추가된 또 다른 차원이며, 결국 저희는 각각 초장 컨텍스트를 가지고 각각 매우 강한 사전을 가진 에이전트 스웜을 갖게 되어 이 전체 에이전트 RL 시스템에서 검색하게 됩니다.
All right. So, we’re going to start from token efficiency. 좋습니다. 그래서 토큰 효율성부터 시작하겠습니다.
So, this is one of the most classical figures in the history of machine learning. 이것은 머신러닝 역사상 가장 고전적인 그림 중 하나입니다.
So, it’s taken from Kaplan et al. Kaplan et al.에서 가져온 것입니다.
And basically says that if we scale proportionally the number of training tokens, the model parameters, and also the amount of compute, we can get lower and lower loss. 그리고 기본적으로 훈련 토큰 수, 모델 파라미터, 그리고 컴퓨트 양을 비례적으로 스케일하면 점점 더 낮은 손실을 얻을 수 있다고 말합니다.
And this is one of the major breakthroughs that the entire community has achieved in the last few years to get better intelligence. 그리고 이것은 지난 몇 년간 전체 커뮤니티가 더 나은 지능을 얻기 위해 달성한 주요 돌파구 중 하나입니다.
But here, what we’re interested in is to have better and better token efficiency. 하지만 여기서 저희가 관심 있는 것은 점점 더 나은 토큰 효율성을 갖는 것입니다.
And here’s the thing. 그리고 여기 포인트가 있습니다.
So, one thing that I would like to emphasize is that token efficiency is not just about efficiency. 제가 강조하고 싶은 한 가지는 토큰 효율성이 단지 효율성에 관한 것이 아니라는 점입니다.
It’s actually also about improving the upper bound of intelligence. 그것은 실제로 지능의 상한선을 개선하는 것에 관한 것이기도 합니다.
So, here’s why. 그 이유는 이렇습니다.
So, suppose you have say 50 trillion tokens, 50 trillion high-quality tokens. 예를 들어 50조 토큰, 50조 고품질 토큰이 있다고 가정해 봅시다.
And then you apply this new optimizer, maybe the Muon optimizer. 그리고 이 새로운 옵티마이저, 아마도 Muon 옵티마이저를 적용합니다.
And then all of a sudden you have a two-times token efficiency. 그러면 갑자기 두 배의 토큰 효율성을 갖게 됩니다.
So, it means that it’s almost like magic that you get equivalently 100 trillion tokens. 즉, 마법처럼 동등하게 100조 토큰을 얻는 것과 같습니다.
And nowadays we’re scaling towards the data wall and we’re hitting the data wall and the amount of high-quality data is quite limited. 그리고 요즘 저희는 데이터 벽을 향해 스케일링하고 있으며 데이터 벽에 부딪히고 있고 고품질 데이터의 양은 상당히 제한적입니다.
And if we suppose that it’s a constant amount, then we increase the token efficiency, it means that we’re going to get better intelligence out of it. 그리고 그것이 상수 양이라고 가정하면, 토큰 효율성을 증가시키면 그로부터 더 나은 지능을 얻게 된다는 의미입니다.
It’s not just about infrastructure efficiency. 그것은 단지 인프라 효율성에 관한 것이 아닙니다.
It’s about better intelligence. 더 나은 지능에 관한 것입니다.
So, this is why we spend a lot of efforts in this aspect because it’s going to push the frontier of intelligence. 그래서 이것이 저희가 이 측면에 많은 노력을 기울이는 이유입니다. 왜냐하면 이것이 지능의 최전선을 밀어낼 것이기 때문입니다.
And Muon optimizer is one of the things that we have heavily invested in since last year. 그리고 Muon 옵티마이저는 작년부터 저희가 크게 투자해 온 것들 중 하나입니다.
So, it’s a second-order optimizer. 그것은 2차 옵티마이저입니다.
And basically every single gradient update is transformed in a way that each entry is orthogonal to each other. 그리고 기본적으로 모든 단일 그래디언트 업데이트가 각 엔트리가 서로 직교하도록 변환됩니다.
And this is very different from the traditional Adam optimizer. 그리고 이것은 전통적인 Adam 옵티마이저와 매우 다릅니다.
And if you implement this optimizer properly, you can get a two-times token efficiency improvement. 그리고 이 옵티마이저를 제대로 구현하면 두 배의 토큰 효율성 개선을 얻을 수 있습니다.
So, we are the first work, we published the first work to demonstrate that Muon optimizer is actually scalable for LLM training. 그래서 저희는 Muon 옵티마이저가 실제로 LLM 훈련에 스케일 가능하다는 것을 입증한 첫 번째 작업을 발표했습니다.
And these are two key techniques that we employed to make it effective for large-scale training. 그리고 대규모 훈련에 효과적으로 만들기 위해 저희가 사용한 두 가지 핵심 기법입니다.
So, one of them is weight decay. 그 중 하나는 weight decay입니다.
It is critical for scaling to larger models. 더 큰 모델로 스케일링하는 데 중요합니다.
And the second is we want to ensure a consistent RMS updates compared to Adam. 그리고 두 번째는 Adam에 비해 일관된 RMS 업데이트를 보장하고 싶다는 것입니다.
So, we have this adjustable coefficients that is applied to each update so that the resulting RMS is going to be comparable to Adam. 그래서 각 업데이트에 적용되는 조정 가능한 계수가 있어 결과 RMS가 Adam과 비교 가능해집니다.
And to make Muon memory efficient across all these Nvidia GPU clusters, we also develop a distributed Muon optimizer implementation that partitions the states across the data parallel group so that we can have a very efficient implementation for the Muon optimizer. 그리고 이 모든 Nvidia GPU 클러스터에서 Muon을 메모리 효율적으로 만들기 위해, 데이터 병렬 그룹에 걸쳐 상태를 분할하는 분산 Muon 옵티마이저 구현도 개발하여 Muon 옵티마이저를 매우 효율적으로 구현할 수 있게 했습니다.
And these are some of the results that were presented in the paper. 그리고 이것이 논문에서 제시된 결과들 중 일부입니다.
So, as you can see, with the same number of parameters and the same number of training tokens, we just replace the original AdamW optimizer with the new Muon optimizer. 보시다시피, 같은 수의 파라미터와 같은 수의 훈련 토큰으로, 원래의 AdamW 옵티마이저를 새로운 Muon 옵티마이저로 교체하기만 하면 됩니다.
It’s going to improve the performance across the board significantly. 전반적으로 성능을 크게 개선할 것입니다.
But there was this new challenge that we encountered when we tried to scale it up further. 하지만 더 나아가 스케일하려고 할 때 새로운 도전에 직면했습니다.
When we tried to scale Muon for a one trillion parameter model, we encountered a new issue about training instability. 1조 파라미터 모델에 Muon을 스케일하려고 할 때, 훈련 불안정성에 관한 새로운 문제를 만났습니다.
So, as you can see on the left figure, we observed that the max logits quickly explodes and quickly exceeds 1,000. 왼쪽 그림에서 보시다시피, max logits가 빠르게 폭발하고 빠르게 1,000을 초과하는 것을 관찰했습니다.
And the typical values for training for this max logits is about say 50 or maybe less than 100. 그리고 이 max logits의 훈련에 대한 전형적인 값은 약 50 또는 100 미만입니다.
But for Muon, it quickly exceeds 1,000. 하지만 Muon의 경우 빠르게 1,000을 초과합니다.
And at the same time, we observe training divergence on the left-hand side. 그리고 동시에 왼쪽에서 훈련 발산을 관찰합니다.
If you look at the training loss, it goes down a bit, but then at the end of the day it explodes and it cannot converge as expected. 훈련 손실을 보면, 조금 내려가다가 결국 폭발하고 예상대로 수렴하지 못합니다.
So, this is one of the technical challenges that we have to address. 이것이 저희가 해결해야 할 기술적 도전 중 하나입니다.
And the solution to this is to introduce this new technique called QK clip. 그리고 이에 대한 해결책은 QK clip이라는 새로운 기법을 도입하는 것입니다.
So, basically what it says is that for each attention head in this entire neural network, we’re going to in the forward pass, we’re going to compute the max logit. 기본적으로 이 전체 신경망의 각 어텐션 헤드에 대해, forward pass에서 max logit을 계산합니다.
And then we’re going to calculate a dividing factor that can be applied to each key projection as well as the query projection so that we can sort of clip the maximum value of the query and the key to constrain it into a given range. 그리고 각 key 프로젝션과 query 프로젝션에 적용할 수 있는 나누는 인자를 계산하여 query와 key의 최대값을 클립하여 주어진 범위로 제약합니다.
So that we’re not going to have explosion anymore. 그래서 더 이상 폭발이 발생하지 않도록 합니다.
So, these are some of the empirical results. 이것이 일부 경험적 결과입니다.
On the left-hand side, there are two curves, but they are strictly overlapped with each other. 왼쪽에는 두 곡선이 있지만, 그것들은 엄격하게 서로 겹쳐 있습니다.
So, these are the training curves before and after applying the clipping technique. 이것이 클리핑 기법을 적용하기 전과 후의 훈련 곡선입니다.
So, you can see the clipping technique does not affect the training loss decrease at all. 클리핑 기법이 훈련 손실 감소에 전혀 영향을 미치지 않는다는 것을 볼 수 있습니다.
But on the right-hand side, if we inspect the intermediate metric, if we inspect the max logit, it’s going to be effectively constrained. 하지만 오른쪽에서 중간 메트릭을 검사하면, max logit을 검사하면 효과적으로 제약됩니다.
So, it first explodes as before, but at the value of 100, it’s going to be clipped at a constant value for a long time. 그래서 처음에는 이전처럼 폭발하지만, 100의 값에서 오랫동안 상수 값으로 클립됩니다.
And then after a certain number of steps, it will just naturally go down. 그리고 일정 수의 스텝 후에 자연스럽게 내려갑니다.
So, the neural network sort of finds a way to constrain the maximum value of the max logit to ensure a stable training process. 그래서 신경망이 안정적인 훈련 과정을 보장하기 위해 max logit의 최대값을 제약하는 방법을 찾습니다.
And at the same time, it doesn’t affect the training convergence as shown in the left figure. 그리고 동시에 왼쪽 그림에서 보듯이 훈련 수렴에 영향을 미치지 않습니다.
So we employed this technique in our K2 model training and successfully scaled it to 1 trillion parameters. 그래서 저희는 이 기법을 K2 모델 훈련에 사용하여 성공적으로 1조 파라미터로 스케일했습니다.
And this is the first example of large-scale Muon training in the history of machine learning. 그리고 이것은 머신러닝 역사상 대규모 Muon 훈련의 첫 번째 사례입니다.
And the second dimension that we’re very interested in is long context. 그리고 저희가 매우 관심 있는 두 번째 차원은 긴 컨텍스트입니다.

And the second dimension that we’re very interested in is long context. 그리고 저희가 매우 관심 있는 두 번째 차원은 긴 컨텍스트입니다.
So this is another figure. 이것은 또 다른 그림입니다.
It’s probably less known. 아마도 덜 알려져 있을 것입니다.
It’s one of the hidden gems in these papers. 이 논문들에 있는 숨겨진 보석 중 하나입니다.
So instead of just pushing down the training laws by training on more tokens, it has some insights from another perspective. 더 많은 토큰으로 훈련하여 훈련 법칙만 밀어내리는 대신, 다른 관점에서 몇 가지 통찰을 가지고 있습니다.
So as we can see, this is a comparison between transformers and LSTMs. 보시다시피, 이것은 트랜스포머와 LSTM의 비교입니다.
So on the left-hand side, we can see that transformers achieve a lower training loss given the same number of parameters and the same number of training tokens as expected. 왼쪽에서 트랜스포머가 같은 수의 파라미터와 같은 수의 훈련 토큰으로 예상대로 더 낮은 훈련 손실을 달성하는 것을 볼 수 있습니다.
And this is why transformers become the de facto architecture that people are using right now. 그리고 이것이 트랜스포머가 현재 사람들이 사용하고 있는 사실상의 아키텍처가 된 이유입니다.
But on the right-hand side, it’s really interesting to see that transformers are actually better because it can improve through the whole context. 하지만 오른쪽에서 트랜스포머가 실제로 더 나은 이유는 전체 컨텍스트를 통해 개선될 수 있기 때문이라는 점이 정말 흥미롭습니다.
So the x-axis is the token index in context. x축은 컨텍스트 내 토큰 인덱스입니다.
And if you increase the token index, you can see that the training loss of transformers actually dropped by a lot. 토큰 인덱스를 증가시키면 트랜스포머의 훈련 손실이 실제로 많이 떨어지는 것을 볼 수 있습니다.
If you just continually increase the context length, the loss just continuously drops down. 컨텍스트 길이를 계속 증가시키기만 하면 손실이 지속적으로 떨어집니다.
But if you look at the curve of LSTM, it just is saturated after a certain number of tokens. 하지만 LSTM의 곡선을 보면 일정 수의 토큰 이후에 그냥 포화됩니다.
It means that transformers have this better capability of capturing longer context, and this is what makes it better, 이것은 트랜스포머가 더 긴 컨텍스트를 포착하는 더 나은 능력을 가지고 있다는 의미이며, 이것이 그것을 더 낫게 만듭니다.
because if you go back to like 10 years ago, people used LSTM for tasks like machine translation. 왜냐하면 약 10년 전으로 돌아가면 사람들이 기계 번역 같은 작업에 LSTM을 사용했기 때문입니다.
But it is not good for, for example, understanding entire code base or running a super long agent trajectory to solve, for example, writing Linux kernels from scratch. 하지만 예를 들어 전체 코드 베이스를 이해하거나 처음부터 Linux 커널을 작성하는 것과 같은 초장 에이전트 궤적을 실행하는 데는 좋지 않습니다.
It’s not going to be accomplished by LSTMs. LSTM으로는 달성되지 않을 것입니다.
So this is a very much needed capability in the era of agents, because tasks are becoming harder and harder, and we need longer and longer contexts. 그래서 이것은 에이전트 시대에 매우 필요한 능력입니다. 왜냐하면 작업이 점점 더 어려워지고 있으며 점점 더 긴 컨텍스트가 필요하기 때문입니다.
So the research idea here is to develop a better architecture so that we can efficiently scale to a longer context length and at the same time achieve a lower per token loss at larger token indices. 그래서 여기서의 연구 아이디어는 더 긴 컨텍스트 길이로 효율적으로 스케일하고 동시에 더 큰 토큰 인덱스에서 더 낮은 토큰당 손실을 달성할 수 있도록 더 나은 아키텍처를 개발하는 것입니다.
And this is the motivation for which we introduced this new architecture called Kimi linear. 그리고 이것이 저희가 Kimi linear라는 새로운 아키텍처를 도입한 동기입니다.
And it contains this new linear attention variant called Kimi delta attention, which improves the original gated delta rule, GDR, by improved recurrent memory. 그리고 이것은 개선된 순환 메모리로 원래의 gated delta rule, GDR을 개선하는 Kimi delta attention이라는 새로운 선형 어텐션 변형을 포함합니다.
I will show the details later. 자세한 내용은 나중에 보여드리겠습니다.
And at the same time, we’re going to mix linear attention layers with full attention layers using a 1-2-3 ratio so that you can balance between this long context capabilities and at the same time having a more efficient implementation. 그리고 동시에 1:2:3 비율로 선형 어텐션 레이어와 풀 어텐션 레이어를 혼합하여 이 긴 컨텍스트 능력과 동시에 더 효율적인 구현 사이의 균형을 맞출 수 있습니다.
So this is some of the formulation. 이것이 일부 공식입니다.
The basic idea is simple. 기본 아이디어는 간단합니다.
If you look at linear attention, in the original formulation, the memory is going to be global. 선형 어텐션을 보면, 원래 공식에서 메모리는 전역적입니다.
So there is a global single decay factor that is applied along the way. 그래서 전체에 적용되는 전역 단일 감쇠 인자가 있습니다.
So it means that basically, there are only two cases. 즉, 기본적으로 두 가지 경우만 있습니다.
In one case, you’re going to forget basically everything, and you’re not going to retain any information. 한 경우에는 기본적으로 모든 것을 잊어버리고 어떤 정보도 유지하지 않습니다.
In the second case, you can choose to retain almost everything, but at the same time, you don’t have the capability to leave out some of the unnecessary information in this long context. 두 번째 경우에는 거의 모든 것을 유지하도록 선택할 수 있지만, 동시에 이 긴 컨텍스트에서 불필요한 정보 일부를 제외하는 능력이 없습니다.
So we introduced this key idea of having a fine-grained decay factor as shown in this highlighted alpha term. 그래서 이 강조된 알파 항에서 보듯이 세분화된 감쇠 인자를 갖는 이 핵심 아이디어를 도입했습니다.
So instead of being a scalar, it’s going to be a diagonal matrix which controls the decay rate for each channel so that we can have two possibilities. 스칼라 대신 각 채널의 감쇠율을 제어하는 대각 행렬이 되어 두 가지 가능성을 가질 수 있습니다.
For some of the channels, we can decay really, really slow, meaning that we can retain this long-context information across a very long range. 일부 채널의 경우 정말 정말 천천히 감쇠하여 매우 긴 범위에 걸쳐 이 긴 컨텍스트 정보를 유지할 수 있습니다.
And at the same time, for the other channels, we can quickly forget the information from the past indices to refresh it and observe new information. 그리고 동시에 다른 채널의 경우 과거 인덱스의 정보를 빠르게 잊어버리고 새로고침하여 새로운 정보를 관찰할 수 있습니다.
And this is to increase the expressivity of this model. 그리고 이것은 이 모델의 표현력을 증가시키기 위함입니다.
And of course, to leverage modern GPUs, we have to use this chunkwise formulations so that we can parallelize the computation of modern GPUs. 그리고 물론 현대 GPU를 활용하기 위해 이 청크 단위 공식을 사용하여 현대 GPU의 계산을 병렬화할 수 있어야 합니다.
So the first equation here is the chunkwise formulation of Kimi linear. 여기 첫 번째 방정식이 Kimi linear의 청크 단위 공식입니다.
But as you can see, this is going to bring massive infrastructure challenges because of this newly introduced alpha term. 하지만 보시다시피, 이 새로 도입된 알파 항 때문에 대규모 인프라 도전을 가져올 것입니다.
Because now it is a matrix instead of a scalar, it cannot easily be factored out. 이제 스칼라 대신 행렬이기 때문에 쉽게 인수분해될 수 없습니다.
So to achieve an efficient implementation, we rewrite the entire equation into the bottom three equations. 효율적인 구현을 달성하기 위해 전체 방정식을 아래 세 방정식으로 다시 작성합니다.
So we introduced this matrix inversion operation, as well as introducing the cumulative decay factor, so that we can implement this entire thing in parallel, without sacrificing any efficiency. 그래서 이 행렬 역연산과 누적 감쇠 인자를 도입하여 효율성을 전혀 희생하지 않고 이 전체를 병렬로 구현할 수 있습니다.
And more importantly, this is not an approximation, it’s an exact mathematically equivalent formulation, so that we can achieve much efficient implementation without sacrificing any loss in terms of performance. 그리고 더 중요하게는, 이것은 근사가 아니라 정확한 수학적으로 동등한 공식이므로 성능 측면에서 손실을 전혀 희생하지 않고 훨씬 효율적인 구현을 달성할 수 있습니다.
So it’s going to be as efficient as previous linear attention variants, but at the same time, much more expressive. 이전 선형 어텐션 변형만큼 효율적이면서 동시에 훨씬 더 표현력이 높습니다.
So these are some of the results that we obtained using a fair comparison. 이것이 공정한 비교를 사용하여 얻은 결과 중 일부입니다.
So on the left-hand side, we see the performance on two different types of tasks. 왼쪽에서 두 가지 다른 유형의 작업에 대한 성능을 볼 수 있습니다.
So MMMU is a short context task. MMMU는 짧은 컨텍스트 작업입니다.
So for short context tasks, Kimi linear achieved a better performance compared to MLA and GDN. 짧은 컨텍스트 작업에서 Kimi linear는 MLA와 GDN에 비해 더 나은 성능을 달성했습니다.
And at the same time, for longer context tasks such as RULER, Kimi linear is also better than the other variants while being much more efficient compared to MLA. 그리고 동시에 RULER와 같은 더 긴 컨텍스트 작업에서 Kimi linear는 다른 변형보다 더 나으면서 MLA에 비해 훨씬 더 효율적입니다.
And when we scale the context further to, for example, 1 million tokens or even longer, it’s going to be much more efficient compared to the baselines. 그리고 컨텍스트를 예를 들어 100만 토큰이나 그 이상으로 더 스케일하면 베이스라인에 비해 훨씬 더 효율적일 것입니다.
And this is also the first architecture that can outperform full attention across the board, including short context tasks, long input tasks, and long output tasks. 그리고 이것은 짧은 컨텍스트 작업, 긴 입력 작업, 긴 출력 작업을 포함하여 전반적으로 풀 어텐션을 능가할 수 있는 첫 번째 아키텍처이기도 합니다.
So these are two key dimensions that we are interested in. 이것이 저희가 관심 있는 두 가지 핵심 차원입니다.
And the third dimension is the agent swarms. 그리고 세 번째 차원은 에이전트 스웜입니다.
So here is a diagram to showcase how we designed this agent swarm paradigm to solve some of the more complex tasks compared to single agent paradigms. 여기 단일 에이전트 패러다임에 비해 더 복잡한 작업을 해결하기 위해 이 에이전트 스웜 패러다임을 어떻게 설계했는지 보여주는 다이어그램이 있습니다.
So here we have an orchestrator. 여기 오케스트레이터가 있습니다.
Or you can call it a main agent. 또는 메인 에이전트라고 부를 수 있습니다.
It’s responsible for orchestrating tasks. 작업을 오케스트레이션하는 책임이 있습니다.
It has different options. 다양한 옵션이 있습니다.
For example, we can spawn a group of sub-agents and assign new tasks to these sub-agents. 예를 들어, 하위 에이전트 그룹을 생성하고 이 하위 에이전트에 새로운 작업을 할당할 수 있습니다.
Or you can collect the results from the return of these sub-agents. 또는 이 하위 에이전트의 반환에서 결과를 수집할 수 있습니다.
And you can sort of perform this process in an iterative way. 그리고 이 과정을 반복적인 방식으로 수행할 수 있습니다.
And at the end of the day, you can accomplish a more complex task compared to using one single agent. 그리고 결국 단일 에이전트를 사용하는 것보다 더 복잡한 작업을 수행할 수 있습니다.
And it’s analogous to human society. 그리고 그것은 인간 사회와 유사합니다.
For example, if we build a company, we need different roles and we need, for example, an orchestrator or maybe we need a CEO to compose and assign the tasks to different roles. 예를 들어, 회사를 구축하면 다른 역할이 필요하고, 예를 들어 오케스트레이터나 CEO가 작업을 구성하고 다른 역할에 할당해야 합니다.
And at the end of the day, the entire organization is going to have to move towards this same goal. 그리고 결국 전체 조직이 이 동일한 목표를 향해 움직여야 합니다.
And here, for example, in this case, we have, maybe you have the AI researchers, you have the web developers, you have physical researchers, and they can study different topics. 그리고 여기 예를 들어 이 경우 AI 연구자, 웹 개발자, 물리 연구자가 있을 수 있으며 그들은 다른 주제를 연구할 수 있습니다.
And at the end of the day, you just collect the results and spawn a group of fact-checkers and web developers and file downloaders to assemble the results into a single report. 그리고 결국 결과를 수집하고 팩트체커, 웹 개발자, 파일 다운로더 그룹을 생성하여 결과를 단일 보고서로 조립합니다.
And this is another perspective to look at this new paradigm. 이것이 이 새로운 패러다임을 보는 또 다른 관점입니다.
So the x-axis is the complexity of the task. x축은 작업의 복잡성입니다.
And the y-axis is the execution time. 그리고 y축은 실행 시간입니다.
And the complexity of the task is measured by the accuracy of a group of models on such tasks. 작업의 복잡성은 그러한 작업에 대한 모델 그룹의 정확도로 측정됩니다.
So we can see with Agent Swarms, it’s going to substantially reduce the execution time compared to single agents. Agent Swarms로 단일 에이전트에 비해 실행 시간을 상당히 줄일 수 있다는 것을 볼 수 있습니다.
It’s going to be more effective, and this means that we can scale this Agent Swarm paradigm. 더 효과적일 것이며, 이것은 이 Agent Swarm 패러다임을 스케일할 수 있다는 의미입니다.
For example, if you run Agent Swarms with 100 or maybe even 1,000 sub-agents, you can accomplish a complex task within a certain period of time that is tolerable for producing real economic value. 예를 들어, 100개 또는 심지어 1,000개의 하위 에이전트로 Agent Swarms를 실행하면 실제 경제적 가치를 생산하는 데 허용 가능한 일정 기간 내에 복잡한 작업을 수행할 수 있습니다.
And we can certainly scale it in different dimensions. 그리고 확실히 다른 차원으로 스케일할 수 있습니다.
We can scale the inputs. 입력을 스케일할 수 있습니다.
For example, we can download and read hundreds of sources or even maybe thousands of sources in parallel. 예를 들어, 수백 개의 소스 또는 심지어 수천 개의 소스를 병렬로 다운로드하고 읽을 수 있습니다.
Or we can output, write a 100-page literature review in parallel. 또는 출력을 스케일하여 100페이지 문헌 검토를 병렬로 작성할 수 있습니다.
Or we can take actions at scale. 또는 대규모로 행동을 취할 수 있습니다.
We can perform data analysis for 10 different tasks. 10개의 다른 작업에 대한 데이터 분석을 수행할 수 있습니다.
And also it is orchestration at scale. 그리고 대규모 오케스트레이션이기도 합니다.
You have to learn to design subtasks and aggregate the results. 하위 작업을 설계하고 결과를 집계하는 법을 배워야 합니다.

And technically, we define some new objective functions to guide the learning process of our agent swarm system. 그리고 기술적으로, 에이전트 스웜 시스템의 학습 과정을 안내하기 위해 몇 가지 새로운 목적 함수를 정의합니다.
So there are three reward functions, reward objectives that are considered here, compared to the conventional single agent RL learning. 그래서 여기에는 전통적인 단일 에이전트 RL 학습과 비교하여 고려되는 세 가지 보상 함수, 보상 목표가 있습니다.
So the first term is what we call the instantiation reward. 첫 번째 항은 저희가 인스턴시에이션 보상이라고 부르는 것입니다.
It incentivizes sub-agent instantiation to prevent this serial collapse phenomenon from happening. 하위 에이전트 인스턴시에이션을 장려하여 이 직렬 붕괴 현상이 발생하는 것을 방지합니다.
So basically, we don’t want it to default to single-agent execution. 기본적으로, 단일 에이전트 실행으로 기본 설정되는 것을 원하지 않습니다.
We want to encourage the parallel executions, especially when it’s early stage in training. 특히 훈련 초기 단계에서 병렬 실행을 장려하고 싶습니다.
And of course we can decay the weight for this instantiation reward term over training course because when it learns parallel execution we can reduce the weight. 그리고 물론 훈련 과정에서 이 인스턴시에이션 보상 항의 가중치를 감소시킬 수 있습니다. 왜냐하면 병렬 실행을 학습하면 가중치를 줄일 수 있기 때문입니다.
And the second term here is Finish Reward. 그리고 여기 두 번째 항은 Finish Reward입니다.
And it is used because we observe one of the things in training that some of these subtasks are just created but never finished. 훈련에서 관찰한 것 중 하나는 이러한 하위 작업 중 일부가 생성되기만 하고 절대 완료되지 않는다는 것이기 때문에 사용됩니다.
So it’s almost like it’s going to hack the first term. 그래서 거의 첫 번째 항을 해킹하는 것과 같습니다.
By just spawning a bunch of sub-agents, and the task might be too complex, or maybe the task just doesn’t make sense. 그냥 하위 에이전트 무리를 생성하기만 하면, 작업이 너무 복잡하거나 작업이 의미가 없을 수 있습니다.
And here, we use this Finish Reward to basically encourage that each of the sub-tasks should have a relatively high ratio of completion. 그리고 여기서 이 Finish Reward를 사용하여 기본적으로 각 하위 작업이 상대적으로 높은 완료 비율을 가져야 한다고 장려합니다.
Instead of just spawning a bunch of pseudo-tasks, we need it to be meaningful. 그냥 의사 작업 무리를 생성하는 대신, 의미 있어야 합니다.
So this is the second term that we use and of course we use the same decay strategy. 이것이 저희가 사용하는 두 번째 항이며, 물론 같은 감쇠 전략을 사용합니다.
We use the relative high weight at the beginning of training and we decay it to a relatively low weight at the end of training. 훈련 초기에 상대적으로 높은 가중치를 사용하고 훈련 끝에는 상대적으로 낮은 가중치로 감쇠합니다.
And of course the third term is the standard term. 그리고 물론 세 번째 항은 표준 항입니다.
It’s the outcome reward. 그것은 결과 보상입니다.
It’s going to measure whether the entire task is completed and then we’re going to add these three terms in our reinforcement learning. 전체 작업이 완료되었는지 측정하고, 그런 다음 이 세 항을 강화 학습에 추가합니다.
And of course, we have to build the entire infrastructure because right now you need to support the parallel execution and they will need to support different reward functions and to maximize the efficiency of the entire agent swarm IO system. 그리고 물론 전체 인프라를 구축해야 합니다. 왜냐하면 지금은 병렬 실행을 지원해야 하고, 다른 보상 함수를 지원해야 하며, 전체 에이전트 스웜 IO 시스템의 효율성을 최대화해야 하기 때문입니다.
So here are three different things that we have tried scaling. 여기 저희가 스케일링을 시도한 세 가지 다른 것들이 있습니다.
The Muon Clip Optimizer improves token efficiency. Muon Clip Optimizer는 토큰 효율성을 개선합니다.
And Kimi Delta attention in the Kimi linear architecture improves long context. 그리고 Kimi linear 아키텍처의 Kimi Delta attention은 긴 컨텍스트를 개선합니다.
And we also have the agent swarms paradigm to further create a new dimension of scaling. 그리고 에이전트 스웜 패러다임을 통해 스케일링의 새로운 차원을 추가로 만듭니다.
And all of this put together, we created Kimi K2.5, a new model that we just released over one month ago. 그리고 이 모든 것을 합쳐서, 한 달 조금 전에 출시한 새로운 모델인 Kimi K2.5를 만들었습니다.
Here’s a short video to demonstrate some of its capabilities. 그 능력 일부를 보여주는 짧은 비디오가 여기 있습니다.
So yeah, there are a lot of interesting things, capabilities that we discover from the model. 네, 모델에서 발견한 흥미로운 것들, 능력들이 많이 있습니다.
For example, it merges the visual capabilities with coding capabilities. 예를 들어, 시각 능력과 코딩 능력을 병합합니다.
So a lot of new things just emerge out of it. 그래서 많은 새로운 것들이 그냥 출현합니다.
It can read a video and then produce a website that sort of replicates or style transfer the original video. 비디오를 읽고 원본 비디오를 복제하거나 스타일 전이하는 웹사이트를 생성할 수 있습니다.
And all of this are due to successful and stable training at the pre-training stage. 그리고 이 모든 것은 사전 훈련 단계에서의 성공적이고 안정적인 훈련 덕분입니다.
So this is also one of the most beautiful curves that I observed in my life. 이것은 또한 제가 인생에서 관찰한 가장 아름다운 곡선 중 하나입니다.
So this is the training curve of the K2.5 based model. 이것이 K2.5 기반 모델의 훈련 곡선입니다.
So as you can see, it went through over 15 trillion tokens and of course in K2.5 we additionally trained another 15 trillion tokens and the entire training process is just so stable, there’s no loss spike, especially when we introduced this new Muon optimizer, we didn’t observe any spike and this smooth, stable training process produces a very stable outcome. 보시다시피, 15조 이상의 토큰을 거쳤고, 물론 K2.5에서는 추가로 15조 토큰을 더 훈련했으며, 전체 훈련 과정이 매우 안정적이고 손실 스파이크가 없으며, 특히 이 새로운 Muon 옵티마이저를 도입했을 때 어떤 스파이크도 관찰하지 않았고, 이 부드럽고 안정적인 훈련 과정이 매우 안정적인 결과를 만들어냅니다.
A very strong base model that we can fine-tune on top of it to achieve new capabilities as we introduced and saw in the video. 비디오에서 소개하고 본 것처럼 새로운 능력을 달성하기 위해 그 위에 파인튜닝할 수 있는 매우 강한 베이스 모델입니다.
And this is also, of course, a trend on NVIDIA H800 GPUs, and you should know in this H800 cluster contains two TB RAM and 8 GPUs. 그리고 이것은 물론 NVIDIA H800 GPU에서의 트렌드이며, 이 H800 클러스터에는 2TB RAM과 8개의 GPU가 있다는 것을 알아야 합니다.
They are connected by NVLink. 그것들은 NVLink로 연결되어 있습니다.
And another key innovation of Kimi K2.5 is that it is the first open model with native joint vision text capabilities. 그리고 Kimi K2.5의 또 다른 핵심 혁신은 네이티브 조인트 비전 텍스트 능력을 가진 첫 번째 오픈 모델이라는 점입니다.
So if you look at previous open models, usually their visual capabilities are added on top of a text base, meaning that, for example, if you train the text models for 20 trillion tokens, and then on top of it, you do another 2 trillion. 이전 오픈 모델을 보면, 보통 시각 능력이 텍스트 베이스 위에 추가됩니다. 예를 들어, 텍스트 모델을 20조 토큰으로 훈련한 다음 그 위에 추가로 2조를 하는 식입니다.
Sort of a post-training process to add additional visual capabilities on top of it. 그 위에 추가 시각 능력을 더하는 일종의 포스트 트레이닝 과정입니다.
But for K2.5, it’s different in the sense that we fuse the training process of vision and text from day one. 하지만 K2.5의 경우, 처음부터 비전과 텍스트의 훈련 과정을 융합한다는 점에서 다릅니다.
So it’s called early fusion here. 그래서 여기서 이를 early fusion이라고 부릅니다.
We start from 0% of the progress. 진행의 0%부터 시작합니다.
So from day one, we’re going to merge the vision and text tokens and as shown in our preliminary experiments, it outperforms our late fusion. 그래서 첫날부터 비전과 텍스트 토큰을 병합하며, 예비 실험에서 보듯이 late fusion을 능가합니다.
And some of the new capabilities that we observe also come from this training recipe. 그리고 저희가 관찰한 일부 새로운 능력도 이 훈련 레시피에서 나옵니다.
For example, if you want to do vision to code, you really have to merge vision and text into a single brain to achieve that. 예를 들어, vision to code를 하려면 정말로 비전과 텍스트를 단일 두뇌로 병합해야 달성할 수 있습니다.
If you separate these two brains, it’s not going to happen. 이 두 두뇌를 분리하면 일어나지 않을 것입니다.
You have to align these two modalities into a shared embedding space, a shared representation space so as to achieve this. 이 두 모달리티를 공유 임베딩 공간, 공유 표현 공간으로 정렬해야 이를 달성할 수 있습니다.
And another interesting thing that we observe is that these two modalities can actually enhance each other. 그리고 저희가 관찰한 또 다른 흥미로운 점은 이 두 모달리티가 실제로 서로를 강화할 수 있다는 것입니다.
So that’s been long been a challenge that if you add vision capabilities into a text model, it’s going to somewhat hurt the text performance. 텍스트 모델에 시각 능력을 추가하면 텍스트 성능이 다소 손상되는 것이 오랫동안 도전이었습니다.
But here we found that if you train it properly, these two modalities can actually enhance each other. 하지만 여기서 제대로 훈련하면 이 두 모달리티가 실제로 서로를 강화할 수 있다는 것을 발견했습니다.
So this is one of the key findings that we observe in our training, so first, vision improves text. 이것이 저희 훈련에서 관찰한 핵심 발견 중 하나입니다. 먼저, 비전이 텍스트를 개선합니다.
So this is so interesting, so before Vision RL, the performance in the first column, and then we have the performance after Vision RL. 이것이 정말 흥미롭습니다. Vision RL 이전의 첫 번째 열 성능과, Vision RL 이후의 성능입니다.
So here, Vision RL refers to a process that we only use vision tasks. 여기서 Vision RL은 비전 작업만 사용하는 과정을 의미합니다.
So there is no text involved here. 여기서 텍스트는 관여하지 않습니다.
We only have vision tasks. 비전 작업만 있습니다.
For example, we teach the model how to count, how to answer some of these visual QA problems without any, for example, math and coding problems in this space. 예를 들어, 모델에게 세는 법, 이러한 시각 QA 문제에 답하는 법을 가르치며, 이 공간에서 수학과 코딩 문제 같은 것은 없습니다.
But we observe that it’s going to improve the performance for even reasonably heavy text tasks. 하지만 상당히 무거운 텍스트 작업의 성능까지 개선한다는 것을 관찰합니다.
And on the other hand, text also improves vision. 그리고 다른 한편으로, 텍스트도 비전을 개선합니다.
If you have a very strong text base, you actually don’t need any vision SFT data in the training process. 매우 강한 텍스트 베이스가 있으면 실제로 훈련 과정에서 어떤 비전 SFT 데이터도 필요하지 않습니다.
And this is the approach that we adopt. 그리고 이것이 저희가 채택한 접근법입니다.
So it’s called Zero Vision SFT. Zero Vision SFT라고 부릅니다.
Basically, we don’t have… We have basically zero vision SFT data, and the only SFT data that we have is the text SFT data. 기본적으로 비전 SFT 데이터가 거의 없으며, 저희가 가진 유일한 SFT 데이터는 텍스트 SFT 데이터입니다.
And then we do a joint RL over text and vision, and you can see that we can achieve almost state-of-the-art performance across the board on vision tasks without any vision data. 그리고 텍스트와 비전에 대한 조인트 RL을 수행하면, 어떤 비전 데이터 없이도 비전 작업 전반에서 거의 최첨단 성능을 달성할 수 있다는 것을 볼 수 있습니다.
So it’s clear that if you have a strong text base, it’s also going to improve the vision if you align these two modalities into a shared space in your pre-training. 강한 텍스트 베이스가 있으면 사전 훈련에서 이 두 모달리티를 공유 공간으로 정렬하면 비전도 개선된다는 것이 분명합니다.
And also, these are some of the examples of, as I’ve shown in the video, it demonstrates strong capabilities of visual design and front-end coding, and this also emerges from our vision text through training. 그리고 또한, 비디오에서 보여드린 것처럼, 시각 디자인과 프론트엔드 코딩의 강한 능력을 보여주며, 이것 또한 저희의 비전-텍스트 훈련을 통해 출현합니다.
So after all this, so this is all about Kimi K2.5. 이 모든 것 후에, 이것이 Kimi K2.5에 관한 전부입니다.
And as you probably know, we released our new architecture yesterday in our tech report. 그리고 아마 아시다시피, 어제 기술 보고서에서 저희의 새로운 아키텍처를 발표했습니다.
It’s called Attention Residue. Attention Residue라고 합니다.
So here I’m also going to briefly talk about our new work, which serves as a sneak peek into our next generation architecture that we’re probably going to adopt in our later models. 그래서 여기서 저희의 새로운 작업에 대해 간략히 이야기하겠습니다. 이는 저희가 이후 모델에서 채택할 가능성이 높은 차세대 아키텍처의 미리보기로 역할합니다.
So here the motivation is quite simple. 여기서 동기는 꽤 간단합니다.
Can we apply some of our techniques that we use in the temporal dimension, and we just take some of the inspirations and then apply it to the depth dimension? 저희가 시간 차원에서 사용하는 일부 기법을 적용할 수 있을까, 그리고 일부 영감을 받아 깊이 차원에 적용할 수 있을까?
And it starts from this residual connection. 그리고 이것은 이 residual connection에서 시작합니다.
So I still remember listening to Kaiming’s talk at a tutorial in ICML 2016, 10 years ago. 저는 여전히 10년 전 ICML 2016의 튜토리얼에서 Kaiming의 발표를 들었던 것을 기억합니다.
So it was a brilliant idea. 그것은 훌륭한 아이디어였습니다.
So basically, before ResNet, nobody was able to train deep networks. 기본적으로 ResNet 이전에는 아무도 깊은 네트워크를 훈련할 수 없었습니다.
If you increase the depth, if you increase the number of layers for neural networks, nobody was able to train it because you observe this gradient explosion and gradient vanishing, all these stability issues. 깊이를 증가시키고, 신경망의 레이어 수를 증가시키면, 그래디언트 폭발과 그래디언트 소실 같은 이러한 안정성 문제 때문에 아무도 훈련할 수 없었습니다.
But then after the introduction of ResNet, we can train an arbitrarily large number of layers. 하지만 ResNet 도입 이후, 임의로 많은 수의 레이어를 훈련할 수 있게 되었습니다.
You can stack as many layers as you want, and you don’t have to worry about the training stability issue and stuff. 원하는 만큼 레이어를 쌓을 수 있으며, 훈련 안정성 문제 같은 것을 걱정할 필요가 없습니다.
And as discussed in Ilya’s talk two years ago, it basically says that residual connection is a variant of LSTM, but just rotated 90 degrees. 그리고 2년 전 Ilya의 발표에서 논의된 바와 같이, 기본적으로 residual connection은 LSTM의 변형이지만 90도 회전된 것이라고 말합니다.
So how do you understand this? 이것을 어떻게 이해할 수 있을까요?
If you look at LSTM, it’s a variant of recurrent net, right? LSTM을 보면, 그것은 순환 네트워크의 변형이죠?
And it’s a recurrent model process. 그리고 그것은 순환 모델 과정입니다.

If you look at LSTM, it’s a variant of recurrent net, right? LSTM을 보면, 그것은 순환 네트워크의 변형이죠?
And it’s a recurrent model process. 그리고 그것은 순환 모델 과정입니다.
So we’re going to take the hidden states from the last step, and then we’re going to have some gating mechanism, some function to produce the current states. 그래서 마지막 스텝의 히든 상태를 가져오고, 그런 다음 현재 상태를 생성하기 위한 일부 게이팅 메커니즘, 일부 함수를 가질 것입니다.
And if you look at the depth dimension, the residual connection is basically the same. 그리고 깊이 차원을 보면, residual connection은 기본적으로 동일합니다.
We’re going to take the outputs from the last layer, and then we’re going to apply some sort of function on top of it to produce the current outputs of the current layer. 마지막 레이어의 출력을 가져오고, 그런 다음 그 위에 일종의 함수를 적용하여 현재 레이어의 현재 출력을 생성합니다.
It’s just the formulation is different. 단지 공식이 다를 뿐입니다.
For example, for a residual connection, we’re going to use a fixed addition. 예를 들어, residual connection의 경우 고정된 덧셈을 사용합니다.
We’re going to add the previous hidden state with the current output. 이전 히든 상태를 현재 출력에 더합니다.
It’s just the formulation that’s different. 단지 공식이 다를 뿐입니다.
But the basic idea is the same. 하지만 기본 아이디어는 동일합니다.
It’s a recurrent that applies in the dimension of depth. 그것은 깊이 차원에 적용되는 순환입니다.
But on the other hand, we can think about reformulating this function, instead of having an LSTM, can we have an attention in the dimension of depth, and it’s going to create new possibilities because attention has been demonstrated to be so successful in the transformer era. 하지만 다른 한편으로, 이 함수를 재구성하는 것을 생각할 수 있습니다. LSTM 대신 깊이 차원에 어텐션을 가질 수 있을까요? 그리고 어텐션이 트랜스포머 시대에 매우 성공적이라는 것이 입증되었기 때문에 새로운 가능성을 만들 것입니다.
So what we’re going to do is not just to take the last hidden state, but we’re going to consider all the previous hidden states and use the attention operation, the attention mechanism, to assemble and aggregate all of these previous hidden states to compute the current state. 그래서 저희가 할 것은 마지막 히든 상태만 가져오는 것이 아니라, 모든 이전 히든 상태를 고려하고 어텐션 연산, 어텐션 메커니즘을 사용하여 이 모든 이전 히든 상태를 조립하고 집계하여 현재 상태를 계산하는 것입니다.
So this is exactly attention rotated by 90 degrees. 이것이 정확히 90도 회전된 어텐션입니다.
We view it as a natural generalization of residual connections in the LSTM analogy. 저희는 이를 LSTM 비유에서 residual connection의 자연스러운 일반화로 봅니다.
Okay, and here is the detailed formulation. 좋습니다, 그리고 여기 상세한 공식이 있습니다.
So, on the left-hand side is a standard residue. 왼쪽은 표준 residual입니다.
As I said, it is basically LSTM rotated by 90 degrees. 말씀드린 대로, 기본적으로 90도 회전된 LSTM입니다.
And the second figure is attention rotated by 90 degrees. 그리고 두 번째 그림은 90도 회전된 어텐션입니다.
So, what we did is to collect all the previous hidden states and have a simple attention operation on top of it to produce the current layer’s outcome. 그래서 저희가 한 것은 모든 이전 히든 상태를 수집하고 그 위에 간단한 어텐션 연산을 하여 현재 레이어의 결과를 생성하는 것입니다.
And of course, to increase the efficiency, to reduce the infrastructure, for example, communication and memory overhead, we also designed a new variant called block attention residue on the right-hand side. 그리고 물론 효율성을 높이고 인프라, 예를 들어 통신 및 메모리 오버헤드를 줄이기 위해, 오른쪽의 block attention residue라는 새로운 변형도 설계했습니다.
So basically, the idea is also simple. 기본적으로 아이디어도 간단합니다.
We’re going to divide all the layers in the neural networks into multiple blocks. 신경망의 모든 레이어를 여러 블록으로 나눕니다.
For example, each block can contain, say, 16 layers, or it can contain maybe four layers. 예를 들어, 각 블록은 16개 레이어를 포함하거나, 어쩌면 4개 레이어를 포함할 수 있습니다.
And then for each block, we’re going to apply this attention residue only on the output of each block. 그리고 각 블록에 대해 이 attention residue를 각 블록의 출력에만 적용합니다.
But within each block, we also still adopt this standard residue. 하지만 각 블록 내에서는 여전히 이 표준 residual을 채택합니다.
So this is going to reduce a lot of overhead while having minimal loss in terms of training accuracy. 그래서 훈련 정확도 측면에서 최소한의 손실로 많은 오버헤드를 줄일 것입니다.
And these are some of the impressive results that we achieved on this new architecture. 그리고 이것이 이 새로운 아키텍처에서 달성한 인상적인 결과 중 일부입니다.
So, on the scaling law, we can improve the token efficiency by 24%, meaning that if you have 50 trillion high-quality tokens, now you just magically have over 60 trillion tokens. 스케일링 법칙에서 토큰 효율성을 24% 개선할 수 있습니다. 즉, 50조 고품질 토큰이 있다면 이제 마법처럼 60조 이상의 토큰을 갖게 됩니다.
For the validation loss, you can also observe that it’s consistently lower than the original curve, demonstrating this stability across optimisation, and also achieves the best improvement on some of these coding, math, and reasoning heavy tasks, as shown in the benchmark results of GPQA, Math, and HumanEval. 검증 손실의 경우, 원래 곡선보다 일관되게 낮다는 것을 관찰할 수 있으며, 최적화 전반의 이 안정성을 보여주고, GPQA, Math, HumanEval의 벤치마크 결과에서 보듯이 코딩, 수학, 추론 중심 작업에서 최고의 개선을 달성합니다.
So the entire community keeps moving forward and we’re happy that we can, we’re able to contribute to the community with new technologies and some of these technologies have been sort of standard and de facto for a long time, but as you can see, we still see a lot of opportunities to improve it, to have revolutionary new design to achieve better performance. 전체 커뮤니티가 계속 앞으로 나아가고 있으며, 저희는 새로운 기술로 커뮤니티에 기여할 수 있어 기쁩니다. 이러한 기술 중 일부는 오랫동안 표준이자 사실상이었지만, 보시다시피 여전히 개선할 많은 기회가 있으며, 더 나은 성능을 달성하기 위한 혁명적인 새로운 설계를 할 수 있습니다.
If you multiply all these gains together, you can actually have a much better model. 이 모든 이득을 곱하면 실제로 훨씬 더 나은 모델을 가질 수 있습니다.
So Adam was invented in 2014, and now we scale an open-source MuonClip, a drop-in replacement for Adam. Adam은 2014년에 발명되었고, 이제 저희는 Adam의 드롭인 대체인 오픈소스 MuonClip을 스케일합니다.
And I’m sure that if you’re training Transformer LLM, it’s going to be much better if you use MuonClip instead of Adam. Transformer LLM을 훈련하고 있다면 MuonClip을 Adam 대신 사용하면 훨씬 더 나을 것이라고 확신합니다.
And Attention was invented over eight years ago. 그리고 Attention은 8년 이상 전에 발명되었습니다.
And then now we have Kimi linear, which is a linear version. 그리고 이제 선형 버전인 Kimi linear가 있습니다.
We don’t have to use full attention across all layers. 모든 레이어에 걸쳐 풀 어텐션을 사용할 필요가 없습니다.
We can have linear attention that performs better on short, long contexts at the same time. 짧은 컨텍스트와 긴 컨텍스트에서 동시에 더 잘 수행하는 선형 어텐션을 가질 수 있습니다.
And also, residual connections are now also a challenge, we scaled an open-source attention residue. 그리고 residual connection도 이제 도전이며, 저희는 오픈소스 attention residue를 스케일했습니다.
So I think one of the interesting things about our era is that we sort of adopt a different mindset for doing research. 저희 시대의 흥미로운 점 중 하나는 연구하는 데 다른 마음가짐을 채택한다는 점이라고 생각합니다.
So if we go back to 10 years ago, it’s mostly about publishing a new idea, but then I think the lack of the rigor of the experiments. 10년 전으로 돌아가면, 주로 새로운 아이디어를 발표하는 것이었지만, 실험의 엄격성이 부족했다고 생각합니다.
It’s very hard to produce solid experimental results. 견고한 실험 결과를 내는 것이 매우 어렵습니다.
But now we have this scaling ladder. 하지만 이제 이 스케일링 사다리가 있습니다.
We have enough resources to train the model and run it at different scales. 모델을 훈련하고 다른 스케일에서 실행할 충분한 자원이 있습니다.
We can have a whole set of benchmarks to measure the progress. 진행을 측정하기 위한 전체 벤치마크 세트를 가질 수 있습니다.
So it becomes easier to make confident and solid conclusion out of it. 그래서 그로부터 확신 있고 견고한 결론을 내리기가 더 쉬워집니다.
And this is one of the reasons why we are observing new progress on these ancient techniques, and I’m sure that we’ll see more and more, especially in the open source community. 그리고 이것이 이러한 고대 기술에서 새로운 진전을 관찰하는 이유 중 하나이며, 특히 오픈소스 커뮤니티에서 점점 더 많이 볼 것이라고 확신합니다.
I think we’re going to have more and more even better architectural and optimization improvement in the next few years. 앞으로 몇 년 안에 더 나은 아키텍처 및 최적화 개선이 점점 더 많이 있을 것이라고 생각합니다.
All right, so to summarize, we’re going to keep scaling our models. 좋습니다, 요약하자면, 저희는 계속 모델을 스케일할 것입니다.
And so these are three dimensions. 그리고 이것이 세 가지 차원입니다.
For example, we see different architectures and optimizers that optimize all three dimensions. 예를 들어, 세 가지 차원을 모두 최적화하는 다른 아키텍처와 옵티마이저를 봅니다.
And we’ll keep. We see new dimensions for scaling, agent swarms is not the end, and we are glad that we can move forward with the entire open source community to achieve better and better intelligence. 그리고 계속할 것입니다. 스케일링을 위한 새로운 차원을 봅니다. 에이전트 스웜이 끝이 아니며, 전체 오픈소스 커뮤니티와 함께 더 나은 지능을 달성하기 위해 앞으로 나아갈 수 있어 기쁩니다.
Thank you so much. 정말 감사합니다.


==============

Zhilin Yang(Kimi / Moonshot AI 공동창업자 겸 CEO)의 GTC 2026 키노트 “How We Scaled Kimi K2.5” 전체 내용 완전 요약
이 키노트는 약 39분 분량으로, 오픈 모델의 스케일링을 세 가지 핵심 차원에서 체계적으로 설명하고, 이를 통해 탄생한 Kimi K2.5의 기술과 성과, 그리고 차세대 아키텍처인 Attention Residue까지 공개한 발표입니다. 발표자는 “오픈 모델이 단순히 열려 있는 것만으로는 부족하며, 훌륭해야 한다”는 전제 아래, 지능의 민주화와 스케일링의 새로운 방향을 제시합니다.
1. 서론: 오픈 모델의 가치와 스케일링의 세 가지 차원
발표는 Jensen의 CES 발표 슬라이드를 인용하며 시작됩니다. 오픈 모델이 독점 모델과의 격차를 빠르게 좁히고 있으며, 프론티어에 도달하고 있다는 점을 강조합니다. 오픈 모델의 핵심 가치는 어디에나 배포할 수 있고(로컬 서버든 클라우드든), 가중치의 모든 비트를 직접 접근할 수 있다는 점입니다. 블랙박스가 아닌 완전한 투명성을 통해 전 세계 어디서든, 누구에게나 지능을 접근 가능하게 만드는 것이 목표입니다.
스케일링이 최근 몇 년간의 주요 AI 발전을 이끈 핵심 동인이라는 점을 확인한 뒤, 단순히 토큰 수를 늘리는 것을 넘어 세 가지 차원으로 스케일링을 확장해야 한다고 주장합니다.
• 토큰 효율성(Token Efficiency): 같은 토큰 수로 더 낮은 손실을 달성하도록 스케일링 곡선을 왼쪽으로 이동시키는 것. 이는 단순한 인프라 효율이 아니라 지능의 상한선 자체를 끌어올리는 일입니다. 고품질 데이터가 제한된 ‘데이터 벽’ 시대에 특히 중요합니다.
• 긴 컨텍스트(Long Context): 컨텍스트 길이를 늘려 더 복잡한 작업을 수행할 수 있게 하는 것. 에이전트가 수일, 수주, 수개월에 걸쳐 실행될 수 있도록 합니다.
• 에이전트 스웜(Agent Swarms): 단일 에이전트가 아닌 여러 에이전트를 오케스트레이션하여 병렬로 하위 작업을 수행함으로써 작업 복잡도(용량)를 높이는 새로운 학습 패러다임.
이 세 차원을 에이전트 관점에서 해석하면, 토큰 효율성은 더 강한 prior를 제공해 RL 탐색을 효율적으로 만들고, 긴 컨텍스트는 장기 실행 에이전트를 가능하게 하며, 에이전트 스웜은 다중 에이전트 협업이라는 새로운 스케일링 축을 추가합니다. 결국 각각 초장 컨텍스트와 강한 prior를 가진 에이전트들의 스웜을 구성하게 됩니다.
2. 첫 번째 차원: 토큰 효율성 – Muon Optimizer와 QK Clip
고전적인 Kaplan 등의 스케일링 법칙을 인용하며, 토큰 수·파라미터·컴퓨트를 비례적으로 늘리면 손실이 낮아진다는 점을 확인합니다. 그러나 발표의 초점은 토큰 효율성 자체를 높이는 것입니다.
Muon Optimizer는 2차 옵티마이저로, 모든 그래디언트 업데이트를 서로 직교하도록 변환합니다. Adam과 근본적으로 다르며, 제대로 구현하면 약 2배의 토큰 효율성을 얻을 수 있습니다. 대규모 학습을 위해 두 가지 핵심 기법을 적용했습니다.
• Weight Decay: 큰 모델로 스케일할 때 필수적.
• Consistent RMS Update: Adam과 비슷한 RMS를 유지하도록 조정 가능한 계수를 적용.
또한 NVIDIA GPU 클러스터 전반에서 메모리 효율을 위해 데이터 병렬 그룹에 걸쳐 상태를 분할하는 분산 Muon 구현을 개발했습니다. 같은 파라미터·같은 토큰 수로 AdamW를 Muon으로 교체하기만 해도 전반적인 성능이 크게 향상되는 실험 결과를 보여줍니다.
그러나 1조(1T) 파라미터 모델로 스케일할 때 훈련 불안정성이 발생했습니다. Max Logit이 빠르게 폭발하여 1,000을 넘고(일반적 값은 50~100 이하), 손실이 잠시 내려가다 폭발하여 수렴하지 못하는 현상이 나타났습니다.
해결책으로 QK Clip을 도입했습니다. 각 어텐션 헤드에서 Forward Pass 중 Max Logit을 계산하고, Query와 Key 프로젝션에 나누는 인자를 적용하여 최대값을 특정 범위로 제약합니다. 실험 결과, 클리핑은 손실 감소 곡선에 거의 영향을 주지 않으면서 Max Logit을 효과적으로 안정화시켰습니다. 처음에는 폭발하다가 100 근처에서 오랫동안 클립되고, 이후 자연스럽게 내려가는 패턴을 보였습니다. 신경망이 스스로 안정적인 방식을 찾아가는 모습입니다.
이 기법을 적용하여 K2 모델을 1조 파라미터로 성공적으로 스케일했으며, 이는 머신러닝 역사상 대규모 Muon 훈련의 첫 사례라고 강조합니다.
3. 두 번째 차원: 긴 컨텍스트 – Kimi Linear와 Kimi Delta Attention
Transformer와 LSTM을 비교하는 덜 알려진 그림을 제시합니다. 같은 파라미터·토큰 수에서 Transformer가 더 낮은 손실을 보이는 것은 예상된 결과이지만, 더 중요한 점은 컨텍스트 전체에서 손실이 지속적으로 감소한다는 것입니다. 토큰 인덱스가 커질수록 Transformer 손실은 계속 떨어지지만, LSTM은 일정 지점에서 포화됩니다. 이것이 Transformer가 긴 컨텍스트를 더 잘 포착하는 이유이며, 10년 전 기계번역에 쓰이던 LSTM으로는 전체 코드베이스 이해나 Linux 커널을 처음부터 작성하는 초장 에이전트 궤적 같은 작업을 수행할 수 없습니다. 에이전트 시대에 점점 더 긴 컨텍스트가 필수적이라는 점을 강조합니다.
이를 위해 Kimi Linear 아키텍처를 도입했습니다. 핵심은 Kimi Delta Attention으로, 기존의 Gated Delta Rule(GDR)을 개선한 선형 어텐션 변형입니다. 개선된 순환 메모리를 제공합니다. 동시에 선형 어텐션 레이어와 풀 어텐션 레이어를 1:2:3 비율로 혼합하여 긴 컨텍스트 능력과 효율성 사이의 균형을 맞춥니다.
기존 선형 어텐션의 문제는 전역 단일 감쇠 인자(스칼라)만 사용한다는 점입니다. 이 경우 “거의 모든 것을 잊거나” “거의 모든 것을 유지하되 불필요한 정보를 걸러내지 못하는” 두 가지 극단만 가능합니다. Kimi Delta Attention은 **세분화된 감쇠 인자(α)**를 도입하여, 이를 대각 행렬로 만듭니다. 채널별로 감쇠율을 다르게 제어할 수 있게 되어, 일부 채널은 매우 느리게 감쇠해 초장거리 정보를 유지하고, 다른 채널은 빠르게 잊어 새로운 정보를 받아들일 수 있습니다. 이로써 표현력이 크게 높아집니다.
GPU 효율을 위해 청크 단위(chunkwise) 공식을 사용합니다. 새로 도입된 α 행렬 때문에 기존처럼 쉽게 인수분해할 수 없어 인프라 도전이 커지지만, 행렬 역연산과 누적 감쇠 인자를 도입해 전체 식을 재작성함으로써 수학적으로 정확한 동등한 공식으로 병렬 구현을 가능하게 했습니다. 근사가 아니라 정확한 등가이므로 성능 손실 없이 기존 선형 어텐션 수준의 효율성을 유지하면서 표현력은 더 높습니다.
공정 비교 결과, 짧은 컨텍스트 작업(MMMU)과 긴 컨텍스트 작업(RULER) 모두에서 MLA·GDN 대비 우수하며, 특히 MLA보다 훨씬 효율적입니다. 100만 토큰 이상으로 스케일할 때도 베이스라인 대비 효율이 크게 뛰어납니다. 발표자는 이것이 짧은 컨텍스트·긴 입력·긴 출력 작업 전반에서 풀 어텐션을 능가한 첫 번째 아키텍처라고 평가합니다.
4. 세 번째 차원: 에이전트 스웜(Agent Swarms)
단일 에이전트 패러다임의 한계를 넘어, 오케스트레이터(메인 에이전트)가 하위 에이전트 그룹을 생성·할당하고, 결과를 수집하며, 이를 반복적으로 수행하는 구조를 제시합니다. 인간 사회·회사 조직과 유사합니다. CEO(오케스트레이터)가 역할을 나누고 목표를 향해 전체 조직을 움직이게 하는 것과 같습니다. AI 연구자, 웹 개발자, 물리 연구자 등 다양한 역할을 병렬로 실행한 뒤 팩트체커·파일 다운로더 등으로 결과를 조립하는 예시가 나옵니다.
작업 복잡도(x축, 모델 그룹의 정확도로 측정) 대비 실행 시간(y축) 그래프에서, 에이전트 스웜이 단일 에이전트 대비 실행 시간을 크게 줄이는 것을 보여줍니다. 100개 또는 1,000개 하위 에이전트로 스케일하면 실제 경제적 가치를 낼 수 있는 시간 내에 복잡한 작업을 완료할 수 있습니다.
스케일링 축은 다양합니다.
• 입력 스케일: 수백~수천 개의 소스를 병렬로 다운로드·읽기
• 출력 스케일: 100페이지 문헌 리뷰를 병렬로 작성
• 행동 스케일: 여러 작업의 데이터 분석 동시 수행
• 오케스트레이션 스케일: 하위 작업 설계와 결과 집계를 학습
학습을 위해 세 가지 보상 함수를 새로 정의했습니다(기존 단일 에이전트 RL과 차별화).
1 Instantiation Reward: 하위 에이전트 생성을 장려하여 “직렬 붕괴(serial collapse)”를 방지. 초기에 가중치를 높게 두고, 병렬 실행을 학습하면 감쇠.
2 Finish Reward: 생성된 하위 작업이 실제로 완료되도록 유도. 의미 없는 의사 작업만 대량 생성하는 해킹을 막고, 완료 비율을 높임. 마찬가지로 초기 고가중치 → 후기 저가중치로 감쇠.
3 Outcome Reward: 전체 작업 완료 여부를 측정하는 표준 보상.
이 세 항을 합쳐 강화 학습에 사용합니다. 병렬 실행과 다중 보상 함수를 지원하는 전체 인프라 구축이 필수적이며, 에이전트 스웜 IO 시스템의 효율성을 극대화해야 합니다.
5. Kimi K2.5: 세 차원의 결합과 성과
Muon Clip(토큰 효율성) + Kimi Delta Attention(긴 컨텍스트) + Agent Swarms를 결합하여 Kimi K2.5를 만들었습니다. 발표 시점 기준 약 한 달 전 출시된 모델입니다.
훈련 곡선은 발표자가 “인생에서 본 가장 아름다운 곡선 중 하나”라고 표현할 정도로 안정적입니다. 15조 토큰을 거치고, K2.5에서는 추가로 15조 토큰을 더 훈련했으며, Muon 도입 후 손실 스파이크가 전혀 없었습니다. NVIDIA H800 클러스터(각 노드 2TB RAM, 8 GPU, NVLink 연결)에서 수행되었습니다.
핵심 혁신은 네이티브 조인트 비전-텍스트 능력입니다. 이전 오픈 모델들은 보통 텍스트 베이스를 먼저 훈련(예: 20조 토큰)한 뒤 포스트 트레이닝으로 비전을 추가하는 Late Fusion 방식이었지만, K2.5는 Early Fusion을 채택해 0% 진행부터 비전과 텍스트 토큰을 병합합니다. 예비 실험에서 Late Fusion을 능가했습니다.
두 모달리티가 서로를 강화하는 현상이 관찰되었습니다.
• 비전이 텍스트를 개선: Vision RL(비전 작업만 사용, 수학·코딩 등 텍스트 작업 없음) 후에도 비교적 무거운 텍스트 작업 성능이 향상.
• 텍스트가 비전을 개선: 강한 텍스트 베이스가 있으면 Zero Vision SFT(비전 SFT 데이터 거의 없음, 텍스트 SFT만 사용) + 조인트 RL로도 비전 작업에서 거의 최첨단 성능을 달성.
이로 인해 비전-투-코드, 시각 디자인, 프론트엔드 코딩, 비디오를 읽고 웹사이트를 생성하거나 스타일 전이하는 등의 능력이 emergent하게 나타났습니다. 이 모든 것은 사전 훈련 단계의 성공적이고 안정적인 훈련 덕분입니다.
6. 차세대 미리보기: Attention Residue
K2.5 이후의 다음 세대 아키텍처로 Attention Residue를 공개했습니다(기술 보고서 발표 직후).
동기: 시간 차원에서 성공한 기법을 깊이(depth) 차원에 적용할 수 있는가? Residual Connection을 LSTM이 90도 회전된 것으로 보는 Ilya의 관점을 확장합니다. Residual은 이전 레이어 출력을 고정 덧셈으로 현재에 더하는 방식이지만, 본질적으로 깊이 방향의 순환입니다.
Attention Residue는 이전 히든 상태 전체를 모아 어텐션 연산으로 현재 레이어 상태를 계산합니다. “어텐션을 90도 회전한 것”이며, Residual의 자연스러운 일반화입니다.
효율을 위해 Block Attention Residue 변형을 도입했습니다. 전체 레이어를 여러 블록(예: 4~16 레이어)으로 나누고, 블록 출력에만 Attention Residue를 적용하며, 블록 내부는 표준 Residual을 유지합니다. 통신·메모리 오버헤드를 크게 줄이면서 훈련 정확도 손실은 최소화합니다.
결과: 토큰 효율성 24% 향상(50조 고품질 토큰이 마법처럼 60조 이상으로 증가). 검증 손실이 일관되게 낮고, GPQA·Math·HumanEval 등 코딩·수학·추론 중심 벤치마크에서 최고의 개선을 보였습니다.
7. 결론과 전망
Adam(2014년 발명)을 대체하는 오픈소스 MuonClip, 8년 이상 된 Attention을 개선한 Kimi Linear, Residual을 발전시킨 Attention Residue 등, “오래된 기술”에서도 여전히 혁신의 여지가 크다는 점을 강조합니다. 과거에는 새로운 아이디어 발표 위주였고 실험의 엄밀성이 부족했지만, 지금은 스케일링 사다리와 충분한 자원·벤치마크가 있어 더 확신 있는 결론을 도출할 수 있는 시대입니다.
세 차원(토큰 효율성·긴 컨텍스트·에이전트 스웜)을 계속 최적화하는 아키텍처와 옵티마이저를 발전시키고, 에이전트 스웜이 끝이 아닌 새로운 스케일링 차원도 탐색할 것입니다. 오픈소스 커뮤니티 전체와 함께 더 나은 지능을 향해 나아가겠다는 메시지로 발표를 마칩니다.
이 키노트는 단순히 특정 모델의 성과를 자랑하는 수준을 넘어, 스케일링의 패러다임 자체를 재정의하고, 오픈 모델이 어떻게 실질적으로 프론티어를 밀어낼 수 있는지를 기술적 디테일과 함께 제시한 발표로 평가할 수 있습니다.