https://youtu.be/3P1XG2BMzS8?si=vg2xYQR9juK9TSbS
That was right when I started my studies, but after that, there were two new waves.
그게 내가 공부를 막 시작했을 때였는데, 그 이후에 두 가지 새로운 물결이 있었어.
The first of these was the emergence of the Transformer architecture around 2016, and then starting in 2017, the concept of pre-training began to gain popularity.
그 중 첫 번째는 2016년경 Transformer 아키텍처의 등장이고, 그다음 2017년부터 pre-training 개념이 인기를 얻기 시작했어.
So, uh in these past few years, things have progressed extremely quickly.
그래서, 어, 지난 몇 년 동안 일이 매우 빠르게 진행됐어.
Uh I feel like I was really lucky to be able to do some work during this wave of innovation.
어, 나는 이 혁신의 물결 동안 어떤 일을 할 수 있어서 정말 운이 좋았다고 느껴.
So, your advisor is Russ, right? Professor Russ, who is also in AI and the head of AI at Apple.
그래서, 당신 지도교수가 Russ죠, 맞죠? AI에 계시고 Apple의 AI 헤드이신 Professor Russ.
Yes.
네.
Right. So, what did you learn from him?
맞아요. 그래서 그에게서 무엇을 배웠나요?
I think one thing I really felt from him is that he provides us with an open environment.
나는 그에게서 정말 느낀 한 가지는 그가 우리에게 열린 환경을 제공한다는 거야.
He’s very open to all kinds of ideas, and that gave me a lot of resources, including the environment, to do the things I wanted to do.
그는 모든 종류의 아이디어에 매우 개방적이고, 그게 내가 하고 싶은 일을 할 수 있는 환경을 포함해서 많은 자원을 줬어.
Okay. Okay. So, then later on in your career, you also had the chance to work closely with some brilliant people from Facebook AI research and the team over at Google Brain, like Coco Lee and Jason Weston, for example.
알겠어요. 알겠어요. 그래서 나중에 경력에서 Facebook AI research와 Google Brain 팀의 뛰어난 사람들, 예를 들어 Coco Lee와 Jason Weston과 가까이 일할 기회도 있었죠.
So, you were in academia and hadn’t worked in industry before, while they were in industry. How do you view this process and what did you learn from them?
당신은 학계에 있었고 이전에 산업계에서 일하지 않았는데, 그들은 산업계에 있었죠. 이 과정을 어떻게 보고, 그들에게서 무엇을 배웠나요?
I think Jason is a very problem-driven person.
나는 Jason이 매우 problem-driven한 사람이라고 생각해.
For example, right now, the issue he most wants to solve is dialogue. So, all of his work revolves around that focus.
예를 들어, 지금 그가 가장 해결하고 싶은 문제는 dialogue야. 그래서 그의 모든 일이 그 초점에 맞춰져 있어.
Mhm. And then, Cook is different from him. Cook focuses more on general methods. He prefers to use a universal approach to solve all problems.
음. 그리고 Cook은 그와 달라. Cook은 더 general methods에 초점을 맞춰. 그는 모든 문제를 해결하기 위해 보편적인 접근법을 선호해.
For example, the AutoML work he did before and our recent work on ExcelNet both reflect this idea, trying to solve problems using a general framework.
예를 들어, 그가 이전에 한 AutoML 작업과 우리의 최근 ExcelNet 작업 둘 다 이 아이디어를 반영해, 일반 프레임워크를 사용해서 문제를 해결하려고 하는 거지.
Basically, to address all kinds of problems like this.
기본적으로, 이런 모든 종류의 문제를 다루기 위해서.
So, he pays more attention to the learning aspect.
그래서 그는 learning 측면에 더 주의를 기울여.
Yes. Yes. Yes. So, actually, I think I am more influenced by Cook, mainly in how to make something first more general, and second, how to scale it up using greater computational power.
네. 네. 네. 그래서 실제로 나는 Cook에게 더 영향을 받았다고 생각해, 주로 무언가를 먼저 더 일반적으로 만드는 방법과, 둘째로 더 큰 컴퓨팅 파워를 사용해서 scale up하는 방법에서.
No. Okay. Okay, right. So, let’s talk about ExcelNet.
아니. 알겠어요. 알겠어요, 맞아요. 그래서 ExcelNet에 대해 이야기해 볼까요.
Why did you decide to work on this project at the time? I remember you wrote two papers before, one on Transformer XL and the other on mixture of softmax maxes, right?
그때 이 프로젝트를 하기로 결정한 이유는 뭐예요? 이전에 두 개의 논문을 썼던 걸로 기억하는데, 하나는 Transformer XL이고 다른 하나는 mixture of softmax maxes 맞죠?
Right.
맞아요.
What was the relationship between these two works and you, Sonar, at the time?
이 두 작업과 당시 당신, Sonar와의 관계는 뭐였나요?
Yes, I think there are two parts to this. There are two issues. One is language modeling, or what we call language modeling, and the other is pre-training.
네, 나는 여기에 두 부분이 있다고 생각해. 두 가지 이슈가 있어. 하나는 language modeling, 우리가 말하는 language modeling이고, 다른 하나는 pre-training이야.
So, it’s like the work we did before was mainly focused on language modeling, and XLNet was essentially about pre-training.
그래서, 우리가 이전에 한 일은 주로 language modeling에 초점을 맞췄고, XLNet은 본질적으로 pre-training에 관한 거였어.
At that time, we submitted Transformer XL to ICLR, right?
그때 우리는 Transformer XL을 ICLR에 제출했어, 맞죠?
The results were really impressive. It achieved state-of-the-art results on all the language modeling benchmarks, but it was rejected at the time because the reviewer challenged us asking why we were doing language modeling if it couldn’t directly improve the performance of downstream tasks.
결과가 정말 인상적이었어. 모든 language modeling 벤치마크에서 state-of-the-art 결과를 냈지만, 당시 리뷰어가 우리에게 도전하면서 language modeling을 왜 하냐고, downstream tasks의 성능을 직접 향상시키지 못하면 왜 하냐고 해서 거절됐어.
So, I started thinking about this issue, and now XLNet is an answer to that problem.
그래서 나는 이 문제에 대해 생각하기 시작했고, 지금 XLNet이 그 문제에 대한 답이야.
We essentially built a bridge between language modeling and pre-training.
우리는 본질적으로 language modeling과 pre-training 사이에 다리를 놓았어.
Right, because there used to be a gap between these two things. Language modeling only has unidirectional information, so it’s not necessarily effective for pre-training.
맞아요, 왜냐하면 이 두 가지 사이에 갭이 있었거든. Language modeling은 단방향 정보만 있어서 pre-training에 반드시 효과적이지는 않아.
Yes, exactly. That’s why they proposed BERT. It was actually meant to solve this problem, but it didn’t really solve it.
네, 정확해요. 그래서 그들이 BERT를 제안한 거야. 실제로 이 문제를 해결하려고 한 거지만, 정말 해결하지는 못했어.
The question is, why can’t language modeling be used directly? Because it’s not actually based on the autoregressive language modeling approach. It’s based on an autoencoding approach.
질문은, 왜 language modeling을 직접 사용할 수 없냐는 거야? 왜냐하면 실제로 autoregressive language modeling 접근법에 기반하지 않거든. autoencoding 접근법에 기반해.
Okay, right. And now XLNet answers this question. Why can language modeling be used for pre-training? It essentially unifies the two frameworks.
알겠어요, 맞아요. 그리고 지금 XLNet이 이 질문에 답해. 왜 language modeling을 pre-training에 사용할 수 있냐? 본질적으로 두 프레임워크를 통합해.
Okay, right. So, our current work is inspired by the previous approaches. It’s also an extension, and we’ve done some framework integration. That’s basically it.
알겠어요, 맞아요. 그래서 우리의 현재 작업은 이전 접근법에서 영감을 받았어. 확장도 되고, 프레임워크 통합도 좀 했어. 기본적으로 그게 다야.
Okay, so right now, within Google Brain, is this model currently being applied? Or what’s the situation?
알겠어요, 그래서 지금 Google Brain 내에서 이 모델이 현재 적용되고 있나요? 아니면 상황이 어떤가요?
It’s not just Google Brain. Now, many departments within Google are using it. They’re all very interested in this, and many of them have already started exploring applications of this model.
Google Brain만이 아니야. 지금 Google 내의 많은 부서들이 사용하고 있어. 모두 이것에 매우 관심이 많고, 많은 곳에서 이미 이 모델의 응용을 탐구하기 시작했어.
For example, they’re considering using this to replace the existing bird. Because ours can be directly substituted without any additional cost.
예를 들어, 기존 BERT를 대체하기 위해 이것을 사용하는 걸 고려하고 있어. 왜냐하면 우리의 것은 추가 비용 없이 직접 대체할 수 있거든.
Okay, so in this project, how do you and your collaborators divide the work? Generally speaking, what are each of your main responsibilities? Because I know you’re the lead, right? You’re definitely a key contributor. So, what do the others provide? Do they give you certain resources or something else?
알겠어요, 그래서 이 프로젝트에서 당신과 협력자들은 일을 어떻게 나누나요? 일반적으로 각자의 주요 책임은 뭐예요? 당신이 리드인 걸 알거든요, 맞죠? 분명히 핵심 기여자예요. 그래서 다른 사람들은 무엇을 제공하나요? 특정 자원을 주나요, 아니면 다른 거?
Uh Zihang and I are mainly co-leads, so we are responsible for the entire experiment, the main ideas, the code, and the paper. It’s all done by the two of us.
어 Zihang과 나는 주로 co-lead야, 그래서 우리가 전체 실험, 주요 아이디어, 코드, 논문을 담당해. 우리 둘이 다 해.
The others mainly contribute in ways like um helping with ideas, providing computing resources, and offering some guidance on the direction.
다른 사람들은 주로 아이디어를 돕고, 컴퓨팅 자원을 제공하고, 방향에 대한 가이드를 주는 식으로 기여해.
Right, okay. So, after this, or actually during this specific period of time, you could say, while you were working hard on finishing your PhD, you also simultaneously got involved with a brand new startup company, right? It’s called Recurrent AI, right?
맞아요, 알겠어요. 그래서 이 후에, 또는 실제로 이 특정 기간 동안, 당신이 PhD를 끝내려고 열심히 일하면서 동시에 새로운 스타트업 회사에 관여했다고 할 수 있죠, 맞죠? Recurrent AI라는 거죠?
Yeah.
응.
So, could you talk a bit about what this company does? What problems it aims to solve, and how far along it is now?
그래서 이 회사가 무엇을 하는지 조금 이야기해 줄 수 있나요? 어떤 문제를 해결하려고 하고, 지금 어디까지 왔나요?
What we’re doing now is AI for sales using artificial intelligence technology to empower enterprises in the communication between their sales teams and customers.
우리가 지금 하는 건 sales를 위한 AI야, 인공지능 기술을 사용해서 기업의 sales 팀과 고객 간 커뮤니케이션을 강화하는 거지.
Specifically, what we do is for their phone sales or text-based sales. If it’s phone sales, we first convert the calls to text. And for text-based sales, we can directly use the text.
구체적으로, 전화 sales나 text 기반 sales를 위해 해. 전화 sales면 먼저 통화를 text로 변환하고. text 기반 sales면 text를 직접 사용해.
Then, we perform some analysis on this data to help improve their sales conversion rates.
그다음 이 데이터에 분석을 해서 sales conversion rates를 높이도록 도와.
There are mainly three scenarios where this applies.
주로 세 가지 시나리오에 적용돼.
The first scenario is conducting comprehensive all-channel sales quality inspections. For example, we have large clients who may have thousands of sales representatives, and every day we perform quality checks on all of their phone sales.
첫 번째 시나리오는 종합적인 all-channel sales quality inspections를 하는 거야. 예를 들어, 수천 명의 sales representatives를 가진 큰 클라이언트가 있고, 매일 그들의 모든 전화 sales에 quality checks를 해.
The second scenario is recommending sales leads. In this area, we can help them match with the best leads.
두 번째 시나리오는 sales leads를 추천하는 거야. 이 영역에서 최고의 leads와 매칭하도록 도와.
Um the third scenario is analyzing customer profiles.
어 세 번째 시나리오는 customer profiles를 분석하는 거야.
In the mining industry in this specific area, we can provide essential assistance to help these clients more effectively match their individual customers with the most appropriate and right solutions for their needs.
mining 산업의 이 특정 영역에서, 클라이언트가 개별 고객을 가장 적절하고 맞는 솔루션과 더 효과적으로 매칭하도록 필수적인 도움을 줄 수 있어.
Okay.
알겠어요.
Knowing their profile, you can suggest the right products or services.
그들의 profile을 알면 맞는 products나 services를 제안할 수 있죠.
So, it sounds like a B2B model.
B2B 모델처럼 들리네요.
Yes, we’re a business provider.
네, 우리는 business provider야.
So, the lead recommendation is actually lead generation, right? Do you mine from public information or from existing market activities and so on? What does that process look like?
그래서 lead recommendation이 실제로 lead generation이죠? public information에서 mine하거나 기존 market activities 등에서 하나요? 그 과정은 어떻게 생겼나요?
We make list recommendations based on phone calls and their historical communication data.
우리는 전화 통화와 그들의 historical communication data에 기반해서 list recommendations를 해.
That is, the list already exists within these calls or text records.
즉, 리스트가 이미 이 통화나 text records 안에 존재해.
Is this process currently feasible in practice?
이 과정이 현재 실제로 실행 가능한가요?
Uh what we have can be fully automated. That is, as long as all the calls and text data are there, we can make recommendations directly based on that content without any manual intervention.
어 우리가 가진 건 완전히 자동화될 수 있어. 즉, 모든 통화와 text data가 있으면, 수동 개입 없이 그 content에 기반해서 직접 recommendations를 할 수 있어.
This is different from traditional methods. Uh traditionally, there would be people offline specifically using Excel spreadsheets to organize that information, but that approach is hard to scale and not particularly accurate.
이건 전통적인 방법과 달라. 어 전통적으로는 offline에서 사람들이 특별히 Excel spreadsheets를 사용해서 그 정보를 정리했는데, 그 접근법은 scale하기 어렵고 특별히 정확하지도 않아.
In the future, we can fully automate this process. And in terms of real-time capability, right now, we update once a day. That’s because the features change a lot every day. And at this scale, we’ve been able to observe pretty good results.
미래에 우리는 이 과정을 완전히 자동화할 수 있어. 그리고 real-time capability 측면에서, 지금은 하루에 한 번 update해. features가 매일 많이 변하기 때문이야. 그리고 이 scale에서 꽤 좋은 결과를 관찰할 수 있었어.
Okay.
알겠어요.
Yes.
네.
How is the application of your solution among your clients now? What is the current progress?
지금 클라이언트들 사이에서 당신의 솔루션 적용은 어떤가요? 현재 진행 상황은?
Yes, right now, we have uh in the fields of internet, education, and finance, we have leading clients in all three sectors.
네, 지금 우리는 어 internet, education, finance 분야에서 세 섹터 모두에 leading clients가 있어.
Okay, so does that mean um for these internet, education, and finance clients, their scenarios involve having a large amount of data or are more suited to this kind of tele sales model, right? So, these kinds of clients are
알겠어요, 그래서 그게 의미하는 건 어 이 internet, education, finance 클라이언트들의 시나리오가 대량의 데이터를 가지거나 이런 tele sales 모델에 더 적합하다는 거죠, 맞죠? 그래서 이런 클라이언트들이
You uh probably more suitable.
당신 어 아마 더 적합할 거야.
Yes, exactly. They have a large amount of communication data.
네, 정확해요. 그들은 대량의 communication data를 가지고 있어.
Mhm.
음.
Yes, and what we do is take this kind of uh this kind of data and turn it into insights that can actually generate value.
네, 그리고 우리가 하는 건 이런 어 이런 데이터를 가져와서 실제로 value를 생성할 수 있는 insights로 바꾸는 거야.
Okay, right.
알겠어요, 맞아요.
Because I think this is the difference between labeled and unlabeled data. The data they originally have is unlabeled.
왜냐하면 이게 labeled와 unlabeled data의 차이라고 생각해. 그들이 원래 가진 데이터는 unlabeled야.
Right, but what we do is we now have tens of thousands of hours of annotated phone call recordings in vertical domains. And on top of that, we’ve also done semantic annotation of the text.
맞아요, 하지만 우리가 하는 건 지금 vertical domains에서 수만 시간의 annotated phone call recordings를 가지고 있어. 그리고 그 위에 text의 semantic annotation도 했어.
It’s precisely through this large volume of labeled data that we’re able to truly turn it into value. Moreover, this annotated data can be transferred between different industries or even between different companies within the same industry, so it’s quite scalable.
바로 이 대량의 labeled data를 통해 진정으로 value로 바꿀 수 있어. 게다가 이 annotated data는 다른 산업 간, 또는 같은 산업 내 다른 회사 간에도 전이될 수 있어서 꽤 scalable해.
So, this actually addresses an efficiency issue in sales, right? In other words, it improves efficiency. And then
그래서 이게 실제로 sales의 efficiency 이슈를 다루죠, 맞죠? 즉, efficiency를 높여요. 그리고
Cut costs, save time.
비용을 줄이고, 시간을 절약하고.
Right, reducing costs and increasing efficiency. So, what are the technical challenges you’ve encountered in this process?
맞아요, 비용을 줄이고 efficiency를 높이는 거죠. 그래서 이 과정에서 마주친 technical challenges는 뭐예요?
I think actually the hardest part is uh you finding a scenario where we can apply these uh state-of-the-art technologies.
나는 실제로 가장 힘든 부분이 어 이런 state-of-the-art technologies를 적용할 수 있는 시나리오를 찾는 거라고 생각해.
Mhm.
음.
Right, so this scenario needs to meet uh several conditions. The first is that it must create significant value for the customer. The second is that the technology you use in this scenario must be something critical, something absolutely essential. Otherwise, even if you use uh state-of-the-art technology and improve by 10 points, if the customer doesn’t perceive it, uh then it’s useless.
맞아요, 그래서 이 시나리오는 어 여러 조건을 충족해야 해. 첫 번째는 고객에게 상당한 value를 만들어야 한다는 거야. 두 번째는 이 시나리오에서 사용하는 technology가 critical하고 절대적으로 필수적인 것이어야 해. 그렇지 않으면 state-of-the-art technology를 사용해서 10 points 향상시켜도 고객이 인식하지 못하면 어 쓸모없어.
Uh the third point is that it should be a supervised learning problem because we encounter many customer needs that are, for example, unsupervised learning problems. At this point, we need to find a way to uh turn it from an unsupervised problem into a supervised learning problem.
어 세 번째 포인트는 supervised learning 문제여야 한다는 거야, 왜냐하면 우리가 마주치는 많은 고객 needs가 예를 들어 unsupervised learning 문제거든. 이때 우리는 어 unsupervised 문제에서 supervised learning 문제로 바꾸는 방법을 찾아야 해.
So, um I think the scenarios we’re working on now all meet these criteria pretty well. That’s why we’ve already applied research results like Transformer XL to our online systems. In these production environments, we’ve indeed achieved significant improvements. And because they meet the criteria mentioned earlier, customers can perceive the value and are willing to pay for it.
그래서, 어 우리가 지금 작업하는 시나리오들이 이 기준들을 꽤 잘 충족한다고 생각해. 그래서 이미 Transformer XL 같은 research results를 online systems에 적용했어. 이 production environments에서 실제로 상당한 improvements를 달성했어. 그리고 앞서 언급한 기준을 충족하기 때문에 고객들이 value를 인식하고 돈을 지불할 의사가 있어.
Okay.
알겠어요.
Yes.
네.
I understand the first two points pretty well. The first is that it’s a real necessity, right? And the second one also has a critical impact, right? What you’re doing can actually bring about substantial improvement, correct?
첫 두 포인트는 꽤 잘 이해해요. 첫 번째는 실제 필요성이라는 거죠, 맞죠? 그리고 두 번째도 critical impact가 있고요, 맞죠? 당신이 하는 일이 실제로 상당한 improvement를 가져올 수 있다는 거죠, 맞죠?
Yes.
네.
But I don’t quite understand the third point. Why do you emphasize turning unsupervised information into light supervision? Is it because your methods are more learning-based, or is there another reason?
하지만 세 번째 포인트는 잘 이해가 안 돼요. 왜 unsupervised information을 light supervision으로 바꾸는 걸 강조하나요? 당신의 methods가 더 learning-based여서인가요, 아니면 다른 이유가 있나요?
I think it’s like this. There are two aspects here. Uh the first aspect is that I think for deep learning, including the whole NLP framework, to achieve results
나는 이렇게 생각해. 여기에 두 측면이 있어. 어 첫 번째 측면은 deep learning, 전체 NLP 프레임워크를 포함해서 결과를 내기 위해서는
Mhm.
음.
Now, uh it has to be done in a supervised learning scenario, because if you’re just doing clustering or something like that, actually uh using traditional methods or using an embedding for clustering yields about the same results.
지금, 어 supervised learning 시나리오에서 해야 해, 왜냐하면 그냥 clustering 같은 걸 하면, 실제로 어 traditional methods를 사용하거나 embedding을 사용해서 clustering하는 게 거의 같은 결과를 내거든.
Of course, the second point is that compared to the past, our current supervised learning requires fewer labels because we have this kind of unsupervised approach.
물론, 두 번째 포인트는 과거와 비교해서 우리의 현재 supervised learning이 더 적은 labels를 필요로 한다는 거야, 이런 unsupervised approach를 가지고 있기 때문에.
Yes, right. Like ExcelNet, which is an example of pre-training, so we can effectively reduce the labeling cost for supervised learning.
네, 맞아요. ExcelNet처럼, 그게 pre-training의 예시니까, supervised learning의 labeling cost를 효과적으로 줄일 수 있어.
Uh-huh.
응.
Yeah, so these two things are not contradictory.
응, 그래서 이 두 가지는 모순되지 않아.
Uh that is to say, in the end, the problem still needs to be defined as a supervised learning task, whether it’s classification or something like structured prediction.
어 즉, 결국 문제는 여전히 supervised learning task로 정의되어야 해, classification이든 structured prediction 같은 것이든.
But ultimately, we hope to use unsupervised learning to improve this process. So, one aspect is the definition of the problem, and the other is the method itself. Okay, these two are different.
하지만 궁극적으로 우리는 unsupervised learning을 사용해서 이 과정을 향상시키고 싶어. 그래서 한 측면은 문제의 정의이고, 다른 측면은 method 자체야. 알겠어, 이 두 가지는 달라.
When you first chose this direction, uh what was your reason? Because NLP has a lot of tasks, a lot of problems it can solve, right?
이 방향을 처음 선택했을 때, 어 이유가 뭐였나요? NLP에는 많은 tasks가 있고, 해결할 수 있는 많은 문제가 있잖아요, 맞죠?
Right.
맞아요.
I mean, there are there are there are dialogues and all kinds of things, so why did you choose this field as the direction for your startup?
내 말은, dialogue도 있고 온갖 것들이 있는데, 왜 이 분야를 스타트업의 방향으로 선택했나요?
What we actually want to do, our ultimate vision is to use AI to improve the efficiency of all kinds of human communication in society.
우리가 실제로 하고 싶은 건, 우리의 궁극적인 vision은 AI를 사용해서 사회에서 모든 종류의 인간 커뮤니케이션의 efficiency를 향상시키는 거야.
Yep.
응.
Right. And there are many types of communication. Right now, we’re focusing on the communication between sales and customers within enterprise services, because I think this area the value that can be brought by reducing costs and increasing efficiency, as well as the scenarios involved, better fit the three criteria I just mentioned.
맞아요. 그리고 많은 종류의 커뮤니케이션이 있어. 지금 우리는 enterprise services 내에서 sales와 고객 간 커뮤니케이션에 초점을 맞추고 있어, 왜냐하면 이 영역이 비용을 줄이고 efficiency를 높여서 가져올 수 있는 value와 관련된 시나리오가 내가 방금 언급한 세 가지 기준에 더 잘 맞는다고 생각하거든.
Okay.
알겠어요.
Right. So, I think in this area, our technology can potentially create even greater value.
맞아요. 그래서 이 영역에서 우리의 technology가 잠재적으로 더 큰 value를 만들 수 있다고 생각해.
Okay. Your main market is China right now, correct?
알겠어요. 지금 주요 시장이 중국이죠, 맞죠?
Yes, that’s right.
네, 맞아요.
Okay. So, um could you tell me a bit more about your other co-founders? Because I actually know that a few of them are also from CMU as well, right?
알겠어요. 그래서, 어 다른 co-founders에 대해 조금 더 이야기해 줄 수 있나요? 사실 그들 중 몇 명도 CMU 출신인 걸 알거든요, 맞죠?
Uh
어
Yes.
네.
Yes, that is exactly right. We currently have four partners in total.
네, 정확히 맞아요. 우리는 현재 총 네 명의 partners가 있어.
Edward, Yutao, and I were all in the very same research lab when I was an undergraduate at Tsinghua. We were all working in that same lab together back then. Uh so, that’s how we met back then.
Edward, Yutao, 그리고 나는 내가 Tsinghua에서 학부생일 때 같은 research lab에 있었어. 그때 모두 같은 lab에서 함께 일했어. 어 그래서 그때 그렇게 만났어.
Uh later Edward went to CMU for his master’s, and Yutao stayed at Tsinghua for his PhD. Yeah, and then we have another partner, Jeff.
어 나중에 Edward는 master’s를 위해 CMU에 갔고, Yutao는 PhD를 위해 Tsinghua에 남았어. 응, 그리고 또 다른 partner인 Jeff가 있어.
So, now there are a few of us, and I think we maybe each of us has a slightly different background and focus.
그래서 지금 우리 몇 명이 있고, 우리 각자가 약간 다른 background와 focus를 가지고 있다고 생각해.
Uh for example, Yutao, he was previously a core developer, probably the most central developer, of a big data analytics platform produced by Tsinghua called a minor.
어 예를 들어, Yutao, 그는 이전에 Tsinghua에서 만든 big data analytics platform인 a minor의 core developer, 아마 가장 중심적인 developer였어.
That platform had millions of unique IP visits. So, the traffic was huge, and the amount of data behind it was also massive, with billions of paper records.
그 platform은 수백만 unique IP visits가 있었어. 그래서 traffic이 엄청났고, 그 뒤의 데이터양도 방대해서 수십억 paper records가 있었어.
Right. So, aside from his academic research abilities, he also has particularly strong engineering skills in this area.
맞아요. 그래서 그의 academic research 능력 외에, 이 영역에서 특히 강한 engineering skills도 가지고 있어.
Then there’s Edward. He’s a serial entrepreneur and has had some entrepreneurial experience in the US before. So, when it comes to products and various other aspects, he’s been very helpful to us.
그다음 Edward가 있어. 그는 serial entrepreneur이고 이전에 US에서 일부 entrepreneurial experience가 있어. 그래서 products와 다양한 다른 측면에서 우리에게 매우 도움이 됐어.
And then there’s Jeff, who is also a serial entrepreneur. Previously at another company, he once he created this kind of achievement, taking the company from zero to nine figures in revenue. He led the team to reach that goal.
그리고 Jeff가 있는데, 그도 serial entrepreneur야. 이전에 다른 회사에서 이런 성과를 낸 적이 있어, 회사를 zero에서 nine figures revenue로 가져갔어. 그 목표에 도달하도록 팀을 이끌었어.
So, now in terms of skill sets, we’re relatively well-rounded in all aspects.
그래서 지금 skill sets 측면에서 우리는 모든 측면에서 상대적으로 well-rounded해.
Okay. Okay. So, during your entrepreneurial journey, this is your first time starting a business, right? So, up to now, what has been your biggest gain or takeaway?
알겠어요. 알겠어요. 그래서 당신의 entrepreneurial journey 동안, 이게 처음으로 비즈니스를 시작하는 거죠, 맞죠? 그래서 지금까지 가장 큰 gain이나 takeaway는 뭐예요?
I think the biggest learning is I used to think that if you were really good at one particular technology, you could build a successful startup and change the world.
나는 가장 큰 learning이, 예전에 한 특정 technology에 정말 잘하면 성공적인 스타트업을 만들고 세상을 바꿀 수 있다고 생각했어.
But later, I realized that’s not the case. For example, your your your scenario needs to meet the three conditions I just mentioned.
하지만 나중에 그게 아니라는 걸 깨달았어. 예를 들어, 당신의 시나리오가 내가 방금 언급한 세 조건을 충족해야 해.
And when it comes to AI, uh another very important lesson I’ve learned is that for AI startups, I think there are three essential production factors.
그리고 AI에 관해서, 어 내가 배운 또 다른 매우 중요한 lesson은 AI startups에 대해 세 가지 필수 production factors가 있다고 생각해.
One is the scenario I just mentioned. Another is people, uh and the third is data.
하나는 내가 방금 언급한 시나리오. 다른 하나는 people, 어 그리고 세 번째는 data.
Uh so, right now, we’re working hard on all three of these aspects. Uh I think what’s really important, and these two points actually illustrate the same thing, is that a key realization for me is this.
어 그래서 지금 우리는 이 세 측면 모두에 열심히 일하고 있어. 어 정말 중요한 건, 그리고 이 두 포인트가 실제로 같은 걸 설명하는데, 나에게 key realization은 이거야.
Previously, I might have been absolutely focused on uh academic or technical pursuits.
이전에는 어 academic이나 technical pursuits에 절대적으로 초점을 맞췄을 수도 있어.
But now, it’s shifted to pursuing things like the entire ecosystem, everything uh within the company from top to bottom, including all the factors of production.
하지만 지금은 entire ecosystem, 회사 내 top부터 bottom까지 모든 것, 모든 production factors를 포함해서 추구하는 것으로 바뀌었어.
I think I’ve really undergone a change in perspective.
나는 정말 perspective의 변화를 겪었다고 생각해.
But ultimately, you still have to build a product, right?
하지만 궁극적으로 여전히 product를 만들어야 하죠, 맞죠?
Right.
맞아요.
Of course, with a product, you need to solve a problem, right? So, my understanding of a product is basically the problem that the customer is facing, right?
물론, product와 함께 문제를 해결해야 하죠, 맞죠? 그래서 product에 대한 내 이해는 기본적으로 고객이 직면한 문제예요, 맞죠?
Right. The problem the customer is facing. How do you define this product? And then, from a technical perspective, how do you turn this product into a a definition of a problem and then use your tools to solve it?
맞아요. 고객이 직면한 문제. 이 product를 어떻게 정의하나요? 그리고 technical perspective에서 이 product를 문제의 정의로 바꾸고 당신의 tools를 사용해서 해결하는 건 어떻게 하나요?
Ideally, you can use your strengths in the process.
이상적으로는 과정에서 당신의 strengths를 사용할 수 있어.
Yes.
네.
So, the technologies you are best at or the ones you are currently doing best at.
그래서 당신이 가장 잘하는 technologies나 현재 가장 잘하고 있는 것들.
Right. Yes.
맞아요. 네.
That way, you’ll be in a leading position in the competition.
그렇게 하면 competition에서 leading position에 있을 거야.
And you’ll have a barrier. After all, you always need some kind of barrier, right? For us, using technology as a barrier might be a good choice.
그리고 barrier를 가질 거야. 결국, 항상 어떤 barrier가 필요하잖아요, 맞죠? 우리에게는 technology를 barrier로 사용하는 게 좋은 선택일 수 있어.
Right. Right. Right. Exactly. So, the company is still
ㅇㅇㅇㅇㅇㅇㅇㅇㅇ
이 영상은 Kimi(문샷 AI) CEO 양지린(Yang Zhilin)이 학계에서 산업계로 넘어가는 초기 경력, 특히 Transformer-XL와 XLNet 연구, 그리고 박사 과정 중에 공동 창업한 스타트업 Recurrent AI에 대해 이야기하는 인터뷰이다. 인터뷰어는 그의 지도교수, 협력자, 연구 동기, 창업 이유와 기술적·사업적 통찰을 차례로 질문한다.
양지린은 자신이 연구를 시작할 무렵을 회상하며 시작한다. 그 직후 AI 분야에 두 가지 큰 물결이 왔다고 말한다. 첫 번째는 2016년경 Transformer 아키텍처의 등장이고, 두 번째는 2017년부터 본격적으로 주목받기 시작한 pre-training 개념이다. 이후 몇 년 동안 분야가 극도로 빠르게 발전했으며, 자신은 이 혁신의 물결 속에서 직접 연구에 참여할 수 있었던 것을 큰 행운으로 여긴다고 강조한다.
지도교수에 대한 질문이 이어진다. 인터뷰어는 그의 지도교수가 Russ(러스, Apple AI 헤드)임을 확인하고, 그에게서 무엇을 배웠는지 묻는다. 양지린은 Russ가 매우 개방적인 환경을 제공했다고 답한다. 교수는 모든 종류의 아이디어에 열린 태도를 보였고, 덕분에 자신이 원하는 연구를 수행할 수 있는 자원과 환경이 충분히 주어졌다고 회상한다.
이후 경력에서 Facebook AI Research와 Google Brain의 뛰어난 연구자들과 가까이 일할 기회가 있었다고 언급된다. 특히 Jason Weston과 Coco Lee(Cook으로 불림)가 예시로 나온다. 양지린은 당시 자신은 학계에만 있었고 산업 경험이 없었던 반면, 그들은 이미 산업계에 있었음을 지적한다. 이들로부터 배운 점을 묻자, 양지린은 두 사람의 스타일을 대비해 설명한다.
Jason Weston은 매우 problem-driven한 사람이라고 한다. 현재 그가 가장 집중하는 문제는 dialogue(대화)이며, 모든 연구가 그 한 가지 초점에 맞춰져 있다고 설명한다. 반면 Coco Lee(Cook)는 그와 다르게 일반적 방법론에 더 집중한다. 그는 보편적 프레임워크로 모든 문제를 해결하려는 접근을 선호한다. 그가 이전에 한 AutoML 작업과 양지린 팀이 최근 한 XLNet 작업이 바로 이 철학을 반영한다고 말한다. 양지린은 자신이 Cook의 영향을 더 많이 받았다고 밝힌다. 구체적으로는 “무언가를 먼저 더 일반화시키는 방법”과 “더 큰 컴퓨팅 파워를 사용해 스케일업하는 방법”을 배웠다고 한다.
다음으로 XLNet 프로젝트에 대한 질문이 이어진다. 인터뷰어는 이전에 Transformer-XL와 mixture of softmax에 관한 두 편의 논문을 썼던 것을 상기시키며, 왜 XLNet을 하게 되었는지, 이 작업들과의 관계가 무엇인지 묻는다. 양지린은 두 가지 이슈가 있었다고 설명한다. 하나는 language modeling이고, 다른 하나는 pre-training이다. 이전 작업은 주로 language modeling에 집중했지만, XLNet은 본질적으로 pre-training을 다루었다.
당시 Transformer-XL를 ICLR에 제출했을 때 결과는 매우 인상적이었다. 모든 주요 language modeling 벤치마크에서 state-of-the-art를 달성했다. 그러나 리뷰어가 “language modeling을 하는 것이 downstream task 성능을 직접 향상시키지 못한다면 왜 하는가?”라고 도전하며 논문을 거절했다. 이 거절이 계기가 되어 양지린은 language modeling과 pre-training 사이의 간극을 bridging하는 방법을 고민하기 시작했고, XLNet이 바로 그 문제에 대한 답이라고 말한다.
그는 language modeling이 단방향 정보만 가지기 때문에 pre-training에 반드시 효과적이지 않다는 기존의 갭을 지적한다. BERT가 이 문제를 해결하려고 제안되었지만, 실제로 완전한 해결책이 되지 못했다고 평가한다. BERT는 autoregressive language modeling이 아니라 autoencoding 접근에 기반하기 때문이다. XLNet은 “왜 language modeling이 pre-training에 사용될 수 있는가”라는 질문에 답하며, 두 프레임워크를 본질적으로 통합한다. 현재 작업은 이전 접근법에서 영감을 받은 확장과 프레임워크 통합이라고 정리한다.
Google Brain 내에서의 적용 상황을 묻자, 양지린은 Google Brain뿐만 아니라 Google 내 여러 부서가 이미 관심을 가지고 있으며, 기존 BERT를 대체하는 용도로 탐색 중이라고 답한다. 자신들의 모델은 추가 비용 없이 직접 대체할 수 있다는 장점이 있다고 강조한다.
프로젝트 내 역할 분담에 대해 묻자, 양지린은 Zihang과 자신이 주로 co-lead라고 말한다. 두 사람이 전체 실험, 주요 아이디어, 코드, 논문 작성까지 담당했고, 다른 협력자들은 아이디어 제안, 컴퓨팅 자원 제공, 방향성에 대한 가이드 정도로 기여했다고 설명한다.
박사 과정을 마무리하는 동시에 새로운 스타트업 Recurrent AI에 참여하게 된 경위를 묻는다. 양지린은 이 회사가 AI for sales를 하고 있다고 설명한다. 인공지능 기술로 기업의 세일즈 팀과 고객 간 커뮤니케이션을 강화하는 것이 목표다. 구체적으로는 전화 세일즈나 텍스트 기반 세일즈 데이터를 다룬다. 전화 세일즈의 경우 통화를 먼저 텍스트로 변환하고, 텍스트 세일즈는 바로 텍스트를 사용한다. 그 후 이 데이터를 분석해 세일즈 전환율을 높이는 인사이트를 제공한다.
적용 시나리오는 크게 세 가지다.
1 전 채널 세일즈 품질 검사: 수천 명의 세일즈 담당자를 둔 대형 고객사를 위해 매일 모든 전화 세일즈를 품질 점검한다.
2 세일즈 리드 추천: 최적의 리드와 매칭을 돕는다.
3 고객 프로필 분석: 특히 마이닝 산업 등에서 개별 고객에게 가장 적합한 솔루션을 매칭하도록 돕는다.
이는 B2B 모델이며, 리드 추천은 기존 통화·텍스트 기록 안에 이미 존재하는 리스트를 기반으로 한다. 현재 프로세스는 완전 자동화가 가능하다고 한다. 통화와 텍스트 데이터만 있으면 수동 개입 없이 바로 추천할 수 있다. 전통적인 방식은 오프라인에서 사람들이 엑셀로 정리하는 것이었는데, 스케일하기 어렵고 정확도도 낮았다. 실시간성은 현재 하루 한 번 업데이트하는 수준이며, 피처가 매일 많이 변하기 때문이다. 이 규모에서도 꽤 좋은 결과를 관찰했다고 한다.
클라이언트 적용 현황을 묻자, 인터넷·교육·금융 세 분야에서 모두 선도 기업 고객을 확보했다고 답한다. 이들 고객은 대량의 커뮤니케이션 데이터를 가지고 있어 텔레세일즈 모델에 적합하다. 핵심은 unlabeled 데이터를 labeled 데이터로 바꿔 가치를 만드는 것이다. 이미 vertical domain에서 수만 시간의 annotated 전화 녹음과 텍스트 semantic annotation을 확보했고, 이 대량의 labeled 데이터 덕분에 실제로 가치를 창출할 수 있게 되었다. 이 annotated 데이터는 산업 간, 같은 산업 내 회사 간에도 전이 가능해 확장성이 높다.
이 작업이 세일즈의 효율성 문제를 해결하는 것이라고 인터뷰어가 정리하자, 양지린은 비용 절감과 효율성 향상이라고 동의한다. 기술적 도전을 묻자, 가장 어려운 부분은 state-of-the-art 기술을 적용할 수 있는 적절한 시나리오를 찾는 것이라고 답한다. 그 시나리오는 세 가지 조건을 충족해야 한다.
1 고객에게 상당한 가치를 만들어야 한다.
2 사용하는 기술이 절대적으로 필수적이고 critical해야 한다. 그렇지 않으면 10점 향상시켜도 고객이 체감하지 못하면 무의미하다.
3 supervised learning 문제여야 한다. 많은 고객 니즈가 unsupervised 문제인데, 이를 supervised로 전환하는 방법을 찾아야 한다.
현재 작업 중인 시나리오들은 이 조건을 잘 충족한다고 말한다. 그래서 Transformer-XL 같은 연구 결과를 이미 온라인 시스템에 적용했고, production 환경에서 유의미한 성능 향상을 달성했다. 고객이 가치를 체감하고 비용을 지불할 의사가 있다고 한다.
세 번째 조건(unsupervised를 supervised로 바꾸는 것)에 대해 추가 질문이 나온다. 양지린은 두 가지 측면이 있다고 설명한다. 첫째, deep learning과 NLP 프레임워크가 실질적인 결과를 내려면 supervised learning 시나리오여야 한다. 단순 clustering은 전통 방법이나 embedding을 써도 비슷한 결과가 나온다. 둘째, 과거보다 현재 supervised learning은 더 적은 라벨로도 가능하다. 왜냐하면 unsupervised pre-training(XLNet 같은)이 있기 때문이다. 결국 문제는 supervised task로 정의되어야 하지만, unsupervised learning으로 그 과정을 개선하는 것이 목표라고 정리한다. 문제 정의와 방법론은 서로 다른 차원이라고 강조한다.
왜 이 방향(세일즈 커뮤니케이션)을 창업 분야로 선택했는지 묻자, 양지린은 궁극적 비전이 “AI로 사회 속 모든 종류의 인간 커뮤니케이션 효율을 향상시키는 것”이라고 답한다. 그중에서도 기업 서비스 내 세일즈-고객 커뮤니케이션에 먼저 집중한 이유는, 비용 절감·효율 향상으로 가져올 수 있는 가치와 시나리오가 앞서 말한 세 가지 조건에 가장 잘 맞기 때문이라고 설명한다. 이 영역에서 기술이 더 큰 가치를 만들 수 있다고 본다. 현재 주요 시장은 중국이다.
공동창업자들에 대한 질문이 이어진다. 양지린은 현재 네 명의 파트너가 있다고 말한다. Edward, Yutao, 자신은 모두 칭화대 학부 시절 같은 연구실에서 일하며 만났다. 이후 Edward는 CMU에서 석사를 했고, Yutao는 칭화대에서 박사 과정을 이어갔다. 또 다른 파트너 Jeff가 있다.
각자의 강점을 설명한다. Yutao는 이전에 칭화대에서 만든 빅데이터 분석 플랫폼(a minor)의 핵심 개발자였으며, 수백만 unique IP, 수십억 건의 논문 데이터를 다룬 경험이 있어 엔지니어링 역량이 특히 강하다. Edward는 시리얼 엔터프레너로 미국에서 창업 경험이 있어 제품과 다양한 측면에서 도움을 준다. Jeff 역시 시리얼 엔터프레너로, 이전 회사에서 매출을 0에서 9자리 숫자까지 끌어올린 경험이 있다. 덕분에 팀의 skill set이 전반적으로 균형 잡혀 있다고 평가한다.
처음으로 창업을 하는 입장에서 지금까지의 가장 큰 배움이 무엇인지 묻자, 양지린은 중요한 인식 변화를 이야기한다. 예전에는 “특정 기술 하나만 정말 잘하면 성공적인 스타트업을 만들고 세상을 바꿀 수 있다”고 생각했다. 그러나 나중에 그게 아니라는 것을 깨달았다. 시나리오가 앞서 말한 세 조건을 충족해야 하고, AI 스타트업에는 세 가지 필수 생산요소가 있다고 본다.
1 시나리오
2 사람
3 데이터
현재 세 가지 모두에 집중하고 있다. 가장 중요한 깨달음은, 이전에는 학문적·기술적 추구에만 절대적으로 집중했다면, 이제는 회사 전체 생태계(위에서 아래까지 모든 생산요소 포함)를 추구하는 것으로 관점이 바뀌었다는 점이다.
인터뷰어가 “결국 제품을 만들어야 하지 않느냐”고 묻자, 양지린은 동의한다. 제품은 고객이 직면한 문제를 해결하는 것이라고 이해한다. 이상적으로는 자신의 강점(가장 잘하는 기술)을 활용해 문제를 정의하고 해결함으로써 경쟁에서 선도적 위치를 차지하고, 기술 자체를 barrier로 삼는 것이 좋은 선택이라고 말한다.
이 인터뷰는 양지린이 순수 연구자에서 창업자로 전환하는 과정의 사고 변화를 생생하게 보여준다. 기술 자체보다 “어떤 시나리오에 어떤 기술을 어떻게 적용해 고객이 체감할 수 있는 가치를 만들 것인가”를 핵심으로 삼게 된 과정, 그리고 language modeling과 pre-training을 연결한 XLNet의 탄생 배경이 특히 자세히 다루어져 있다.