Yi Tay introduces himself as co-founder of Reka AI, a startup building foundation models. He sets the context for his talk: sharing insights and lessons from training large language models (LLMs) as a small startup. He highlights the novelty of pre-training LLMs from scratch, especially under startup constraints like limited compute, small team size, fewer tries, and money constraints. He contrasts his experience at Google Brain with the startup environment.
Yi Tay giới thiệu bản thân là đồng sáng lập của Reka AI, một startup xây dựng các mô hình nền tảng. Ông đặt bối cảnh cho bài nói chuyện: chia sẻ những hiểu biết và bài học từ việc huấn luyện các mô hình ngôn ngữ lớn (LLM) với tư cách là một startup nhỏ. Ông nhấn mạnh tính mới lạ của việc tiền huấn luyện LLM từ đầu, đặc biệt trong các ràng buộc của startup như tài nguyên tính toán hạn chế, đội ngũ nhỏ, ít cơ hội thử nghiệm và hạn chế về tài chính. Ông so sánh trải nghiệm của mình tại Google Brain với môi trường startup.
Music hello and welcome back we are sorry for the short tea break uh it's supposed to be 15 minutes but uh due to uh our running late so we hope uh that uh you understand and thank you so much our next speaker here is y co-founder and chief scientist at Rea AI a globally distributed setup startup building multimodel large language models and advancing AI research to dat
[Nhạc] Xin chào và chào mừng trở lại. Chúng tôi xin lỗi vì thời gian nghỉ trà ngắn, lẽ ra là 15 phút nhưng vì chúng tôi chạy chậm nên hy vọng các bạn hiểu. Cảm ơn rất nhiều. Diễn giả tiếp theo của chúng tôi là đồng sáng lập và giám đốc khoa học tại Rea AI, một startup phân tán toàn cầu đang xây dựng các mô hình ngôn ngữ lớn đa phương thức và thúc đẩy nghiên cứu AI cho đến nay.
Rea has raised more than 100 million us with a recent 1 billion valuation previously he was a senior research scientist at Google Brain before Google he obtained his PhD from NTU Singapore please welcome him onto the stage Applause Music hello tting one two three uh how do I yeah okay I'll just yeah hello everybody uh good morning everybody so today I'm very happy to uh
Rea đã huy động hơn 100 triệu đô la Mỹ với định giá gần đây là 1 tỷ đô la. Trước đây, anh ấy là nhà nghiên cứu khoa học cấp cao tại Google Brain. Trước Google, anh ấy lấy bằng Tiến sĩ từ NTU Singapore. Hãy chào đón anh ấy lên sân khấu. [Vỗ tay] [Nhạc] Xin chào, kiểm tra một hai ba... Ừm, tôi sẽ... Xin chào mọi người, chào buổi sáng. Hôm nay tôi rất vui được...
give a talk today uh so my name is e and uh I am one of the co-founders of uh reca AI uh one of the startups that are building Foundation models uh so today my talk will be uh I will try to make it interesting uh like basically I will share like the insights and lessons I've learned uh building I uh as a small startup so okay so to to start off like the context right so like uh training and
Thuyết trình hôm nay. Tên tôi là E và tôi là một trong những đồng sáng lập của Reca AI, một trong những startup đang xây dựng các mô hình nền tảng. Bài nói hôm nay của tôi sẽ cố gắng làm cho nó thú vị, về cơ bản tôi sẽ chia sẻ những hiểu biết và bài học tôi đã học được khi xây dựng AI với tư cách là một startup nhỏ. Để bắt đầu, bối cảnh là việc huấn luyện và tiền huấn luyện LM từ đầu là một trải nghiệm tương đối mới lạ.
and pre-training like LM from scratch is like a relatively noal experience I think if you want to do this these days you kind of have to be at one of the frontier Labs or like Google or entropic or something like that uh so it's kind of like a experience that like uh pretty novel uh a small amount of the human population will actually get to pre-train and post-train large language
Và tiền huấn luyện LM từ đầu là một trải nghiệm tương đối mới lạ. Tôi nghĩ nếu bạn muốn làm điều này ngày nay, bạn phải ở một trong các phòng thí nghiệm tiên tiến như Google hay Anthropic hay gì đó. Vì vậy, đó là một trải nghiệm khá mới lạ mà một phần nhỏ dân số thực sự có được để tiền huấn luyện và hậu huấn luyện các mô hình ngôn ngữ lớn.
models uh and then like the next point which are which will add constraints to the first uh statement is basically like if you train LMS when considering startup constraints is even uh much more harder and over experience uh for example like instead of like uh you know um you have to kind of like procure your own compute you have to you know negotiate the prices for compute find
Và điểm tiếp theo, điều này sẽ thêm ràng buộc cho phát biểu đầu tiên, về cơ bản là nếu bạn huấn luyện LM với các ràng buộc của startup thì thậm chí còn khó khăn hơn nhiều và là một trải nghiệm quá mức. Ví dụ, thay vì... bạn phải tự mua sắm compute, thương lượng giá compute, tìm nguồn, bạn cũng có những ràng buộc về đội ngũ nhỏ.
sources you also have small T constraints uh for example like instead of I mean you you might not have like teams of 100 people but you might just have like 10 people or 20 people uh and this was an interesting constraint that we Face uh natur this kind of natural to uh how a startup functions and then uh Point C is like you kind of have significantly fewer tries then uh
Ví dụ, thay vì có đội 100 người, bạn chỉ có 10 hoặc 20 người. Và đây là một ràng buộc thú vị mà chúng tôi phải đối mặt, điều này khá tự nhiên đối với cách một startup hoạt động. Và điểm C là bạn có ít cơ hội thử nghiệm hơn đáng kể so với khi bạn có nhiều compute hơn.
as compared to when you have like a large more compute uh so I guess you could have like less much lesser ablations and stuff like this right and then lastly is like also money constraints uh so you could uh I mean uh uh even if you could raise a lot money you kind of have to make business sense because like if you could raise like 1 billion or2 billion like you kind of
So với khi bạn có nhiều compute hơn, tôi đoán bạn có thể có ít ablation hơn nhiều và những thứ tương tự. Và cuối cùng là ràng buộc về tiền bạc. Bạn có thể huy động nhiều tiền nhưng bạn phải kinh doanh có lãi vì nếu bạn huy động được 1 tỷ hay 2 tỷ, bạn phải kiếm lại số tiền đó để biện minh cho khoản đầu tư.
have to make the money back uh to justify the investment so I think like uh the research question here is like can we build really good LMS with uh in the most capital and resource efficient and we try to push the limit of how we are able to do that with relatively uh uh relatively being GPU poor uh so I think not many people get this experience to build LMS in the W like
Vì vậy, câu hỏi nghiên cứu ở đây là liệu chúng ta có thể xây dựng các LM thực sự tốt với hiệu quả vốn và tài nguyên tối đa không? Và chúng tôi cố gắng đẩy giới hạn về cách chúng tôi có thể làm điều đó với tương đối ít GPU. Tôi nghĩ không nhiều người có được trải nghiệm xây dựng LM theo cách này và tôi nghĩ trải nghiệm của tôi có thể thú vị với mọi người.
this and uh and I just thought like my experience could be interesting to people uh so also for context that I used to train LM at Google research and Google brand so like uh it's inevitable that I will make some comparisons and it's kind of like uh a natural like there's some contrast with how the experience has been uh so more about who we are I think we are kind of like one of the uh more
Cũng để bối cảnh, tôi từng huấn luyện LM tại Google Research và Google Brain, vì vậy tôi chắc chắn sẽ so sánh và có sự tương phản tự nhiên với trải nghiệm hiện tại. Về chúng tôi, tôi nghĩ chúng tôi là một trong những công ty mô hình nền tảng khiêm tốn hơn.
lowkey uh Foundation model companies so we are like about 1.5 to two years old uh model startup we are globally distributed uh and uh co-founders and most of our employees are all around the world like UK California Singapore and stuff like that so I think uh we raised about $120 million USD in total uh a 60 million series a and a 60 million recently um it's uh it may seem
Chúng tôi khoảng 1,5 đến 2 năm tuổi, là startup mô hình, phân bố toàn cầu, các đồng sáng lập và hầu hết nhân viên ở khắp nơi trên thế giới như Anh, California, Singapore, v.v. Chúng tôi đã huy động tổng cộng khoảng 120 triệu đô la Mỹ, 60 triệu series A và 60 triệu gần đây. Có vẻ nhiều tiền nhưng thực ra rất ít so với những người khác đã huy động hàng tỷ.
like a lot of money but I mean it's actually very little compared to like what uh other people have raised uh like in billions and stuff uh so three to four months ago we released two models uh record call and record flash that uh we kind of when we debuted we were rank seven on lmes uh and uh so I guess at that point if you remove duplicate models from the ch Arena uh so for
Ba đến bốn tháng trước, chúng tôi đã phát hành hai mô hình, record call và record flash, khi ra mắt chúng tôi đứng thứ 7 trên lmes. Và tôi nghĩ tại thời điểm đó, nếu loại bỏ các mô hình trùng lặp khỏi ch Arena...
context chat arena is like basically like the one of the primary evals that uh like the community cast about to when evaluating like language models uh so uh I think that kind of like I mean as Jeff has spoke earlier like LMC is basically like you get two comparisons and then voters will choose between and then you could compute ELO score and then there's a leaderboard and something like
Để bối cảnh, chat arena là một trong những đánh giá chính mà cộng đồng quan tâm khi đánh giá các mô hình ngôn ngữ. Như Jeff đã nói trước đó, LMC về cơ bản là bạn có hai so sánh và cử tri sẽ chọn giữa chúng, sau đó bạn có thể tính điểm ELO và có bảng xếp hạng.
that for those who are not that familiar with uh the chatboard arena uh um so at this was kind of like I think three or four months ago um and uh when we debuted we were like seven and so that if you remove duplicate models that was like a fifth uh or uh so I think we were pretty proud of this achievement and uh I'm just going to share like how like what was the process like and how uh
Đối với những ai chưa quen với chatboard arena, điều này cách đây ba hoặc bốn tháng. Khi ra mắt, chúng tôi đứng thứ 7, và nếu loại bỏ các mô hình trùng lặp thì là thứ 5. Chúng tôi khá tự hào về thành tích này và tôi sẽ chia sẻ quá trình, thách thức và khó khăn chúng tôi gặp phải.
the challenges and struggles we Face uh so it also kind of like one notable thing is that like uh we uh it took four months from obtaining our cluster which I'll talk about the interesting details about like uh procuring h100s in last year which was a very interesting time um and then uh we uh so basically to we we raised 100 plus million dollar but we only spent like 20 million so far on
Một điều đáng chú ý là chúng tôi mất bốn tháng để có được cụm máy tính, tôi sẽ nói về chi tiết thú vị về việc mua h100 vào năm ngoái, một thời điểm rất thú vị. Chúng tôi đã huy động hơn 100 triệu đô la nhưng chỉ chi khoảng 20 triệu cho các mô hình, với ít hơn 15 người.
the models and then we had less than 15 people and then uh it will also be able to we were able to outperform the early versions of gp4 I think 0613 on Cheo Arena when we deut so I guess that was a score like uh kind of like uh a couple of months ago but this is more recent scores uh and we we are kind of like uh you know I think it's not so easy to keep to keep up so
Chúng tôi đã có thể vượt trội so với các phiên bản đầu của gp4, tôi nghĩ là 0613 trên Cheo Arena khi ra mắt. Đó là điểm số cách đây vài tháng, nhưng đây là điểm số gần đây hơn. Thật không dễ để theo kịp.
we just basically do use this comparison to uh share some uh comparisons for example like uh flash we and call 5T tokens we spend 3 million to train it the IES elow is kind of comparable to Lama 370b and you can see the cost like like uh difference right so I think this is basically like a technical thing that we are like quite uh proud of and uh yeah so I think it's been quite a quite
Chúng tôi sử dụng so sánh này để chia sẻ một số so sánh, ví dụ như flash we và call 5T tokens, chúng tôi chi 3 triệu để huấn luyện nó. Điểm IES elow tương đương với Lama 370b và bạn có thể thấy sự khác biệt về chi phí. Đây là một điều kỹ thuật mà chúng tôi khá tự hào.
quite a fun right uh and then so we have two models that are on LM is like 21b model and a 67b model uh and then recently we also did a spar Up Cycle model which is like a uh converting the dance model toe and uh we we spend another like $5 million on that so I think this kind of interesting technical choice that we kind of uh I'll talk about more subsequently so trdr and outline of the
Đã khá vui. Chúng tôi có hai mô hình trên LM là 21b và 67b, và gần đây chúng tôi cũng thực hiện mô hình spar Up Cycle, chuyển đổi mô hình dance, và chi thêm 5 triệu đô la cho việc đó. Đây là một lựa chọn kỹ thuật thú vị mà tôi sẽ nói thêm sau.