N
NatiStream
Transcript
Hiển thị

Limitations of Bag-of-Words and Introduction to RNNs

This chapter introduces the limitations of traditional bag-of-words text analysis, which loses word order and structure. It explains how recurrent neural networks (RNNs) preserve sequence information by reading inputs step by step. Techniques like embeddings and gating units are introduced to improve learning. The chapter also covers the instability of simple RNNs and how gating units like LSTM enhance stability and power.

Chương này giới thiệu những hạn chế của phương pháp phân tích văn bản bag-of-words truyền thống, vốn làm mất thứ tự và cấu trúc từ. Nó giải thích cách mạng nơ-ron hồi tiếp (RNN) bảo toàn thông tin trình tự bằng cách đọc từng bước đầu vào. Các kỹ thuật như embedding và cổng (gating) được giới thiệu để cải thiện việc học. Chương cũng đề cập đến sự bất ổn của RNN đơn giản và cách các cổng như LSTM tăng cường độ ổn định và sức mạnh.

hi everyone this is Alec and I'm going to be talking to you today about using recurrent Neal networks for text analysis um and to get an understanding of why this is a potential um tool to use uh it's good to look at a bit of a history of how text analysis has been done um particularly from a machine learning perspective so in machine learning typically we're used to Vector

Xin chào mọi người, tôi là Alec và hôm nay tôi sẽ nói với các bạn về việc sử dụng mạng nơ-ron hồi tiếp cho phân tích văn bản. Để hiểu tại sao đây là một công cụ tiềm năng, chúng ta hãy xem xét một chút lịch sử về cách phân tích văn bản đã được thực hiện, đặc biệt là từ góc độ học máy. Trong học máy, chúng ta thường quen với biểu diễn vector.

representation so we know how to deal with numbers for categories we would use a one hot vectorization model but when we move to trying to understand and class ify and regress on sequences for instance it becomes much less clear because our tools are typically based on Vector approaches and so the way this is typically dealt with is by Computing some uh hard-coded feature

Chúng ta biết cách xử lý số cho các danh mục bằng cách sử dụng mô hình vector one-hot. Nhưng khi chuyển sang hiểu, phân loại và hồi quy trên các chuỗi, mọi thứ trở nên kém rõ ràng hơn vì các công cụ của chúng ta thường dựa trên phương pháp vector. Cách thường được xử lý là tính toán một số biến đổi đặc trưng được mã hóa cứng.

Transformations for instance using tfid effector risers um some sort of compression model like LSA um and then plugging a linear model such as support Vector machine or a softmax classifier on top of that and the premise of this talk today is what happens if we cut out those techniques and instead replace them with an RNN uh and to get an understanding of why this might be an

Ví dụ như sử dụng bộ biến đổi tf-idf, một số mô hình nén như LSA, và sau đó ghép một mô hình tuyến tính như máy vector hỗ trợ hoặc bộ phân loại softmax lên trên. Tiền đề của bài nói hôm nay là: điều gì sẽ xảy ra nếu chúng ta loại bỏ các kỹ thuật đó và thay thế chúng bằng một RNN? Để hiểu tại sao điều này có thể là một lợi thế, cấu trúc là khó.

advantage is structure is hard so engrams are the typical way of preserving some structure so this would be we take our sentence for instance the cat sat on the mat and we re represent it as the occurrence of any individual word or combinations of words and these combinations of words begin to get us the way to see a little bit of structure by structure I mean preserving

N-gram là cách điển hình để bảo toàn một số cấu trúc. Ví dụ, chúng ta lấy câu 'the cat sat on the mat' và biểu diễn lại nó dưới dạng sự xuất hiện của từng từ hoặc tổ hợp từ. Các tổ hợp từ này bắt đầu cho chúng ta thấy một chút cấu trúc. Ý tôi về cấu trúc là bảo toàn thứ tự của các từ.

The Ordering of words um and the problem with this is because once we have byrams are try diagrams you know any combination of two or three words it quickly becomes huge possibilities you can have easily 10 million plus features and this become begins to get cumbersome require lots of memory and slows things down in and of itself um structure although it's difficult is also very

Vấn đề là khi chúng ta có bigram, trigram, bất kỳ tổ hợp nào của hai hoặc ba từ, nó nhanh chóng trở nên khổng lồ. Bạn có thể dễ dàng có hơn 10 triệu đặc trưng, điều này trở nên cồng kềnh, đòi hỏi nhiều bộ nhớ và làm chậm mọi thứ. Mặc dù cấu trúc khó, nó cũng rất quan trọng cho một số tác vụ như hài hước hoặc mỉa mai.

important for certain tasks such as humor or sarcasm looking at a collection of the word you know cat appeared or the word dog appeared isn't going to cut it um and that's what a lot of our models today do so to understand though why many models today are based on this and do quite successful engrams can get you a long way for many tasks so specific words are often very strong indicators

Nhìn vào một tập hợp các từ như 'cat' xuất hiện hay 'dog' xuất hiện là không đủ, và đó là điều mà nhiều mô hình hiện nay làm. Tuy nhiên, để hiểu tại sao nhiều mô hình hiện nay dựa trên điều này và khá thành công, n-gram có thể giúp ích rất nhiều cho nhiều tác vụ. Các từ cụ thể thường là chỉ báo rất mạnh.

useless in the case of negative sentiment and fantastic in the case of positive sentiment and if you're for instance trying to classify whether a document is about the stock market or about a as a recipe you know you don't see the word green tea come up very much in the stock market conversation you don't see the word NASDAQ come up very much in a recipe so you can very much

Vô dụng trong trường hợp cảm xúc tiêu cực và tuyệt vời trong trường hợp cảm xúc tích cực. Ví dụ, nếu bạn đang cố gắng phân loại xem một tài liệu nói về thị trường chứng khoán hay công thức nấu ăn, bạn sẽ không thấy từ 'green tea' xuất hiện nhiều trong cuộc trò chuyện về thị trường chứng khoán, và bạn sẽ không thấy từ 'NASDAQ' xuất hiện nhiều trong công thức nấu ăn. Vì vậy, bạn có thể nhanh chóng phân tách mọi thứ bằng mô hình tuyến tính trong những loại tác vụ đó.

quickly separate things VI a linear model in those kinds of um task and so it's often a question of knowing what's right for your task at hand so if you're trying to get a more qualitative understanding of like what's going on in a body of Texs this is where structure may be very important whereas if you're just trying to separate out um you know something that may be very indicative on

Vì vậy, thường là vấn đề biết cái gì phù hợp với tác vụ của bạn. Nếu bạn đang cố gắng có được sự hiểu biết định tính hơn về những gì đang diễn ra trong một khối văn bản, thì cấu trúc có thể rất quan trọng. Trong khi nếu bạn chỉ cố gắng phân tách một cái gì đó rất chỉ báo ở cấp độ từ, thì mô hình túi từ có thể khá mạnh.

like a word level um then in many ways uh a bag of wordss model can be quite strong okay so how an RNN work to understand its potential advantages over a back Awards model what an RNN does is it reads through a sequence tively which is really nice because it's how um it's how people do it as well so it's able to preserve some of the structure of the of the model so what it does is it goes

Được rồi, làm thế nào RNN hoạt động? Để hiểu lợi thế tiềm năng của nó so với mô hình túi từ, RNN đọc qua một chuỗi một cách tuần tự, điều này thực sự tốt vì đó cũng là cách con người làm. Nó có thể bảo toàn một số cấu trúc của mô hình. Những gì nó làm là đi qua từng từ và cập nhật biểu diễn ẩn của nó dựa trên từ đó và đầu vào từ trạng thái ẩn trước đó.

through each word and updates its hidden representation based on that word and the input from the previous hidden State at Time Zero where we have no previous hidden State we feed in a fixed um you know either a bunch of zeros or we treat it as another parameter to be learned in initial hidden representation and it just continues to do this all the way through the sequence

Tại thời điểm 0, khi không có trạng thái ẩn trước đó, chúng ta đưa vào một giá trị cố định, ví dụ như một loạt số 0 hoặc coi nó như một tham số khác cần học trong biểu diễn ẩn ban đầu. Và nó tiếp tục làm điều này trong suốt chuỗi.

so at each time step we have a 512 in this case if we had 512 hidden units dimensional Vector representation of our sequence so it's a way of kind of taking this sequence of words for instance and at each time step converting it into a fix length representation uh as a bit of notation arrows would be projections dot products um and boxes would represent activities

Vì vậy, tại mỗi bước thời gian, chúng ta có một biểu diễn vector 512 chiều (nếu chúng ta có 512 đơn vị ẩn) của chuỗi. Đây là một cách để lấy chuỗi từ này và tại mỗi bước thời gian chuyển đổi nó thành một biểu diễn có độ dài cố định. Về ký hiệu, mũi tên sẽ là các phép chiếu, tích vô hướng, và hộp sẽ là các hoạt động, vector giá trị.

vectors of values so for instance the activation of each hidden unit um would be this box and it just proceeds it early through it's important to note these um projections are deterministic largely they are shared across all time sequences so this projection with this arrow is shared for all inputs across all time SE steps and this um hidden to Hidden unit connections um are preserved as well

Ví dụ, kích hoạt của mỗi đơn vị ẩn sẽ là hộp này và nó chỉ tiến triển. Điều quan trọng cần lưu ý là các phép chiếu này phần lớn là xác định và được chia sẻ trên tất cả các chuỗi thời gian. Phép chiếu với mũi tên này được chia sẻ cho tất cả các đầu vào qua tất cả các bước thời gian, và các kết nối từ ẩn đến ẩn cũng được bảo toàn qua chuỗi.

across the sequence so this is what makes learning tractable in these models so at the end of this iterating through this sequence we've got now a learned representation of the sequence a vector form of our sequence which can then be used by slapping on a traditional output classifier so in this toy example what we've done is we've read in an input sentence and for

Đây là điều làm cho việc học trở nên khả thi trong các mô hình này. Vào cuối quá trình lặp qua chuỗi, chúng ta có một biểu diễn đã học của chuỗi dưới dạng vector, sau đó có thể được sử dụng bằng cách thêm một bộ phân loại đầu ra truyền thống. Trong ví dụ đồ chơi này, chúng ta đã đọc một câu đầu vào và đang cố gắng dạy mô hình phân loại chủ đề của câu.

instance we're trying to teach a model to classify the subject of the sentence you can also stack them so just as well as your RNN can go through and an input sequence and return its you know internal representation of that sequence you can then train another RNN on top of it or you can jointly train both um the structure is actually quite flexible one final note to note is the

Bạn cũng có thể xếp chồng chúng. Cũng như RNN có thể đi qua một chuỗi đầu vào và trả về biểu diễn nội bộ của chuỗi đó, bạn có thể huấn luyện một RNN khác lên trên nó hoặc huấn luyện cả hai cùng nhau. Cấu trúc thực sự khá linh hoạt. Một lưu ý cuối cùng là cách chúng ta thực hiện việc truyền thẳng từ đầu vào đến ẩn ban đầu.

way that we do um this original input to Hidden feed forward is this is typically represented either as a traditional one hot which doesn't really get us too much Advantage but what's really exciting is we can represent this as what's called an embedding Matrix so these words the cat sat on the map would be represented as indexes into a matrix and what they would do is the would be represented as

Điều này thường được biểu diễn dưới dạng one-hot truyền thống, nhưng điều thực sự thú vị là chúng ta có thể biểu diễn nó dưới dạng ma trận nhúng. Các từ 'the cat sat on the mat' sẽ được biểu diễn dưới dạng chỉ số vào một ma trận. Khi đọc qua chuỗi, chúng ta tra cứu hàng của ma trận, ví dụ hàng 100, và trả về đầu vào được đưa vào RNN là biểu diễn đã học trong ma trận nhúng đó.

like index 100 so when we read through this sequence we would look up the row of the on The Matrix you know row 100 and we would return as the input being fed into the RNN the Learned representation of in that embedding Matrix so let's say we had 128 Dimensions to be learned as an input representation for our words it would be equalent to a 1828 by like let's say

Giả sử chúng ta có 128 chiều để học làm biểu diễn đầu vào cho các từ, nó sẽ tương đương với ma trận 128 x 10.000 nếu chúng ta biết 10.000 từ. Sau đó, chúng ta đưa cái này vào làm đầu vào. Điều này thực sự tuyệt vời vì chúng ta có thể coi nó như một cách học để học biểu diễn của các từ. Chúng ta sẽ xem xét sau trong bài thuyết trình những biểu diễn đó thực sự trông như thế nào và chúng mang lại cho mô hình rất nhiều sức mạnh.

10,000 Matrix if we knew 10,000 words and we would then feed this in as input and that's really cool because we can treat it as a learned way to learn representations of our words and we'll look later in this presentation at what those actually look like and they give the model a lot of power all right so the big thing in the literature is rnn's have a reputation

Được rồi, vấn đề lớn trong tài liệu là RNN có tiếng là rất khó học. Chúng thường được biết đến là không ổn định và các RNN đơn giản được huấn luyện với gradient descent ngẫu nhiên thông thường thực sự rất không ổn định và khó học. Nhưng những gì đã xảy ra trong tài liệu nghiên cứu trong vài năm qua là có một loạt các thủ thuật khác nhau đã được phát triển giúp chúng ổn định hơn, mạnh mẽ hơn và đáng tin cậy hơn.

for being very difficult to learn they are often known to be unstable and simple rnn's trained with generic stochastic gradient descent are actually very unstable and difficult to learn but what has happened in the research literature over the last few years is there's a bunch of various tricks that have been developed that help them be much more stable much more powerful and

Để hiểu những điều này, chúng ta sẽ nhanh chóng đi qua tất cả các thủ thuật khác nhau này. Đầu tiên là các đơn vị gating. Để hiểu đơn vị gating là gì, trước tiên chúng ta cần xem xét chi tiết hơn một chút về cách một RNN đơn giản hoạt động. Những gì xảy ra là chúng ta có trạng thái ẩn từ bước thời gian trước đó (tại bước thời gian ban đầu, nó có thể là số 0 hoặc tham số học) và chúng ta nhận đầu vào tại bước thời gian t.

much more reliable effectively um and to get an understanding of these we're going to go quickly through all of these various tricks um the first of these is gating units so to understand what a gating unit is we first need to look a little bit more into detail of how a simple RNN works and so what happens is we have our hidden state from our previous time step again at the original

Chúng ta lấy đầu vào từ trạng thái ẩn H của t-1 và đầu vào của t, cộng chúng lại với nhau thông qua phép chiếu tích vô hướng, sau đó áp dụng một hàm kích hoạt theo từng phần tử như tanh, và chúng ta có một trạng thái ẩn mới. Tại bước thời gian tiếp theo, chúng ta nhận thêm đầu vào, cộng lại, áp dụng hàm kích hoạt, và quá trình này tiếp tục. Vấn đề là thông tin luôn được cập nhật mỗi bước thời gian, do đó thông tin khó tồn tại qua mô hình này.

time setep this can just be zeros or learn parameters and we receive input at time step T so we take input uh from the hidden state of H of T minus t minus one and we take input of T we um just add them for together for instance um via a DOT product projection and then we apply an element wise activation function like T for instance and then we have a new hidden State at the next time step we

Bạn có thể nghĩ về điều này như một dạng suy giảm hàm mũ. Nếu chúng ta có một giá trị, giả sử là 1, và qua quá trình này chúng ta nhân giá trị đó với 0.5, sau một vài bước thời gian, giá trị đó sẽ suy giảm theo hàm mũ về 0. Vì vậy, thông tin khó lan truyền qua một cấu trúc như thế này. Đã có nhiều thay đổi được gọi là đơn vị gating để làm cho điều này hoạt động tốt hơn.

1 / 6