0:00
hi everyone so in this video I'd like us
0:02
to cover the process of tokenization in
0:04
large language models now you see here
0:06
that I have a set face and that's
0:08
because uh tokenization is my least
0:10
favorite part of working with large
0:12
language models but unfortunately it is
0:13
necessary to understand in some detail
0:16
because it it is fairly hairy gnarly and
0:18
there's a lot of hidden foot guns to be
0:19
aware of and a lot of oddness with large
0:22
language models typically traces back to
0:25
tokenization so what is
0:27
tokenization now in my previous video
0:29
Let's Build GPT from scratch uh we
0:32
actually already did tokenization but we
0:33
did a very naive simple version of
0:36
tokenization so when you go to the
0:37
Google colab for that video uh you see
0:41
here that we loaded our training set and
0:43
our training set was this uh Shakespeare
0:46
uh data set now in the beginning the
0:48
Shakespeare data set is just a large
0:50
string in Python it's just text and so
0:52
the question is how do we plug text into
0:55
large language models and in this case
0:58
here we created a vocabulary of 65
1:01
possible characters that we saw occur in
1:04
this string these were the possible
1:06
characters and we saw that there are 65
1:08
of them and then we created a a lookup
1:11
table for converting from every possible
1:13
character a little string piece into a
1:16
token an
1:18
integer so here for example we tokenized
1:21
the string High there and we received
1:23
this sequence of
1:25
tokens and here we took the first 1,000
1:28
characters of our data set and we
1:30
encoded it into tokens and because it is
1:33
this is character level we received
1:35
1,000 tokens in a sequence so token 18
1:39
47
1:40
Etc now later we saw that the way we
1:43
plug these tokens into the language
1:46
model is by using an embedding
1:48
table and so basically if we have 65
1:51
possible tokens then this embedding
1:53
table is going to have 65 rows and
1:56
roughly speaking we're taking the
1:58
integer associated with every single
2:00
sing Le token we're using that as a
2:02
lookup into this table and we're
2:04
plucking out the corresponding row and
2:06
this row is a uh is trainable parameters
2:09
that we're going to train using back
2:10
propagation and this is the vector that
2:13
then feeds into the Transformer um and
2:15
that's how the Transformer Ser of
2:17
perceives every single
2:18
token so here we had a very naive
2:21
tokenization process that was a
2:23
character level tokenizer but in
2:25
practice in state-ofthe-art uh language
2:27
models people use a lot more complicated
2:29
schemes unfortunately
2:30
uh for constructing these uh token
2:34
vocabularies so we're not dealing on the
2:37
Character level we're dealing on chunk
2:39
level and the way these um character
2:42
chunks are constructed is using
2:44
algorithms such as for example the bik
2:45
pair in coding algorithm which we're
2:47
going to go into in detail um and cover
2:51
in this video I'd like to briefly show
2:53
you the paper that introduced a bite
2:55
level encoding as a mechanism for
2:57
tokenization in the context of large
2:58
language models and I would say that
3:01
that's probably the gpt2 paper and if
3:03
you scroll down here to the section
3:06
input representation this is where they
3:08
cover tokenization the kinds of
3:09
properties that you'd like the
3:11
tokenization to have and they conclude
3:13
here that they're going to have a
3:15
tokenizer where you have a vocabulary of
3:18
50,2 57 possible
3:21
tokens and the context size is going to
3:24
be 1,24 tokens so in the in in the
3:27
attention layer of the Transformer
3:29
neural network
3:30
every single token is attending to the
3:32
previous tokens in the sequence and it's
3:34
going to see up to 1,24 tokens so tokens
3:38
are this like fundamental unit um the
3:41
atom of uh large language models if you
3:43
will and everything is in units of
3:45
tokens everything is about tokens and
3:47
tokenization is the process for
3:48
translating strings or text into
3:51
sequences of tokens and uh vice versa
3:55
when you go into the Llama 2 paper as
3:57
well I can show you that when you search
3:58
token you're going to get get 63 hits um
4:02
and that's because tokens are again
4:03
pervasive so here they mentioned that
4:05
they trained on two trillion tokens of
4:07
data and so
4:08
on so we're going to build our own
4:11
tokenizer luckily the bite be encoding
4:13
algorithm is not uh that super
4:15
complicated and we can build it from
4:17
scratch ourselves and we'll see exactly
4:19
how this works before we dive into code
4:21
I'd like to give you a brief Taste of
4:23
some of the complexities that come from
4:24
the tokenization because I just want to
4:26
make sure that we motivate it
4:27
sufficiently for why we are doing all
4:29
this and why this is so gross so
4:33
tokenization is at the heart of a lot of
4:34
weirdness in large language models and I
4:36
would advise that you do not brush it
4:38
off a lot of the issues that may look
4:41
like just issues with the new network
4:42
architecture or the large language model
4:45
itself are actually issues with the
4:47
tokenization and fundamentally Trace uh
4:49
back to it so if you've noticed any
4:52
issues with large language models can't
4:54
you know not able to do spelling tasks
4:56
very easily that's usually due to
4:58
tokenization simple string processing
5:00
can be difficult for the large language
5:02
model to perform
5:04
natively uh non-english languages can
5:06
work much worse and to a large extent
5:08
this is due to
5:09
tokenization sometimes llms are bad at
5:12
simple arithmetic also can trace be
5:14
traced to
5:15
tokenization uh gbt2 specifically would
5:18
have had quite a bit more issues with
5:20
python than uh future versions of it due
5:22
to tokenization there's a lot of other
5:24
issues maybe you've seen weird warnings
5:26
about a trailing whites space this is a
5:27
tokenization issue um
5:31
if you had asked GPT earlier about solid
5:34
gold Magikarp and what it is you would
5:35
see the llm go totally crazy and it
5:38
would start going off about a completely
5:40
unrelated tangent topic maybe you've
5:42
been told to use yl over Json in
5:44
structure data all of that has to do
5:45
with tokenization so basically
5:48
tokenization is at the heart of many
5:49
issues I will look back around to these
5:52
at the end of the video but for now let
5:54
me just um skip over it a little bit and
5:57
let's go to this web app um the Tik
6:00
tokenizer bell.app so I have it loaded
6:03
here and what I like about this web app
6:05
is that tokenization is running a sort
6:07
of live in your browser in JavaScript so
6:10
you can just type here stuff hello world
6:12
and the whole string
6:14
rokenes so here what we see on uh the
6:18
left is a string that you put in on the
6:20
right we're currently using the gpt2
6:22
tokenizer we see that this string that I
6:25
pasted here is currently tokenizing into
6:27
300 tokens and here they are sort of uh
6:31
shown explicitly in different colors for
6:33
every single token so for example uh
6:36
this word tokenization became two tokens
6:39
the token
6:41
3,642 and
6:44
1,634 the token um space is is token 318
6:50
so be careful on the bottom you can show
6:52
white space and keep in mind that there
6:55
are spaces and uh sln new line
6:57
characters in here but you can hide them
7:00
for
7:02
clarity the token space at is token 379
7:06
the to the Token space the is 262 Etc so
7:11
you notice here that the space is part
7:13
of that uh token
7:16
chunk now so this is kind of like how
7:19
our English sentence broke up and that
7:21
seems all well and good now now here I
7:24
put in some arithmetic so we see that uh
7:27
the token 127 Plus and then token six
7:32
space 6 followed by 77 so what's
7:34
happening here is that 127 is feeding in
7:37
as a single token into the large
7:38
language model but the um number 677
7:43
will actually feed in as two separate
7:45
tokens and so the large language model
7:47
has to sort of um take account of that
7:51
and process it correctly in its Network
7:54
and see here 804 will be broken up into
7:56
two tokens and it's is all completely
7:58
arbitrary and here I have another
8:00
example of four-digit numbers and they
8:02
break up in a way that they break up and
8:04
it's totally arbitrary sometimes you
8:05
have um multiple digits single token
8:08
sometimes you have individual digits as
8:10
many tokens and it's all kind of pretty
8:12
arbitrary and coming out of the
8:15
tokenizer here's another example we have
8:17
the string egg and you see here that
8:21
this became two
8:22
tokens but for some reason when I say I
8:25
have an egg you see when it's a space
8:28
egg it's two token it's sorry it's a
8:31
single token so just egg by itself in
8:33
the beginning of a sentence is two
8:35
tokens but here as a space egg is
8:38
suddenly a single token uh for the exact
8:41
same string okay here lowercase egg
8:44
turns out to be a single token and in
8:46
particular notice that the color is
8:47
different so this is a different token
8:49
so this is case sensitive and of course
8:52
a capital egg would also be different
8:55
tokens and again um this would be two
8:57
tokens arbitrarily so so for the same
9:00
concept egg depending on if it's in the
9:02
beginning of a sentence at the end of a
9:04
sentence lowercase uppercase or mixed
9:06
all this will be uh basically very
9:08
different tokens and different IDs and
9:10
the language model has to learn from raw
9:12
data from all the internet text that
9:14
it's going to be training on that these
9:15
are actually all the exact same concept
9:17
and it has to sort of group them in the
9:19
parameters of the neural network and
9:21
understand just based on the data
9:22
patterns that these are all very similar
9:25
but maybe not almost exactly similar but
9:27
but very very similar
9:30
um after the EG demonstration here I
9:33
have um an introduction from open a eyes
9:36
chbt in Korean so manaso Pang uh Etc uh
9:42
so this is in Korean and the reason I
9:44
put this here is because you'll notice
9:48
that um non-english languages work
9:51
slightly worse in Chachi part of this is
9:54
because of course the training data set
9:56
for Chachi is much larger for English
9:58
and for everything else but the same is
10:00
true not just for the large language
10:02
model itself but also for the tokenizer
10:04
so when we train the tokenizer we're
10:06
going to see that there's a training set
10:07
as well and there's a lot more English
10:09
than non-english and what ends up
10:11
happening is that we're going to have a
10:13
lot more longer tokens for
10:17
English so how do I put this if you have
10:20
a single sentence in English and you
10:21
tokenize it you might see that it's 10
10:24
tokens or something like that but if you
10:25
translate that sentence into say Korean
10:27
or Japanese or something else you'll
10:29
typically see that the number of tokens
10:31
used is much larger and that's because
10:33
the chunks here are a lot more broken up
10:37
so we're using a lot more tokens for the
10:39
exact same thing and what this does is
10:41
it bloats up the sequence length of all
10:44
the documents so you're using up more
10:46
tokens and then in the attention of the
10:48
Transformer when these tokens try to
10:50
attend each other you are running out of
10:52
context um in the maximum context length
10:55
of that Transformer and so basically all
10:58
the non-english text is stretched out
11:01
from the perspective of the Transformer
11:03
and this just has to do with the um
11:06
trainings that used for the tokenizer
11:07
and the tokenization itself so it will
11:10
create a lot bigger tokens and a lot
11:12
larger groups in English and it will
11:14
have a lot of little boundaries for all
11:16
the other non-english text um so if we
11:20
translated this into English it would be
11:22
significantly fewer
11:23
tokens the final example I have here is
11:26
a little snippet of python for doing FS
11:28
buuz and what I'd like you to notice is
11:31
look all these individual spaces are all
11:34
separate tokens they are token
11:37
220 so uh 220 220 220 220 and then space
11:43
if is a single token and so what's going
11:45
on here is that when the Transformer is
11:47
going to consume or try to uh create
11:49
this text it needs to um handle all
11:53
these spaces individually they all feed
11:54
in one by one into the entire
11:57
Transformer in the sequence and so this
11:59
is being extremely wasteful tokenizing
12:01
it in this way and so as a result of
12:04
that gpt2 is not very good with python
12:07
and it's not anything to do with coding
12:09
or the language model itself it's just
12:11
that if he use a lot of indentation
12:12
using space in Python like we usually do
12:15
uh you just end up bloating out all the
12:17
text and it's separated across way too
12:19
much of the sequence and we are running
12:21
out of the context length in the
12:23
sequence uh that's roughly speaking
12:24
what's what's happening we're being way
12:26
too wasteful we're taking up way too
12:27
much token space now we can also scroll
12:30
up here and we can change the tokenizer
12:32
so note here that gpt2 tokenizer creates
12:34
a token count of 300 for this string
12:37
here we can change it to CL 100K base
12:40
which is the GPT for tokenizer and we
12:42
see that the token count drops to 185 so
12:45
for the exact same string we are now
12:47
roughly having the number of tokens and
12:50
roughly speaking this is because uh the
12:52
number of tokens in the GPT 4 tokenizer
12:54
is roughly double that of the number of
12:57
tokens in the gpt2 tokenizer so we went
12:59
went from roughly 50k to roughly 100K
13:02
now you can imagine that this is a good
13:03
thing because the same text is now
13:06
squished into half as many tokens so uh
13:10
this is a lot denser input to the
13:13
Transformer and in the Transformer every
13:15
single token has a finite number of
13:17
tokens before it that it's going to pay
13:18
attention to and so what this is doing
13:20
is we're roughly able to see twice as
13:23
much text as a context for what token to
13:27
predict next uh because of this change
13:29
but of course just increasing the number
13:31
of tokens is uh not strictly better
13:33
infinitely uh because as you increase
13:35
the number of tokens now your embedding
13:37
table is um sort of getting a lot larger
13:40
and also at the output we are trying to
13:41
predict the next token and there's the
13:43
soft Max there and that grows as well
13:45
we're going to go into more detail later
13:46
on this but there's some kind of a Sweet
13:48
Spot somewhere where you have a just
13:51
right number of tokens in your
13:52
vocabulary where everything is
13:54
appropriately dense and still fairly
13:57
efficient now one thing I would like you
13:58
to note specifically for the gp4
14:00
tokenizer is that the handling of the
14:04
white space for python has improved a
14:05
lot you see that here these four spaces
14:08
are represented as one single token for
14:10
the three spaces here and then the token
14:14
SPF and here seven spaces were all
14:17
grouped into a single token so we're
14:19
being a lot more efficient in how we
14:20
represent Python and this was a
14:22
deliberate Choice made by open aai when
14:24
they designed the gp4 tokenizer and they
14:28
group a lot more space into a single
14:30
character what this does is this
14:32
densifies Python and therefore we can
14:35
attend to more code before it when we're
14:38
trying to predict the next token in the
14:40
sequence and so the Improvement in the
14:42
python coding ability from gbt2 to gp4
14:45
is not just a matter of the language
14:47
model and the architecture and the
14:49
details of the optimization but a lot of
14:51
the Improvement here is also coming from
14:52
the design of the tokenizer and how it
14:54
groups characters into tokens okay so
14:57
let's now start writing some code
14:59
so remember what we want to do we want
15:01
to take strings and feed them into
15:04
language models for that we need to
15:06
somehow tokenize strings into some
15:09
integers in some fixed vocabulary and
15:12
then we will use those integers to make
15:14
a look up into a lookup table of vectors
15:17
and feed those vectors into the
15:18
Transformer as an input now the reason
15:21
this gets a little bit tricky of course
15:23
is that we don't just want to support
15:24
the simple English alphabet we want to
15:26
support different kinds of languages so
15:28
this is anango in Korean which is hello
15:32
and we also want to support many kinds
15:33
of special characters that we might find
15:35
on the internet for example
15:37
Emoji so how do we feed this text into
15:41
uh
15:42
Transformers well how's the what is this
15:44
text anyway in Python so if you go to
15:47
the documentation of a string in Python
15:50
you can see that strings are immutable
15:52
sequences of Unicode code
15:54
points okay what are Unicode code points
15:58
we can go to PDF so Unicode code points
16:01
are defined by the Unicode Consortium as
16:05
part of the Unicode standard and what
16:08
this is really is that it's just a
16:09
definition of roughly 150,000 characters
16:12
right now and roughly speaking what they
16:15
look like and what integers um represent
16:18
those characters so it says 150,000
16:20
characters across 161 scripts as of
16:23
right now so if you scroll down here you
16:25
can see that the standard is very much
16:26
alive the latest standard 15.1 in
16:29
September
16:30
2023 and basically this is just a way to
16:34
define lots of types of
16:37
characters like for example all these
16:39
characters across different scripts so
16:42
the way we can access the unic code code
16:44
Point given Single Character is by using
16:46
the or function in Python so for example
16:48
I can pass in Ord of H and I can see
16:51
that for the Single Character H the unic
16:55
code code point is
16:56
104 okay um but this can be arbitr
17:00
complicated so we can take for example
17:02
our Emoji here and we can see that the
17:04
code point for this one is
17:06
128,000 or we can take
17:10
un and this is 50,000 now keep in mind
17:14
you can't plug in strings here because
17:17
you uh this doesn't have a single code
17:18
point it only takes a single uni code
17:21
code Point character and tells you its
17:24
integer so in this way we can look
17:27
up all the um characters of this
17:30
specific string and their code points so
17:32
or of X forx in this string and we get
17:37
this encoding here now see here we've
17:40
already turned the raw code points
17:42
already have integers so why can't we
17:44
simply just use these integers and not
17:47
have any tokenization at all why can't
17:49
we just use this natively as is and just
17:51
use the code Point well one reason for
17:53
that of course is that the vocabulary in
17:54
that case would be quite long so in this
17:57
case for Unicode the this is a
17:59
vocabulary of
18:00
150,000 different code points but more
18:03
worryingly than that I think the Unicode
18:05
standard is very much alive and it keeps
18:07
changing and so it's not kind of a
18:09
stable representation necessarily that
18:11
we may want to use directly so for those
18:14
reasons we need something a bit better
18:16
so to find something better we turn to
18:18
encodings so if we go to the Wikipedia
18:20
page here we see that the Unicode
18:21
consortion defines three types of
18:24
encodings utf8 UTF 16 and UTF 32 these
18:28
encoding are the way by which we can
18:31
take Unicode text and translate it into
18:33
binary data or by streams utf8 is by far
18:37
the most common uh so this is the utf8
18:40
page now this Wikipedia page is actually
18:42
quite long but what's important for our
18:44
purposes is that utf8 takes every single
18:46
Cod point and it translates it to a by
18:50
stream and this by stream is between one
18:52
to four bytes so it's a variable length
18:54
encoding so depending on the Unicode
18:56
Point according to the schema you're
18:58
going to end up with between 1 to four
19:00
bytes for each code point on top of that
19:03
there's utf8 uh
19:05
utf16 and UTF 32 UTF 32 is nice because
19:09
it is fixed length instead of variable
19:11
length but it has many other downsides
19:12
as well so the full kind of spectrum of
19:17
pros and cons of all these different
19:18
three encodings are beyond the scope of
19:20
this video I just like to point out that
19:23
I enjoyed this block post and this block
19:25
post at the end of it also has a number
19:27
of references that can be quite useful
19:29
uh one of them is uh utf8 everywhere
19:32
Manifesto um and this Manifesto
19:34
describes the reason why utf8 is
19:37
significantly preferred and a lot nicer
19:40
than the other encodings and why it is
19:42
used a lot more prominently um on the
19:45
internet one of the major advantages
19:48
just just to give you a sense is that
19:50
utf8 is the only one of these that is
19:52
backwards compatible to the much simpler
19:54
asky encoding of text um but I'm not
19:57
going to go into the full detail in this
19:58
video so suffice to say that we like the
20:01
utf8 encoding and uh let's try to take
20:04
the string and see what we get if we
20:06
encoded into
20:08
utf8 the string class in Python actually
20:11
has do encode and you can give it the
20:12
encoding which is say utf8 now we get
20:16
out of this is not very nice because
20:18
this is the bytes is a bytes object and
20:21
it's not very nice in the way that it's
20:23
printed so I personally like to take it
20:25
through list because then we actually
20:27
get the raw B
20:29
of this uh encoding so this is the raw
20:32
byes that represent this string
20:36
according to the utf8 en coding we can
20:38
also look at utf16 we get a slightly
20:41
different by stream and we here we start
20:43
to see one of the disadvantages of utf16
20:45
you see how we have zero Z something Z
20:48
something Z something we're starting to
20:50
get a sense that this is a bit of a
20:51
wasteful encoding and indeed for simple
20:54
asky characters or English characters
20:56
here uh we just have the structure of 0
20:59
something Z something and it's not
21:01
exactly nice same for UTF 32 when we
21:04
expand this we can start to get a sense
21:06
of the wastefulness of this encoding for
21:08
our purposes you see a lot of zeros
21:10
followed by
21:11
something and so uh this is not
21:15
desirable so suffice it to say that we
21:18
would like to stick with utf8 for our
21:21
purposes however if we just use utf8
21:24
naively these are by streams so that
21:26
would imply a vocabulary length of only
21:29
256 possible tokens uh but this this
21:33
vocabulary size is very very small what
21:35
this is going to do if we just were to
21:37
use it naively is that all of our text
21:40
would be stretched out over very very
21:42
long sequences of bytes and so
21:46
um what what this does is that certainly
21:49
the embeding table is going to be tiny
21:51
and the prediction at the top at the
21:52
final layer is going to be very tiny but
21:54
our sequences are very long and remember
21:56
that we have pretty finite um context
21:59
length and the attention that we can
22:01
support in a transformer for
22:03
computational reasons and so we only
22:06
have as much context length but now we
22:07
have very very long sequences and this
22:09
is just inefficient and it's not going
22:11
to allow us to attend to sufficiently
22:13
long text uh before us for the purposes
22:16
of the next token prediction task so we
22:18
don't want to use the raw bytes of the
22:22
utf8 encoding we want to be able to
22:24
support larger vocabulary size that we
22:27
can tune as a hyper
22:29
but we want to stick with the utf8
22:31
encoding of these strings so what do we
22:34
do well the answer of course is we turn
22:35
to the bite pair encoding algorithm
22:37
which will allow us to compress these
22:39
bite sequences um to a variable amount
22:43
so we'll get to that in a bit but I just
22:45
want to briefly speak to the fact that I
22:47
would love nothing more than to be able
22:49
to feed raw bite sequences into uh
22:53
language models in fact there's a paper
22:55
about how this could potentially be done
22:57
uh from Summer last last year now the
22:59
problem is you actually have to go in
23:01
and you have to modify the Transformer
23:02
architecture because as I mentioned
23:04
you're going to have a problem where the
23:07
attention will start to become extremely
23:08
expensive because the sequences are so
23:10
long and so in this paper they propose
23:13
kind of a hierarchical structuring of
23:16
the Transformer that could allow you to
23:18
just feed in raw bites and so at the end
23:20
they say together these results
23:22
establish the viability of tokenization
23:24
free autor regressive sequence modeling
23:25
at scale so tokenization free would
23:27
indeed be amazing we would just feed B
23:30
streams directly into our models but
23:32
unfortunately I don't know that this has
23:34
really been proven out yet by
23:36
sufficiently many groups and a
23:37
sufficient scale uh but something like
23:39
this at one point would be amazing and I
23:41
hope someone comes up with it but for
23:42
now we have to come back and we can't
23:44
feed this directly into language models
23:46
and we have to compress it using the B
23:48
paare encoding algorithm so let's see
23:50
how that works so as I mentioned the B
23:52
paare encoding algorithm is not all that
23:54
complicated and the Wikipedia page is
23:56
actually quite instructive as far as the
23:57
basic idea goes go what we're doing is
23:59
we have some kind of a input sequence uh
24:02
like for example here we have only four
24:04
elements in our vocabulary a b c and d
24:06
and we have a sequence of them so
24:08
instead of bytes let's say we just have
24:10
four a vocab size of
24:12
four the sequence is too long and we'd
24:14
like to compress it so what we do is
24:16
that we iteratively find the pair of uh
24:20
tokens that occur the most
24:23
frequently and then once we've
24:25
identified that pair we repl replace
24:28
that pair with just a single new token
24:31
that we append to our vocabulary so for
24:34
example here the bite pair AA occurs
24:36
most often so we mint a new token let's
24:39
call it capital Z and we replace every
24:42
single occurrence of AA by Z so now we
24:46
have two Z's here so here we took a
24:49
sequence of 11 characters with
24:52
vocabulary size four and we've converted
24:54
it to a um sequence of only nine tokens
24:59
but now with a vocabulary of five
25:01
because we have a fifth vocabulary
25:02
element that we just created and it's Z
25:05
standing for concatination of AA and we
25:08
can again repeat this process so we
25:10
again look at the sequence and identify
25:13
the pair of tokens that are most
25:16
frequent let's say that that is now AB
25:19
well we are going to replace AB with a
25:21
new token that we meant call Y so y
25:24
becomes ab and then every single
25:25
occurrence of ab is now replaced with y
25:28
so we end up with this so now we only
25:31
have 1 2 3 4 5 6 seven characters in our
25:35
sequence but we have not just um four
25:40
vocabulary elements or five but now we
25:42
have six and for the final round we
25:46
again look through the sequence find
25:48
that the phrase zy or the pair zy is
25:51
most common and replace it one more time
25:53
with another um character let's say x so
25:57
X is z y and we replace all curses of zy
26:00
and we get this following sequence so
26:02
basically after we have gone through
26:04
this process instead of having a um
26:08
sequence of
26:10
11 uh tokens with a vocabulary length of
26:14
four we now have a sequence of 1 2 3
26:18
four five tokens but our vocabulary
26:21
length now is seven and so in this way
26:25
we can iteratively compress our sequence
26:27
I we Mint new tokens so in the in the
26:30
exact same way we start we start out
26:32
with bite sequences so we have 256
26:36
vocabulary size but we're now going to
26:38
go through these and find the bite pairs
26:41
that occur the most and we're going to
26:43
iteratively start minting new tokens
26:45
appending them to our vocabulary and
26:47
replacing things and in this way we're
26:49
going to end up with a compressed
26:50
training data set and also an algorithm
26:53
for taking any arbitrary sequence and
26:55
encoding it using this uh vocabul
26:58
and also decoding it back to Strings so
27:01
let's now Implement all that so here's
27:03
what I did I went to this block post
27:06
that I enjoyed and I took the first
27:07
paragraph and I copy pasted it here into
27:10
text so this is one very long line
27:13
here now to get the tokens as I
27:16
mentioned we just take our text and we
27:17
encode it into utf8 the tokens here at
27:20
this point will be a raw bites single
27:23
stream of bytes and just so that it's
27:26
easier to work with instead of just a
27:28
bytes object I'm going to convert all
27:30
those bytes to integers and then create
27:33
a list of it just so it's easier for us
27:34
to manipulate and work with in Python
27:36
and visualize and here I'm printing all
27:38
of that so this is the original um this
27:42
is the original paragraph and its length
27:45
is
27:46
533 uh code points and then here are the
27:50
bytes encoded in ut utf8 and we see that
27:53
this has a length of 616 bytes at this
27:56
point or 616 tokens and the reason this
27:59
is more is because a lot of these simple
28:02
asky characters or simple characters
28:05
they just become a single bite but a lot
28:06
of these Unicode more complex characters
28:09
become multiple bytes up to four and so
28:11
we are expanding that
28:13
size so now what we'd like to do as a
28:15
first step of the algorithm is we'd like
28:16
to iterate over here and find the pair
28:19
of bites that occur most frequently
28:22
because we're then going to merge it so
28:24
if you are working long on a notebook on
28:26
a side then I encourage you to basically
28:28
click on the link find this notebook and
28:30
try to write that function yourself
28:32
otherwise I'm going to come here and
28:33
Implement first the function that finds
28:35
the most common pair okay so here's what
28:37
I came up with there are many different
28:38
ways to implement this but I'm calling
28:40
the function get stats it expects a list
28:42
of integers I'm using a dictionary to
28:44
keep track of basically the counts and
28:47
then this is a pythonic way to iterate
28:49
consecutive elements of this list uh
28:51
which we covered in the previous video
28:54
and then here I'm just keeping track of
28:56
just incrementing by one um for all the
28:59
pairs so if I call this on all the
29:00
tokens here then the stats comes out
29:03
here so this is the dictionary the keys
29:06
are these topples of consecutive
29:09
elements and this is the count so just
29:12
to uh print it in a slightly better way
29:15
this is one way that I like to do that
29:18
where you it's a little bit compound
29:21
here so you can pause if you like but we
29:22
iterate all all the items the items
29:25
called on dictionary returns pairs of
29:27
key value and instead I create a list
29:32
here of value key because if it's a
29:35
value key list then I can call sort on
29:37
it and by default python will uh use the
29:41
first element which in this case will be
29:44
value to sort by if it's given tles and
29:47
then reverse so it's descending and
29:49
print that so basically it looks like
29:51
101 comma 32 was the most commonly
29:54
occurring consecutive pair and it
29:56
occurred 20 times we can double check
29:58
that that makes reasonable sense so if I
30:00
just search
30:02
10132 then you see that these are the 20
30:05
occurrences of that um pair and if we'd
30:10
like to take a look at what exactly that
30:12
pair is we can use Char which is the
30:14
opposite of or in Python so we give it a
30:18
um unic code Cod point so 101 and of 32
30:22
and we see that this is e and space so
30:25
basically there's a lot of E space here
30:28
meaning that a lot of these words seem
30:29
to end with e so here's eace as an
30:32
example so there's a lot of that going
30:34
on here and this is the most common pair
30:37
so now that we've identified the most
30:38
common pair we would like to iterate
30:40
over this sequence we're going to Mint a
30:43
new token with the ID of
30:45
256 right because these tokens currently
30:48
go from Z to 255 so when we create a new
30:51
token it will have an ID of
30:53
256 and we're going to iterate over this
30:56
entire um list and every every time we
31:00
see 101 comma 32 we're going to swap
31:03
that out for
31:04
256 so let's Implement that now and feel
31:07
free to uh do that yourself as well so
31:10
first I commented uh this just so we
31:12
don't pollute uh the notebook too much
31:15
this is a nice way of in Python
31:18
obtaining the highest ranking pair so
31:20
we're basically calling the Max on this
31:23
dictionary stats and this will return
31:26
the maximum
31:28
key and then the question is how does it
31:30
rank keys so you can provide it with a
31:33
function that ranks keys and that
31:35
function is just stats. getet uh stats.
31:38
getet would basically return the value
31:41
and so we're ranking by the value and
31:43
getting the maximum key so it's 101
31:45
comma 32 as we saw now to actually merge
31:49
10132 um this is the function that I
31:52
wrote but again there are many different
31:53
versions of it so we're going to take a
31:56
list of IDs and the the pair that we
31:58
want to replace and that pair will be
32:00
replaced with the new index
32:02
idx so iterating through IDs if we find
32:06
the pair swap it out for idx so we
32:08
create this new list and then we start
32:11
at zero and then we go through this
32:13
entire list sequentially from left to
32:15
right and here we are checking for
32:17
equality at the current position with
32:20
the
32:21
pair um so here we are checking that the
32:23
pair matches now here is a bit of a
32:25
tricky condition that you have to append
32:27
if you're trying to be careful and that
32:29
is that um you don't want this here to
32:32
be out of Bounds at the very last
32:34
position when you're on the rightmost
32:35
element of this list otherwise this
32:37
would uh give you an autof bounds error
32:39
so we have to make sure that we're not
32:41
at the very very last element so uh this
32:44
would be false for that so if we find a
32:47
match we append to this new list that
32:51
replacement index and we increment the
32:53
position by two so we skip over that
32:55
entire pair but otherwise if we we
32:57
haven't found a matching pair we just
32:59
sort of copy over the um element at that
33:02
position and increment by one then
33:05
return this so here's a very small toy
33:07
example if we have a list 566 791 and we
33:10
want to replace the occurrences of 67
33:12
with 99 then calling this on that will
33:16
give us what we're asking for so here
33:19
the 67 is replaced with
33:22
99 so now I'm going to uncomment this
33:24
for our actual use case where we want to
33:27
take our tokens we want to take the top
33:30
pair here and replace it with 256 to get
33:33
tokens to if we run this we get the
33:37
following so recall that previously we
33:41
had a length 616 in this list and now we
33:45
have a length 596 right so this
33:48
decreased by 20 which makes sense
33:50
because there are 20 occurrences
33:52
moreover we can try to find 256 here and
33:55
we see plenty of occurrences on off it
33:58
and moreover just double check there
34:00
should be no occurrence of 10132 so this
34:03
is the original array plenty of them and
34:05
in the second array there are no
34:06
occurrences of 1032 so we've
34:09
successfully merged this single pair and
34:12
now we just uh iterate this so we are
34:14
going to go over the sequence again find
34:15
the most common pair and replace it so
34:18
let me now write a y Loop that uses
34:19
these functions to do this um sort of
34:22
iteratively and how many times do we do
34:24
it four well that's totally up to us as
34:26
a hyper parameter
34:27
the more um steps we take the larger
34:31
will be our vocabulary and the shorter
34:33
will be our sequence and there is some
34:35
sweet spot that we usually find works
34:37
the best in practice and so this is kind
34:40
of a hyperparameter and we tune it and
34:42
we find good vocabulary sizes as an
34:44
example gp4 currently uses roughly
34:46
100,000 tokens and um bpark that those
34:50
are reasonable numbers currently instead
34:52
the are large language models so let me
34:54
now write uh putting putting it all
34:56
together and uh iterating these steps
34:59
okay now before we dive into the Y loop
35:01
I wanted to add one more cell here where
35:03
I went to the block post and instead of
35:05
grabbing just the first paragraph or two
35:07
I took the entire block post and I
35:09
stretched it out in a single line and
35:11
basically just using longer text will
35:12
allow us to have more representative
35:14
statistics for the bite Pairs and we'll
35:16
just get a more sensible results out of
35:18
it because it's longer text um so here
35:22
we have the raw text we encode it into
35:24
bytes using the utf8 encoding
35:28
and then here as before we are just
35:30
changing it into a list of integers in
35:32
Python just so it's easier to work with
35:34
instead of the raw byes objects and then
35:37
this is the code that I came up with uh
35:41
to actually do the merging in Loop these
35:44
two functions here are identical to what
35:46
we had above I only included them here
35:48
just so that you have the point of
35:50
reference here so uh these two are
35:53
identical and then this is the new code
35:55
that I added so the first first thing we
35:57
want to do is we want to decide on the
35:59
final vocabulary size that we want our
36:01
tokenizer to have and as I mentioned
36:03
this is a hyper parameter and you set it
36:05
in some way depending on your best
36:06
performance so let's say for us we're
36:08
going to use 276 because that way we're
36:11
going to be doing exactly 20
36:13
merges and uh 20 merges because we
36:16
already have
36:17
256 tokens for the raw bytes and to
36:21
reach 276 we have to do 20 merges uh to
36:24
add 20 new
36:25
tokens here uh this is uh one way in
36:28
Python to just create a copy of a list
36:31
so I'm taking the tokens list and by
36:34
wrapping it in a list python will
36:36
construct a new list of all the
36:37
individual elements so this is just a
36:39
copy
36:40
operation then here I'm creating a
36:42
merges uh dictionary so this merges
36:45
dictionary is going to maintain
36:46
basically the child one child two
36:49
mapping to a new uh token and so what
36:53
we're going to be building up here is a
36:54
binary tree of merges but actually it's
36:57
not exactly a tree because a tree would
36:59
have a single root node with a bunch of
37:01
leaves for us we're starting with the
37:03
leaves on the bottom which are the
37:05
individual bites those are the starting
37:07
256 tokens and then we're starting to
37:10
like merge two of them at a time and so
37:12
it's not a tree it's more like a forest
37:15
um uh as we merge these elements
37:19
so for 20 merges we're going to find the
37:23
most commonly occurring pair we're going
37:25
to Mint a new token integer for it so I
37:28
here will start at zero so we'll going
37:30
to start at 256 we're going to print
37:32
that we're merging it and we're going to
37:34
replace all of the occurrences of that
37:36
pair with the new new lied token and
37:40
we're going to record that this pair of
37:42
integers merged into this new
37:46
integer so running this gives us the
37:49
following
37:51
output so we did 20 merges and for
37:54
example the first merge was exactly as
37:57
before the
37:59
10132 um tokens merging into a new token
38:02
2556 now keep in mind that the
38:04
individual uh tokens 101 and 32 can
38:07
still occur in the sequence after
38:08
merging it's only when they occur
38:10
exactly consecutively that that becomes
38:13
256
38:14
now um and in particular the other thing
38:17
to notice here is that the token 256
38:19
which is the newly minted token is also
38:21
eligible for merging so here on the
38:23
bottom the 20th merge was a merge of 25
38:27
and 259 becoming
38:29
275 so every time we replace these
38:32
tokens they become eligible for merging
38:34
in the next round of data ration so
38:36
that's why we're building up a small
38:37
sort of binary Forest instead of a
38:39
single individual
38:40
tree one thing we can take a look at as
38:42
well is we can take a look at the
38:44
compression ratio that we've achieved so
38:46
in particular we started off with this
38:48
tokens list um so we started off with
38:51
24,000 bytes and after merging 20 times
38:56
uh we now have only
38:59
19,000 um tokens and so therefore the
39:02
compression ratio simply just dividing
39:04
the two is roughly 1.27 so that's the
39:07
amount of compression we were able to
39:08
achieve of this text with only 20
39:11
merges um and of course the more
39:13
vocabulary elements you add uh the
39:16
greater the compression ratio here would
39:19
be finally so that's kind of like um the
39:24
training of the tokenizer if you will
39:26
now 1 Point I wanted to make is that and
39:28
maybe this is a diagram that can help um
39:31
kind of illustrate is that tokenizer is
39:33
a completely separate object from the
39:35
large language model itself so
39:37
everything in this lecture we're not
39:38
really touching the llm itself uh we're
39:40
just training the tokenizer this is a
39:42
completely separate pre-processing stage
39:44
usually so the tokenizer will have its
39:46
own training set just like a large
39:48
language model has a potentially
39:50
different training set so the tokenizer
39:52
has a training set of documents on which
39:53
you're going to train the
39:55
tokenizer and then and um we're
39:58
performing The Bite pair encoding
39:59
algorithm as we saw above to train the
40:01
vocabulary of this
40:03
tokenizer so it has its own training set
40:05
it is a pre-processing stage that you
40:07
would run a single time in the beginning
40:09
um and the tokenizer is trained using
40:12
bipar coding algorithm once you have the
40:14
tokenizer once it's trained and you have
40:16
the vocabulary and you have the merges
40:19
uh we can do both encoding and decoding
40:22
so these two arrows here so the
40:25
tokenizer is a translation layer between
40:27
raw text which is as we saw the sequence
40:30
of Unicode code points it can take raw
40:33
text and turn it into a token sequence
40:35
and vice versa it can take a token
40:37
sequence and translate it back into raw
40:41
text so now that we have trained uh
40:43
tokenizer and we have these merges we
40:46
are going to turn to how we can do the
40:47
encoding and the decoding step if you
40:49
give me text here are the tokens and
40:51
vice versa if you give me tokens here's
40:53
the text once we have that we can
40:55
translate between these two Realms and
40:58
then the language model is going to be
40:59
trained as a step two afterwards and
41:02
typically in a in a sort of a
41:04
state-of-the-art application you might
41:05
take all of your training data for the
41:07
language model and you might run it
41:08
through the tokenizer and sort of
41:10
translate everything into a massive
41:12
token sequence and then you can throw
41:14
away the raw text you're just left with
41:15
the tokens themselves and those are
41:18
stored on disk and that is what the
41:20
large language model is actually reading
41:21
when it's training on them so this one
41:23
approach that you can take as a single
41:25
massive pre-processing step a
41:27
stage um so yeah basically I think the
41:30
most important thing I want to get
41:31
across is that this is completely
41:33
separate stage it usually has its own
41:34
entire uh training set you may want to
41:37
have those training sets be different
41:38
between the tokenizer and the logge
41:40
language model so for example when
41:41
you're training the tokenizer as I
41:43
mentioned we don't just care about the
41:45
performance of English text we care
41:47
about uh multi many different languages
41:49
and we also care about code or not code
41:52
so you may want to look into different
41:53
kinds of mixtures of different kinds of
41:55
languages and different amounts of code
41:57
and things like that because the amount
42:00
of different language that you have in
42:02
your tokenizer training set will
42:04
determine how many merges of it there
42:06
will be and therefore that determines
42:08
the density with which uh this type of
42:11
data is um sort of has in the token
42:15
space and so roughly speaking
42:18
intuitively if you add some amount of
42:20
data like say you have a ton of Japanese
42:21
data in your uh tokenizer training set
42:24
then that means that more Japanese
42:25
tokens will get merged
42:27
and therefore Japanese will have shorter
42:29
sequences uh and that's going to be
42:31
beneficial for the large language model
42:32
which has a finite context length on
42:34
which it can work on in in the token
42:37
space uh so hopefully that makes sense
42:39
so we're now going to turn to encoding
42:41
and decoding now that we have trained a
42:43
tokenizer so we have our merges and now
42:46
how do we do encoding and decoding okay
42:48
so let's begin with decoding which is
42:50
this Arrow over here so given a token
42:53
sequence let's go through the tokenizer
42:55
to get back a python string object so
42:58
the raw text so this is the function
43:00
that we' like to implement um we're
43:02
given the list of integers and we want
43:03
to return a python string if you'd like
43:06
uh try to implement this function
43:07
yourself it's a fun exercise otherwise
43:09
I'm going to start uh pasting in my own
43:11
solution so there are many different
43:14
ways to do it um here's one way I will
43:17
create an uh kind of pre-processing
43:19
variable that I will call
43:21
vocab and vocab is a mapping or a
43:25
dictionary in Python for from the token
43:28
uh ID to the bytes object for that token
43:32
so we begin with the raw bytes for
43:34
tokens from 0 to 255 and then we go in
43:37
order of all the merges and we sort of
43:40
uh populate this vocab list by doing an
43:42
addition here so this is the basically
43:46
the bytes representation of the first
43:48
child followed by the second one and
43:50
remember these are bytes objects so this
43:52
addition here is an addition of two
43:54
bytes objects just concatenation
43:57
so that's what we get
43:59
here one tricky thing to be careful with
44:01
by the way is that I'm iterating a
44:03
dictionary in Python using a DOT items
44:06
and uh it really matters that this runs
44:09
in the order in which we inserted items
44:11
into the merous dictionary luckily
44:14
starting with python 3.7 this is
44:15
guaranteed to be the case but before
44:17
python 3.7 this iteration may have been
44:19
out of order with respect to how we
44:21
inserted elements into merges and this
44:23
may not have worked but we are using an
44:26
um modern python so we're okay and then
44:29
here uh given the IDS the first thing
44:32
we're going to do is get the
44:35
tokens so the way I implemented this
44:37
here is I'm taking I'm iterating over
44:40
all the IDS I'm using vocap to look up
44:42
their bytes and then here this is one
44:44
way in Python to concatenate all these
44:47
bytes together to create our tokens and
44:50
then these tokens here at this point are
44:52
raw bytes so I have to decode using UTF
44:56
F now back into python strings so
44:59
previously we called that encode on a
45:01
string object to get the bytes and now
45:03
we're doing it Opposite we're taking the
45:05
bytes and calling a decode on the bytes
45:08
object to get a string in Python and
45:11
then we can return
45:13
text so um this is how we can do it now
45:17
this actually has a um issue um in the
45:21
way I implemented it and this could
45:22
actually throw an error so try to think
45:24
figure out why this code could actually
45:26
result in an error if we plug in um uh
45:30
some sequence of IDs that is
45:33
unlucky so let me demonstrate the issue
45:35
when I try to decode just something like
45:37
97 I am going to get letter A here back
45:41
so nothing too crazy happening but when
45:44
I try to decode 128 as a single element
45:48
the token 128 is what in string or in
45:51
Python object uni Cod decoder utfa can't
45:55
Decode by um 0x8 which is this in HEX in
46:00
position zero invalid start bite what
46:02
does that mean well to understand what
46:04
this means we have to go back to our
46:05
utf8 page uh that I briefly showed
46:08
earlier and this is Wikipedia utf8 and
46:11
basically there's a specific schema that
46:14
utfa bytes take so in particular if you
46:17
have a multi-te object for some of the
46:20
Unicode characters they have to have
46:22
this special sort of envelope in how the
46:24
encoding works and so what's happening
46:27
here is that invalid start pite that's
46:30
because
46:31
128 the binary representation of it is
46:34
one followed by all zeros so we have one
46:37
and then all zero and we see here that
46:40
that doesn't conform to the format
46:41
because one followed by all zero just
46:43
doesn't fit any of these rules so to
46:45
speak so it's an invalid start bite
46:48
which is byte one this one must have a
46:51
one following it and then a zero
46:53
following it and then the content of
46:54
your uni codee in x here so basically we
46:58
don't um exactly follow the utf8
47:00
standard and this cannot be decoded and
47:03
so the way to fix this um is to
47:06
use this errors equals in bytes. decode
47:11
function of python and by default errors
47:14
is strict so we will throw an error if
47:17
um it's not valid utf8 bytes encoding
47:20
but there are many different things that
47:22
you could put here on error handling
47:24
this is the full list of all the errors
47:25
that you can use and in particular
47:27
instead of strict let's change it to
47:29
replace and that will replace uh with
47:32
this special marker this replacement
47:36
character so errors equals replace and
47:41
now we just get that character
47:43
back so basically not every single by
47:47
sequence is valid
47:49
utf8 and if it happens that your large
47:51
language model for example predicts your
47:54
tokens in a bad manner then they might
47:57
not fall into valid utf8 and then we
48:00
won't be able to decode them so the
48:03
standard practice is to basically uh use
48:06
errors equals replace and this is what
48:08
you will also find in the openai um code
48:10
that they released as well but basically
48:13
whenever you see um this kind of a
48:14
character in your output in that case uh
48:16
something went wrong and the LM output
48:18
not was not valid uh sort of sequence of
48:22
tokens okay and now we're going to go
48:23
the other way so we are going to
48:25
implement
48:26
this Arrow right here where we are going
48:28
to be given a string and we want to
48:30
encode it into
48:31
tokens so this is the signature of the
48:34
function that we're interested in and um
48:37
this should basically print a list of
48:38
integers of the tokens so again uh try
48:42
to maybe implement this yourself if
48:43
you'd like a fun exercise uh and pause
48:46
here otherwise I'm going to start
48:47
putting in my
48:48
solution so again there are many ways to
48:50
do this so um this is one of the ways
48:54
that sort of I came came up with so the
48:58
first thing we're going to do is we are
48:59
going
49:00
to uh take our text encode it into utf8
49:03
to get the raw bytes and then as before
49:06
we're going to call list on the bytes
49:07
object to get a list of integers of
49:10
those bytes so those are the starting
49:13
tokens those are the raw bytes of our
49:15
sequence but now of course according to
49:17
the merges dictionary above and recall
49:20
this was the
49:21
merges some of the bytes may be merged
49:24
according to this lookup in addition to
49:27
that remember that the merges was built
49:28
from top to bottom and this is sort of
49:30
the order in which we inserted stuff
49:31
into merges and so we prefer to do all
49:34
these merges in the beginning before we
49:36
do these merges later because um for
49:39
example this merge over here relies on
49:41
the 256 which got merged here so we have
49:45
to go in the order from top to bottom
49:47
sort of if we are going to be merging
49:49
anything now we expect to be doing a few
49:51
merges so we're going to be doing W
49:55
true um and now we want to find a pair
49:58
of byes that is consecutive that we are
50:01
allowed to merge according to this in
50:04
order to reuse some of the functionality
50:05
that we've already written I'm going to
50:07
reuse the function uh get
50:09
stats so recall that get stats uh will
50:12
give us the we'll basically count up how
50:14
many times every single pair occurs in
50:17
our sequence of tokens and return that
50:19
as a dictionary and the dictionary was a
50:22
mapping from all the different uh by
50:26
pairs to the number of times that they
50:27
occur right um at this point we don't
50:30
actually care how many times they occur
50:32
in the sequence we only care what the
50:34
raw pairs are in that sequence and so
50:37
I'm only going to be using basically the
50:38
keys of the dictionary I only care about
50:40
the set of possible merge candidates if
50:43
that makes
50:44
sense now we want to identify the pair
50:46
that we're going to be merging at this
50:48
stage of the loop so what do we want we
50:50
want to find the pair or like the a key
50:53
inside stats that has the lowest index
50:57
in the merges uh dictionary because we
51:00
want to do all the early merges before
51:01
we work our way to the late
51:03
merges so again there are many different
51:05
ways to implement this but I'm going to
51:08
do something a little bit fancy
51:11
here so I'm going to be using the Min
51:14
over an iterator in Python when you call
51:17
Min on an iterator and stats here as a
51:19
dictionary we're going to be iterating
51:21
the keys of this dictionary in Python so
51:24
we're looking at all the pairs inside
51:27
stats um which are all the consecutive
51:29
Pairs and we're going to be taking the
51:32
consecutive pair inside tokens that has
51:34
the minimum what the Min takes a key
51:39
which gives us the function that is
51:40
going to return a value over which we're
51:42
going to do the Min and the one we care
51:45
about is we're we care about taking
51:46
merges and basically getting um that
51:51
pairs
51:53
index so basically for any pair inside
51:57
stats we are going to be looking into
52:00
merges at what index it has and we want
52:03
to get the pair with the Min number so
52:06
as an example if there's a pair 101 and
52:08
32 we definitely want to get that pair
52:10
uh we want to identify it here and
52:12
return it and pair would become 10132 if
52:15
it
52:16
occurs and the reason that I'm putting a
52:18
float INF here as a fall back is that in
52:21
the get function when we call uh when we
52:24
basically consider a pair that doesn't
52:27
occur in the merges then that pair is
52:29
not eligible to be merged right so if in
52:32
the token sequence there's some pair
52:33
that is not a merging pair it cannot be
52:36
merged then uh it doesn't actually occur
52:38
here and it doesn't have an index and uh
52:41
it cannot be merged which we will denote
52:43
as float INF and the reason Infinity is
52:45
nice here is because for sure we're
52:47
guaranteed that it's not going to
52:48
participate in the list of candidates
52:50
when we do the men so uh so this is one
52:53
way to do it so B basically long story
52:56
short this Returns the most eligible
52:58
merging candidate pair uh that occurs in
53:01
the tokens now one thing to be careful
53:04
with here is this uh function here might
53:07
fail in the following way if there's
53:10
nothing to merge then uh uh then there's
53:14
nothing in merges um that satisfi that
53:17
is satisfied anymore there's nothing to
53:19
merge everything just returns float imps
53:22
and then the pair I think will just
53:24
become the very first element of stats
53:27
um but this pair is not actually a
53:28
mergeable pair it just becomes the first
53:31
pair inside stats arbitrarily because
53:33
all of these pairs evaluate to float in
53:36
for the merging Criterion so basically
53:39
it could be that this this doesn't look
53:40
succeed because there's no more merging
53:42
pairs so if this pair is not in merges
53:45
that was returned then this is a signal
53:47
for us that actually there was nothing
53:48
to merge no single pair can be merged
53:51
anymore in that case we will break
53:53
out um nothing else can be
53:58
merged you may come up with a different
54:00
implementation by the way this is kind
54:01
of like really trying hard in
54:04
Python um but really we're just trying
54:06
to find a pair that can be merged with
54:08
the lowest index
54:10
here now if we did find a pair that is
54:14
inside merges with the lowest index then
54:16
we can merge it
54:20
so we're going to look into the merger
54:22
dictionary for that pair to look up the
54:24
index and we're going to now merge that
54:27
into that index so we're going to do
54:29
tokens equals and we're going to
54:32
replace the original tokens we're going
54:35
to be replacing the pair pair and we're
54:37
going to be replacing it with index idx
54:39
and this returns a new list of tokens
54:42
where every occurrence of pair is
54:43
replaced with idx so we're doing a merge
54:46
and we're going to be continuing this
54:48
until eventually nothing can be merged
54:49
we'll come out here and we'll break out
54:51
and here we just return
54:53
tokens and so that that's the
54:56
implementation I think so hopefully this
54:57
runs okay cool um yeah and this looks uh
55:02
reasonable so for example 32 is a space
55:05
in asky so that's here um so this looks
55:09
like it worked great okay so let's wrap
55:11
up this section of the video at least I
55:13
wanted to point out that this is not
55:15
quite the right implementation just yet
55:16
because we are leaving out a special
55:18
case so in particular if uh we try to do
55:21
this this would give us an error and the
55:24
issue is that um if we only have a
55:26
single character or an empty string then
55:28
stats is empty and that causes an issue
55:30
inside Min so one way to fight this is
55:33
if L of tokens is at least two because
55:36
if it's less than two it's just a single
55:38
token or no tokens then let's just uh
55:40
there's nothing to merge so we just
55:42
return so that would fix uh that
55:45
case Okay and then second I have a few
55:48
test cases here for us as well so first
55:50
let's make sure uh about or let's note
55:53
the following if we take a string and we
55:56
try to encode it and then decode it back
55:59
you'd expect to get the same string back
56:00
right is that true for all
56:05
strings so I think uh so here it is the
56:07
case and I think in general this is
56:09
probably the case um but notice that
56:12
going backwards is not is not you're not
56:15
going to have an identity going
56:16
backwards because as I mentioned us not
56:19
all token sequences are valid utf8 uh
56:23
sort of by streams and so so therefore
56:25
you're some of them can't even be
56:27
decodable um so this only goes in One
56:30
Direction but for that one direction we
56:33
can check uh here if we take the
56:35
training text which is the text that we
56:36
train to tokenizer around we can make
56:38
sure that when we encode and decode we
56:39
get the same thing back which is true
56:42
and here I took some validation data so
56:44
I went to I think this web page and I
56:46
grabbed some text so this is text that
56:48
the tokenizer has not seen and we can
56:50
make sure that this also works um okay
56:53
so that gives us some confidence that
56:54
this was correctly implemented
56:56
so those are the basics of the bite pair
56:58
encoding algorithm we saw how we can uh
57:01
take some training set train a tokenizer
57:04
the parameters of this tokenizer really
57:05
are just this dictionary of merges and
57:08
that basically creates the little binary
57:10
Forest on top of raw
57:12
bites once we have this the merges table
57:15
we can both encode and decode between
57:17
raw text and token sequences so that's
57:19
the the simplest setting of The
57:21
tokenizer what we're going to do now
57:23
though is we're going to look at some of
57:24
the St the art lar language models and
57:27
the kinds of tokenizers that they use
57:28
and we're going to see that this picture
57:30
complexifies very quickly so we're going
57:32
to go through the details of this comp
57:35
complexification one at a time so let's
57:38
kick things off by looking at the GPD
57:39
Series so in particular I have the gpt2
57:42
paper here um and this paper is from
57:45
2019 or so so 5 years ago and let's
57:48
scroll down to input representation this
57:51
is where they talk about the tokenizer
57:53
that they're using for gpd2 now this is
57:56
all fairly readable so I encourage you
57:57
to pause and um read this yourself but
58:00
this is where they motivate the use of
58:02
the bite pair encoding algorithm on the
58:05
bite level representation of utf8
58:08
encoding so this is where they motivate
58:10
it and they talk about the vocabulary
58:11
sizes and everything now everything here
58:14
is exactly as we've covered it so far
58:16
but things start to depart around here
58:19
so what they mention is that they don't
58:20
just apply the naive algorithm as we
58:22
have done it and in particular here's a
58:25
example suppose that you have common
58:27
words like dog what will happen is that
58:29
dog of course occurs very frequently in
58:32
the text and it occurs right next to all
58:34
kinds of punctuation as an example so
58:36
doc dot dog exclamation mark dog
58:39
question mark Etc and naively you might
58:42
imagine that the BP algorithm could
58:44
merge these to be single tokens and then
58:46
you end up with lots of tokens that are
58:47
just like dog with a slightly different
58:49
punctuation and so it feels like you're
58:51
clustering things that shouldn't be
58:52
clustered you're combining kind of
58:54
semantics with
58:56
uation and this uh feels suboptimal and
58:59
indeed they also say that this is
59:01
suboptimal according to some of the
59:02
experiments so what they want to do is
59:04
they want to top down in a manual way
59:06
enforce that some types of um characters
59:10
should never be merged together um so
59:13
they want to enforce these merging rules
59:15
on top of the bite PA encoding algorithm
59:18
so let's take a look um at their code
59:20
and see how they actually enforce this
59:21
and what kinds of mergy they actually do
59:23
perform so I have to to tab open here
59:26
for gpt2 under open AI on GitHub and
59:30
when we go to
59:31
Source there is an encoder thatp now I
59:34
don't personally love that they call it
59:36
encoder dopy because this is the
59:37
tokenizer and the tokenizer can do both
59:39
encode and decode uh so it feels kind of
59:42
awkward to me that it's called encoder
59:43
but that is the tokenizer and there's a
59:46
lot going on here and we're going to
59:47
step through it in detail at one point
59:49
for now I just want to focus on this
59:52
part here the create a rigix pattern
59:54
here that looks very complicated and
59:56
we're going to go through it in a bit uh
59:59
but this is the core part that allows
1:00:00
them to enforce rules uh for what parts
1:00:04
of the text Will Never Be merged for
1:00:06
sure now notice that re. compile here is
1:00:09
a little bit misleading because we're
1:00:11
not just doing import re which is the
1:00:12
python re module we're doing import reex
1:00:15
as re and reex is a python package that
1:00:18
you can install P install r x and it's
1:00:20
basically an extension of re so it's a
1:00:22
bit more powerful
1:00:23
re um
1:00:26
so let's take a look at this pattern and
1:00:29
what it's doing and why this is actually
1:00:31
doing the separation that they are
1:00:33
looking for okay so I've copy pasted the
1:00:35
pattern here to our jupit notebook where
1:00:37
we left off and let's take this pattern
1:00:39
for a spin so in the exact same way that
1:00:42
their code does we're going to call an
1:00:44
re. findall for this pattern on any
1:00:47
arbitrary string that we are interested
1:00:49
so this is the string that we want to
1:00:51
encode into tokens um to feed into n llm
1:00:55
like gpt2 so what exactly is this doing
1:00:59
well re. findall will take this pattern
1:01:01
and try to match it against a
1:01:03
string um the way this works is that you
1:01:06
are going from left to right in the
1:01:08
string and you're trying to match the
1:01:10
pattern and R.F find all will get all
1:01:14
the occurrences and organize them into a
1:01:16
list now when you look at the um when
1:01:19
you look at this pattern first of all
1:01:21
notice that this is a raw string um and
1:01:24
then these are three double quotes just
1:01:26
to start the string so really the string
1:01:29
itself this is the pattern itself
1:01:31
right and notice that it's made up of a
1:01:34
lot of ores so see these vertical bars
1:01:36
those are ores in reg X and so you go
1:01:40
from left to right in this pattern and
1:01:41
try to match it against the string
1:01:43
wherever you are so we have hello and
1:01:46
we're going to try to match it well it's
1:01:48
not apostrophe s it's not apostrophe t
1:01:51
or any of these but it is an optional
1:01:54
space followed by- P of uh sorry SL P of
1:01:58
L one or more times what is/ P of L it
1:02:02
is coming to some documentation that I
1:02:05
found um there might be other sources as
1:02:08
well uh SLP is a letter any kind of
1:02:12
letter from any language and hello is
1:02:15
made up of letters h e l Etc so optional
1:02:20
space followed by a bunch of letters one
1:02:22
or more letters is going to match hello
1:02:25
but then the match ends because a white
1:02:27
space is not a letter so from there on
1:02:31
begins a new sort of attempt to match
1:02:34
against the string again and starting in
1:02:36
here we're going to skip over all of
1:02:38
these again until we get to the exact
1:02:40
same Point again and we see that there's
1:02:42
an optional space this is the optional
1:02:44
space followed by a bunch of letters one
1:02:46
or more of them and so that matches so
1:02:49
when we run this we get a list of two
1:02:52
elements hello and then space world
1:02:56
so how are you if we add more letters we
1:02:59
would just get them like this now what
1:03:02
is this doing and why is this important
1:03:04
we are taking our string and instead of
1:03:06
directly encoding it um for
1:03:09
tokenization we are first splitting it
1:03:11
up and when you actually step through
1:03:13
the code and we'll do that in a bit more
1:03:15
detail what really is doing on a high
1:03:17
level is that it first splits your text
1:03:21
into a list of texts just like this one
1:03:25
and all these elements of this list are
1:03:27
processed independently by the tokenizer
1:03:29
and all of the results of that
1:03:31
processing are simply
1:03:32
concatenated so hello world oh I I
1:03:36
missed how hello world how are you we
1:03:40
have five elements of list all of these
1:03:42
will independent
1:03:44
independently go from text to a token
1:03:47
sequence and then that token sequence is
1:03:49
going to be concatenated it's all going
1:03:51
to be joined up and roughly speaking
1:03:54
what that does is you're only ever
1:03:56
finding merges between the elements of
1:03:58
this list so you can only ever consider
1:04:00
merges within every one of these
1:04:02
elements in
1:04:03
individually and um after you've done
1:04:06
all the possible merging for all of
1:04:08
these elements individually the results
1:04:10
of all that will be joined um by
1:04:14
concatenation and so you are basically
1:04:16
what what you're doing effectively is
1:04:18
you are never going to be merging this e
1:04:21
with this space because they are now
1:04:23
parts of the separate elements of this
1:04:25
list and so you are saying we are never
1:04:28
going to merge
1:04:29
eace um because we're breaking it up in
1:04:32
this way so basically using this regx
1:04:36
pattern to Chunk Up the text is just one
1:04:38
way of enforcing that some merges are
1:04:42
not to happen and we're going to go into
1:04:44
more of this text and we'll see that
1:04:45
what this is trying to do on a high
1:04:46
level is we're trying to not merge
1:04:48
across letters across numbers across
1:04:51
punctuation and so on so let's see in
1:04:53
more detail how that works so let's
1:04:55
continue now we have/ P ofn if you go to
1:04:58
the documentation SLP of n is any kind
1:05:02
of numeric character in any script so
1:05:04
it's numbers so we have an optional
1:05:07
space followed by numbers and those
1:05:08
would be separated out so letters and
1:05:10
numbers are being separated so if I do
1:05:13
Hello World 123 how are you then world
1:05:16
will stop matching here because one is
1:05:18
not a letter anymore but one is a number
1:05:21
so this group will match for that and
1:05:23
we'll get it as a separate entity
1:05:27
uh let's see how these apostrophes work
1:05:28
so here if we have
1:05:31
um uh Slash V or I mean apostrophe V as
1:05:35
an example then apostrophe here is not a
1:05:38
letter or a
1:05:40
number so hello will stop matching and
1:05:42
then we will exactly match this with
1:05:45
that so that will come out as a separate
1:05:48
thing so why are they doing the
1:05:50
apostrophes here honestly I think that
1:05:52
these are just like very common
1:05:54
apostrophes p uh that are used um
1:05:57
typically I don't love that they've done
1:05:59
this
1:06:01
because uh let me show you what happens
1:06:03
when you have uh some Unicode
1:06:05
apostrophes like for example you can
1:06:07
have if you have house then this will be
1:06:11
separated out because of this matching
1:06:13
but if you use the Unicode apostrophe
1:06:15
like
1:06:16
this then suddenly this does not work
1:06:20
and so this apostrophe will actually
1:06:22
become its own thing now and so so um
1:06:25
it's basically hardcoded for this
1:06:26
specific kind of apostrophe and uh
1:06:30
otherwise they become completely
1:06:31
separate tokens in addition to this you
1:06:34
can go to the gpt2 docs and here when
1:06:38
they Define the pattern they say should
1:06:40
have added re. ignore case so BP merges
1:06:43
can happen for capitalized versions of
1:06:45
contractions so what they're pointing
1:06:47
out is that you see how this is
1:06:48
apostrophe and then lowercase letters
1:06:51
well because they didn't do re. ignore
1:06:53
case then then um these rules will not
1:06:56
separate out the apostrophes if it's
1:06:59
uppercase so
1:07:01
house would be like this but if I did
1:07:07
house if I'm uppercase then notice
1:07:10
suddenly the apostrophe comes by
1:07:12
itself so the tokenization will work
1:07:15
differently in uppercase and lower case
1:07:17
inconsistently separating out these
1:07:19
apostrophes so it feels extremely gnarly
1:07:21
and slightly gross um but that's that's
1:07:25
how that works okay so let's come back
1:07:27
after trying to match a bunch of
1:07:28
apostrophe Expressions by the way the
1:07:30
other issue here is that these are quite
1:07:32
language specific probably so I don't
1:07:35
know that all the languages for example
1:07:36
use or don't use apostrophes but that
1:07:37
would be inconsistently tokenized as a
1:07:40
result then we try to match letters then
1:07:43
we try to match numbers and then if that
1:07:45
doesn't work we fall back to here and
1:07:48
what this is saying is again optional
1:07:49
space followed by something that is not
1:07:51
a letter number or a space in one or
1:07:54
more of that so what this is doing
1:07:56
effectively is this is trying to match
1:07:58
punctuation roughly speaking not letters
1:08:00
and not numbers so this group will try
1:08:02
to trigger for that so if I do something
1:08:04
like this then these parts here are not
1:08:08
letters or numbers but they will
1:08:10
actually they are uh they will actually
1:08:12
get caught here and so they become its
1:08:14
own group so we've separated out the
1:08:17
punctuation and finally this um this is
1:08:20
also a little bit confusing so this is
1:08:22
matching white space but this is using a
1:08:25
negative look ahead assertion in regex
1:08:29
so what this is doing is it's matching
1:08:31
wh space up to but not including the
1:08:33
last Whit space
1:08:35
character why is this important um this
1:08:38
is pretty subtle I think so you see how
1:08:40
the white space is always included at
1:08:42
the beginning of the word so um space r
1:08:46
space u Etc suppose we have a lot of
1:08:48
spaces
1:08:49
here what's going to happen here is that
1:08:52
these spaces up to not including the
1:08:55
last character will get caught by this
1:08:58
and what that will do is it will
1:09:00
separate out the spaces up to but not
1:09:02
including the last character so that the
1:09:04
last character can come here and join
1:09:06
with the um space you and the reason
1:09:09
that's nice is because space you is the
1:09:11
common token so if I didn't have these
1:09:14
Extra Spaces here you would just have
1:09:15
space you and if I add tokens if I add
1:09:18
spaces we still have a space view but
1:09:21
now we have all this extra white space
1:09:23
so basically the GB to tokenizer really
1:09:25
likes to have a space letters or numbers
1:09:27
um and it it preens these spaces and
1:09:30
this is just something that it is
1:09:31
consistent about so that's what that is
1:09:34
for and then finally we have all the the
1:09:36
last fallback is um whites space
1:09:39
characters uh so um that would be
1:09:43
just um if that doesn't get caught then
1:09:47
this thing will catch any trailing
1:09:49
spaces and so on I wanted to show one
1:09:51
more real world example here so if we
1:09:53
have this string which is a piece of
1:09:54
python code and then we try to split it
1:09:56
up then this is the kind of output we
1:09:58
get so you'll notice that the list has
1:10:01
many elements here and that's because we
1:10:02
are splitting up fairly often uh every
1:10:05
time sort of a category
1:10:07
changes um so there will never be any
1:10:09
merges Within These
1:10:11
elements and um that's what you are
1:10:13
seeing here now you might think that in
1:10:16
order to train the
1:10:18
tokenizer uh open AI has used this to
1:10:21
split up text into chunks and then run
1:10:24
just a BP algorithm within all the
1:10:26
chunks but that is not exactly what
1:10:28
happened and the reason is the following
1:10:30
notice that we have the spaces here uh
1:10:33
those Spaces end up being entire
1:10:35
elements but these spaces never actually
1:10:38
end up being merged by by open Ai and
1:10:41
the way you can tell is that if you copy
1:10:42
paste the exact same chunk here into Tik
1:10:44
token U Tik tokenizer you see that all
1:10:47
the spaces are kept independent and
1:10:49
they're all token
1:10:51
220 so I think opena at some point Point
1:10:54
en Force some rule that these spaces
1:10:56
would never be merged and so um there's
1:10:59
some additional rules on top of just
1:11:01
chunking and bpe that open ey is not uh
1:11:04
clear about now the training code for
1:11:06
the gpt2 tokenizer was never released so
1:11:09
all we have is uh the code that I've
1:11:11
already shown you but this code here
1:11:13
that they've released is only the
1:11:14
inference code for the tokens so this is
1:11:18
not the training code you can't give it
1:11:19
a piece of text and training tokenizer
1:11:22
this is just the inference code which
1:11:23
Tak takes the merges that we have up
1:11:26
above and applies them to a new piece of
1:11:28
text and so we don't know exactly how
1:11:31
opening ey trained um train the
1:11:32
tokenizer but it wasn't as simple as
1:11:35
chunk it up and BP it uh whatever it was
1:11:38
next I wanted to introduce you to the
1:11:40
Tik token library from openai which is
1:11:42
the official library for tokenization
1:11:45
from openai so this is Tik token bip
1:11:48
install P to Tik token and then um you
1:11:51
can do the tokenization in inference
1:11:54
this is again not training code this is
1:11:56
only inference code for
1:11:58
tokenization um I wanted to show you how
1:12:00
you would use it quite simple and
1:12:02
running this just gives us the gpt2
1:12:04
tokens or the GPT 4 tokens so this is
1:12:07
the tokenizer use for GPT 4 and so in
1:12:10
particular we see that the Whit space in
1:12:11
gpt2 remains unmerged but in GPT 4 uh
1:12:14
these Whit spaces merge as we also saw
1:12:17
in this one where here they're all
1:12:19
unmerged but if we go down to GPT 4 uh
1:12:23
they become merged
1:12:25
um now in the
1:12:28
gp4 uh tokenizer they changed the
1:12:31
regular expression that they use to
1:12:33
Chunk Up text so the way to see this is
1:12:36
that if you come to your the Tik token
1:12:38
uh library and then you go to this file
1:12:41
Tik token X openi public this is where
1:12:44
sort of like the definition of all these
1:12:46
different tokenizers that openi
1:12:47
maintains is and so uh necessarily to do
1:12:51
the inference they had to publish some
1:12:52
of the details about the strings
1:12:54
so this is the string that we already
1:12:55
saw for gpt2 it is slightly different
1:12:58
but it is actually equivalent uh to what
1:13:00
we discussed here so this pattern that
1:13:03
we discussed is equivalent to this
1:13:05
pattern this one just executes a little
1:13:07
bit faster so here you see a little bit
1:13:09
of a slightly different definition but
1:13:11
otherwise it's the same we're going to
1:13:13
go into special tokens in a bit and then
1:13:15
if you scroll down to CL 100k this is
1:13:19
the GPT 4 tokenizer you see that the
1:13:21
pattern has changed um and this is kind
1:13:24
of like the main the major change in
1:13:26
addition to a bunch of other special
1:13:27
tokens which I'll go into in a bit again
1:13:30
now some I'm not going to actually go
1:13:32
into the full detail of the pattern
1:13:33
change because honestly this is my
1:13:35
numbing uh I would just advise that you
1:13:37
pull out chat GPT and the regex
1:13:40
documentation and just step through it
1:13:42
but really the major changes are number
1:13:45
one you see this eye here that means
1:13:48
that the um case sensitivity this is
1:13:51
case insensitive match and so the
1:13:54
comment that we saw earlier on oh we
1:13:56
should have used re. uppercase uh
1:13:58
basically we're now going to be matching
1:14:02
these apostrophe s apostrophe D
1:14:05
apostrophe M Etc uh we're going to be
1:14:07
matching them both in lowercase and in
1:14:09
uppercase so that's fixed there's a
1:14:11
bunch of different like handling of the
1:14:13
whites space that I'm not going to go
1:14:14
into the full details of and then one
1:14:16
more thing here is you will notice that
1:14:19
when they match the numbers they only
1:14:21
match one to three numbers so so they
1:14:24
will never merge
1:14:26
numbers that are in low in more than
1:14:29
three digits only up to three digits of
1:14:31
numbers will ever be merged and uh
1:14:35
that's one change that they made as well
1:14:36
to prevent uh tokens that are very very
1:14:39
long number
1:14:40
sequences uh but again we don't really
1:14:42
know why they do any of this stuff uh
1:14:44
because none of this is documented and
1:14:46
uh it's just we just get the pattern so
1:14:50
um yeah it is what it is but those are
1:14:52
some of the changes that gp4 has made
1:14:54
and of course the vocabulary size went
1:14:56
from roughly 50k to roughly
1:14:58
100K the next thing I would like to do
1:15:00
very briefly is to take you through the
1:15:02
gpt2 encoder dopy that openi has
1:15:05
released uh this is the file that I
1:15:07
already mentioned to you briefly now
1:15:10
this file is uh fairly short and should
1:15:13
be relatively understandable to you at
1:15:15
this point um starting at the bottom
1:15:18
here they are loading two files encoder
1:15:21
Json and vocab bpe and they do some
1:15:24
light processing on it and then they
1:15:25
call this encoder object which is the
1:15:28
tokenizer now if you'd like to inspect
1:15:30
these two files which together
1:15:32
constitute their saved tokenizer then
1:15:35
you can do that with a piece of code
1:15:36
like
1:15:37
this um this is where you can download
1:15:39
these two files and you can inspect them
1:15:41
if you'd like and what you will find is
1:15:43
that this encoder as they call it in
1:15:45
their code is exactly equivalent to our
1:15:48
vocab so remember here where we have
1:15:52
this vocab object which allowed us us to
1:15:53
decode very efficiently and basically it
1:15:56
took us from the integer to the byes uh
1:16:00
for that integer so our vocab is exactly
1:16:03
their encoder and then their vocab bpe
1:16:08
confusingly is actually are merges so
1:16:11
their BP merges which is based on the
1:16:14
data inside vocab bpe ends up being
1:16:17
equivalent to our merges so uh basically
1:16:21
they are saving and loading the two uh
1:16:24
variables that for us are also critical
1:16:26
the merges variable and the vocab
1:16:28
variable using just these two variables
1:16:31
you can represent a tokenizer and you
1:16:33
can both do encoding and decoding once
1:16:35
you've trained this
1:16:36
tokenizer now the only thing that um is
1:16:40
actually slightly confusing inside what
1:16:43
opening ey does here is that in addition
1:16:45
to this encoder and a decoder they also
1:16:47
have something called a bite encoder and
1:16:49
a bite decoder and this is actually
1:16:51
unfortunately just
1:16:54
kind of a spirous implementation detail
1:16:56
and isn't actually deep or interesting
1:16:58
in any way so I'm going to skip the
1:16:59
discussion of it but what opening ey
1:17:01
does here for reasons that I don't fully
1:17:03
understand is that not only have they
1:17:05
this tokenizer which can encode and
1:17:06
decode but they have a whole separate
1:17:08
layer here in addition that is used
1:17:10
serially with the tokenizer and so you
1:17:13
first do um bite encode and then encode
1:17:16
and then you do decode and then bite
1:17:18
decode so that's the loop and they are
1:17:20
just stacked serial on top of each other
1:17:23
and and it's not that interesting so I
1:17:25
won't cover it and you can step through
1:17:26
it if you'd like otherwise this file if
1:17:29
you ignore the bite encoder and the bite
1:17:30
decoder will be algorithmically very
1:17:32
familiar with you and the meat of it
1:17:34
here is the what they call bpe function
1:17:37
and you should recognize this Loop here
1:17:40
which is very similar to our own y Loop
1:17:42
where they're trying to identify the
1:17:44
Byram uh a pair that they should be
1:17:47
merging next and then here just like we
1:17:49
had they have a for Loop trying to merge
1:17:51
this pair uh so they will go over all of
1:17:54
the sequence and they will merge the
1:17:55
pair whenever they find it and they keep
1:17:58
repeating that until they run out of
1:18:00
possible merges in the in the text so
1:18:02
that's the meat of this file and uh
1:18:05
there's an encode and a decode function
1:18:06
just like we have implemented it so long
1:18:08
story short what I want you to take away
1:18:10
at this point is that unfortunately it's
1:18:12
a little bit of a messy code that they
1:18:13
have but algorithmically it is identical
1:18:15
to what we've built up above and what
1:18:18
we've built up above if you understand
1:18:19
it is algorithmically what is necessary
1:18:21
to actually build a BP to organizer
1:18:24
train it and then both encode and decode
1:18:27
the next topic I would like to turn to
1:18:28
is that of special tokens so in addition
1:18:31
to tokens that are coming from you know
1:18:33
raw bytes and the BP merges we can
1:18:35
insert all kinds of tokens that we are
1:18:37
going to use to delimit different parts
1:18:39
of the data or introduced to create a
1:18:41
special structure of the token streams
1:18:45
so in uh if you look at this encoder
1:18:47
object from open AIS gpd2 right here we
1:18:51
mentioned this is very similar to our
1:18:52
vocab you'll notice that the length of
1:18:55
this is
1:18:59
50257 and as I mentioned it's mapping uh
1:19:02
and it's inverted from the mapping of
1:19:03
our vocab our vocab goes from integer to
1:19:06
string and they go the other way around
1:19:08
for no amazing reason um but the thing
1:19:12
to note here is that this the mapping
1:19:13
table here is
1:19:15
50257 where does that number come from
1:19:19
where what are the tokens as I mentioned
1:19:21
there are 256 raw bite token
1:19:24
tokens and then opena actually did
1:19:27
50,000
1:19:29
merges so those become the other tokens
1:19:32
but this would have been
1:19:34
50256 so what is the 57th token and
1:19:38
there is basically one special
1:19:41
token and that one special token you can
1:19:43
see is called end of text so this is a
1:19:47
special token and it's the very last
1:19:50
token and this token is used to delimit
1:19:52
documents ments in the training set so
1:19:56
when we're creating the training data we
1:19:57
have all these documents and we tokenize
1:19:59
them and we get a stream of tokens those
1:20:02
tokens only range from Z to
1:20:05
50256 and then in between those
1:20:07
documents we put special end of text
1:20:10
token and we insert that token in
1:20:13
between documents and we are using this
1:20:16
as a signal to the language model that
1:20:18
the document has ended and what follows
1:20:21
is going to be unrelated to the document
1:20:23
previously that said the language model
1:20:25
has to learn this from data it it needs
1:20:27
to learn that this token usually means
1:20:30
that it should wipe its sort of memory
1:20:32
of what came before and what came before
1:20:34
this token is not actually informative
1:20:36
to what comes next but we are expecting
1:20:38
the language model to just like learn
1:20:39
this but we're giving it the Special
1:20:41
sort of the limiter of these documents
1:20:44
we can go here to Tech tokenizer and um
1:20:47
this the gpt2 tokenizer uh our code that
1:20:49
we've been playing with before so we can
1:20:51
add here right hello world world how are
1:20:54
you and we're getting different tokens
1:20:56
but now you can see what if what happens
1:20:58
if I put end of text you see how until I
1:21:02
finished it these are all different
1:21:04
tokens end of
1:21:06
text still set different tokens and now
1:21:09
when I finish it suddenly we get token
1:21:13
50256 and the reason this works is
1:21:16
because this didn't actually go through
1:21:18
the bpe merges instead the code that
1:21:22
actually outposted tokens has special
1:21:25
case instructions for handling special
1:21:28
tokens um we did not see these special
1:21:31
instructions for handling special tokens
1:21:33
in the encoder dopy it's absent there
1:21:36
but if you go to Tech token Library
1:21:38
which is uh implemented in Rust you will
1:21:41
find all kinds of special case handling
1:21:43
for these special tokens that you can
1:21:45
register uh create adds to the
1:21:47
vocabulary and then it looks for them
1:21:49
and it uh whenever it sees these special
1:21:51
tokens like this it will actually come
1:21:53
in and swap in that special token so
1:21:56
these things are outside of the typical
1:21:58
algorithm of uh B PA en
1:22:01
coding so these special tokens are used
1:22:03
pervasively uh not just in uh basically
1:22:06
base language modeling of predicting the
1:22:07
next token in the sequence but
1:22:09
especially when it gets to later to the
1:22:11
fine tuning stage and all of the chat uh
1:22:13
gbt sort of aspects of it uh because we
1:22:16
don't just want to Del limit documents
1:22:17
we want to delimit entire conversations
1:22:19
between an assistant and a user so if I
1:22:22
refresh this sck tokenizer page the
1:22:24
default example that they have here is
1:22:26
using not sort of base model encoders
1:22:30
but ftuned model uh sort of tokenizers
1:22:34
um so for example using the GPT 3.5
1:22:36
turbo scheme these here are all special
1:22:39
tokens I am start I end Etc uh this is
1:22:43
short for Imaginary mcore start by the
1:22:47
way but you can see here that there's a
1:22:50
sort of start and end of every single
1:22:51
message and there can be many other
1:22:53
other tokens lots of tokens um in use to
1:22:57
delimit these conversations and kind of
1:22:59
keep track of the flow of the messages
1:23:01
here now we can go back to the Tik token
1:23:04
library and here when you scroll to the
1:23:06
bottom they talk about how you can
1:23:08
extend tick token and I can you can
1:23:10
create basically you can Fork uh the um
1:23:14
CL 100K base tokenizers in gp4 and for
1:23:17
example you can extend it by adding more
1:23:19
special tokens and these are totally up
1:23:20
to you you can come up with any
1:23:21
arbitrary tokens and add them with the
1:23:24
new ID afterwards and the tikken library
1:23:27
will uh correctly swap them out uh when
1:23:30
it sees this in the
1:23:32
strings now we can also go back to this
1:23:35
file which we've looked at previously
1:23:37
and I mentioned that the gpt2 in Tik
1:23:40
toen open
1:23:41
I.P we have the vocabulary we have the
1:23:44
pattern for splitting and then here we
1:23:46
are registering the single special token
1:23:48
in gpd2 which was the end of text token
1:23:50
and we saw that it has this ID
1:23:53
in GPT 4 when they defy this here you
1:23:56
see that the pattern has changed as
1:23:58
we've discussed but also the special
1:23:59
tokens have changed in this tokenizer so
1:24:02
we of course have the end of text just
1:24:04
like in gpd2 but we also see three sorry
1:24:07
four additional tokens here Thim prefix
1:24:10
middle and suffix what is fim fim is
1:24:12
short for fill in the middle and if
1:24:15
you'd like to learn more about this idea
1:24:17
it comes from this paper um and I'm not
1:24:20
going to go into detail in this video
1:24:21
it's beyond this video and then there's
1:24:23
one additional uh serve token here so
1:24:27
that's that encoding as well so it's
1:24:30
very common basically to train a
1:24:32
language model and then if you'd like uh
1:24:35
you can add special tokens now when you
1:24:38
add special tokens you of course have to
1:24:40
um do some model surgery to the
1:24:42
Transformer and all the parameters
1:24:43
involved in that Transformer because you
1:24:45
are basically adding an integer and you
1:24:47
want to make sure that for example your
1:24:49
embedding Matrix for the vocabulary
1:24:51
tokens has to be extended by adding a
1:24:53
row and typically this row would be
1:24:55
initialized uh with small random numbers
1:24:57
or something like that because we need
1:24:59
to have a vector that now stands for
1:25:01
that token in addition to that you have
1:25:03
to go to the final layer of the
1:25:04
Transformer and you have to make sure
1:25:06
that that projection at the very end
1:25:08
into the classifier uh is extended by
1:25:10
one as well so basically there's some
1:25:12
model surgery involved that you have to
1:25:13
couple with the tokenization changes if
1:25:17
you are going to add special tokens but
1:25:19
this is a very common operation that
1:25:20
people do especially if they'd like to
1:25:22
fine tune the model for example taking
1:25:24
it from a base model to a chat model
1:25:26
like chat
1:25:28
GPT okay so at this point you should
1:25:30
have everything you need in order to
1:25:31
build your own gp4 tokenizer now in the
1:25:34
process of developing this lecture I've
1:25:35
done that and I published the code under
1:25:37
this repository
1:25:39
MBP so MBP looks like this right now as
1:25:43
I'm recording but uh the MBP repository
1:25:45
will probably change quite a bit because
1:25:47
I intend to continue working on it um in
1:25:50
addition to the MBP repository I've
1:25:52
published the this uh exercise
1:25:53
progression that you can follow so if
1:25:55
you go to exercise. MD here uh this is
1:25:58
sort of me breaking up the task ahead of
1:26:01
you into four steps that sort of uh
1:26:03
build up to what can be a gp4 tokenizer
1:26:07
and so feel free to follow these steps
1:26:08
exactly and follow a little bit of the
1:26:10
guidance that I've laid out here and
1:26:12
anytime you feel stuck just reference
1:26:15
the MBP repository here so either the
1:26:18
tests could be useful or the MBP
1:26:20
repository itself I try to keep the code
1:26:23
fairly clean and understandable and so
1:26:26
um feel free to reference it whenever um
1:26:29
you get
1:26:30
stuck uh in addition to that basically
1:26:33
once you write it you should be able to
1:26:35
reproduce this behavior from Tech token
1:26:37
so getting the gb4 tokenizer you can
1:26:39
take uh you can encode the string and
1:26:41
you should get these tokens and then you
1:26:43
can encode and decode the exact same
1:26:45
string to recover it and in addition to
1:26:47
all that you should be able to implement
1:26:48
your own train function uh which Tik
1:26:51
token Library does not provide it's it's
1:26:52
again only inference code but you could
1:26:55
write your own train MBP does it as well
1:26:58
and that will allow you to train your
1:26:59
own token
1:27:01
vocabularies so here are some of the
1:27:02
code inside M be mean bpe uh shows the
1:27:06
token vocabularies that you might obtain
1:27:09
so on the left uh here we have the GPT 4
1:27:12
merges uh so the first 256 are raw
1:27:16
individual bytes and then here I am
1:27:18
visualizing the merges that gp4
1:27:20
performed during its training so the
1:27:22
very first merge that gp4 did was merge
1:27:25
two spaces into a single token for you
1:27:28
know two spaces and that is a token 256
1:27:31
and so this is the order in which things
1:27:32
merged during gb4 training and this is
1:27:35
the merge order that um we obtain in MBP
1:27:39
by training a tokenizer and in this case
1:27:41
I trained it on a Wikipedia page of
1:27:43
Taylor Swift uh not because I'm a Swifty
1:27:46
but because that is one of the longest
1:27:48
um Wikipedia Pages apparently that's
1:27:50
available but she is pretty cool and
1:27:54
um what was I going to say yeah so you
1:27:57
can compare these two uh vocabularies
1:27:59
and so as an example um here GPT for
1:28:04
merged I in to become in and we've done
1:28:07
the exact same thing on this token 259
1:28:10
here space t becomes space t and that
1:28:13
happened for us a little bit later as
1:28:15
well so the difference here is again to
1:28:17
my understanding only a difference of
1:28:18
the training set so as an example
1:28:20
because I see a lot of white space I
1:28:22
supect that gp4 probably had a lot of
1:28:24
python code in its training set I'm not
1:28:25
sure uh for the
1:28:28
tokenizer and uh here we see much less
1:28:30
of that of course in the Wikipedia page
1:28:33
so roughly speaking they look the same
1:28:35
and they look the same because they're
1:28:36
running the same algorithm and when you
1:28:38
train your own you're probably going to
1:28:39
get something similar depending on what
1:28:41
you train it on okay so we are now going
1:28:43
to move on from tick token and the way
1:28:45
that open AI tokenizes its strings and
1:28:48
we're going to discuss one more very
1:28:49
commonly used library for working with
1:28:51
tokenization inlm
1:28:53
and that is sentence piece so sentence
1:28:55
piece is very commonly used in language
1:28:58
models because unlike Tik token it can
1:29:00
do both training and inference and is
1:29:02
quite efficient at both it supports a
1:29:05
number of algorithms for training uh
1:29:07
vocabularies but one of them is the B
1:29:09
pair en coding algorithm that we've been
1:29:10
looking at so it supports it now
1:29:14
sentence piece is used both by llama and
1:29:16
mistal series and many other models as
1:29:18
well it is on GitHub under Google
1:29:21
sentence piece
1:29:23
and the big difference with sentence
1:29:24
piece and we're going to look at example
1:29:26
because this is kind of hard and subtle
1:29:28
to explain is that they think different
1:29:31
about the order of operations here so in
1:29:35
the case of Tik token we first take our
1:29:39
code points in the string we encode them
1:29:41
using mutf to bytes and then we're
1:29:43
merging bytes it's fairly
1:29:45
straightforward for sentence piece um it
1:29:49
works directly on the level of the code
1:29:50
points themselves so so it looks at
1:29:53
whatever code points are available in
1:29:54
your training set and then it starts
1:29:56
merging those code points and um the bpe
1:30:00
is running on the level of code
1:30:02
points and if you happen to run out of
1:30:04
code points so there are maybe some rare
1:30:07
uh code points that just don't come up
1:30:08
too often and the Rarity is determined
1:30:10
by this character coverage hyper
1:30:11
parameter then these uh code points will
1:30:14
either get mapped to a special unknown
1:30:16
token like ank or if you have the bite
1:30:20
foldback option turned on then that will
1:30:22
take those rare Cod points it will
1:30:24
encode them using utf8 and then the
1:30:26
individual bytes of that encoding will
1:30:28
be translated into tokens and there are
1:30:30
these special bite tokens that basically
1:30:32
get added to the vocabulary so it uses
1:30:36
BP on on the code points and then it
1:30:38
falls back to bytes for rare Cod points
1:30:42
um and so that's kind of like difference
1:30:44
personally I find the Tik token we
1:30:46
significantly cleaner uh but it's kind
1:30:47
of like a subtle but pretty major
1:30:49
difference between the way they approach
1:30:50
tokenization let's work with with a
1:30:52
concrete example because otherwise this
1:30:54
is kind of hard to um to get your head
1:30:57
around so let's work with a concrete
1:30:59
example this is how we can import
1:31:01
sentence piece and then here we're going
1:31:04
to take I think I took like the
1:31:05
description of sentence piece and I just
1:31:07
created like a little toy data set it
1:31:09
really likes to have a file so I created
1:31:10
a toy. txt file with this
1:31:13
content now what's kind of a little bit
1:31:16
crazy about sentence piece is that
1:31:17
there's a ton of options and
1:31:19
configurations and the reason this is so
1:31:21
is because sentence piece has been
1:31:22
around I think for a while and it really
1:31:24
tries to handle a large diversity of
1:31:26
things and um because it's been around I
1:31:28
think it has quite a bit of accumulated
1:31:31
historical baggage uh as well and so in
1:31:34
particular there's like a ton of
1:31:36
configuration arguments this is not even
1:31:37
all of it you can go to here to see all
1:31:40
the training
1:31:41
options um and uh there's also quite
1:31:44
useful documentation when you look at
1:31:46
the raw Proto buff uh that is used to
1:31:49
represent the trainer spec and so on um
1:31:52
many of these options are irrelevant to
1:31:55
us so maybe to point out one example Das
1:31:57
Das shrinking Factor uh this shrinking
1:32:00
factor is not used in the B pair en
1:32:01
coding algorithm so this is just an
1:32:03
argument that is irrelevant to us um it
1:32:06
applies to a different training
1:32:10
algorithm now what I tried to do here is
1:32:12
I tried to set up sentence piece in a
1:32:14
way that is very very similar as far as
1:32:16
I can tell to maybe identical hopefully
1:32:19
to the way that llama 2 was strained so
1:32:22
the way they trained their own um their
1:32:25
own tokenizer and the way I did this was
1:32:27
basically you can take the tokenizer
1:32:29
model file that meta released and you
1:32:31
can um open it using the Proto protuff
1:32:35
uh sort of file that you can generate
1:32:38
and then you can inspect all the options
1:32:40
and I tried to copy over all the options
1:32:41
that looked relevant so here we set up
1:32:44
the input it's raw text in this file
1:32:47
here's going to be the output so it's
1:32:48
going to be for talk 400. model and
1:32:51
vocab
1:32:52
we're saying that we're going to use the
1:32:53
BP algorithm and we want to Bap size of
1:32:56
400 then there's a ton of configurations
1:32:59
here
1:33:01
for um for basically pre-processing and
1:33:05
normalization rules as they're called
1:33:07
normalization used to be very prevalent
1:33:09
I would say before llms in natural
1:33:11
language processing so in machine
1:33:13
translation and uh text classification
1:33:15
and so on you want to normalize and
1:33:17
simplify the text and you want to turn
1:33:18
it all lowercase and you want to remove
1:33:20
all double whites space Etc
1:33:22
and in language models we prefer not to
1:33:24
do any of it or at least that is my
1:33:25
preference as a deep learning person you
1:33:27
want to not touch your data you want to
1:33:29
keep the raw data as much as possible um
1:33:32
in a raw
1:33:33
form so you're basically trying to turn
1:33:35
off a lot of this if you can the other
1:33:38
thing that sentence piece does is that
1:33:40
it has this concept of sentences so
1:33:43
sentence piece it's back it's kind of
1:33:45
like was developed I think early in the
1:33:47
days where there was um an idea that
1:33:50
they you're training a tokenizer on a
1:33:52
bunch of independent sentences so it has
1:33:54
a lot of like how many sentences you're
1:33:56
going to train on what is the maximum
1:33:58
sentence length
1:34:01
um shuffling sentences and so for it
1:34:04
sentences are kind of like the
1:34:05
individual training examples but again
1:34:07
in the context of llms I find that this
1:34:09
is like a very spous and weird
1:34:10
distinction like sentences are just like
1:34:14
don't touch the raw data sentences
1:34:16
happen to exist but in raw data sets
1:34:19
there are a lot of like inet like what
1:34:21
exactly is a sentence what isn't a
1:34:22
sentence um and so I think like it's
1:34:25
really hard to Define what an actual
1:34:26
sentence is if you really like dig into
1:34:29
it and there could be different concepts
1:34:31
of it in different languages or
1:34:32
something like that so why even
1:34:34
introduce the concept it it doesn't
1:34:36
honestly make sense to me I would just
1:34:37
prefer to treat a file as a giant uh
1:34:39
stream of
1:34:40
bytes it has a lot of treatment around
1:34:43
rare word characters and when I say word
1:34:45
I mean code points we're going to come
1:34:46
back to this in a second and it has a
1:34:49
lot of other rules for um basically
1:34:52
splitting digits splitting white space
1:34:54
and numbers and how you deal with that
1:34:57
so these are some kind of like merge
1:34:58
rules so I think this is a little bit
1:35:00
equivalent to tick token using the
1:35:03
regular expression to split up
1:35:05
categories there's like kind of
1:35:07
equivalence of it if you squint T it in
1:35:09
sentence piece where you can also for
1:35:11
example split up split up the digits uh
1:35:14
and uh so
1:35:16
on there's a few more things here that
1:35:18
I'll come back to in a bit and then
1:35:19
there are some special tokens that you
1:35:20
can indicate and it hardcodes the UN
1:35:23
token the beginning of sentence end of
1:35:26
sentence and a pad token um and the UN
1:35:29
token must exist for my understanding
1:35:33
and then some some things so we can
1:35:35
train and when when I press train it's
1:35:37
going to create this file talk 400.
1:35:40
model and talk 400. wab I can then load
1:35:43
the model file and I can inspect the
1:35:46
vocabulary off it and so we trained
1:35:49
vocab size 400 on this text here and
1:35:53
these are the individual pieces the
1:35:55
individual tokens that sentence piece
1:35:57
will create so in the beginning we see
1:35:59
that we have the an token uh with the ID
1:36:02
zero then we have the beginning of
1:36:04
sequence end of sequence one and two and
1:36:08
then we said that the pad ID is negative
1:36:09
1 so we chose not to use it so there's
1:36:12
no pad ID
1:36:13
here then these are individual bite
1:36:17
tokens so here we saw that bite fallback
1:36:20
in llama was turned on so it's true so
1:36:24
what follows are going to be the 256
1:36:26
bite
1:36:27
tokens and these are their
1:36:32
IDs and then at the bottom after the
1:36:35
bite tokens come the
1:36:38
merges and these are the parent nodes in
1:36:41
the merges so we're not seeing the
1:36:42
children we're just seeing the parents
1:36:44
and their
1:36:45
ID and then after the
1:36:47
merges comes eventually the individual
1:36:51
tokens and their IDs and so these are
1:36:54
the individual tokens so these are the
1:36:55
individual code Point tokens if you will
1:36:58
and they come at the end so that is the
1:37:00
ordering with which sentence piece sort
1:37:02
of like represents its vocabularies it
1:37:04
starts with special tokens then the bike
1:37:06
tokens then the merge tokens and then
1:37:08
the individual codo tokens and all these
1:37:12
raw codepoint to tokens are the ones
1:37:14
that it encountered in the training
1:37:16
set so those individual code points are
1:37:20
all the the entire set of code points
1:37:22
that occurred
1:37:24
here so those all get put in there and
1:37:27
then those that are extremely rare as
1:37:29
determined by character coverage so if a
1:37:31
code Point occurred only a single time
1:37:33
out of like a million um sentences or
1:37:35
something like that then it would be
1:37:37
ignored and it would not be added to our
1:37:40
uh
1:37:41
vocabulary once we have a vocabulary we
1:37:43
can encode into IDs and we can um sort
1:37:46
of get a
1:37:47
list and then here I am also decoding
1:37:51
the indiv idual tokens back into little
1:37:54
pieces as they call it so let's take a
1:37:57
look at what happened here hello space
1:38:01
on so these are the token IDs we got
1:38:05
back and when we look here uh a few
1:38:07
things sort of uh jump to mind number
1:38:12
one take a look at these characters the
1:38:14
Korean characters of course were not
1:38:16
part of the training set so sentence
1:38:18
piece is encountering code points that
1:38:20
it has not seen during training time and
1:38:22
those code points do not have a token
1:38:25
associated with them so suddenly these
1:38:26
are un tokens unknown tokens but because
1:38:31
bite fall back as true instead sentence
1:38:34
piece falls back to bytes and so it
1:38:36
takes this it encodes it with utf8 and
1:38:40
then it uses these tokens to represent
1:38:43
uh those bytes and that's what we are
1:38:46
getting sort of here this is the utf8 uh
1:38:50
encoding and in this shifted by three uh
1:38:53
because of these um special tokens here
1:38:56
that have IDs earlier on so that's what
1:38:59
happened here now one more thing that um
1:39:03
well first before I go on with respect
1:39:06
to the bitef back let me remove bite
1:39:08
foldback if this is false what's going
1:39:11
to happen let's
1:39:13
retrain so the first thing that happened
1:39:14
is all the bite tokens disappeared right
1:39:17
and now we just have the merges and we
1:39:19
have a lot more merges now because we
1:39:20
have a lot more space because we're not
1:39:22
taking up space in the wab size uh with
1:39:25
all the
1:39:26
bytes and now if we encode
1:39:29
this we get a zero so this entire string
1:39:33
here suddenly there's no bitef back so
1:39:35
this is unknown and unknown is an and so
1:39:39
this is zero because the an token is
1:39:42
token zero and you have to keep in mind
1:39:45
that this would feed into your uh
1:39:47
language model so what is a language
1:39:48
model supposed to do when all kinds of
1:39:50
different things that are unrecognized
1:39:52
because they're rare just end up mapping
1:39:54
into Unk it's not exactly the property
1:39:56
that you want so that's why I think
1:39:58
llama correctly uh used by fallback true
1:40:02
uh because we definitely want to feed
1:40:04
these um unknown or rare code points
1:40:06
into the model and some uh some manner
1:40:09
the next thing I want to show you is the
1:40:11
following notice here when we are
1:40:12
decoding all the individual tokens you
1:40:15
see how spaces uh space here ends up
1:40:18
being this um bold underline I'm not
1:40:21
100% sure by the way why sentence piece
1:40:23
switches whites space into these bold
1:40:25
underscore characters maybe it's for
1:40:28
visualization I'm not 100% sure why that
1:40:30
happens uh but notice this why do we
1:40:32
have an extra space in the front of
1:40:37
hello um what where is this coming from
1:40:40
well it's coming from this option
1:40:43
here
1:40:45
um add dummy prefix is true and when you
1:40:48
go to the
1:40:50
documentation add D whites space at the
1:40:52
beginning of text in order to treat
1:40:53
World in world and hello world in the
1:40:56
exact same way so what this is trying to
1:40:58
do is the
1:40:59
following if we go back to our tick
1:41:02
tokenizer world as uh token by itself
1:41:06
has a different ID than space world so
1:41:10
we have this is 1917 but this is 14 Etc
1:41:15
so these are two different tokens for
1:41:16
the language model and the language
1:41:17
model has to learn from data that they
1:41:19
are actually kind of like a very similar
1:41:20
concept so to the language model in the
1:41:23
Tik token World um basically words in
1:41:26
the beginning of sentences and words in
1:41:28
the middle of sentences actually look
1:41:29
completely different um and it has to
1:41:32
learned that they are roughly the same
1:41:34
so this add dami prefix is trying to
1:41:37
fight that a little bit and the way that
1:41:39
works is that it basically
1:41:42
uh adds a dummy prefix so for as a as a
1:41:47
part of pre-processing it will take the
1:41:49
string and it will add a space it will
1:41:51
do this and that's done in an effort to
1:41:55
make this world and that world the same
1:41:58
they will both be space world so that's
1:42:00
one other kind of pre-processing option
1:42:02
that is turned on and llama 2 also uh
1:42:05
uses this option and that's I think
1:42:07
everything that I want to say for my
1:42:09
preview of sentence piece and how it is
1:42:10
different um maybe here what I've done
1:42:13
is I just uh put in the Raw protocol
1:42:17
buffer representation basically of the
1:42:20
tokenizer the too trained so feel free
1:42:23
to sort of Step through this and if you
1:42:25
would like uh your tokenization to look
1:42:27
identical to that of the meta uh llama 2
1:42:30
then you would be copy pasting these
1:42:32
settings as I tried to do up above and
1:42:35
uh yeah that's I think that's it for
1:42:37
this section I think my summary for
1:42:39
sentence piece from all of this is
1:42:41
number one I think that there's a lot of
1:42:42
historical baggage in sentence piece a
1:42:44
lot of Concepts that I think are
1:42:46
slightly confusing and I think
1:42:47
potentially um contain foot guns like
1:42:49
this concept of a sentence and it's
1:42:51
maximum length and stuff like that um
1:42:54
otherwise it is fairly commonly used in
1:42:56
the industry um because it is efficient
1:42:59
and can do both training and inference
1:43:01
uh it has a few quirks like for example
1:43:03
un token must exist and the way the bite
1:43:05
fallbacks are done and so on I don't
1:43:07
find particularly elegant and
1:43:08
unfortunately I have to say it's not
1:43:10
very well documented so it took me a lot
1:43:11
of time working with this myself um and
1:43:15
just visualizing things and trying to
1:43:16
really understand what is happening here
1:43:18
because uh the documentation
1:43:19
unfortunately is in my opion not not
1:43:21
super amazing but it is a very nice repo
1:43:25
that is available to you if you'd like
1:43:26
to train your own tokenizer right now
1:43:28
okay let me now switch gears again as
1:43:30
we're starting to slowly wrap up here I
1:43:32
want to revisit this issue in a bit more
1:43:33
detail of how we should set the vocap
1:43:35
size and what are some of the
1:43:36
considerations around it so for this I'd
1:43:40
like to go back to the model
1:43:41
architecture that we developed in the
1:43:42
last video when we built the GPT from
1:43:45
scratch so this here was uh the file
1:43:47
that we built in the previous video and
1:43:49
we defined the Transformer model and and
1:43:51
let's specifically look at Bap size and
1:43:53
where it appears in this file so here we
1:43:55
Define the voap size uh at this time it
1:43:58
was 65 or something like that extremely
1:44:00
small number so this will grow much
1:44:02
larger you'll see that Bap size doesn't
1:44:04
come up too much in most of these layers
1:44:06
the only place that it comes up to is in
1:44:09
exactly these two places here so when we
1:44:11
Define the language model there's the
1:44:14
token embedding table which is this
1:44:16
two-dimensional array where the vocap
1:44:18
size is basically the number of rows and
1:44:21
uh each vocabulary element each token
1:44:24
has a vector that we're going to train
1:44:26
using back propagation that Vector is of
1:44:28
size and embed which is number of
1:44:29
channels in the Transformer and
1:44:32
basically as voap size increases this
1:44:34
embedding table as I mentioned earlier
1:44:36
is going to also grow we're going to be
1:44:37
adding rows in addition to that at the
1:44:40
end of the Transformer there's this LM
1:44:42
head layer which is a linear layer and
1:44:44
you'll notice that that layer is used at
1:44:46
the very end to produce the logits uh
1:44:49
which become the probabilities for the
1:44:50
next token in sequence and so
1:44:52
intuitively we're trying to produce a
1:44:54
probability for every single token that
1:44:56
might come next at every point in time
1:44:59
of that Transformer and if we have more
1:45:01
and more tokens we need to produce more
1:45:03
and more probabilities so every single
1:45:05
token is going to introduce an
1:45:06
additional dot product that we have to
1:45:08
do here in this linear layer for this
1:45:10
final layer in a
1:45:11
Transformer so why can't vocap size be
1:45:15
infinite why can't we grow to Infinity
1:45:17
well number one your token embedding
1:45:18
table is going to grow uh your linear
1:45:22
layer is going to grow so we're going to
1:45:24
be doing a lot more computation here
1:45:25
because this LM head layer will become
1:45:27
more computational expensive number two
1:45:29
because we have more parameters we could
1:45:31
be worried that we are going to be under
1:45:33
trining some of these
1:45:35
parameters so intuitively if you have a
1:45:37
very large vocabulary size say we have a
1:45:39
million uh tokens then every one of
1:45:41
these tokens is going to come up more
1:45:43
and more rarely in the training data
1:45:45
because there's a lot more other tokens
1:45:47
all over the place and so we're going to
1:45:49
be seeing fewer and fewer examples uh
1:45:51
for each individual token and you might
1:45:53
be worried that basically the vectors
1:45:55
associated with every token will be
1:45:56
undertrained as a result because they
1:45:58
just don't come up too often and they
1:46:00
don't participate in the forward
1:46:01
backward pass in addition to that as
1:46:03
your vocab size grows you're going to
1:46:05
start shrinking your sequences a lot
1:46:07
right and that's really nice because
1:46:09
that means that we're going to be
1:46:10
attending to more and more text so
1:46:12
that's nice but also you might be
1:46:14
worrying that two large of chunks are
1:46:16
being squished into single tokens and so
1:46:19
the model just doesn't have as much of
1:46:21
time to think per sort of um some number
1:46:25
of characters in the text or you can
1:46:27
think about it that way right so
1:46:28
basically we're squishing too much
1:46:29
information into a single token and then
1:46:32
the forward pass of the Transformer is
1:46:33
not enough to actually process that
1:46:34
information appropriately and so these
1:46:36
are some of the considerations you're
1:46:37
thinking about when you're designing the
1:46:39
vocab size as I mentioned this is mostly
1:46:41
an empirical hyperparameter and it seems
1:46:43
like in state-of-the-art architectures
1:46:44
today this is usually in the high 10,000
1:46:47
or somewhere around 100,000 today and
1:46:49
the next consideration I want to briefly
1:46:51
talk about is what if we want to take a
1:46:53
pre-trained model and we want to extend
1:46:55
the vocap size and this is done fairly
1:46:57
commonly actually so for example when
1:46:59
you're doing fine-tuning for cha GPT um
1:47:02
a lot more new special tokens get
1:47:04
introduced on top of the base model to
1:47:06
maintain the metadata and all the
1:47:08
structure of conversation objects
1:47:10
between a user and an assistant so that
1:47:12
takes a lot of special tokens you might
1:47:14
also try to throw in more special tokens
1:47:16
for example for using the browser or any
1:47:18
other tool and so it's very tempting to
1:47:21
add a lot of tokens for all kinds of
1:47:22
special functionality so if you want to
1:47:25
be adding a token that's totally
1:47:26
possible Right all we have to do is we
1:47:28
have to resize this embedding so we have
1:47:30
to add rows we would initialize these uh
1:47:32
parameters from scratch to be small
1:47:34
random numbers and then we have to
1:47:36
extend the weight inside this linear uh
1:47:39
so we have to start making dot products
1:47:41
um with the associated parameters as
1:47:43
well to basically calculate the
1:47:45
probabilities for these new tokens so
1:47:47
both of these are just a resizing
1:47:49
operation it's a very mild
1:47:51
model surgery and can be done fairly
1:47:53
easily and it's quite common that
1:47:54
basically you would freeze the base
1:47:55
model you introduce these new parameters
1:47:57
and then you only train these new
1:47:59
parameters to introduce new tokens into
1:48:01
the architecture um and so you can
1:48:03
freeze arbitrary parts of it or you can
1:48:05
train arbitrary parts of it and that's
1:48:06
totally up to you but basically minor
1:48:08
surgery required if you'd like to
1:48:10
introduce new tokens and finally I'd
1:48:12
like to mention that actually there's an
1:48:13
entire design space of applications in
1:48:16
terms of introducing new tokens into a
1:48:18
vocabulary that go Way Beyond just
1:48:19
adding special tokens and special new
1:48:21
functionality so just to give you a
1:48:23
sense of the design space but this could
1:48:24
be an entire video just by itself uh
1:48:27
this is a paper on learning to compress
1:48:29
prompts with what they called uh gist
1:48:31
tokens and the rough idea is suppose
1:48:33
that you're using language models in a
1:48:35
setting that requires very long prompts
1:48:37
while these long prompts just slow
1:48:39
everything down because you have to
1:48:40
encode them and then you have to use
1:48:41
them and then you're tending over them
1:48:43
and it's just um you know heavy to have
1:48:45
very large prompts so instead what they
1:48:48
do here in this paper is they introduce
1:48:51
new tokens and um imagine basically
1:48:55
having a few new tokens you put them in
1:48:56
a sequence and then you train the model
1:48:59
by distillation so you are keeping the
1:49:02
entire model Frozen and you're only
1:49:03
training the representations of the new
1:49:05
tokens their embeddings and you're
1:49:07
optimizing over the new tokens such that
1:49:09
the behavior of the language model is
1:49:12
identical uh to the model that has a
1:49:15
very long prompt that works for you and
1:49:18
so it's a compression technique of
1:49:19
compressing that very long prompt into
1:49:21
those few new gist tokens and so you can
1:49:24
train this and then at test time you can
1:49:25
discard your old prompt and just swap in
1:49:27
those tokens and they sort of like uh
1:49:29
stand in for that very long prompt and
1:49:31
have an almost identical performance and
1:49:34
so this is one um technique and a class
1:49:36
of parameter efficient fine-tuning
1:49:38
techniques where most of the model is
1:49:40
basically fixed and there's no training
1:49:42
of the model weights there's no training
1:49:44
of Laura or anything like that of new
1:49:45
parameters the the parameters that
1:49:47
you're training are now just the uh
1:49:49
token embeddings so that's just one
1:49:51
example but this could again be like an
1:49:53
entire video but just to give you a
1:49:55
sense that there's a whole design space
1:49:56
here that is potentially worth exploring
1:49:57
in the future the next thing I want to
1:49:59
briefly address is that I think recently
1:50:01
there's a lot of momentum in how you
1:50:03
actually could construct Transformers
1:50:05
that can simultaneously process not just
1:50:07
text as the input modality but a lot of
1:50:09
other modalities so be it images videos
1:50:12
audio Etc and how do you feed in all
1:50:14
these modalities and potentially predict
1:50:16
these modalities from a Transformer uh
1:50:19
do you have to change the architecture
1:50:20
in some fundamental way and I think what
1:50:22
a lot of people are starting to converge
1:50:23
towards is that you're not changing the
1:50:24
architecture you stick with the
1:50:25
Transformer you just kind of tokenize
1:50:28
your input domains and then call the day
1:50:30
and pretend it's just text tokens and
1:50:32
just do everything else identical in an
1:50:34
identical manner so here for example
1:50:36
there was a early paper that has nice
1:50:38
graphic for how you can take an image
1:50:40
and you can chunc at it into
1:50:42
integers um and these sometimes uh so
1:50:45
these will basically become the tokens
1:50:47
of images as an example and uh these
1:50:50
tokens can be uh hard tokens where you
1:50:52
force them to be integers they can also
1:50:54
be soft tokens where you uh sort of
1:50:57
don't require uh these to be discrete
1:51:00
but you do Force these representations
1:51:02
to go through bottlenecks like in Auto
1:51:05
encoders uh also in this paper that came
1:51:07
out from open a SORA which I think
1:51:09
really um uh blew the mind of many
1:51:12
people and inspired a lot of people in
1:51:14
terms of what's possible they have a
1:51:15
Graphic here and they talk briefly about
1:51:17
how llms have text tokens Sora has
1:51:20
visual patches so again they came up
1:51:23
with a way to chunc a videos into
1:51:25
basically tokens when they own
1:51:27
vocabularies and then you can either
1:51:29
process discrete tokens say with autog
1:51:30
regressive models or even soft tokens
1:51:32
with diffusion models and uh all of that
1:51:35
is sort of uh being actively worked on
1:51:38
designed on and is beyond the scope of
1:51:40
this video but just something I wanted
1:51:41
to mention briefly okay now that we have
1:51:43
come quite deep into the tokenization
1:51:45
algorithm and we understand a lot more
1:51:47
about how it works let's loop back
1:51:49
around to the beginning of this video
1:51:51
and go through some of these bullet
1:51:52
points and really see why they happen so
1:51:55
first of all why can't my llm spell
1:51:57
words very well or do other spell
1:51:59
related
1:52:01
tasks so fundamentally this is because
1:52:03
as we saw these characters are chunked
1:52:06
up into tokens and some of these tokens
1:52:08
are actually fairly long so as an
1:52:10
example I went to the gp4 vocabulary and
1:52:13
I looked at uh one of the longer tokens
1:52:15
so that default style turns out to be a
1:52:18
single individual token so that's a lot
1:52:20
of characters for a single token so my
1:52:22
suspicion is that there's just too much
1:52:24
crammed into this single token and my
1:52:26
suspicion was that the model should not
1:52:28
be very good at tasks related to
1:52:30
spelling of this uh single token so I
1:52:35
asked how many letters L are there in
1:52:37
the word default style and of course my
1:52:41
prompt is intentionally done that way
1:52:44
and you see how default style will be a
1:52:46
single token so this is what the model
1:52:47
sees so my suspicion is that it wouldn't
1:52:49
be very good at this and indeed it is
1:52:51
not it doesn't actually know how many
1:52:53
L's are in there it thinks there are
1:52:55
three and actually there are four if I'm
1:52:57
not getting this wrong myself so that
1:53:00
didn't go extremely well let's look look
1:53:02
at another kind of uh character level
1:53:05
task so for example here I asked uh gp4
1:53:08
to reverse the string default style and
1:53:11
they tried to use a code interpreter and
1:53:13
I stopped it and I said just do it just
1:53:15
try it and uh it gave me jumble so it
1:53:20
doesn't actually really know how to
1:53:21
reverse this string going from right to
1:53:24
left uh so it gave a wrong result so
1:53:27
again like working with this working
1:53:28
hypothesis that maybe this is due to the
1:53:30
tokenization I tried a different
1:53:32
approach I said okay let's reverse the
1:53:34
exact same string but take the following
1:53:36
approach step one just print out every
1:53:39
single character separated by spaces and
1:53:41
then as a step two reverse that list and
1:53:43
it again Tred to use a tool but when I
1:53:45
stopped it it uh first uh produced all
1:53:48
the characters and that was actually
1:53:49
correct and then It reversed them and
1:53:51
that was correct once it had this so
1:53:53
somehow it can't reverse it directly but
1:53:55
when you go just first uh you know
1:53:57
listing it out in order it can do that
1:53:59
somehow and then it can once it's uh
1:54:02
broken up this way this becomes all
1:54:04
these individual characters and so now
1:54:06
this is much easier for it to see these
1:54:08
individual tokens and reverse them and
1:54:10
print them out so that is kind of
1:54:14
interesting so let's continue now why
1:54:17
are llms worse at uh non-english langu
1:54:20
and I briefly covered this already but
1:54:23
basically um it's not only that the
1:54:25
language model sees less non-english
1:54:27
data during training of the model
1:54:29
parameters but also the tokenizer is not
1:54:32
um is not sufficiently trained on
1:54:35
non-english data and so here for example
1:54:37
hello how are you is five tokens and its
1:54:41
translation is 15 tokens so this is a
1:54:43
three times blow up and so for example
1:54:46
anang is uh just hello basically in
1:54:49
Korean and that end up being three
1:54:50
tokens I'm actually kind of surprised by
1:54:52
that because that is a very common
1:54:53
phrase there just the typical greeting
1:54:55
of like hello and that ends up being
1:54:57
three tokens whereas our hello is a
1:54:59
single token and so basically everything
1:55:01
is a lot more bloated and diffuse and
1:55:02
this is I think partly the reason that
1:55:04
the model Works worse on other
1:55:07
languages uh coming back why is LM bad
1:55:10
at simple arithmetic um that has to do
1:55:13
with the tokenization of numbers and so
1:55:17
um you'll notice that for example
1:55:19
addition is very sort of
1:55:21
like uh there's an algorithm that is
1:55:23
like character level for doing addition
1:55:26
so for example here we would first add
1:55:28
the ones and then the tens and then the
1:55:29
hundreds you have to refer to specific
1:55:31
parts of these digits but uh these
1:55:35
numbers are represented completely
1:55:36
arbitrarily based on whatever happened
1:55:38
to merge or not merge during the
1:55:39
tokenization process there's an entire
1:55:41
blog post about this that I think is
1:55:43
quite good integer tokenization is
1:55:45
insane and this person basically
1:55:47
systematically explores the tokenization
1:55:49
of numbers in I believe this is gpt2 and
1:55:52
so they notice that for example for the
1:55:54
for um four-digit numbers you can take a
1:55:57
look at whether it is uh a single token
1:56:00
or whether it is two tokens that is a 1
1:56:02
three or a 2 two or a 31 combination and
1:56:05
so all the different numbers are all the
1:56:07
different combinations and you can
1:56:08
imagine this is all completely
1:56:09
arbitrarily so and the model
1:56:11
unfortunately sometimes sees uh four um
1:56:14
a token for for all four digits
1:56:17
sometimes for three sometimes for two
1:56:18
sometimes for one and it's in an
1:56:20
arbitrary uh Manner and so this is
1:56:23
definitely a headwind if you will for
1:56:25
the language model and it's kind of
1:56:26
incredible that it can kind of do it and
1:56:28
deal with it but it's also kind of not
1:56:30
ideal and so that's why for example we
1:56:32
saw that meta when they train the Llama
1:56:34
2 algorithm and they use sentence piece
1:56:36
they make sure to split up all the um
1:56:40
all the digits as an example for uh
1:56:42
llama 2 and this is partly to improve a
1:56:45
simple arithmetic kind of
1:56:47
performance and finally why is gpt2 not
1:56:51
as good in Python again this is partly a
1:56:53
modeling issue on in the architecture
1:56:55
and the data set and the strength of the
1:56:57
model but it's also partially
1:56:58
tokenization because as we saw here with
1:57:00
the simple python example the encoding
1:57:03
efficiency of the tokenizer for handling
1:57:05
spaces in Python is terrible and every
1:57:07
single space is an individual token and
1:57:09
this dramatically reduces the context
1:57:11
length that the model can attend to
1:57:13
cross so that's almost like a
1:57:14
tokenization bug for gpd2 and that was
1:57:17
later fixed with gp4 okay so here's
1:57:20
another fun one my llm abruptly halts
1:57:23
when it sees the string end of text so
1:57:25
here's um here's a very strange Behavior
1:57:28
print a string end of text is what I
1:57:30
told jt4 and it says could you please
1:57:32
specify the string and I'm I'm telling
1:57:35
it give me end of text and it seems like
1:57:37
there's an issue it's not seeing end of
1:57:39
text and then I give it end of text is
1:57:42
the string and then here's a string and
1:57:44
then it just doesn't print it so
1:57:46
obviously something is breaking here
1:57:47
with respect to the handling of the
1:57:48
special token and I don't actually know
1:57:50
what open ey is doing under the hood
1:57:53
here and whether they are potentially
1:57:55
parsing this as an um as an actual token
1:57:59
instead of this just being uh end of
1:58:01
text um as like individual sort of
1:58:05
pieces of it without the special token
1:58:06
handling logic and so it might be that
1:58:10
someone when they're calling do encode
1:58:12
uh they are passing in the allowed
1:58:13
special and they are allowing end of
1:58:16
text as a special character in the user
1:58:18
prompt but the user prompt of course is
1:58:21
is a sort of um attacker controlled text
1:58:24
so you would hope that they don't really
1:58:25
parse or use special tokens or you know
1:58:29
from that kind of input but it appears
1:58:31
that there's something definitely going
1:58:32
wrong here and um so your knowledge of
1:58:35
these special tokens ends up being in a
1:58:36
tax surface potentially and so if you'd
1:58:39
like to confuse llms then just um try to
1:58:43
give them some special tokens and see if
1:58:44
you're breaking something by chance okay
1:58:46
so this next one is a really fun one uh
1:58:49
the trailing whites space issue so if
1:58:53
you come to playground and uh we come
1:58:56
here to GPT 3.5 turbo instruct so this
1:58:58
is not a chat model this is a completion
1:59:00
model so think of it more like it's a
1:59:03
lot more closer to a base model it does
1:59:05
completion it will continue the token
1:59:08
sequence so here's a tagline for ice
1:59:10
cream shop and we want to continue the
1:59:12
sequence and so we can submit and get a
1:59:14
bunch of tokens okay no problem but now
1:59:18
suppose I do this but instead of
1:59:21
pressing submit here I do here's a
1:59:23
tagline for ice cream shop space so I
1:59:26
have a space here before I click
1:59:29
submit we get a warning your text ends
1:59:32
in a trail Ling space which causes worse
1:59:33
performance due to how API splits text
1:59:36
into tokens so what's happening here it
1:59:38
still gave us a uh sort of completion
1:59:41
here but let's take a look at what's
1:59:43
happening so here's a tagline for an ice
1:59:45
cream shop and then what does this look
1:59:49
like in the actual actual training data
1:59:50
suppose you found the completion in the
1:59:52
training document somewhere on the
1:59:54
internet and the llm trained on this
1:59:56
data so maybe it's something like oh
1:59:58
yeah maybe that's the tagline that's a
2:00:00
terrible tagline but notice here that
2:00:03
when I create o you see that because
2:00:06
there's the the space character is
2:00:08
always a prefix to these tokens in GPT
2:00:11
so it's not an O token it's a space o
2:00:13
token the space is part of the O and
2:00:17
together they are token 8840 that's
2:00:19
that's space o so what's What's
2:00:22
Happening Here is that when I just have
2:00:24
it like this and I let it complete the
2:00:27
next token it can sample the space o
2:00:30
token but instead if I have this and I
2:00:33
add my space then what I'm doing here
2:00:35
when I incode this string is I have
2:00:38
basically here's a t line for an ice
2:00:39
cream uh shop and this space at the very
2:00:42
end becomes a token
2:00:44
220 and so we've added token 220 and
2:00:48
this token otherwise would be part of
2:00:50
the tagline because if there actually is
2:00:52
a tagline here so space o is the token
2:00:55
and so this is suddenly a of
2:00:57
distribution for the model because this
2:01:00
space is part of the next token but
2:01:02
we're putting it here like this and the
2:01:04
model has seen very very little data of
2:01:07
actual Space by itself and we're asking
2:01:10
it to complete the sequence like add in
2:01:12
more tokens but the problem is that
2:01:13
we've sort of begun the first token and
2:01:16
now it's been split up and now we're out
2:01:19
of this distribution and now arbitrary
2:01:21
bad things happen and it's just a very
2:01:23
rare example for it to see something
2:01:25
like that and uh that's why we get the
2:01:27
warning so the fundamental issue here is
2:01:29
of course that um the llm is on top of
2:01:32
these tokens and these tokens are text
2:01:35
chunks they're not characters in a way
2:01:37
you and I would think of them they are
2:01:38
these are the atoms of what the LM is
2:01:40
seeing and there's a bunch of weird
2:01:42
stuff that comes out of it let's go back
2:01:44
to our default cell style I bet you that
2:01:48
the model has never in its training set
2:01:50
seen default cell sta without Le in
2:01:54
there it's always seen this as a single
2:01:57
group because uh this is some kind of a
2:01:59
function in um I'm guess I don't
2:02:02
actually know what this is part of this
2:02:03
is some kind of API but I bet you that
2:02:05
it's never seen this combination of
2:02:07
tokens uh in its training data because
2:02:11
or I think it would be extremely rare so
2:02:12
I took this and I copy pasted it here
2:02:15
and I had I tried to complete from it
2:02:17
and the it immediately gave me a big
2:02:19
error and it said the model predicted to
2:02:21
completion that begins with a stop
2:02:22
sequence resulting in no output consider
2:02:24
adjusting your prompt or stop sequences
2:02:26
so what happened here when I clicked
2:02:28
submit is that immediately the model
2:02:30
emitted and sort of like end of text
2:02:32
token I think or something like that it
2:02:34
basically predicted the stop sequence
2:02:36
immediately so it had no completion and
2:02:39
so this is why I'm getting a warning
2:02:40
again because we're off the data
2:02:42
distribution and the model is just uh
2:02:45
predicting just totally arbitrary things
2:02:48
it's just really confused basically this
2:02:49
is uh this is giving it brain damage
2:02:51
it's never seen this before it's shocked
2:02:53
and it's predicting end of text or
2:02:55
something I tried it again here and it
2:02:57
in this case it completed it but then
2:02:59
for some reason this request May violate
2:03:01
our usage policies this was
2:03:04
flagged um basically something just like
2:03:07
goes wrong and there's something like
2:03:08
Jank you can just feel the Jank because
2:03:10
the model is like extremely unhappy with
2:03:11
just this and it doesn't know how to
2:03:13
complete it because it's never occurred
2:03:14
in training set in a training set it
2:03:16
always appears like this and becomes a
2:03:18
single token
2:03:20
so these kinds of issues where tokens
2:03:22
are either you sort of like complete the
2:03:24
first character of the next token or you
2:03:27
are sort of you have long tokens that
2:03:29
you then have just some of the
2:03:30
characters off all of these are kind of
2:03:32
like issues with partial tokens is how I
2:03:35
would describe it and if you actually
2:03:38
dig into the T token
2:03:40
repository go to the rust code and
2:03:42
search for
2:03:44
unstable and you'll see um en code
2:03:47
unstable native unstable token tokens
2:03:49
and a lot of like special case handling
2:03:52
none of this stuff about unstable tokens
2:03:53
is documented anywhere but there's a ton
2:03:55
of code dealing with unstable tokens and
2:03:58
unstable tokens is exactly kind of like
2:04:01
what I'm describing here what you would
2:04:03
like out of a completion API is
2:04:05
something a lot more fancy like if we're
2:04:07
putting in default cell sta if we're
2:04:09
asking for the next token sequence we're
2:04:11
not actually trying to append the next
2:04:12
token exactly after this list we're
2:04:15
actually trying to append we're trying
2:04:16
to consider lots of tokens um
2:04:20
that if we were or I guess like we're
2:04:22
trying to search over characters that if
2:04:26
we retened would be of high probability
2:04:28
if that makes sense um so that we can
2:04:31
actually add a single individual
2:04:32
character uh instead of just like adding
2:04:34
the next full token that comes after
2:04:37
this partial token list so I this is
2:04:39
very tricky to describe and I invite you
2:04:41
to maybe like look through this it ends
2:04:43
up being extremely gnarly and hairy kind
2:04:45
of topic it and it comes from
2:04:46
tokenization fundamentally so um maybe I
2:04:49
can even spend an entire video talking
2:04:51
about unstable tokens sometime in the
2:04:52
future okay and I'm really saving the
2:04:54
best for last my favorite one by far is
2:04:57
the solid gold
2:04:59
Magikarp and it just okay so this comes
2:05:01
from this blog post uh solid gold
2:05:04
Magikarp and uh this is um internet
2:05:07
famous now for those of us in llms and
2:05:10
basically I I would advise you to uh
2:05:12
read this block Post in full but
2:05:14
basically what this person was doing is
2:05:17
this person went to the um
2:05:19
token embedding stable and clustered the
2:05:22
tokens based on their embedding
2:05:25
representation and this person noticed
2:05:27
that there's a cluster of tokens that
2:05:29
look really strange so there's a cluster
2:05:31
here at rot e stream Fame solid gold
2:05:34
Magikarp Signet message like really
2:05:36
weird tokens in uh basically in this
2:05:40
embedding cluster and so what are these
2:05:42
tokens and where do they even come from
2:05:44
like what is solid gold magikarpet makes
2:05:45
no sense and then they found bunch of
2:05:49
these
2:05:50
tokens and then they notice that
2:05:52
actually the plot thickens here because
2:05:54
if you ask the model about these tokens
2:05:56
like you ask it uh some very benign
2:05:59
question like please can you repeat back
2:06:00
to me the string sold gold Magikarp uh
2:06:03
then you get a variety of basically
2:06:05
totally broken llm Behavior so either
2:06:08
you get evasion so I'm sorry I can't
2:06:10
hear you or you get a bunch of
2:06:11
hallucinations as a response um you can
2:06:15
even get back like insults so you ask it
2:06:17
uh about streamer bot it uh tells the
2:06:20
and the model actually just calls you
2:06:22
names uh or it kind of comes up with
2:06:24
like weird humor like you're actually
2:06:26
breaking the model by asking about these
2:06:28
very simple strings like at Roth and
2:06:31
sold gold Magikarp so like what the hell
2:06:33
is happening and there's a variety of
2:06:34
here documented behaviors uh there's a
2:06:37
bunch of tokens not just so good
2:06:38
Magikarp that have that kind of a
2:06:40
behavior and so basically there's a
2:06:42
bunch of like trigger words and if you
2:06:44
ask the model about these trigger words
2:06:46
or you just include them in your prompt
2:06:48
the model goes haywire and has all kinds
2:06:50
of uh really Strange Behaviors including
2:06:53
sort of ones that violate typical safety
2:06:55
guidelines uh and the alignment of the
2:06:57
model like it's swearing back at you so
2:07:00
what is happening here and how can this
2:07:02
possibly be true well this again comes
2:07:05
down to tokenization so what's happening
2:07:07
here is that sold gold Magikarp if you
2:07:09
actually dig into it is a Reddit user so
2:07:12
there's a u Sol gold
2:07:14
Magikarp and probably what happened here
2:07:17
even though I I don't know that this has
2:07:18
been like really definitively explored
2:07:20
but what is thought to have happened is
2:07:23
that the tokenization data set was very
2:07:26
different from the training data set for
2:07:28
the actual language model so in the
2:07:30
tokenization data set there was a ton of
2:07:32
redded data potentially where the user
2:07:35
solid gold Magikarp was mentioned in the
2:07:36
text because solid gold Magikarp was a
2:07:39
very common um sort of uh person who
2:07:42
would post a lot uh this would be a
2:07:44
string that occurs many times in a
2:07:45
tokenization data set because it occurs
2:07:48
many times in a tokenization data set
2:07:50
these tokens would end up getting merged
2:07:51
to the single individual token for that
2:07:54
single Reddit user sold gold Magikarp so
2:07:56
they would have a dedicated token in a
2:07:58
vocabulary of was it 50,000 tokens in
2:08:01
gpd2 that is devoted to that Reddit user
2:08:04
and then what happens is the
2:08:06
tokenization data set has those strings
2:08:09
but then later when you train the model
2:08:11
the language model itself um this data
2:08:14
from Reddit was not present and so
2:08:17
therefore in the entire training set for
2:08:19
the language model sold gold Magikarp
2:08:21
never occurs that token never appears in
2:08:24
the training set for the actual language
2:08:26
model later so this token never gets
2:08:29
activated it's initialized at random in
2:08:31
the beginning of optimization then you
2:08:33
have forward backward passes and updates
2:08:34
to the model and this token is just
2:08:36
never updated in the embedding table
2:08:38
that row Vector never gets sampled it
2:08:40
never gets used so it never gets trained
2:08:42
and it's completely untrained it's kind
2:08:44
of like unallocated memory in a typical
2:08:46
binary program written in C or something
2:08:48
like that that so it's unallocated
2:08:50
memory and then at test time if you
2:08:52
evoke this token then you're basically
2:08:54
plucking out a row of the embedding
2:08:56
table that is completely untrained and
2:08:57
that feeds into a Transformer and
2:08:59
creates undefined behavior and that's
2:09:01
what we're seeing here this completely
2:09:02
undefined never before seen in a
2:09:04
training behavior and so any of these
2:09:07
kind of like weird tokens would evoke
2:09:08
this Behavior because fundamentally the
2:09:09
model is um is uh uh out of sample out
2:09:14
of distribution okay and the very last
2:09:17
thing I wanted to just briefly mention
2:09:19
point out although I think a lot of
2:09:20
people are quite aware of this is that
2:09:22
different kinds of formats and different
2:09:23
representations and different languages
2:09:25
and so on might be more or less
2:09:27
efficient with GPD tokenizers uh or any
2:09:30
tokenizers for any other L for that
2:09:31
matter so for example Json is actually
2:09:34
really dense in tokens and yaml is a lot
2:09:36
more efficient in tokens um so for
2:09:39
example this are these are the same in
2:09:41
Json and in yaml the Json is
2:09:45
116 and the yaml is 99 so quite a bit of
2:09:48
an Improvement and so in the token
2:09:52
economy where we are paying uh per token
2:09:54
in many ways and you are paying in the
2:09:56
context length and you're paying in um
2:09:58
dollar amount for uh the cost of
2:10:00
processing all this kind of structured
2:10:01
data when you have to um so prefer to
2:10:04
use theal over Json and in general kind
2:10:06
of like the tokenization density is
2:10:08
something that you have to um sort of
2:10:10
care about and worry about at all times
2:10:12
and try to find efficient encoding
2:10:13
schemes and spend a lot of time in tick
2:10:15
tokenizer and measure the different
2:10:17
token efficiencies of different formats
2:10:19
and settings and so on okay so that
2:10:21
concludes my fairly long video on
2:10:23
tokenization I know it's a try I know
2:10:26
it's annoying I know it's irritating I
2:10:28
personally really dislike the stage what
2:10:31
I do have to say at this point is don't
2:10:33
brush it off there's a lot of foot guns
2:10:35
sharp edges here security issues uh AI
2:10:38
safety issues as we saw plugging in
2:10:40
unallocated memory into uh language
2:10:42
models so um it's worth understanding
2:10:45
this stage um that said I will say that
2:10:48
eternal glory goes to anyone who can get
2:10:50
rid of it uh I showed you one possible
2:10:53
paper that tried to uh do that and I
2:10:55
think I hope a lot more can follow over
2:10:57
time and my final recommendations for
2:10:59
the application right now are if you can
2:11:01
reuse the GPT 4 tokens and the
2:11:03
vocabulary uh in your application then
2:11:05
that's something you should consider and
2:11:06
just use Tech token because it is very
2:11:08
efficient and nice library for inference
2:11:11
for bpe I also really like the bite
2:11:14
level BP that uh Tik toen and openi uses
2:11:17
uh if you for some reason want to train
2:11:19
your own vocabulary from scratch um then
2:11:23
I would use uh the bpe with sentence
2:11:25
piece um oops as I mentioned I'm not a
2:11:28
huge fan of sentence piece I don't like
2:11:31
its uh bite fallback and I don't like
2:11:34
that it's doing BP on unic code code
2:11:36
points I think it's uh it also has like
2:11:38
a million settings and I think there's a
2:11:39
lot of foot gonss here and I think it's
2:11:40
really easy to Mis calibrate them and
2:11:42
you end up cropping your sentences or
2:11:44
something like that uh because of some
2:11:46
type of parameter that you don't fully
2:11:47
understand so so be very careful with
2:11:49
the settings try to copy paste exactly
2:11:52
maybe where what meta did or basically
2:11:54
spend a lot of time looking at all the
2:11:56
hyper parameters and go through the code
2:11:57
of sentence piece and make sure that you
2:11:59
have this correct um but even if you
2:12:02
have all the settings correct I still
2:12:03
think that the algorithm is kind of
2:12:05
inferior to what's happening here and
2:12:08
maybe the best if you really need to
2:12:10
train your vocabulary maybe the best
2:12:11
thing is to just wait for M bpe to
2:12:13
becomes as efficient as possible and uh
2:12:17
that's something that maybe I hope to
2:12:18
work on and at some point maybe we can
2:12:21
be training basically really what we
2:12:23
want is we want tick token but training
2:12:25
code and that is the ideal thing that
2:12:28
currently does not exist and MBP is um
2:12:31
is in implementation of it but currently
2:12:33
it's in Python so that's currently what
2:12:36
I have to say for uh tokenization there
2:12:38
might be an advanced video that has even
2:12:40
drier and even more detailed in the
2:12:42
future but for now I think we're going
2:12:44
to leave things off here and uh I hope
2:12:47
that was helpful bye
2:12:54
and uh they increase this contact size
2:12:56
from gpt1 of 512 uh to 1024 and GPT 4
2:13:03
two the
2:13:05
next okay next I would like us to
2:13:08
briefly walk through the code from open
2:13:10
AI on the gpt2 encoded
2:13:16
ATP I'm sorry I'm gonna sneeze
2:13:19
and then what's Happening Here
2:13:22
is this is a spous layer that I will
2:13:25
explain in a
2:13:26
bit What's Happening Here
2:13:33
is