I did manage to write a back-propogation algorithm, at this point I don't fully understand the math behind back-propogation. Generally back-propogation algorithms take the activation, calculate the delta(?) with the activation and the target output (only on last layer). I don't know where tokens come in. From your comment it sounds like it has to do something in a unsupervised learning network. I am also not a professional. Sorry if I didn't really understand your comment.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: