Embedding Policy Results

For the new Embedding you need to train stories with those chitchats and corrections. So, where is now the advantage/improvement compared to normal LSTM? Is it that you need way less stories to write such unccooperative stories, because attention layer learns not to pay attention to this part and will generalise to stories not trained?