Advice on generating the right amount of NLU training data

Fair point! I fully understand what you are saying and believe that looking at the actual data/results and ideally taking enduser experience into account is the best way to go. However in this case, I was just looking for inspiration, if there is a way to get a first view on the performance already (but then a bit more than a general F1 score :smiley:) But will think about a sort of “A/B test-setup” then maybe!

FYI: found this one when doing some research [2005.04118] Beyond Accuracy: Behavioral Testing of NLP models with CheckList . Comes with a github repo.
Not fully suitable for (end-2-end) Conversational AI, but interesting to read.