# Adding Larger Lookup Table Causes Ill Defined F-Scores

**URL:** <https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501>\
**Category:** Rasa Open Source\
**Created:** [May 8, 2019, 5:16pm UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501 "2019-05-08T17:16:11Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![akshay2000](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akshay2000/32/415_2.png) [@akshay2000](https://forum.rasa.com/u/akshay2000)\
**Post date:** [May 8, 2019, 5:16pm UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/1 "2019-05-08T17:16:11Z")

</div>

I have an entity that is supposed to be extracted by `ner_crf`. Earlier I had about 50-60 examples in the lookup table for this entity. Everything seemed to be working as expected.

Now, I have added about 1000 entries in my lookup table. Suddenly, `intent_classifier_sklearn` component gives following warning: UndefinedMetricWarning: F-score is ill-defined and being set to 0.0 in labels with no predicted samples.

What is happening here? Shouldn’t intent classifier work independently of the lookup table - which is entity extraction construct?

---

<div class="post-metadata">

**Author:** ![damao](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/damao/32/2465_2.png) [@damao](https://forum.rasa.com/u/damao)\
**Post date:** [May 9, 2019, 6:20am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/2 "2019-05-09T06:20:12Z")

</div>

I always have this warning during the training. Also can’t figure it out…

---

<div class="post-metadata">

**Author:** ![akshay2000](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akshay2000/32/415_2.png) [@akshay2000](https://forum.rasa.com/u/akshay2000)\
**Post date:** [May 13, 2019, 8:18am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/4 "2019-05-13T08:18:28Z")

</div>

Can you share the data and the logs? I suspect the either one of the intents have too little information or it is too diverse. Have you tried calculating the confusion matrix using `evaluate` module?

While these methods aren’t concrete, they should give you rough idea of where the problem lies.

---

<div class="post-metadata">

**Author:** ![damao](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/damao/32/2465_2.png) [@damao](https://forum.rasa.com/u/damao)\
**Post date:** [May 14, 2019, 1:40am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/5 "2019-05-14T01:40:36Z")

</div>

Sorry, it is a business project so I can’t share any data

---

<div class="post-metadata">

**Author:** ![souvikg10](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/souvikg10/32/93_2.png) [@souvikg10](https://forum.rasa.com/u/souvikg10)\
**Post date:** [May 17, 2019, 9:36pm UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/6 "2019-05-17T21:36:50Z")

</div>

The ill defined f-score is based on lack of enough training examples to evaluate a certain intent or entities, when you are using lookup tables, it uses one of the features of ner\_crf , which then enforces training and tries to evaluate based on the training set, however we don’t put all the examples in the training set and hence could be the reason for this warning

---

<div class="post-metadata">

**Author:** ![akshay2000](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akshay2000/32/415_2.png) [@akshay2000](https://forum.rasa.com/u/akshay2000)\
**Post date:** [May 18, 2019, 5:47am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/7 "2019-05-18T05:47:33Z")

</div>

Isn’t the very purpose of a lookup table that we shouldn’t have to include all the examples in training set? Secondly, I know it is just a warning, but it looks like it might affect predictions. Is there a guideline on how training should look to avoid it?

---

<div class="post-metadata">

**Author:** ![souvikg10](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/souvikg10/32/93_2.png) [@souvikg10](https://forum.rasa.com/u/souvikg10)\
**Post date:** [May 18, 2019, 7:43am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/8 "2019-05-18T07:43:00Z")

</div>

Indeed, but the lookup table is using one of the feature of ner\_crf - pattern feature

take a look at this blog - [Entity extraction with the new lookup table feature in Rasa NLU | The Rasa Blog | Rasa](https://blog.rasa.com/improving-entity-extraction/)

> However, because the training set is still so small, you’d likely need a few hundred more examples to push this score to above 80% in practice.

One of the sentences from the blog

Lookup tables are means to improve entity extraction of NER\_CRF unlike a phrase matcher. It depends on the entity itself

---

<div class="post-metadata">

**Author:** ![akshay2000](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akshay2000/32/415_2.png) [@akshay2000](https://forum.rasa.com/u/akshay2000)\
**Post date:** [May 19, 2019, 6:48am UTC](https://forum.rasa.com/t/adding-larger-lookup-table-causes-ill-defined-f-scores/9501/9 "2019-05-19T06:48:29Z")

</div>

I’m sorry, I am a bit confused. Here’s what I understand:

Lookup tables use `pattern` feature from `ner_crf`. So, for entity extraction to work properly, we need enough (not all, but enough) data in the intent itself. So, essentially, if I add more values from lookup table to my training samples, performance should improve.

Is that correct?
