# Rasa classifies random input as intents with high probability

**URL:** <https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331>\
**Category:** Rasa Open Source\
**Created:** [July 30, 2021, 12:24pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331 "2021-07-30T12:24:03Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jakub](https://avatars.discourse-cdn.com/v4/letter/j/ecb155/32.png) [@Jakub](https://forum.rasa.com/u/Jakub)\
**Post date:** [July 30, 2021, 12:24pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/1 "2021-07-30T12:24:03Z")

</div>

Hello, I have trained my model using pipeline:

- name: “DucklingHTTPExtractor” url: “[http://localhost:8000](http://localhost:8000)” dimensions: [“duration”]
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: char\_wb min\_ngram: 1 max\_ngram: 4
- name: DIETClassifier epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector epochs: 100
- name: FallbackClassifier threshold: 0.3 ambiguity\_threshold: 0.1

However random words like ‘asdf’, ‘qwerty’ and even single letters are classified as intents (‘greet’, ‘neutral’ respectively). I could increase FallbackClassifier threshold, but some examples have confidence over 0.9. I’m using Rasa 2.8. How to deal with it?

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [July 30, 2021, 2:02pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/2 "2021-07-30T14:02:15Z")

</div>

@Jakub Hi! Can you please share some example or file or screenshot?

@Jakub Why you are using `ambiguity_threshold: 0.1` inside FallbackClassifier or share the link from where you get this idea?

@Jakub Please share complete config.yml or update the above one.

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [July 30, 2021, 2:20pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/3 "2021-07-30T14:20:34Z")

</div>

> [@nik202](#):
>
> share some example or file or screenshot?

Hi! Same on my site. No Matter what kind of input. There is always a intent recognition. Eg “123456”, “dfjjaspfj” or something like “i want pizza”.

I’m looking forward for some help. Thanks!

Using diffrent versions with docker:

- rasa/rasa:main-spacy-de
- rasa/rasa:2.8.1-spacy-de

pipeline: language: de

```
pipeline:
    # # No configuration for the NLU pipeline was provided. The following default pipeline was used to train your model.
    # # If you'd like to customize it, uncomment and adjust the pipeline.
    # # See https://rasa.com/docs/rasa/tuning-your-model for more information.
    - name: SpacyNLP
      model: de_core_news_sm
    - name: SpacyTokenizer
    - name: SpacyFeaturizer
    - name: RegexFeaturizer
    - name: LexicalSyntacticFeaturizer
    - name: CountVectorsFeaturizer
    - name: CountVectorsFeaturizer
      analyzer: char_wb
      min_ngram: 1
      max_ngram: 4
    - name: DIETClassifier
      epochs: 100
      constrain_similarities: true
    - name: EntitySynonymMapper
    - name: FallbackClassifier
      threshold: 0.3
      ambiguity_threshold: 0.1

```

nlu.yml: version: “2.0”

```
nlu:

- intent: greet

  examples: |

    - hey

    - hallo

    - hi

    - hallo du

    - guten morgen

    - guten abend

    - morgen

    - guten tag

- intent: goodbye

  examples: |

    - tschüss

    - wiederhören

    - ciao

    - bis dann

    - bye

```

result:

```
{

    "text": "12345",

    "intent": {

        "id": 8761605359927853060,

        "name": "goodbye",

        "confidence": 0.9780691862106323

    },

    "entities": [],

    "intent_ranking": [

        {

            "id": 8761605359927853060,

            "name": "goodbye",

            "confidence": 0.9780691862106323

        },

        {

            "id": -3994034142759590561,

            "name": "greet",

            "confidence": 0.02193082869052887

        }

    ]

}

```

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [July 30, 2021, 2:25pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/4 "2021-07-30T14:25:21Z")

</div>

> [@thorty](#):
>
> `12345`

@thorty Means whatever input i.e “I want pizza” , “I need drinks” do such examples are in your nlu.yml file?; it’s giving you goodbye output?

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [July 30, 2021, 2:48pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/5 "2021-07-30T14:48:33Z")

</div>

> [@nik202](#):
>
> @thorty Means whatever input i.e “I want pizza” , “I need drinks” do such examples are in your nlu.yml file?; it’s giving you goodbye output?

No they are **not** into my nlu.yml as examples. In this case I’m expecting a result like: intent: none. But what I’ m getting here is something like intent: goodbye with a conf of over 90.

my nlu.yml is exactly like in the post above. the pipeline also. and the result is an example with this config.

thank you @nik202 !

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [July 30, 2021, 3:05pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/6 "2021-07-30T15:05:41Z")

</div>

> [@thorty](#):
>
> `ambiguity_threshold: 0.1`

@thorty Why this? any significance? @thorty Link please, I guess its used for two-stage fallback not for FallbackClassifier

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [July 30, 2021, 3:12pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/7 "2021-07-30T15:12:21Z")

</div>

@Nik202 No. It’s from the doc or from the example of the init pipeline. Should I use another threshold? I will try this out later.

---

<div class="post-metadata">

**Author:** ![rctatman](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/rctatman/32/6734_2.png) [@rctatman](https://forum.rasa.com/u/rctatman)\
**Post date:** [July 30, 2021, 4:52pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/8 "2021-07-30T16:52:46Z")

</div>

@thorty All your inputs will be given an intent. If you have very high confidence and a lot of incorrect classifications it’s often the case that this is due to your data not being a good representation of what your assistant is seeing in production.

You should create an out of scope intent if you want to be able to capture things your assistant _can’t_ do. This page goes over how to do it: [Fallback and Human Handoff](https://rasa.com/docs/rasa/fallback-handoff/#handling-out-of-scope-messages)

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [July 31, 2021, 6:29am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/9 "2021-07-31T06:29:25Z")

</div>

@Nik202: Thanks, but this does not change enything.

@rctatman: Thanks for your explanation and the Link to the right Step in Documentation. This helps a lot in understanding rasa.

In my case I do not really know the utterances from users. So I will start with a small set of examples and train the NLU after going live with real-world examples. To handle the “non” trained utterances the right way it would be better to get **no** intent rather then the **wrong** intent. Especially when utterances are non-sense like in my example. Is there any chance to achieve this?

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [July 31, 2021, 12:27pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/10 "2021-07-31T12:27:37Z")

</div>

@thorty ok

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [July 31, 2021, 12:29pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/11 "2021-07-31T12:29:06Z")

</div>

@Jakub Any progress about your error?

---

<div class="post-metadata">

**Author:** ![Jakub](https://avatars.discourse-cdn.com/v4/letter/j/ecb155/32.png) [@Jakub](https://forum.rasa.com/u/Jakub)\
**Post date:** [August 2, 2021, 8:40am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/12 "2021-08-02T08:40:43Z")

</div>

@nik202 sorry for late response. I removed ambiguity\_threshold: 0.1 from my config file (the one I sent in first message is my whole config.yml, I have default policies), but there’s only a slight diffrence (a little bit more of inputs are classified as nlu\_fallback). I upload examples and neutral intent below.

![example1](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/c/2/c2ceecbfeb19b877b26701552a36c89f5f68e08a.png) ![example2](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/5/b/5ba1ec77514c7443fd9305892680290d3f76addf.png) ![example3](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/e/0/e0ce2e5ff0d6519cc7c1a899fca2f668a6690a41.png) ![example4](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/8/0/80528192cb77cdcb94e577c58a6dbd60fb852be8.png)

 ![example5](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/9/e/9e1fcf327897408da5d9d03db4c8f8e65f531e65.png)

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [August 2, 2021, 1:51pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/13 "2021-08-02T13:51:01Z")

</div>

Short Update from my side:

when I also use the KeywordIntentClassifier in my pipline the NLU response as follows:

```
"intent": {
    "name": null,
    "confidence": 0.0
},

```

pipeline `- name: KeywordIntentClassifier`

But I think that makes the DIET Classifier with all its advantages useless. ?!?

---

<div class="post-metadata">

**Author:** ![thorty](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/thorty/32/16303_2.png) [@thorty](https://forum.rasa.com/u/thorty)\
**Post date:** [August 4, 2021, 2:15pm UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/14 "2021-08-04T14:15:39Z")

</div>

Short Update:

Like @rctatman explained. With more trainingdata I become better confidence values for my utterences. Also for “nonsense” imput like “askjfnqiurz”.

---

<div class="post-metadata">

**Author:** ![SowmyaBalam](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@SowmyaBalam](https://forum.rasa.com/u/SowmyaBalam)\
**Post date:** [October 12, 2021, 6:24am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/15 "2021-10-12T06:24:08Z")

</div>

hi, i am also facing the same issue, for whatever nonsense query "abcdefghijkl’ it is going to greet with high confidence of 0.99 in diet classifier. can anybody help me resolve this issue.

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [October 12, 2021, 7:30am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/16 "2021-10-12T07:30:54Z")

</div>

@SowmyaBalam delete the older trained model first, some times it help and did you mention the default fallback, please share some supporting files, so that we can see.

---

<div class="post-metadata">

**Author:** ![SowmyaBalam](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@SowmyaBalam](https://forum.rasa.com/u/SowmyaBalam)\
**Post date:** [October 12, 2021, 7:34am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/17 "2021-10-12T07:34:18Z")

</div>

@nik202this is the config file. i am using:

language: en

pipeline:

# # No configuration for the NLU pipeline was provided. The following default pipeline was used to train your model.

# # If you’d like to customize it, uncomment and adjust the pipeline.

# # See [https://rasa.com/docs/rasa/tuning-your-model](https://rasa.com/docs/rasa/tuning-your-model) for more information.

- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: char\_wb min\_ngram: 1 max\_ngram: 4
- name: DIETClassifier epochs: 300 evaluate\_on\_number\_of\_examples: 80 evaluate\_every\_number\_of\_epochs: 5 tensorboard\_log\_directory: “tensorboard” tensorboard\_log\_level: “epoch” checkpoint\_model: True constrain\_similarities: True
- name: EntitySynonymMapper

# - name: ResponseSelector

# epochs: 100

# constrain\_similarities: true

# - name: FallbackClassifier

# threshold: 0.3

# ambiguity\_threshold: 0.1

# Configuration for Rasa Core.

# [https://rasa.com/docs/rasa/core/policies/](https://rasa.com/docs/rasa/core/policies/)

policies:

# # No configuration for policies was provided. The following default policies were used to train your model.

# # If you’d like to customize them, uncomment and adjust the policies.

# # See [Policy Overview | Rasa Documentation](https://rasa.com/docs/rasa/policies) for more information.

- name: MemoizationPolicy
- name: TEDPolicy #max\_history: 5 epochs: 300 evaluate\_on\_number\_of\_examples: 80 evaluate\_every\_number\_of\_epochs: 5 tensorboard\_log\_directory: “tensorboard” tensorboard\_log\_level: “epoch” checkpoint\_model: True constrain\_similarities: True
- name: RulePolicy

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [October 12, 2021, 7:36am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/18 "2021-10-12T07:36:41Z")

</div>

@SowmyaBalam can you please format the code using “\</\>” or “”" “”" ?

---

<div class="post-metadata">

**Author:** ![SowmyaBalam](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@SowmyaBalam](https://forum.rasa.com/u/SowmyaBalam)\
**Post date:** [October 12, 2021, 7:40am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/19 "2021-10-12T07:40:57Z")

</div>

language: en

pipeline:

- name: WhitespaceTokenizer\
- name: RegexFeaturizer\
- name: LexicalSyntacticFeaturizer\
- name: CountVectorsFeaturizer\
- name: CountVectorsFeaturizer  
analyzer: char\_wb  
min\_ngram: 1  
max\_ngram: 4\
- name: DIETClassifier  
epochs: 300  
evaluate\_on\_number\_of\_examples: 80  
evaluate\_every\_number\_of\_epochs: 5  
tensorboard\_log\_directory: “tensorboard”  
tensorboard\_log\_level: “epoch”  
checkpoint\_model: True  
constrain\_similarities: True\
- name: EntitySynonymMapper\

policies:

- name: MemoizationPolicy\
- name: TEDPolicy  
max\_history: 5  
epochs: 300  
evaluate\_on\_number\_of\_examples: 80  
evaluate\_every\_number\_of\_epochs: 5  
tensorboard\_log\_directory: “tensorboard”  
tensorboard\_log\_level: “epoch”  
checkpoint\_model: True  
constrain\_similarities: True\
- name: RulePolicy

---

<div class="post-metadata">

**Author:** ![nik202](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/nik202/32/16098_2.png) [@nik202](https://forum.rasa.com/u/nik202)\
**Post date:** [October 12, 2021, 7:54am UTC](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331/20 "2021-10-12T07:54:22Z")

</div>

@SowmyaBalam Hey! why there is a \ (backward slash) ?infront of the pipelines and policies ?

[Next page](https://forum.rasa.com/t/rasa-classifies-random-input-as-intents-with-high-probability/46331.md?page=2)
