# Similar Entity Extraction

**URL:** <https://forum.rasa.com/t/similar-entity-extraction/1159>\
**Category:** Rasa Open Source\
**Created:** [September 17, 2018, 11:13am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159 "2018-09-17T11:13:09Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [September 17, 2018, 11:13am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/1 "2018-09-17T11:13:09Z")

</div>

I am facing problem in extracting the entity for following use case. User:Hi Bot:hello! How may I help you User:I want my tax reciept Bot: Sure! Could you provide me your ID User:59(he can also say like:it is 59 or ID is 59 or ID=59 or my ID is 59) Bot: Your User number? User: it is 9154 Bot:Your transaction number user: 745

I want to extract 59,9154,745. How to proceed with this problem? I have trained the data but when I provide ID and user number of the same length. the entity is not extracted. However intent is working fine and the the conversation is happening as per the story.

---

<div class="post-metadata">

**Author:** ![akelad](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akelad/32/69_2.png) [@akelad](https://forum.rasa.com/u/akelad)\
**Post date:** [September 17, 2018, 11:15am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/2 "2018-09-17T11:15:19Z")

</div>

have you tried using the `ner_duckling_http` entity extractor?

---

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [September 17, 2018, 11:17am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/3 "2018-09-17T11:17:01Z")

</div>

I am using ner\_crf container

---

<div class="post-metadata">

**Author:** ![akelad](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akelad/32/69_2.png) [@akelad](https://forum.rasa.com/u/akelad)\
**Post date:** [September 18, 2018, 3:57pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/4 "2018-09-18T15:57:29Z")

</div>

for numbers, please try the `ner_duckling_https` extractor

---

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [September 18, 2018, 4:11pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/5 "2018-09-18T16:11:37Z")

</div>

Actually installing the Duckling is a tedious task. I m getting too many errors while installing it.

---

<div class="post-metadata">

**Author:** ![akelad](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akelad/32/69_2.png) [@akelad](https://forum.rasa.com/u/akelad)\
**Post date:** [September 19, 2018, 4:19pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/6 "2018-09-19T16:19:38Z")

</div>

you don’t need to install it. you can run it with docker [https://rasa.com/docs/nlu/master/components/#ner-duckling-http](https://rasa.com/docs/nlu/master/components/#ner-duckling-http)

---

<div class="post-metadata">

**Author:** ![andrewbain](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/andrewbain/32/570_2.png) [@andrewbain](https://forum.rasa.com/u/andrewbain)\
**Post date:** [September 24, 2018, 10:26am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/7 "2018-09-24T10:26:50Z")

</div>

Hi there, I have it the duckling server running in docker and it seems to be working however it seems to have slowed down my training to snail pace. I am also unsure how to enter training data as duckling creates it’s own entities, do I just enter the intent and add the duckling entity names to my rasa\_core config?

---

<div class="post-metadata">

**Author:** ![akelad](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/akelad/32/69_2.png) [@akelad](https://forum.rasa.com/u/akelad)\
**Post date:** [September 26, 2018, 10:00am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/8 "2018-09-26T10:00:25Z")

</div>

hmmm, duckling shouldn’t affect your training at all. and duckling just extracts the entities, no need to label them

---

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [September 26, 2018, 2:50pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/9 "2018-09-26T14:50:00Z")

</div>

What if the inputs are alphanumeric for e.g. my ID is NP\_45680780. And how to correct order in case user replies in a different order. E.g. BOT-Can I have Ur ID? USER- my user number is NA\_56098

And one more doubt,how to extract desired entity from multiple entities As : my transaction number is QR%56873578 for user number 987\_AT. I have to extract QR%56873578

---

<div class="post-metadata">

**Author:** ![ccelotto](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ccelotto/32/333_2.png) [@ccelotto](https://forum.rasa.com/u/ccelotto)\
**Post date:** [October 1, 2018, 6:49pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/10 "2018-10-01T18:49:59Z")

</div>

I’m wondering the same thing here. It doesn’t seem like Duckling is the solution for custom entities such as alphanumeric codes like you and I are working with.

Does anyone have any suggestions for extracting custom alphanumeric codes? Duckling and entity lookup tables haven’t seemed to work after some testing. Maybe creating some Regex patterns would work?

Thanks!

---

<div class="post-metadata">

**Author:** ![souvikg10](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/souvikg10/32/93_2.png) [@souvikg10](https://forum.rasa.com/u/souvikg10)\
**Post date:** [October 1, 2018, 8:28pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/11 "2018-10-01T20:28:01Z")

</div>

> **[NLU Training Data](https://rasa.com/docs/rasa/nlu-training-data/)**
>
> Read more about how to format training data with Rasa NLU for open source natural language processing.

You have in this page at the bottom a description about Regex patterns.

---

<div class="post-metadata">

**Author:** ![ccelotto](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ccelotto/32/333_2.png) [@ccelotto](https://forum.rasa.com/u/ccelotto)\
**Post date:** [October 1, 2018, 8:37pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/12 "2018-10-01T20:37:19Z")

</div>

Thank you. I’m familiar with the Regex documentation, but zip codes and phone numbers seem to be a bit different thank custom alphanumeric codes since zip codes and phone numbers always have the same format while unique alphanumeric codes do not.

I was more so looking for some thoughts on the effectiveness (and if it is possible) to use Regex patterns for a situation such as this.

---

<div class="post-metadata">

**Author:** ![souvikg10](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/souvikg10/32/93_2.png) [@souvikg10](https://forum.rasa.com/u/souvikg10)\
**Post date:** [October 2, 2018, 7:38am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/14 "2018-10-02T07:38:28Z")

</div>

you can also have a regex pattern for alphanumeric codes as well

I take the example of userID

let’s say it is 6 digits and starts with G

```auto
G([1-9]\d{4})

```

Then I should provide examples such as

my id is **G15367** [ID] …

In another way, you can also add regex entity extractor, that takes a regular expression pattern as rule and find entities from a given token (similar to duckling)

also FYI, in duckling you can add custom rules if you have a hang on Haskell. They have recently added a new feature to add custom dimension.

---

<div class="post-metadata">

**Author:** ![ccelotto](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ccelotto/32/333_2.png) [@ccelotto](https://forum.rasa.com/u/ccelotto)\
**Post date:** [October 2, 2018, 4:56pm UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/15 "2018-10-02T16:56:42Z")

</div>

Awesome! Thank you. I can confirm for @ashukrishna100 that this does indeed work. I used a few guides online that provided regex variable charts to come up with the regex pattern suited for our use case. Performance is great after implementing a regex pattern!

---

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [October 9, 2018, 5:52am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/16 "2018-10-09T05:52:56Z")

</div>

Yeah it works and with optimum performance for sure. Thank you so much @ccelotto

---

<div class="post-metadata">

**Author:** ![ashukrishna100](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ashukrishna100/32/516_2.png) [@ashukrishna100](https://forum.rasa.com/u/ashukrishna100)\
**Post date:** [October 9, 2018, 5:56am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/17 "2018-10-09T05:56:45Z")

</div>

Very informative. no doubts you are a RASA star. have tried regex entity extractor and it works fine. Thanks @souvikg10

---

<div class="post-metadata">

**Author:** ![neerajb1](https://avatars.discourse-cdn.com/v4/letter/n/90ced4/32.png) [@neerajb1](https://forum.rasa.com/u/neerajb1)\
**Post date:** [October 24, 2018, 10:17am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/18 "2018-10-24T10:17:50Z")

</div>

> [@ccelotto](#):
>
> le charts to come up with the regex pattern suited for our use c

Hi ,

Could you please share how it works. I am also using it but no luck.

NLU: “regex\_features”: [{ “name”: “Transaction\_ID”, “pattern”: “\[1\]+$” }]

pipeline:

- name: “tokenizer\_whitespace”
- name: “intent\_entity\_featurizer\_regex”
- name: “ner\_crf”
- name: “ner\_synonyms”
- name: “intent\_featurizer\_count\_vectors”
- name: intent\_classifier\_tensorflow\_embedding

Thanks

* * *

1. a-zA-Z0-9

---

<div class="post-metadata">

**Author:** ![ccelotto](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/ccelotto/32/333_2.png) [@ccelotto](https://forum.rasa.com/u/ccelotto)\
**Post date:** [October 25, 2018, 2:37am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/19 "2018-10-25T02:37:40Z")

</div>

This was super helpful for me in creating the regex for my use case [http://www.cbs.dtu.dk/courses/27610/regular-expressions-cheat-sheet-v2.pdf](http://www.cbs.dtu.dk/courses/27610/regular-expressions-cheat-sheet-v2.pdf)

From looking at your regex pattern, it seems like it may be missing some parentheses/brackets.

This is mine: (([A-z]{1})([0-9]{6,7}))

It is used for picking up on alphanumeric codes similar to “A123456”, “G493024”, “F4930294”.

So breaking it down… (**([A-z]{1})**([0-9]{6,7}))… this bolded section states that the first character {1} will be a letter A-Z [A-z].

(([A-z]{1})**([0-9]{6,7})**)… this bolded section states that the next 6-7 characters {6,7} will be numbers 0 through 9 [0-9].

Hopefully you can use this as a reference! If not, tell me what you’re trying to accomplish, and I’ll do my best to help you craft it.

P.S. don’t forget (like mentioned in [Rasa docs for regex](https://rasa.com/docs/nlu/dataformat/)) to provide some training examples using the regex or the model won’t know to pick up on the pattern.

---

<div class="post-metadata">

**Author:** ![neerajb1](https://avatars.discourse-cdn.com/v4/letter/n/90ced4/32.png) [@neerajb1](https://forum.rasa.com/u/neerajb1)\
**Post date:** [October 26, 2018, 4:33am UTC](https://forum.rasa.com/t/similar-entity-extraction/1159/20 "2018-10-26T04:33:05Z")

</div>

> [@ccelotto](#):
>
> 9

Hi Christopher,

Thank you very much for detailed explanation. It worked i for me . I think it might be due to i forgot including intent\_entity\_featurizer\_regex in config file
