# Local LLM with text-generation-webui steps

**URL:** <https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020>\
**Category:** Rasa CALM\
**Created:** [August 14, 2024, 9:00am UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020 "2024-08-14T09:00:46Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![chatz](https://avatars.discourse-cdn.com/v4/letter/c/b77776/32.png) [@chatz](https://forum.rasa.com/u/chatz)\
**Post date:** [August 14, 2024, 9:00am UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/1 "2024-08-14T09:00:46Z")

</div>

Hello everyone I was trying to deploy a local llm with RASA PRO and finally I found the solution here is the details if anyone needs it:

I have installed text-generation-webui → [link](https://github.com/oobabooga/text-generation-webui)

Then:

1. I started the server with ./start\_linux.sh
2. Loaded the model through “Model” tab
3. In “Session” tab I selected openai, api, listen and pressed Apply flags ![image](https://europe1.discourse-cdn.com/flex013/uploads/rasa/original/3X/b/9/b93a94f4b397a53c12de2b98ed4e6bd3dae2b8fe.png)

In rasa endpoints.yml:

```
nlg:
  type: rephrase
  rephrase_all: true
  llm:
    model: 'model_gemma_27b_it'
    model_name : 'model_gemma_27b_it'
    type: "openai"
    openai_api_key: "NULL"
    openai_api_base: http://127.0.0.1:5000/v1
    request_timeout: 800

```

If you have an error

```
AttributeError: module ‘openai’ has no attribute ‘error’

```

you have to install this:

```
  pip install openai==0.28.1

```

---

<div class="post-metadata">

**Author:** ![Sanjukta.bs](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/sanjukta.bs/32/24453_2.png) [@Sanjukta.bs](https://forum.rasa.com/u/Sanjukta.bs)\
**Post date:** [September 12, 2024, 6:53pm UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/2 "2024-09-12T18:53:49Z")

</div>

Hey what changes should I be making to my config for this to work! Also I dont have any access to open api key how can I bypass open api keys. Because it keeps popping up. Any help will be appreciated. Thanks

---

<div class="post-metadata">

**Author:** ![sahibpreetsingh12](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/sahibpreetsingh12/32/14031_2.png) [@sahibpreetsingh12](https://forum.rasa.com/u/sahibpreetsingh12)\
**Post date:** [September 14, 2024, 9:11pm UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/3 "2024-09-14T21:11:55Z")

</div>

Try and use a huggingface model (Mixtral would be fine)

---

<div class="post-metadata">

**Author:** ![Sanjukta.bs](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/sanjukta.bs/32/24453_2.png) [@Sanjukta.bs](https://forum.rasa.com/u/Sanjukta.bs)\
**Post date:** [September 15, 2024, 6:22am UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/4 "2024-09-15T06:22:02Z")

</div>

I want to use local models, I was trying ollama but it is taking a lot of time to generate a reply. Thus I am stuck.

---

<div class="post-metadata">

**Author:** ![sahibpreetsingh12](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/sahibpreetsingh12/32/14031_2.png) [@sahibpreetsingh12](https://forum.rasa.com/u/sahibpreetsingh12)\
**Post date:** [September 15, 2024, 1:46pm UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/5 "2024-09-15T13:46:23Z")

</div>

That’s the only stopping point with HF Try and see if you can use vLLM

---

<div class="post-metadata">

**Author:** ![chatz](https://avatars.discourse-cdn.com/v4/letter/c/b77776/32.png) [@chatz](https://forum.rasa.com/u/chatz)\
**Post date:** [September 16, 2024, 8:51am UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/6 "2024-09-16T08:51:16Z")

</div>

Hello, if the model takes too much time to generate I believe the problem is that your system is struggling to load the model. Maybe you should try to improve it by reducing the characters generated or other configuration variables.

---

<div class="post-metadata">

**Author:** ![Sanjukta.bs](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.rasa.com/sanjukta.bs/32/24453_2.png) [@Sanjukta.bs](https://forum.rasa.com/u/Sanjukta.bs)\
**Post date:** [September 16, 2024, 11:08am UTC](https://forum.rasa.com/t/local-llm-with-text-generation-webui-steps/63020/7 "2024-09-16T11:08:07Z")

</div>

More than taking time, it is predicting wrong flows! Is there any good demo that we can follow that uses local llms instead of open ai?
