Език: English
This lecture is an introduction on how you can host open source Large Language Models (LLMs) entirely locally on old/spare hardware. You will learn what an inference server is, how you can set one up, and how you can use it to chat with an LLM that is entirely contained on your own hardware.
This lecture will show the basic principles of running a Large Language Model (LLM) inference server with an open source inference engine and open source models, all hosted locally on relatively old and available hardware.
We will explain:
- What tools are needed and how they work together to host your local LLM
- What hardware is needed, and what determines the LLMs performance
- What is LLM quantization, why is it important, and when/how to use it
- What are the benefits and limitations of hosting LLMs locally
- The process of setting up an LLM hosting stack step-by-step