<> = ollama = * https://ollama.com/ * https://ollama.com/search * https://github.com/ollama/ollama {{{#!highlight sh cd ~ curl -fsSL https://ollama.com/install.sh | sh # >>> The Ollama API is now available at 127.0.0.1:11434. curl localhost:11434 # Ollama is running ollama run llama3.2:1b ollama show llama3.2:1b ollama list ollama ps ollama serve # start server ollama stop llama3.2:1b ollama rm llama3.2:1b }}} == Setup proxy == {{{#!highlight sh sudo nano /etc/systemd/system/ollama.service }}} {{{#!highlight sh [Unit] Description=Ollama Service After=network-online.target [Service] ExecStart=/usr/local/bin/ollama serve User=ollama Group=ollama Restart=always RestartSec=3 Environment="PATH=/home/vagrant/.dotnet/tools:/home/vagrant/dotnetcore9:/usr/local/bin:/usr/bin:/bin:/home/vagrant/jdk-17.0.7+7/bin/:/home/vagrant/gradle-8.1/bin/" Environment="HTTPS_PROXY=http://192.168.0.123:3128/" Environment="HTTP_PROXY=http://192.168.0.123:3128/" [Install] WantedBy=default.target }}} {{{#!highlight sh sudo systemctl daemon-reload sudo systemctl restart ollama systemctl show ollama ollama run llama3.2:1b }}} == Chatbots - LLM - AI accessible via browser == * Gemini prompt - https://gemini.google.com/ * Copilot prompt - https://copilot.microsoft.com/ * Grok prompt - https://grok.com/ * ChatGPT prompt - https://chatgpt.com/ == Stuff == Does a local run of ollama spends tokens ? No LLMs might generate different versions because they are probabilistic. They predict the next word/token based on probabilities. The reason for results variation is called temperature. High temperature [0.8 - 1.0] the model takes more risks that leads to creative code. Low temperature [0.0 - 0.2] becomes very focused and predictable. At temperature 0 (zero) it will choose the most likely word/token. Seed value, in ollama, if not specified a random one is generated for each request. If you want the same code every time set the same seed and temperature 0. RAG, retrieved augmented generation. Create a domain-specialized assistant. Tools like copilot chat looks at the .md files and source code and perform a RAG workflow. RAG. * R retrieval, the tool searches for the relevant code or markdown text * A augmentation, takes those snippets and pastes them into the prompt and send it to the model * G the model reads the code/docs and generates an answer based only on the specific context A specific project is the domain. Project memory/domain (markdown + code) Local LLM runner (ollama) VS code extension continue RAM local models * 8GB RAM -> llama3:8b * 16GB - 32GB -> codestral Continue talks with ollama via a local REST API