ollama

   1 cd ~
   2 curl -fsSL https://ollama.com/install.sh | sh
   3 # >>> The Ollama API is now available at 127.0.0.1:11434.
   4 curl localhost:11434 
   5 # Ollama is running
   6 ollama run llama3.2:1b
   7 ollama show llama3.2:1b
   8 ollama list
   9 ollama ps
  10 ollama serve # start server 
  11 ollama stop llama3.2:1b
  12 ollama rm llama3.2:1b

Setup proxy

   1 sudo nano  /etc/systemd/system/ollama.service

   1 [Unit]
   2 Description=Ollama Service
   3 After=network-online.target
   4 
   5 [Service]
   6 ExecStart=/usr/local/bin/ollama serve
   7 User=ollama
   8 Group=ollama
   9 Restart=always
  10 RestartSec=3
  11 Environment="PATH=/home/vagrant/.dotnet/tools:/home/vagrant/dotnetcore9:/usr/local/bin:/usr/bin:/bin:/home/vagrant/jdk-17.0.7+7/bin/:/home/vagrant/gradle-8.1/bin/"
  12 Environment="HTTPS_PROXY=http://192.168.0.123:3128/"
  13 Environment="HTTP_PROXY=http://192.168.0.123:3128/"
  14 
  15 [Install]
  16 WantedBy=default.target

   1 sudo systemctl daemon-reload
   2 sudo systemctl restart ollama
   3 systemctl show ollama
   4 ollama run llama3.2:1b

Chatbots - LLM - AI accessible via browser

Stuff

Does a local run of ollama spends tokens ? No

LLMs might generate different versions because they are probabilistic. They predict the next word/token based on probabilities.

The reason for results variation is called temperature.

High temperature [0.8 - 1.0] the model takes more risks that leads to creative code.

Low temperature [0.0 - 0.2] becomes very focused and predictable. At temperature 0 (zero) it will choose the most likely word/token.

Seed value, in ollama, if not specified a random one is generated for each request.

If you want the same code every time set the same seed and temperature 0.

RAG, retrieved augmented generation. Create a domain-specialized assistant.

Tools like copilot chat looks at the .md files and source code and perform a RAG workflow.

RAG.

A specific project is the domain. Project memory/domain (markdown + code)

Local LLM runner (ollama) VS code extension continue RAM local models

Continue talks with ollama via a local REST API

ollama (last edited 2026-08-29 09:22:58 by vitor)