ollama
1 cd ~
2 curl -fsSL https://ollama.com/install.sh | sh
3 # >>> The Ollama API is now available at 127.0.0.1:11434.
4 curl localhost:11434
5 # Ollama is running
6 ollama run llama3.2:1b
7 ollama show llama3.2:1b
8 ollama list
9 ollama ps
10 ollama serve # start server
11 ollama stop llama3.2:1b
12 ollama rm llama3.2:1b
Setup proxy
1 sudo nano /etc/systemd/system/ollama.service
1 [Unit]
2 Description=Ollama Service
3 After=network-online.target
4
5 [Service]
6 ExecStart=/usr/local/bin/ollama serve
7 User=ollama
8 Group=ollama
9 Restart=always
10 RestartSec=3
11 Environment="PATH=/home/vagrant/.dotnet/tools:/home/vagrant/dotnetcore9:/usr/local/bin:/usr/bin:/bin:/home/vagrant/jdk-17.0.7+7/bin/:/home/vagrant/gradle-8.1/bin/"
12 Environment="HTTPS_PROXY=http://192.168.0.123:3128/"
13 Environment="HTTP_PROXY=http://192.168.0.123:3128/"
14
15 [Install]
16 WantedBy=default.target
Chatbots - LLM - AI accessible via browser
Gemini prompt - https://gemini.google.com/
Copilot prompt - https://copilot.microsoft.com/
Grok prompt - https://grok.com/
ChatGPT prompt - https://chatgpt.com/
Stuff
Does a local run of ollama spends tokens ? No
LLMs might generate different versions because they are probabilistic. They predict the next word/token based on probabilities.
The reason for results variation is called temperature.
High temperature [0.8 - 1.0] the model takes more risks that leads to creative code.
Low temperature [0.0 - 0.2] becomes very focused and predictable. At temperature 0 (zero) it will choose the most likely word/token.
Seed value, in ollama, if not specified a random one is generated for each request.
If you want the same code every time set the same seed and temperature 0.
RAG, retrieved augmented generation. Create a domain-specialized assistant.
Tools like copilot chat looks at the .md files and source code and perform a RAG workflow.
RAG.
- R retrieval, the tool searches for the relevant code or markdown text
- A augmentation, takes those snippets and pastes them into the prompt and send it to the model
- G the model reads the code/docs and generates an answer based only on the specific context
A specific project is the domain. Project memory/domain (markdown + code)
Local LLM runner (ollama) VS code extension continue RAM local models
8GB RAM -> llama3:8b
16GB - 32GB -> codestral
Continue talks with ollama via a local REST API
A chatbot just talks. An agent can do things.
Continue is agentic. It has perception (context awareness). It sees the prompt, sees open files, the git git history and terminal errors.
Tool use: can perform actions on my behalf.
Reasoning: if i prompt "fix this error", looks at the stack trace, reasons which file is causing the bug and proposes a fix.
Continue is a "human in the loop" agent, asks my command/prompt, performs complex task, ask for approval.
Ollama is the brain. Continue (private coding agent) is the agentic body lives inside vscode .
Private JARVIS-style assistant
Openclaw + ollama
