ollama

   1 cd ~
   2 curl -fsSL https://ollama.com/install.sh | sh
   3 # >>> The Ollama API is now available at 127.0.0.1:11434.
   4 curl localhost:11434 
   5 # Ollama is running
   6 ollama run llama3.2:1b
   7 ollama show llama3.2:1b
   8 ollama list
   9 ollama ps
  10 ollama serve # start server 
  11 ollama stop llama3.2:1b
  12 ollama rm llama3.2:1b

Setup proxy

   1 sudo nano  /etc/systemd/system/ollama.service

   1 [Unit]
   2 Description=Ollama Service
   3 After=network-online.target
   4 
   5 [Service]
   6 ExecStart=/usr/local/bin/ollama serve
   7 User=ollama
   8 Group=ollama
   9 Restart=always
  10 RestartSec=3
  11 Environment="PATH=/home/vagrant/.dotnet/tools:/home/vagrant/dotnetcore9:/usr/local/bin:/usr/bin:/bin:/home/vagrant/jdk-17.0.7+7/bin/:/home/vagrant/gradle-8.1/bin/"
  12 Environment="HTTPS_PROXY=http://192.168.0.123:3128/"
  13 Environment="HTTP_PROXY=http://192.168.0.123:3128/"
  14 
  15 [Install]
  16 WantedBy=default.target

   1 sudo systemctl daemon-reload
   2 sudo systemctl restart ollama
   3 systemctl show ollama
   4 ollama run llama3.2:1b

Chatbots - LLM - AI accessible via browser

Stuff

Does a local run of ollama spends tokens ? No

LLMs might generate different versions because they are probabilistic. They predict the next word/token based on probabilities.

The reason for results variation is called temperature.

High temperature [0.8 - 1.0] the model takes more risks that leads to creative code.

Low temperature [0.0 - 0.2] becomes very focused and predictable. At temperature 0 (zero) it will choose the most likely word/token.

Seed value, in ollama, if not specified a random one is generated for each request.

If you want the same code every time set the same seed and temperature 0.

RAG, retrieved augmented generation. Create a domain-specialized assistant.

Tools like copilot chat looks at the .md files and source code and perform a RAG workflow.

RAG.

A specific project is the domain. Project memory/domain (markdown + code)

Local LLM runner (ollama) VS code extension continue RAM local models

Continue talks with ollama via a local REST API

A chatbot just talks. An agent can do things.

Continue is agentic. It has perception (context awareness). It sees the prompt, sees open files, the git git history and terminal errors.

Tool use: can perform actions on my behalf.

Reasoning: if i prompt "fix this error", looks at the stack trace, reasons which file is causing the bug and proposes a fix.

Continue is a "human in the loop" agent, asks my command/prompt, performs complex task, ask for approval.

Ollama is the brain. Continue (private coding agent) is the agentic body lives inside vscode .

Private JARVIS-style assistant

Openclaw + ollama

   1 ollama run llama3:8b
   2 # http://localhost:11434
   3 cd ~
   4 git clone https://github.com/openclaw/OpenClaw.git 
   5 cd OpenClaw
   6 npm i 
   7 # .env file 
   8 AI_PROVIDER=ollama
   9 OLLAMA_BASE_URL=http://localhost:11434
  10 MODEL=llama3:8b
  11 API_KEY=ollama
  12 OPENCLAW_WORKSPACE_DIR=/tmp
  13 # launch openclaw
  14 npm start 
  15 # http://localhost:3000
  16 

TextToSPeech (TTS) or SpeechToText (STT)

Private alexa

MCP Servers act as the workshop full of tools (the hammers, wrenches, and drills).

RAG, looks up information in a book Skill, a specific task to be done with that info

A skill is hybrid, it's part human language (instructions) and part code (logic and parameters).

A skill is like a job description:

A skill turns a complex logic into a reusable template

"Persona" in systems instructions tells which parts if the model/brain it is allowed to be used:

Persona also sets the output target audience. Persona is the speaker, target audience the listener.

You are X. Your task is to xyz. You must follow these rules/constraints asdf.

A model is a giant library and the persona tells which shelves to visit. Shelves might also be functional clusters, features or circuits.

Ask LLM to write a skill for itself is called meta-prompting.

Autonomous role discovery, let the LLM analyze the task and choose the most efficient speaker and listener to get the task done.

Meta-prompt template:

A skill represents a single expert. A toolchain represents a whole department.

Meta-prompt

Convert this task into a x.md skill: I want to update the version in a pom.xml. It must use the current date and time in ISO8601 format with only numbers.

Convert this task into a y.md skill: I wnat to bump up the minor version in Java pom files of this project. It must add 1 to the minor version.

In VSCode copilot to run the skill use #SkillName .

ollama (last edited 2026-08-29 16:15:01 by vitor)