Ollama Explained: How to Run Large Language Models Locally for Free

Most people meet AI through the cloud. You type a prompt into a chatbot, or a developer calls a hosted API, and somewhere in a data center a remote machine does the work. That works well, but it comes with recurring costs, usage limits, and the fact that your data leaves your machine.

There is an open-source alternative. Ollama lets you download and run large language models (LLMs) directly on your own computer. It is free, it works offline once a model is downloaded, and it keeps your data on your own hardware.

This guide explains what Ollama is, how it works, how to install it, and how to use it from the terminal and from your own code.

What Is Ollama?

Ollama is a free, open-source tool for downloading, running, and managing large language models locally. It runs on Windows, macOS, and Linux.

Before tools like this existed, running a model yourself meant downloading model weights from repositories such as Hugging Face and configuring an inference environment by hand. Ollama simplifies this into a single command. You can think of it ollama run as a package manager for AI: you name a model, and Ollama fetches it, sets it up, and starts it.

Ollama also provides a local HTTP server, so any application on your machine can use the models, not just the terminal.

Why Run LLMs Locally?

Running models on your own hardware has some clear advantages.

Lower cost. There are no per-token charges or subscriptions. Once a model is downloaded, using it costs only your own electricity and hardware.

Privacy and security. Prompts and responses never leave your machine. This matters for organizations that must keep customer data inside a secure environment, and for anyone working with sensitive documents.

Low latency. With no network round trip, responses begin quickly. Speed still depends on your hardware and the size of the model.

Offline and limited-connectivity use. After the initial download, a local model works without internet access, which suits edge and IoT scenarios.

Developer flexibility. You can integrate models into your own applications without depending on a third-party service.

How Ollama Works

When Ollama runs, it starts a background server on your machine. By default, this is a REST API at http://localhost:11434.

Everything goes through this server. When you chat with a model in the terminal, the CLI sends your prompt to it. When an application built with a framework like LangChain wants a model response, it sends an HTTP POST request to the same address.

This design means Ollama handles the heavy lifting of loading and running the model, while your application treats the model as a simple API. You can also run Ollama on another machine and connect to it remotely, or plug in interfaces such as Open WebUI to build a document-question-answering (RAG) setup on top of it.

How to Install Ollama

  1. Go to ollama.com and click Download.
  2. Choose your operating system (Windows, macOS, or Linux).
  3. Run the installer.

After installation, you can start Ollama in two ways. You can open the desktop application, which starts the background server silently (you won’t see a window, only a tray icon). Or you can open a terminal and type ollama. If a list of commands appears, the installation worked.

Running Your First Model

To run a model, use ollama run followed by the model name:

ollama run llama3.2

If the model isn’t on your machine yet, Ollama downloads it first. Once it is ready, you get an interactive prompt where you can start chatting. To leave the chat, type /bye.

You can run as many models as you like and switch between them by changing the name:

ollama run mistral

Essential Ollama Commands

Command What it does
ollama run <model>                      Downloads (if needed) and starts a model
ollama list Shows all models installed on your machine
ollama rm <model> Removes a model
ollama cp <source> <copy> Copies a model
ollama create <name> -f <file> Builds a custom model from a Modelfile
ollama serve Manually starts the Ollama server

Typing ollama on its own shows the full list of available commands.

Types of Models You Can Use

Ollama’s library includes hundreds of models, which fall into a few broad categories:

  • Language models work with text, either in a conversational style or in an instruction-following style for question answering.
  • Multimodal models can also analyze images, for example,e describing what is happening in a picture.
  • Embedding models convert documents such as PDFs into a form suitable for a vector database, so you can ask questions about your own data.
  • Tool-calling models are fine-tuned to call functions, APIs, and services, which makes them useful for agent-style applications.

Popular families include Meta’s Llama series, Mistral, IBM’s Granite (an enterprise-oriented model that works with RAG and agent workflows), and DeepSeek. Reasoning models, which work through a problem step by step before answering, are also increasingly common. You can browse the full catalog at ollama.com/library.

How to Choose the Right Model and Check Your Hardware

Because models run locally, you need enough disk space to store them and enough RAM to load them. Sizes vary enormously. Some models are only a few gigabytes, while the largest, such as Llama 3.1 with 405 billion parameters, require over 200 GB of storage and far more memory than a typical computer has, even one with 64 GB of RAM.

A common rule of thumb from the Ollama documentation is:

  • About 8 GB of RAM for 7B-parameter models
  • About 16 GB for 13B models
  • About 32 GB for 33B models

Many models are also quantized, meaning compressed so they use fewer resources with only a small loss in quality. This is what makes running capable models on ordinary hardware realistic.

If you’re just starting, pick a smaller model first. You can always move up once you know what your machine can handle.

Using the Ollama HTTP API

Everything you can do in the terminal can also be done through the API. This is what makes Ollama useful for developers: any language that can send HTTP requests can use your local models.

If the desktop app is running, the API is already available. If not, start it manually:

ollama serve

Here is a Python example that sends a chat request and streams the response as it is generated. Install the requests library first with pip install requests.

import requests
import json

url = "http://localhost:11434/api/chat"

payload = {
    "model": "mistral",
    "messages": [
        {"role": "user", "content": "What is Python?"}
    ]
}

with requests.post(url, json=payload, stream=True) as response:
    for line in response.iter_lines():
        if line:
            data = json.loads(line)
            print(data.get("message", {}).get("content", ""), end="", flush=True)

Streaming lets you display the answer word by word instead of waiting for the full response. Besides /api/chatthe API has other endpoints, including ones for generating text and for managing models.

Using the Official Python Library

If you’d rather not handle raw HTTP requests, Ollama offers an official Python package (a JavaScript one is available too). Install it with:

pip install ollama

Then:

from ollama import Client

client = Client(host="http://localhost:11434")

response = client.generate(
    model="mistral",
    prompt="Explain what a REST API is in two sentences."
)

print(response["response"])

This does the same job in a few lines. Any model you have installed, including custom ones, can be used by name.

Customizing Models with a Modelfile

A Modelfile lets you build your own version of a model. The idea is similar to how a Dockerfile describes a container: you start from a base and add your own configuration. You can import models from sources like Hugging Face, or take an existing model and change its parameters and system prompt.

Create a plain text file named. Modelfile (no extension) with contents like this:

FROM llama3.2

PARAMETER temperature 1

SYSTEM
You are Mario from Super Mario Bros. Answer as Mario, the assistant, only.
  • FROM Sets the base model.
  • PARAMETER Adjusts settings such as temperature, which controls how creative or predictable the output is.
  • SYSTEM Defines instructions the model follows in every conversation.

Then, from the folder containing the file, create and run the model:

ollama create mario -f ./Modelfile
ollama run mario

Your custom model now behaves according to the system prompt, and you can call it from code exactly like any other model by using mario as the model name. To delete it later, run ollama rm mario.

This is a practical way to build specialized assistants, such as a customer-support bot, a code reviewer, or a tutor with a fixed personality, without training anything.

Common Use Cases

  • Private chat assistants for confidential documents
  • Coding help without sending source code to an external service
  • Prototyping AI features in apps without API bills
  • Document question answering using embedding models and RAG
  • Offline AI tools for environments with poor connectivity
  • Testing and comparing different open-source models

Limitations to Keep in Mind

Local models are convenient, but they involve trade-offs. The largest, most capable models need serious hardware. Smaller local models may not match the quality of the biggest hosted models on complex tasks. Speed depends on your CPU, GPU, and memory. Ollama is also not the only tool in this space, so it’s worth comparing it with alternatives if you have specific needs.

Frequently Asked Questions

Is Ollama really free?
Yes. Ollama is open source and free to use, and the models in its library are openly available, though each model has its own license terms.

Does Ollama work offline?
Yes, after you have downloaded a model. Downloading requires an internet connection; running does not.

Do I need a GPU?
Not strictly. Smaller models can run on a CPU, but a compatible GPU generally makes responses noticeably faster.

What port does Ollama use?
Port 11434 on localhost by default.

How do I see which models I have installed?
Run ollama list.

Can I use Ollama from my own application?
Yes. Send requests to the local REST API, or use the official Python or JavaScript libraries.

Is my data sent to the cloud?
No. Prompts and responses are processed on your own machine.

Conclusion

Ollama makes running AI models locally straightforward. One command downloads and starts a model, a local API lets you plug it into any application, and Modelfiles let you shape models to your own needs. Whether your goal is to cut cloud costs, keep sensitive data private, or simply experiment with open-source AI, it is a solid place to start. Install it, pull a small model, and try it for yourself.

Leave a Reply

Your email address will not be published. Required fields are marked *