← Back to blog

AI · Ollama · Local LLM · macOS · Developer Tools

Run an LLM Locally on Your Mac with Ollama

A beginner's guide to installing Ollama on macOS, running Llama 3.2 in the terminal, and calling its local API with curl, using tested examples.

Ravi Kumar//9 min read

You can run an LLM on your Mac with Ollama: install the app, download a model, and send it a prompt. This tutorial takes you from installation to a terminal conversation, then makes a request to the same model using curl.

The setup below was tested on a MacBook Pro with an M1 Pro chip and 16 GB of unified memory. Both the terminal command and local API worked. One answer was correct. Another confidently misunderstood the question. We'll look at both, because getting a model to run and getting a useful answer are separate things.

What does running an LLM locally mean?

An LLM, or large language model, generates text from the prompt you give it. Running it locally means the computation that produces the answer happens on your computer.

There are two pieces in this setup:

  • Ollama downloads, loads, and runs models. It also exposes an HTTP API.
  • Llama 3.2 is the model we'll run. Its downloaded files contain the learned parameters used to generate responses.

Think of the path like this:

 
Your prompt → Ollama → Llama 3.2 running on your Mac → response
text

You need an internet connection to install Ollama and download the model. Once downloaded, this model can generate answers without an internet connection. That doesn't give it live knowledge of the web.

Ollama also supports cloud models. This walkthrough uses the local llama3.2:3b model and the API on localhost; no cloud account or API key is needed. See the Ollama quickstart for the distinction.

Check your Mac before downloading

Ollama currently requires macOS Sonoma, version 14, or newer. Apple M-series Macs support CPU and GPU execution; Intel Macs use CPU execution. These requirements come from the macOS documentation.

You can check your macOS version from Terminal:

sw_vers -productVersion
bash

Open Apple menu → About This Mac to check your chip and memory.

The test machine for this post had:

ItemTested setup
ComputerMacBook Pro
ChipApple M1 Pro
Memory16 GB unified memory
macOS27.0.1
Ollama0.40.1

We'll use llama3.2:3b. The model listing gives its download size as roughly 2 GB. Leave additional free space for Ollama and other files.

The 3b tag refers to approximately three billion model parameters. It doesn't mean the model needs exactly 3 GB of memory. Download size and memory use are different, and memory requirements also depend on how much context the model processes.

Install Ollama on macOS

Download Ollama from the official macOS download page. For the manual app installation:

  1. Open the downloaded disk image.
  2. Drag Ollama into Applications.
  3. Open Ollama.
  4. If prompted, allow it to add the ollama terminal command.

This is the preferred manual installation described in the macOS documentation. Open a new Terminal window and check:

ollama --version
bash

The tested installation reported:

ollama version is 0.40.1
text

Your version may be newer.

The test setup used Ollama's official installer script. It installed the app successfully but couldn't create the command link in /usr/local/bin without an administrator password. The link was added to an existing writable Homebrew bin directory instead. If you hit command not found, check the app's CLI setup before assuming the installation failed.

Ollama needs a running server to load models and answer requests. If the app hasn't started it, run:

ollama serve
bash

Keep that terminal open and use a second terminal for the remaining commands. This was the server startup method used during testing. If you see an error saying the address is already in use, Ollama may already be serving; you don't need a second instance.

Download Llama 3.2 and start a conversation

Download the model:

ollama pull llama3.2:3b
bash

This saves the model on disk. The first download can take a few minutes, depending on your connection. Then start a conversation:

ollama run llama3.2:3b
bash

Ollama loads the model and opens an interactive prompt. Try:

Explain what an HTTP request is in two short sentences.
text

Then ask a follow-up:

Give me an example using a web browser.
text

These are suggested prompts; the exact wording of the response will vary. Type /bye to leave the conversation. The quickstart documents this interactive flow.

You can also send a prompt directly from the command line. During testing, this command completed successfully:

ollama run llama3.2:3b 'Explain what a local LLM is in two short sentences.'
bash

The answer was wrong. It interpreted “local” as a geographic region and described a model trained for regional language and culture.

In this tutorial, local means the model runs on your computer. The response sounded confident, but it explained a different concept. Keep that example in mind when checking the model's output: a successful command proves that inference worked, not that the answer is accurate.

Call the same model with curl

The terminal is convenient for chatting. The API lets another program send a prompt and read the answer.

With the server running and the model downloaded, paste this into Terminal:

curl http://localhost:11434/api/chat \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "llama3.2:3b",
    "messages": [
      {
        "role": "user",
        "content": "Explain this Java function to a beginner in at most four sentences.\nMention what happens when the list is empty.\n\nstatic int sum(List<Integer> numbers) {\n    int total = 0;\n    for (int number : numbers) {\n        total += number;\n    }\n    return total;\n}"
      }
    ],
    "stream": false
  }'
bash

This sends an HTTP POST request to Ollama on your own Mac. curl selects POST because we're supplying a request body with -d.

The request has three important fields:

  • model selects the downloaded model.
  • messages contains the conversation. Here, it contains one user message.
  • stream: false asks for a complete JSON response instead of incremental streamed responses.

These fields are documented in the chat API reference. The \n sequences represent line breaks inside the JSON string.

Here's the function we're asking the model to explain:

static int sum(List<Integer> numbers) {
    int total = 0;
    for (int number : numbers) {
        total += number;
    }
    return total;
}
java

You don't need Java installed to try this. We're sending the code as text and asking the model to explain what it does.

The answer comes back in the JSON response's message.content field. Here's what the model returned during testing:

 
This Java function, `sum`, calculates the total of all numbers in a list. It starts with a variable `total` set to 0, then goes through each number in the list one by one, adding each number to the total. If the list is empty, the function will return 0 because `total` was never changed.
text

It got the basics right: add the numbers, return the total, and return zero if the list is empty. It also kept the explanation to three sentences, within the four-sentence limit we asked for. That's useful for this example. For more complex code, you'll still need to check the explanation against what the code actually does.

Ollama reported a total duration of about 2.13 seconds for this request, with the model already loaded. That is one observation on this machine, not an estimate for every prompt or Mac.

If you want a follow-up conversation through the API, include the previous user and assistant messages in the next request's messages array. The model doesn't get the previous exchange just because you call the same endpoint again.

Downloaded models and running models are different

List the models saved on disk:

ollama ls
bash

During testing, llama3.2:3b appeared with a size of 2.0 GB.

To see which models are currently loaded, run:

ollama ps
bash

The tested model showed 2.5 GB, 100% GPU, and a 4096-token context while loaded. Those are Ollama's reported values for the running model, not a measurement of total system memory use.

When you're done, unload it:

ollama stop llama3.2:3b
bash

The downloaded files remain on disk, so you can run it again without downloading it again. If you want to delete those files too, use:

ollama rm llama3.2:3b
bash

You don't need to remove the model after completing the tutorial. These commands are covered in the CLI reference.

Common first-run problems

The terminal says command not found

Open Ollama and complete its terminal-command setup. Then open a new terminal and retry ollama --version. The test installation hit a permissions issue while creating the command link; the app itself was already installed.

curl cannot connect to localhost

The local server may not be running. Open Ollama or run ollama serve in a separate terminal, then retry the request. This happened during testing before the server was started explicitly.

The API says the model was not found

Run ollama ls and compare the installed name with the request's model field. Download the tutorial's model with ollama pull llama3.2:3b if it's missing.

Responses are slow

The first request may include loading the model. Larger models and longer prompts can also need more memory and computation. Start with the small example here, check ollama ps, and close memory-heavy applications before trying a larger model.

What to try next

You now have a model you can chat with and call over HTTP. Try replacing the Java function with a short function from your own work, or ask the model to summarize a paragraph. Keep the task small enough that you can check the answer yourself.

The next post will connect a local Ollama model to Spring AI. For now, the useful foundation is already in place: a downloaded model, a working local endpoint, and a clear way to inspect what it returns.

Further reading