Quick answer: how do you run an LLM locally?
Install LM Studio for a graphical setup or Ollama for a terminal workflow, download a model that fits your computer, then chat with it on your own machine. Start with a small task, such as summarizing a short document, before deciding whether the speed and quality are good enough for daily work.
Best first step: Check your RAM and operating system against the app’s requirements. A local model is not automatically a replacement for a hosted assistant, but it can be useful for offline experiments and sensitive drafts.
A local large language model (LLM) runs on your computer instead of sending each prompt to a hosted AI service. You download the model once, load it in an app, and chat locally. The setup is easier than it used to be, but model size, memory, speed, and answer quality still depend on your hardware.
For a first experiment, use a desktop app and a task you can check yourself. This guide uses LM Studio and Ollama, two ways to run local models. If your goal is building an app with a local coding assistant, see our Dyad review for that separate workflow.
What local AI is good for, and where it falls short
Local models can be useful when you want to experiment without a per-request API bill, work with downloaded models while offline, or process a document on-device. LM Studio’s offline documentation says downloaded models, document chat, and a local server can work without internet; searching for and downloading models still requires a connection.
The trade-off is capability and convenience. A small model may be slower or less reliable than a hosted service on difficult questions, and it will not automatically have current web information. Local use also shifts the work to your CPU, GPU, memory, storage, and battery. Keep your expectations narrow: a local assistant can be handy for drafts, simple summaries, and private experiments, but it may not be the best tool for every job.
Check your computer before downloading a model
Start with the current LM Studio system requirements, then check free disk space. Its documentation recommends 16 GB of RAM on Apple Silicon Macs and Windows PCs, with smaller models and modest context sizes as the practical route on an 8 GB Mac. Windows x64 requires AVX2 support, and LM Studio recommends at least 4 GB of dedicated GPU memory. These are product recommendations, not a promise that every model will run well.
Model downloads often come in different quantized versions. Quantization reduces file size while trading away some fidelity; LM Studio recommends choosing a 4-bit option or higher when your machine can handle it. Begin with a smaller download, close memory-heavy apps, and only move up if the first model leaves enough headroom. LM Studio explains the download options here.
Choose a simple way to run the model
| Tool | Good starting point if you want | First action |
|---|---|---|
| LM Studio | A graphical model browser, chat, and document workflow | Open Discover, download a model, then load it in Chat |
| Ollama | A terminal command or local model service for other tools | Install Ollama and run a model from the command line |
Option 1: LM Studio for a graphical setup
Download LM Studio for your operating system, open the Discover tab, and search for a model. Select a smaller quantized option that fits your available memory. When the download finishes, load it in the Chat tab and send a test prompt. LM Studio’s getting-started guide walks through the same basic sequence.
To try a private document task, open a non-sensitive sample file and ask for a short summary plus three questions the document answers. Check each answer against the source. The goal is not to prove the model is perfect; it is to see whether this setup is accurate and responsive enough for your real use.
Option 2: Ollama for a terminal workflow
Install Ollama for macOS, Windows, or Linux, open a terminal, and run the sample command from the official quickstart:
ollama run gemma4The first run downloads the model, so leave time and storage for that step. Type a prompt in the chat and use `/bye` to exit. If you are new to the terminal, LM Studio’s graphical path may feel simpler; Ollama is useful when you want a repeatable command or plan to connect a local model to another tool.
Test it before trusting it with important work
Use one task you already know how to evaluate. For example, give the model a short product FAQ and ask it to list the cancellation policy, summarize the setup steps, and identify anything the FAQ does not answer. Compare the result with the source. Note how long it takes, whether the answer stays grounded in the file, and whether you had to repeat the prompt.
If it struggles, try a smaller context, a different model, or a shorter prompt before assuming you need a more powerful computer. For critical decisions, verify the answer yourself. A local model can confidently invent details just like a hosted model.
Privacy: local does not mean every workflow is offline
A downloaded model can answer locally, but model discovery, downloads, and app updates need an internet connection. Some apps also offer cloud models or integrations, which are different from local inference. Confirm the selected model and connection mode before sharing confidential material. If you enable an API server, keep its access limited to devices and users you trust; a local-network server is not the same as an offline chat window.
When to use local AI instead of a hosted model
Use local AI when offline access, experimentation, or keeping a specific file on-device matters more than top-end capability and convenience. Use a hosted model when you need stronger reasoning, current web access, or do not want to manage local memory and downloads. Many builders can keep both options: a local model for low-risk, repeatable work and a hosted assistant for tasks where quality or current information matters.
If you are choosing a model for software work rather than general offline chat, compare the trade-offs in our guide to AI coding models.
Bottom line
Start with one small model and one testable task. LM Studio is the friendlier first stop if you want a visual interface; Ollama is a good fit if you prefer commands and integrations. Check requirements before downloading, verify answers against the source, and keep a hosted model available for work your computer cannot handle well.
Frequently Asked Questions
Can I run an LLM locally on a normal laptop?
Often, yes, if you choose a model that fits your computer. LM Studio recommends 16 GB of RAM for common supported setups and says 8 GB Macs may work with smaller models and modest context sizes. Check the current requirements and expect performance to vary by model and hardware.
What is the easiest way to run an LLM locally?
LM Studio is a straightforward graphical starting point: download the app, choose a model in Discover, then load it in Chat. Ollama is another simple option if you are comfortable running a terminal command such as `ollama run gemma4`.
Can I use a local LLM without internet?
Yes, after you have downloaded the app’s runtime and model files. LM Studio documents offline chat, document chat, and local-server use; model search, downloads, and updates still require connectivity.
Is running an LLM locally private?
A model running locally can process prompts and documents on your computer, but confirm that you selected a local model rather than a cloud option or external integration. Also check what data the app sends for updates or model discovery before using sensitive material.
Is local AI better than ChatGPT or Claude?
Not for every task. Local models can be useful for offline work and experimentation, while hosted models can offer stronger capabilities and web access without using your computer’s memory. Test both on a task you can evaluate and choose based on accuracy, speed, privacy needs, and effort.
Advertiser disclosure: some links on this website are affiliate links, meaning No Code MBA may make a commission if you click through and purchase.