🇯🇵 Tokyo is live! 🚀 Launch your VPS and enjoy 2 months off — use code KONNICHIWA50 🎉 Get Started Today →

Which Local LLM Can You Actually Run Per VPS Spec? A Model-to-CPU/RAM Benchmark Table

VPS server with terminal and processor elements representing a local LLM benchmark

A local language model can load successfully and still be a poor fit for a VPS. The model file must fit alongside the operating system, Ollama, the context cache, and every other service. CPU generation speed must also match the workload. A careful benchmark therefore answers two separate questions: does it fit safely? and is it fast enough for this job?

This guide gives you a reproducible CPU-only test for the current public VPS.us KVM tiers. It uses exact model tags and runtime counters instead of fabricated benchmark results. Run the same prompt several times on the exact server you plan to use, record the evidence, and size from your own results.

⚡ Spin up a Premium VPS in 2 minutes
17 locations worldwide
NVMe  ·  Unmetered 1 Gbps  ·  Full root access  ·  From $10/mo

Separate Model Capacity From Inference Performance

Model capacity is a memory question. The registry artifact is only the starting point: runtime state, the context cache, Ubuntu, and background services also need RAM. Ollama documents that longer context increases memory use and that parallel requests multiply context allocation. A download smaller than total RAM is not proof of a reliable fit.

Performance is a measurement question. CPU generation depends on the exact processor, instruction set, model architecture, quantization, context, prompt, and competing load. Do not copy tokens-per-second figures from another host and present them as a promise for yours. Benchmark the final tag on the final VPS image, then repeat after material runtime or model changes.

Screen Current Model Artifacts Against Available RAM

Model weight blocks allocated within available VPS memory alongside reserved system headroom

Start with the exact artifact size listed in the Ollama library. The following examples were checked on September 22, 2026. They are screening values, not measured working-memory requirements or performance guarantees; each tag links to its current registry entry.

Exact Ollama tagPublished artifact sizeFirst capacity check
gemma3:270m292 MBSmall smoke tests and simple extraction
gemma3:1b815 MBLow-memory experimentation
qwen3:1.7b-q4_K_M1.4 GBCompact text tasks with controlled context
qwen3:4b-instruct-2507-q4_K_M2.5 GBInstruction work on a node with several GB of free RAM
gemma3:4b-it-q4_K_M3.3 GBMultimodal-capable model with more memory pressure
qwen3:8b-q4_K_M5.2 GBHigher-memory CPU experimentation

Quantization reduces storage and memory pressure by representing weights at lower precision, but it can also affect quality. The VPS.us self-hosted AI chatbot guide explains the wider stack around the model. Keep enough free space for more than model weights: a chatbot may also need an API layer, logs, retrieval services, and a database.

Prepare a Clean CPU-Only Test Node

Use the same Ubuntu release and VPS tier that will host the workload. Install Ollama using its official Linux instructions, and make sure curl and jq are available. Stop unrelated jobs, but keep normal security and monitoring services running. Record the baseline before testing.

ollama --version
uname -r
lsb_release -ds
lscpu
free -h
swapon --show
nproc
uptime
ollama list

If the load average is already high or free memory is unexpectedly low, fix that condition before testing. Avoid using swap as a way to declare a model successful. Swap may prevent an immediate out-of-memory failure, but it replaces RAM access with storage I/O and can make interactive latency unacceptable. The VPS server optimization guide covers the broader discipline of measuring before tuning.

Run a Reproducible Ollama Benchmark

VPS test environment with storage, tools, and server resources for repeatable LLM benchmarks

Pull one exact tag at a time. Tags may change, so retain the complete model digest from the local model-list API with your results. Keep the prompt, context, output limit, thinking setting, and concurrency constant across tiers. The example disables thinking and caps output at 256 tokens; benchmark your real workload separately if it requires different settings.

ollama pull qwen3:1.7b-q4_K_M
ollama list
curl -fSs http://127.0.0.1:11434/api/tags > model-manifest.json
# Warm the model before measured runs.
curl -fSs http://127.0.0.1:11434/api/generate \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3:1.7b-q4_K_M","prompt":"Summarize three benefits of TLS 1.3.","stream":false,"think":false,"keep_alive":"10m","options":{"temperature":0,"num_ctx":4096,"num_predict":256}}' \
  > warmup.json

The Ollama generate API returns timing and token counters in the final response. Store the complete JSON so you can distinguish model-load time, prompt evaluation, and generation. After the warm-up, run ollama ps and verify that PROCESSOR reports 100% CPU.

for run in 1 2 3; do
  curl -fSs http://127.0.0.1:11434/api/generate \
    -H 'Content-Type: application/json' \
    -d '{"model":"qwen3:1.7b-q4_K_M","prompt":"Summarize three benefits of TLS 1.3.","stream":false,"think":false,"keep_alive":"10m","options":{"temperature":0,"num_ctx":4096,"num_predict":256}}' \
    > "run-$run.json"
done

Run the loop from an idle node, then repeat it while the real supporting services are active. If the production workload requires simultaneous users, perform a separate concurrency test; do not multiply a single-request rate and call it capacity.

Measure Throughput, Latency, and Peak Memory

Ollama reports durations in nanoseconds. Generation throughput is eval_count × 1,000,000,000 ÷ eval_duration. Prompt throughput uses the uncached prompt counters; report it as unavailable when its duration is zero. Keep cold-model loading separate from warm generation and record cache effects rather than treating a repeated prompt as a fresh prompt-processing test.

jq '
  def rate($count; $duration):
    if ($duration // 0) > 0 then $count * 1000000000 / $duration else null end;
  {
    total_seconds: (.total_duration / 1000000000),
    load_seconds: (.load_duration / 1000000000),
    uncached_prompt_tokens_per_second: rate((.prompt_eval_count - (.prompt_eval_cached_count // 0)); .prompt_eval_duration),
    generation_tokens_per_second: rate(.eval_count; .eval_duration),
    prompt_tokens: .prompt_eval_count,
    cached_prompt_tokens: (.prompt_eval_cached_count // 0),
    generated_tokens: .eval_count
  }' run-1.json

Monitor memory throughout the run and keep the samples, not only a before-and-after reading. Record the highest observed use and any swapping; periodic samples can miss brief peaks, so do not present them as an exact process high-water mark. A request that completes after heavy swapping is memory-constrained rather than a clean pass.

ollama ps
free -h
vmstat 1 | tee memory-samples.txt
# Stop vmstat with Ctrl-C after the benchmark.
journalctl -u ollama --since '10 minutes ago' --no-pager

Interpret Results Without Inventing Universal Thresholds

There is no honest universal boundary where a model becomes “interactive.” A short internal classification task and a streaming chat interface have different latency budgets. Define the target before reading the results.

  • Single-user chat: judge time to first visible output, sustained generation, and the longest response users will tolerate.
  • Background automation: judge completed jobs per hour, queue depth, and recovery after failure.
  • Document processing: include prompt-evaluation time because long inputs can dominate the run.
  • Concurrent service: test the real parallel request count and context length while watching RAM.

Use the median of repeated warm runs and retain the slowest result. Investigate large variation rather than averaging it away. CPU steal, another workload, model reloads, or memory pressure can make one run unrepresentative.

Match the Current VPS.us Tiers to a Test Plan

The current public VPS.us LLM hosting page lists four KVM tiers from 1 GB to 8 GB of RAM. The table below is deliberately conservative. “Candidate” means the artifact is worth testing with a controlled context; it does not guarantee a clean load, a particular throughput, or production suitability.

Current public tierPublished resourcesPractical first testImportant limit
KVM1-US1 vCore, 1 GB RAM, 20 GB NVMegemma3:270m smoke testLittle room for a general-purpose LLM service
KVM2-US2 vCores, 2 GB RAM, 25 GB NVMegemma3:1b; optionally screen a 1.7B Q4 tagContext and OS headroom can make larger candidates fail
KVM4-US4 vCores, 4 GB RAM, 40 GB NVMe1.7B Q4; cautiously test one 4B Q4 tagA 4B artifact leaves limited headroom
KVM8-US8 vCores, 8 GB RAM, 80 GB NVMe4B Q4; screen an 8B Q4 tag only with measured headroomConcurrency and long context can exceed memory

If your chosen tag cannot meet its memory or latency target on KVM8-US, do not conceal the problem with swap or an unsupported extrapolation. Choose a smaller model, reduce context or concurrency with an explicit quality test, or ask about a custom resource profile. For containerized deployments, the Docker VPS hosting guide explains the surrounding operational trade-offs.

Operate the Winning Configuration Safely

Keep Ollama on a loopback or private interface unless you have designed authentication, TLS, rate limits, and access policy for a shared endpoint. Pin the exact model tag in your application configuration, record its digest, and rerun the benchmark after changing the model, context length, Ollama version, VPS tier, or concurrency.

  • Leave operating-system and monitoring headroom instead of sizing to the model file alone.
  • Set a queue limit so demand cannot consume unbounded memory.
  • Monitor free memory, swap activity, request duration, and failures.
  • Keep benchmark JSON and environment details with the deployment record.
  • Roll back when a new model or runtime version misses the accepted baseline.

The useful benchmark is not the highest number from an idle lab. It is the repeatable result your actual service can sustain while remaining stable, observable, and within its resource budget.

Frequently Asked Questions

Is the model download size the same as required RAM?

No. The artifact size is only a first screening value. The runtime, context cache, operating system, and concurrent requests also consume memory, so verify peak use on the exact server.

Does adding more vCPUs always increase tokens per second?

No. Scaling depends on the host CPU, model, quantization, memory bandwidth, and runtime. Measure each tier instead of assuming linear gains.

Can swap make a model suitable for production?

Swap can sometimes prevent an immediate failure, but storage is far slower than RAM. Treat swap activity during inference as a warning and test the real latency before accepting the configuration.

How many benchmark runs should I keep?

Keep at least one warm-up and three measured runs for an initial comparison, then add a longer test under representative load. Retain the raw API responses so rates can be recomputed.
Facebook
Twitter
LinkedIn

Table of Contents

Get started today

With VPS.US VPS Hosting you get all the features, tools

Image