A local language model can load successfully and still be a poor fit for a VPS. The model file must fit alongside the operating system, Ollama, the context cache, and every other service. CPU generation speed must also match the workload. A careful benchmark therefore answers two separate questions: does it fit safely? and is it fast enough for this job?
This guide gives you a reproducible CPU-only test for the current public VPS.us KVM tiers. It uses exact model tags and runtime counters instead of fabricated benchmark results. Run the same prompt several times on the exact server you plan to use, record the evidence, and size from your own results.
Separate Model Capacity From Inference Performance
Model capacity is a memory question. The registry artifact is only the starting point: runtime state, the context cache, Ubuntu, and background services also need RAM. Ollama documents that longer context increases memory use and that parallel requests multiply context allocation. A download smaller than total RAM is not proof of a reliable fit.
Performance is a measurement question. CPU generation depends on the exact processor, instruction set, model architecture, quantization, context, prompt, and competing load. Do not copy tokens-per-second figures from another host and present them as a promise for yours. Benchmark the final tag on the final VPS image, then repeat after material runtime or model changes.
Screen Current Model Artifacts Against Available RAM

Start with the exact artifact size listed in the Ollama library. The following examples were checked on September 22, 2026. They are screening values, not measured working-memory requirements or performance guarantees; each tag links to its current registry entry.
| Exact Ollama tag | Published artifact size | First capacity check |
|---|---|---|
gemma3:270m | 292 MB | Small smoke tests and simple extraction |
gemma3:1b | 815 MB | Low-memory experimentation |
qwen3:1.7b-q4_K_M | 1.4 GB | Compact text tasks with controlled context |
qwen3:4b-instruct-2507-q4_K_M | 2.5 GB | Instruction work on a node with several GB of free RAM |
gemma3:4b-it-q4_K_M | 3.3 GB | Multimodal-capable model with more memory pressure |
qwen3:8b-q4_K_M | 5.2 GB | Higher-memory CPU experimentation |
Quantization reduces storage and memory pressure by representing weights at lower precision, but it can also affect quality. The VPS.us self-hosted AI chatbot guide explains the wider stack around the model. Keep enough free space for more than model weights: a chatbot may also need an API layer, logs, retrieval services, and a database.
Prepare a Clean CPU-Only Test Node
Use the same Ubuntu release and VPS tier that will host the workload. Install Ollama using its official Linux instructions, and make sure curl and jq are available. Stop unrelated jobs, but keep normal security and monitoring services running. Record the baseline before testing.
ollama --version uname -r lsb_release -ds lscpu free -h swapon --show nproc uptime ollama list
If the load average is already high or free memory is unexpectedly low, fix that condition before testing. Avoid using swap as a way to declare a model successful. Swap may prevent an immediate out-of-memory failure, but it replaces RAM access with storage I/O and can make interactive latency unacceptable. The VPS server optimization guide covers the broader discipline of measuring before tuning.
Run a Reproducible Ollama Benchmark

Pull one exact tag at a time. Tags may change, so retain the complete model digest from the local model-list API with your results. Keep the prompt, context, output limit, thinking setting, and concurrency constant across tiers. The example disables thinking and caps output at 256 tokens; benchmark your real workload separately if it requires different settings.
ollama pull qwen3:1.7b-q4_K_M
ollama list
curl -fSs http://127.0.0.1:11434/api/tags > model-manifest.json
# Warm the model before measured runs.
curl -fSs http://127.0.0.1:11434/api/generate \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3:1.7b-q4_K_M","prompt":"Summarize three benefits of TLS 1.3.","stream":false,"think":false,"keep_alive":"10m","options":{"temperature":0,"num_ctx":4096,"num_predict":256}}' \
> warmup.json
The Ollama generate API returns timing and token counters in the final response. Store the complete JSON so you can distinguish model-load time, prompt evaluation, and generation. After the warm-up, run ollama ps and verify that PROCESSOR reports 100% CPU.
for run in 1 2 3; do
curl -fSs http://127.0.0.1:11434/api/generate \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3:1.7b-q4_K_M","prompt":"Summarize three benefits of TLS 1.3.","stream":false,"think":false,"keep_alive":"10m","options":{"temperature":0,"num_ctx":4096,"num_predict":256}}' \
> "run-$run.json"
done
Run the loop from an idle node, then repeat it while the real supporting services are active. If the production workload requires simultaneous users, perform a separate concurrency test; do not multiply a single-request rate and call it capacity.
Measure Throughput, Latency, and Peak Memory
Ollama reports durations in nanoseconds. Generation throughput is eval_count × 1,000,000,000 ÷ eval_duration. Prompt throughput uses the uncached prompt counters; report it as unavailable when its duration is zero. Keep cold-model loading separate from warm generation and record cache effects rather than treating a repeated prompt as a fresh prompt-processing test.
jq '
def rate($count; $duration):
if ($duration // 0) > 0 then $count * 1000000000 / $duration else null end;
{
total_seconds: (.total_duration / 1000000000),
load_seconds: (.load_duration / 1000000000),
uncached_prompt_tokens_per_second: rate((.prompt_eval_count - (.prompt_eval_cached_count // 0)); .prompt_eval_duration),
generation_tokens_per_second: rate(.eval_count; .eval_duration),
prompt_tokens: .prompt_eval_count,
cached_prompt_tokens: (.prompt_eval_cached_count // 0),
generated_tokens: .eval_count
}' run-1.json
Monitor memory throughout the run and keep the samples, not only a before-and-after reading. Record the highest observed use and any swapping; periodic samples can miss brief peaks, so do not present them as an exact process high-water mark. A request that completes after heavy swapping is memory-constrained rather than a clean pass.
ollama ps free -h vmstat 1 | tee memory-samples.txt # Stop vmstat with Ctrl-C after the benchmark. journalctl -u ollama --since '10 minutes ago' --no-pager
Interpret Results Without Inventing Universal Thresholds
There is no honest universal boundary where a model becomes “interactive.” A short internal classification task and a streaming chat interface have different latency budgets. Define the target before reading the results.
- Single-user chat: judge time to first visible output, sustained generation, and the longest response users will tolerate.
- Background automation: judge completed jobs per hour, queue depth, and recovery after failure.
- Document processing: include prompt-evaluation time because long inputs can dominate the run.
- Concurrent service: test the real parallel request count and context length while watching RAM.
Use the median of repeated warm runs and retain the slowest result. Investigate large variation rather than averaging it away. CPU steal, another workload, model reloads, or memory pressure can make one run unrepresentative.
Match the Current VPS.us Tiers to a Test Plan
The current public VPS.us LLM hosting page lists four KVM tiers from 1 GB to 8 GB of RAM. The table below is deliberately conservative. “Candidate” means the artifact is worth testing with a controlled context; it does not guarantee a clean load, a particular throughput, or production suitability.
| Current public tier | Published resources | Practical first test | Important limit |
|---|---|---|---|
| KVM1-US | 1 vCore, 1 GB RAM, 20 GB NVMe | gemma3:270m smoke test | Little room for a general-purpose LLM service |
| KVM2-US | 2 vCores, 2 GB RAM, 25 GB NVMe | gemma3:1b; optionally screen a 1.7B Q4 tag | Context and OS headroom can make larger candidates fail |
| KVM4-US | 4 vCores, 4 GB RAM, 40 GB NVMe | 1.7B Q4; cautiously test one 4B Q4 tag | A 4B artifact leaves limited headroom |
| KVM8-US | 8 vCores, 8 GB RAM, 80 GB NVMe | 4B Q4; screen an 8B Q4 tag only with measured headroom | Concurrency and long context can exceed memory |
If your chosen tag cannot meet its memory or latency target on KVM8-US, do not conceal the problem with swap or an unsupported extrapolation. Choose a smaller model, reduce context or concurrency with an explicit quality test, or ask about a custom resource profile. For containerized deployments, the Docker VPS hosting guide explains the surrounding operational trade-offs.
Operate the Winning Configuration Safely
Keep Ollama on a loopback or private interface unless you have designed authentication, TLS, rate limits, and access policy for a shared endpoint. Pin the exact model tag in your application configuration, record its digest, and rerun the benchmark after changing the model, context length, Ollama version, VPS tier, or concurrency.
- Leave operating-system and monitoring headroom instead of sizing to the model file alone.
- Set a queue limit so demand cannot consume unbounded memory.
- Monitor free memory, swap activity, request duration, and failures.
- Keep benchmark JSON and environment details with the deployment record.
- Roll back when a new model or runtime version misses the accepted baseline.
The useful benchmark is not the highest number from an idle lab. It is the repeatable result your actual service can sustain while remaining stable, observable, and within its resource budget.
Frequently Asked Questions
Is the model download size the same as required RAM?
Does adding more vCPUs always increase tokens per second?
Can swap make a model suitable for production?
How many benchmark runs should I keep?