Cheap and local models
Zeltro doesn't sell you AI. It runs whichever agent you point it at, against whichever model you want, including models running on your own machine.
The default agents bill at frontier-model prices. For a lot of Zeltro work (scaffolding, wiring config, routine edits) a much cheaper model is fine, and the difference is large:
| Model | Input / Output per 1M tokens |
|---|---|
| Claude Sonnet 5 (Anthropic API) | $2 / $10 |
| Qwen3 Coder Next (OpenRouter) | ~$0.12 / ~$0.80 |
| Qwen3 Coder 30B A3B Instruct (OpenRouter) | ~$0.07 / ~$0.28 |
| Any model on local Ollama | nothing per token (your hardware and power) |
Prices as of September 2026; OpenRouter prices move, so check the model's page
before you rely on them. In a test in August 2026, Qwen Code with
qwen/qwen3-coder-30b-a3b-instruct built a working Django app end to end for
$0.87.
Zeltro already saves tokens before the model is involved: the environment, networking, database and container plumbing are pre-built, so the agent spends its context on your app instead of deriving Docker setup. Switching models compounds that saving rather than replacing it.
How it works#
zeltro ai-set --agent <agent> --model <model> --api-base <url> --api-key <key>
--api-base is the important one. Zeltro hands it to the agent through whichever
setting that CLI reads:
| Agent | Endpoint passed as | Key passed as | Endpoint it expects |
|---|---|---|---|
qwen |
OPENAI_BASE_URL |
OPENAI_API_KEY |
OpenAI-compatible |
aider |
--openai-api-base |
--api-key <provider>=<key> |
OpenAI-compatible |
claude |
ANTHROPIC_BASE_URL |
ANTHROPIC_API_KEY, only if it starts with sk-ant- |
Anthropic-compatible |
codex |
-c openai_base_url=… (Zeltro passes it for you) |
OPENAI_API_KEY, only if it starts with sk- |
OpenAI Responses API |
gemini |
nothing | nothing | Its own Google sign-in, or a GEMINI_API_KEY you export |
"OpenAI-compatible" covers OpenRouter, DeepInfra, Together, Fireworks, Ollama,
LM Studio and vLLM, so qwen and aider are the agents to use with cheap or local
models.
For one run with a different model, without touching the global setting, use the
per-session overrides (ZELTRO_AI_AGENT,
ZELTRO_AI_MODEL, ZELTRO_AI_API_BASE, ZELTRO_AI_API_KEY_FILE). In the Zeltro app,
an AI profile does the same.
OpenRouter (cheapest hosted)#
zeltro ai-set --agent qwen \
--model qwen/qwen3-coder-30b-a3b-instruct \
--api-base https://openrouter.ai/api/v1 \
--api-key sk-or-v1-...
Pay-as-you-go, no subscription. qwen/qwen3-coder-next is a newer, larger Qwen
coding model for two to three times the price. With Aider, use OpenRouter's own
prefix and no endpoint:
--agent aider --model openrouter/qwen/qwen3-coder-30b-a3b-instruct --api-key sk-or-v1-....
Ollama (local, private)#
Install Ollama, pull a coding model, point Zeltro at it:
ollama pull qwen3-coder:30b
zeltro ai-set --agent qwen \
--model qwen3-coder:30b \
--api-base http://localhost:11434/v1 \
--api-key ollama
The key is a placeholder. Ollama ignores it, but the CLIs expect something.
Nothing leaves your machine, which matters for client work under NDA.
Raise Ollama's context length before you start. Its default depends on your VRAM
and is only 4k tokens under 24 GB, which is far too short for an agent. Ollama
recommends at least 64k for coding tools: start the server with
OLLAMA_CONTEXT_LENGTH=64000 ollama serve, or use the slider in the Ollama app's
settings. More context needs more VRAM.
Model size matters more than you would like, and this is measured rather
than assumed. Tested here against qwen2.5-coder:1.5b, zeltro create --classify-only returned valid, correctly shaped JSON (the plumbing is fine)
but recommended a budgeting app for a guitar pedal tracker, with the reason
"Laravel is great for building budgeting apps". Coherent output, incoherent
thinking.
So small models fail in the worst way: they succeed mechanically and are wrong
on the substance. A ~30B model on 24 GB of VRAM is roughly the floor for agentic
work (qwen3-coder:30b is a 19 GB download); below that, expect plausible nonsense
rather than errors.
Qwen Code also requires Node 22 or newer. It installs and runs on Node 20
(npm warns EBADENGINE and it works anyway), but that is unsupported.
LM Studio or vLLM#
Both expose an OpenAI-compatible server, so the shape is identical. LM Studio listens on port 1234; vLLM defaults to 8000:
zeltro ai-set --agent qwen --model <model> \
--api-base http://localhost:1234/v1 --api-key local
Codex against another endpoint#
Current Codex CLI no longer reads OPENAI_BASE_URL, so Zeltro passes the endpoint
from zeltro ai-set --agent codex --api-base ... (or ZELTRO_AI_API_BASE) as
Codex's own setting, -c openai_base_url="…", on every run. The catch is that
Codex only speaks OpenAI's Responses API, which not every compatible server
implements. For anything more involved, configure Codex itself in
~/.codex/config.toml with a [model_providers] entry. For Ollama or LM Studio,
Codex has its own --oss mode (oss_provider in the same file). See
Codex's advanced configuration.
For cheap or local models, qwen or aider is less work.
Claude Code against a local model#
Claude Code speaks Anthropic's API, not OpenAI's. Ollama 0.14 and later also speak Anthropic's API, so Claude Code can use an Ollama model directly:
zeltro ai-set --agent claude --model qwen3-coder:30b \
--api-base http://localhost:11434
Note the URL has no /v1. Leave --api-key empty: Ollama ignores the key, and
Zeltro only passes keys that start with sk-ant-. If Claude Code asks you to log
in, export ANTHROPIC_AUTH_TOKEN=ollama in your shell first, as
Ollama's guide does.
For a server that only speaks OpenAI's API, put a translating proxy such as
LiteLLM in front and pass its address as
--api-base.
Worth it only if you specifically want Claude Code's interface over a local
model. Otherwise qwen is less machinery.
Going back#
zeltro ai-set --agent claude --api-base none --api-key ""
none clears the endpoint; an empty --api-key clears a stored key. Switching
agents also clears the stored model, so this returns you to Claude Code's own
default model. Check where you stand with zeltro ai-set --json-output.
Choosing honestly#
Cheaper models are genuinely worse at long autonomous work. The realistic split:
- Routine edits, scaffolding, config, boilerplate. A cheap or local model is fine and the savings are real.
zeltro createbuilding a whole app from a sentence, or debugging something subtle. Frontier models still earn their price.
Nothing stops you switching per task. zeltro ai-set takes a second and changes the
global default; the ZELTRO_AI_* variables change it for a single run.
Spotted a mistake? Edit this page on GitHub.