Making a $20 Claude Code plan last: Haiku subagents and a DeepSeek side tool
The problem
I use Claude Code on the $20 per month plan. For the kind of work I do with it (Kubernetes manifests, CI pipelines, a FastAPI backend, docs) it is very good, and that is exactly the problem: I use it a lot, and the quota runs out long before the day does. Then I sit and wait for the usage window to reset.
When I looked at where the usage went, most of it was not the hard part. It was the main model, the expensive one, doing things like:
- searching the repo for where something is defined
- renaming a symbol in six files
- summarising a few hundred lines of logs
- writing boilerplate it had already decided on
None of that needs the best model. So the goal became simple: keep Claude as the primary model for design and debugging, and push the simple work to something cheaper.
Repo for this article: https://github.com/besmirzanaj/claude-cheap-models
What I looked at first, and why I dropped it
The first thing I found was OmniRoute,
a local AI gateway. You point Claude Code at localhost:20128 and it routes
every request to one of a few hundred providers, many of them free. On paper
that is exactly what I wanted.
I decided against it, and against gateways in general, for three reasons:
- Everything goes through it. A gateway in front of Claude Code sees every prompt, every file and every command output. On my machines that includes kubectl output and infrastructure details. With free tiers that content then goes to whichever provider was picked, often one whose terms allow training on it.
- My credentials go through it too. I did not want my Claude login, or any API key, to pass through a third-party proxy. One leak and the damage is not limited to a side project.
- It is built to work around provider rules. Its own README lists TLS fingerprint impersonation and a multi-level proxy for when “AI is blocked”. Routing a Claude subscription through something like that is not a risk I want to take with my account.
There is also a quality argument. Claude Code’s prompts and tools are tuned for Claude. Put a cheaper model behind it and multi-step work gets worse exactly where I need it to be good.
So the rule I set was: nothing sits between Claude Code and Anthropic. The cheap model has to be something Claude calls, not something that replaces it.
The setup
Two pieces, both installed at user level so they work in every project.
1. Haiku subagents
Claude Code lets you define subagents as Markdown files with a model: in
the frontmatter. I have two, both pinned to Haiku:
---
name: search
description: Read-only code/exploration search. Use for "where is X", "which files use Y", naming/convention sweeps, and gathering excerpts — anything that only needs to find and read, not change. Runs on a cheap model; don't use it for design, debugging logic, or edits.
tools: Read, Grep, Glob
model: haiku
---
You are a fast read-only search agent. Find what the caller asked for and report the conclusion, not a file dump.
search can only read. simple-edit makes mechanical changes that have
already been decided (rename, find and replace, version bump) and reports
back instead of guessing when the change turns out to need a decision.
This still counts against the Claude plan, but Haiku is the cheapest tier, and the tool use stays reliable because it is still Claude.
2. A cheap_llm MCP tool
The second piece is a small MCP server, about 80 lines of Python, that
exposes one tool. Claude calls it with a prompt, the server sends that
prompt to deepseek/deepseek-v4-flash on OpenRouter
and returns the answer:
@server.tool()
async def cheap_llm(prompt: str, system: str = "", max_tokens: int = 4000, reasoning: bool = False) -> str:
...
body = {
"model": MODEL,
"messages": messages,
"max_tokens": max_tokens,
"reasoning": {"enabled": reasoning},
# Only providers that don't store or train on prompts
"provider": {"data_collection": "deny"},
}
The important properties:
- It only sees what Claude sends it. The DeepSeek model gets one prompt. It has no access to the repository, the shell or my Claude session.
- No stored or trained-on prompts. The request is restricted to OpenRouter providers that do not keep or train on the data. If none can serve the model the call fails; it does not quietly fall back.
- The key is not in any file. The server reads it at runtime from the
macOS Keychain (or
secret-toolon Linux, orOPENROUTER_API_KEY). - It is cheap. At the time of writing this model costs $0.03 per million input tokens and $1.28 per million output tokens on OpenRouter.
The tool description tells Claude what it is for (summaries, boilerplate, translation, reformatting, field extraction) and what must never go into it: secrets, credentials, customer data, private source code.
The routing rules
The last part is four lines in ~/.claude/CLAUDE.md so every session knows
the rule:
## Cheap models for simple work
- Delegate read-only searches to the `search` subagent and already-decided mechanical edits to `simple-edit` (both Haiku). Keep design, debugging and judgement in the main session.
- The `cheap_llm` MCP tool (DeepSeek V4 Flash via OpenRouter) is for self-contained text tasks: summarising logs or docs, boilerplate, translation, reformatting. Never send it secrets, keys, customer data or private source code; it goes to a third party.
Installing it
You need Claude Code and uv.
git clone https://github.com/besmirzanaj/claude-cheap-models.git
cd claude-cheap-models && ./install.sh
The installer copies the two agents into ~/.claude/agents/, copies the
server into ~/.claude/mcp/cheap-llm/, registers it with
claude mcp add --scope user, and appends the routing rules to
~/.claude/CLAUDE.md if they are not there yet. It is safe to re-run.
Then store the OpenRouter key once per machine:
# macOS
security add-generic-password -U -s openrouter-api-key -a "$USER" -w
# Linux
secret-tool store --label='OpenRouter API key' service openrouter-api-key
Run that in a real terminal. My first attempt was from a prompt that could not show the password prompt, and it happily stored an empty key.
Start a new Claude Code session and /mcp should list cheap-llm.
One thing that bit me: reasoning tokens
The first live call returned HTTP 200 and an empty answer:
[deepseek/deepseek-v4-flash · 10 in / 20 out tokens]
DeepSeek V4 Flash reasons before it answers by default, and the reasoning
counts against max_tokens. With a small cap the whole budget went to
thinking and nothing was left for the answer. For the kind of work this
tool does, thinking is a waste anyway, so reasoning is now off unless the
caller asks for it:
Kubernetes NetworkPolicies use label selectors to restrict pod traffic, and
when combined with a default-deny policy, only explicitly allowed ingress
and egress connections are permitted.
[deepseek/deepseek-v4-flash · 39 in / 37 out tokens]
What this does not do
A few honest limits:
- It does not help once the quota is gone.
cheap_llmis a tool Claude calls. If Claude Code is locked out until the reset, nothing calls it. This setup makes the quota last longer; it is not a replacement model. - The cheap model has no tools. It gets text and returns text. Anything that needs the repo or a command stays with Claude or a Haiku subagent.
- Check what comes back. It is a first pass. Claude is told to verify its output before relying on it.
- I have not measured the savings yet. I will update this post once I have a few weeks of usage to compare.
Conclusions
The pattern I ended up with is simple: Claude for judgement, Haiku for searches and mechanical edits, and a very cheap model for text chores, called as a tool and never placed in the path of my Claude session. No gateway, no credentials passing through someone else’s proxy, and the only thing the third-party model ever sees is the prompt Claude chose to send.
The whole thing is one small repo under the MIT licence:
https://github.com/besmirzanaj/claude-cheap-models. Use it, change the
model with CHEAP_LLM_MODEL, and send a pull request if you improve it.
