The problem

I use Claude Code on the $20 per month plan. For the kind of work I do with it (Kubernetes manifests, CI pipelines, a FastAPI backend, docs) it is very good, and that is exactly the problem: I use it a lot, and the quota runs out long before the day does. Then I sit and wait for the usage window to reset.

When I looked at where the usage went, most of it was not the hard part. It was the main model, the expensive one, doing things like:

  • searching the repo for where something is defined
  • renaming a symbol in six files
  • summarising a few hundred lines of logs
  • writing boilerplate it had already decided on

None of that needs the best model. So the goal became simple: keep Claude as the primary model for design and debugging, and push the simple work to something cheaper.

Repo for this article: https://github.com/besmirzanaj/claude-cheap-models

What I looked at first, and why I dropped it

The first thing I found was OmniRoute, a local AI gateway. You point Claude Code at localhost:20128 and it routes every request to one of a few hundred providers, many of them free. On paper that is exactly what I wanted.

I decided against it, and against gateways in general, for three reasons:

  • Everything goes through it. A gateway in front of Claude Code sees every prompt, every file and every command output. On my machines that includes kubectl output and infrastructure details. With free tiers that content then goes to whichever provider was picked, often one whose terms allow training on it.
  • My credentials go through it too. I did not want my Claude login, or any API key, to pass through a third-party proxy. One leak and the damage is not limited to a side project.
  • It is built to work around provider rules. Its own README lists TLS fingerprint impersonation and a multi-level proxy for when “AI is blocked”. Routing a Claude subscription through something like that is not a risk I want to take with my account.

There is also a quality argument. Claude Code’s prompts and tools are tuned for Claude. Put a cheaper model behind it and multi-step work gets worse exactly where I need it to be good.

So the rule I set was: nothing sits between Claude Code and Anthropic. The cheap model has to be something Claude calls, not something that replaces it.

The setup

Two pieces, both installed at user level so they work in every project.

1. Haiku subagents

Claude Code lets you define subagents as Markdown files with a model: in the frontmatter. I have two, both pinned to Haiku:

---
name: search
description: Read-only code/exploration search. Use for "where is X", "which files use Y", naming/convention sweeps, and gathering excerpts — anything that only needs to find and read, not change. Runs on a cheap model; don't use it for design, debugging logic, or edits.
tools: Read, Grep, Glob
model: haiku
---

You are a fast read-only search agent. Find what the caller asked for and report the conclusion, not a file dump.

search can only read. simple-edit makes mechanical changes that have already been decided (rename, find and replace, version bump) and reports back instead of guessing when the change turns out to need a decision.

This still counts against the Claude plan, but Haiku is the cheapest tier, and the tool use stays reliable because it is still Claude.

2. A cheap_llm MCP tool

The second piece is a small MCP server, about 80 lines of Python, that exposes one tool. Claude calls it with a prompt, the server sends that prompt to deepseek/deepseek-v4-flash on OpenRouter and returns the answer:

@server.tool()
async def cheap_llm(prompt: str, system: str = "", max_tokens: int = 4000, reasoning: bool = False) -> str:
    ...
    body = {
        "model": MODEL,
        "messages": messages,
        "max_tokens": max_tokens,
        "reasoning": {"enabled": reasoning},
        # Only providers that don't store or train on prompts
        "provider": {"data_collection": "deny"},
    }

The important properties:

  • It only sees what Claude sends it. The DeepSeek model gets one prompt. It has no access to the repository, the shell or my Claude session.
  • No stored or trained-on prompts. The request is restricted to OpenRouter providers that do not keep or train on the data. If none can serve the model the call fails; it does not quietly fall back.
  • The key is not in any file. The server reads it at runtime from the macOS Keychain (or secret-tool on Linux, or OPENROUTER_API_KEY).
  • It is cheap. At the time of writing this model costs $0.03 per million input tokens and $1.28 per million output tokens on OpenRouter.

The tool description tells Claude what it is for (summaries, boilerplate, translation, reformatting, field extraction) and what must never go into it: secrets, credentials, customer data, private source code.

The routing rules

The last part is four lines in ~/.claude/CLAUDE.md so every session knows the rule:

## Cheap models for simple work

- Delegate read-only searches to the `search` subagent and already-decided mechanical edits to `simple-edit` (both Haiku). Keep design, debugging and judgement in the main session.
- The `cheap_llm` MCP tool (DeepSeek V4 Flash via OpenRouter) is for self-contained text tasks: summarising logs or docs, boilerplate, translation, reformatting. Never send it secrets, keys, customer data or private source code; it goes to a third party.

Installing it

You need Claude Code and uv.

git clone https://github.com/besmirzanaj/claude-cheap-models.git
cd claude-cheap-models && ./install.sh

The installer copies the two agents into ~/.claude/agents/, copies the server into ~/.claude/mcp/cheap-llm/, registers it with claude mcp add --scope user, and appends the routing rules to ~/.claude/CLAUDE.md if they are not there yet. It is safe to re-run.

Then store the OpenRouter key once per machine:

# macOS
security add-generic-password -U -s openrouter-api-key -a "$USER" -w
# Linux
secret-tool store --label='OpenRouter API key' service openrouter-api-key

Run that in a real terminal. My first attempt was from a prompt that could not show the password prompt, and it happily stored an empty key.

Start a new Claude Code session and /mcp should list cheap-llm.

One thing that bit me: reasoning tokens

The first live call returned HTTP 200 and an empty answer:

[deepseek/deepseek-v4-flash · 10 in / 20 out tokens]

DeepSeek V4 Flash reasons before it answers by default, and the reasoning counts against max_tokens. With a small cap the whole budget went to thinking and nothing was left for the answer. For the kind of work this tool does, thinking is a waste anyway, so reasoning is now off unless the caller asks for it:

Kubernetes NetworkPolicies use label selectors to restrict pod traffic, and
when combined with a default-deny policy, only explicitly allowed ingress
and egress connections are permitted.

[deepseek/deepseek-v4-flash · 39 in / 37 out tokens]

What this does not do

A few honest limits:

  • It does not help once the quota is gone. cheap_llm is a tool Claude calls. If Claude Code is locked out until the reset, nothing calls it. This setup makes the quota last longer; it is not a replacement model.
  • The cheap model has no tools. It gets text and returns text. Anything that needs the repo or a command stays with Claude or a Haiku subagent.
  • Check what comes back. It is a first pass. Claude is told to verify its output before relying on it.
  • I have not measured the savings yet. I will update this post once I have a few weeks of usage to compare.

Conclusions

The pattern I ended up with is simple: Claude for judgement, Haiku for searches and mechanical edits, and a very cheap model for text chores, called as a tool and never placed in the path of my Claude session. No gateway, no credentials passing through someone else’s proxy, and the only thing the third-party model ever sees is the prompt Claude chose to send.

The whole thing is one small repo under the MIT licence: https://github.com/besmirzanaj/claude-cheap-models. Use it, change the model with CHEAP_LLM_MODEL, and send a pull request if you improve it.