THE PROBLEM ISN'T THE PROMPT
LLMLingua, LongLLMLingua, training-free evaluator heads -- there's real work happening in prompt compression. All of it targets the same surface: the user's input.
That's not where the tokens go.
We've been running Claude Code as our primary development environment for months. Dozens of tool calls per task. Each Read returns up to 2,000 lines of source code. Each Grep dumps file matches with surrounding context. Each Bash spits back build output, test results, full stack traces. One agent task can push 200K tokens of tool output into the context window before the model even starts thinking about a solution.
We instrumented our sessions across a month. The breakdown:
| SOURCE | % OF TOTAL TOKENS |
|---|---|
| Tool output (Read, Grep, Bash, etc.) | 62-78% |
| Model responses | 14-22% |
| System prompt + user messages | 8-16% |
The prompt is a rounding error. Tool output is the context window.
WHAT TOOL OUTPUT ACTUALLY LOOKS LIKE
Tool output isn't natural language. That's what makes it both harder and easier to compress than prompts.
Harder because the structure matters. A function signature, an import path, an error line number -- lose any of these and the model makes wrong edits. Easier because tool output is dense with predictable redundancy: whitespace patterns, boilerplate imports, repeated variable names, framework scaffolding, comment blocks the model doesn't need to read.
We categorized the output types we see in practice:
- Source code (Read tool) -- heavily structured, high redundancy in formatting and imports, but semantically dense in signatures and logic
- Search results (Grep tool) -- repeated file paths, line number prefixes, context lines that duplicate nearby reads
- Build/test output (Bash tool) -- verbose by default, 90% noise in passing tests, signal concentrated in failures
- Directory listings (Glob, ls) -- pure structure, highly compressible
- Git output (diff, log, status) -- mixed structure and content, diff headers are pure overhead
Each type has different compression characteristics. A grep result compresses differently than a stack trace. We needed a model that could learn these patterns, not a rule-based system that would break on edge cases.
WREN: THE ARCHITECTURE
Wren is a LoRA fine-tuned 1.5B parameter model trained specifically for tool output compression. It runs locally on Apple Silicon via MLX -- no API calls, no latency penalty beyond local inference.
Three constraints shaped everything:
Local only. Sending tool output to a remote API for compression means spending tokens to save tokens. MLX on Apple Silicon gives sub-second inference. Zero API cost.
Semantic preservation over ratio. A 90% compression that drops a function signature is worse than 50% that keeps it. We'd rather under-compress than lose a single load-bearing token.
Invisible integration. Wren sits in the pipeline as an MCP server. Claude calls tools normally. Output passes through Wren before hitting the context window. You don't change how you work.
Training data came from our own sessions -- hundreds of real tool outputs paired with manually verified compressed versions. Every compressed form had to preserve what the model would need to produce correct edits, correct commands, correct reasoning. We threw out pairs where the compression dropped anything semantic, even if the ratio was impressive.
COMPRESSION THRESHOLDS
Not everything should be compressed. Short outputs have low redundancy and the inference cost isn't justified. We enforce two gates:
minimum input: 300 characters
minimum savings: 20% reduction
If the input is too short or the compression doesn't clear 20%, Wren returns the original unchanged. In practice, this means small tool outputs (single-line grep matches, short git status) pass through untouched, while large code reads and verbose build logs get significant compression.
RESULTS IN PRODUCTION
Across 30 days of daily use in real development workflows (not benchmarks, not synthetic inputs), Wren achieves:
| OUTPUT TYPE | COMPRESSION RATIO | SEMANTIC RETENTION |
|---|---|---|
| Source code (Read) | 55-65% | Full -- signatures, logic, imports preserved |
| Search results (Grep) | 60-75% | Full -- matched lines, paths, line numbers preserved |
| Build output (Bash) | 70-85% | Full -- errors and warnings preserved, passing test noise removed |
| Directory listings | 65-80% | Full -- structure and paths preserved |
| Git diff/log | 50-65% | Full -- hunks and messages preserved, headers compressed |
The aggregate across all output types lands at 60-80% token reduction with no observed degradation in model behavior. Claude produces the same edits, the same commands, the same reasoning -- just with a smaller context window footprint.
TOKEN ECONOMICS
Wren runs locally. Inference cost is your laptop's power draw. The savings are pure:
typical agent task: ~120,000 tokens tool output
after wren: ~35,000 tokens tool output
savings per task: ~85,000 tokens
at opus pricing: ~$1.28 saved per task
across 20 tasks/day: ~$25.60/day, ~$768/month
Output tokens are priced 3-8x higher than input tokens on most APIs. Since tool output gets cached and re-read by the model across multiple turns, the actual savings compound -- compressed output means smaller cache reads on every subsequent turn.
WHY THIS ISN'T PROMPT COMPRESSION
Different problem than prompt compression. Worth being explicit about why.
Prompts are natural language. The redundancy is linguistic -- filler words, syntactic scaffolding, implied context. LLMLingua drops low-information tokens and meaning survives because natural language is fault-tolerant. You can remove "the" and "is" and nothing breaks.
Tool output is structured, machine-generated text. The redundancy is structural -- repeated patterns, formatting overhead, boilerplate. But the high-signal tokens are load-bearing in a way natural language rarely is. A variable name, an error code, a file path. Drop the wrong one from a grep result and the model edits the wrong file.
The model needs to understand code structure, not just token frequency. import React from 'react' appearing in 15 files is compressible. import { useCustomHook } from './hooks' is probably unique and critical. That distinction requires training on real tool outputs, not Wikipedia articles.
INTEGRATION
Wren runs as an MCP server. The setup is one line in your Claude config:
{
"mcpServers": {
"wren": {
"command": "python3",
"args": ["-m", "wren.mcp"]
}
}
}
It also ships built into Flock, our agent multiplexer, where it automatically compresses paste input into Claude sessions when the text exceeds the 300-character threshold.
The code is open source at github.com/Divagation/wren.
WHAT'S NEXT
Compression after the fact is step one. Step two is agents that are compression-aware from the start -- structuring tool calls to minimize redundancy before it exists. Why read an entire file when 40 lines around the target function carry all the signal? The compression model already knows this. The agent should too.
We're also building a benchmark suite for tool output compression specifically. The field lacks a standard evaluation that measures semantic retention across output types -- code, grep, build logs, git. Aggregate compression ratios don't tell you if the model can still make correct edits. We'll publish the suite alongside the next Wren release.