Production agents hit context limits fast. A coding agent that runs tests, reads logs, and pulls documentation can burn through 100k tokens in three turns. RAG pipelines dump entire chunks into the prompt. Tool outputs return verbose JSON. Every token costs money and adds latency.
Headroom is a compression layer that sits between your agent and the LLM. It shrinks tool outputs, logs, RAG chunks, and files before they reach the context window. The project claims 20% token reduction for coding agents and 60–95% for JSON-heavy workflows, with no change to the final answer. It ships as a Python library, a FastAPI proxy, and an MCP server, giving you three deployment patterns with different trade-offs in latency, observability, and integration complexity.
What Gets Compressed and What Stays Intact
Headroom uses a fine-tuned model (kompress-v2-base) to decide what to keep and what to discard. The compression is not blind truncation. The model learns to preserve semantic anchors: error messages, function signatures, stack traces, and critical log lines.
The hero image in the repo shows a 55,957-token agent prompt compressed to 24,340 tokens. The FATAL line at item 67 survives byte-for-byte. This is the key engineering question: how do you compress structured data without breaking parsers downstream?
Headroom handles:
- JSON tool outputs: Strips redundant keys, collapses nested structures, preserves schema-critical fields.
- Log files: Keeps ERROR and FATAL lines, drops repetitive INFO entries, maintains timestamps for correlation.
- Code diffs: Preserves function signatures and changed lines, compresses unchanged context.
- RAG chunks: Keeps sentences with high semantic density, drops boilerplate.
The model does not compress the user query or the assistant’s response. It only touches the context that the agent reads before generating an answer.
Three Deployment Modes
Headroom ships in three forms. Each has different latency, observability, and integration costs.
| Deployment Mode | Latency Overhead | Observability | Integration Effort | Best For |
|---|---|---|---|---|
| Library | 50–200ms per call | Full control, log everything | Low (import and wrap) | Agents you control end-to-end |
| Proxy | 100–300ms per call | Centralized metrics, request logs | Medium (point agent at proxy URL) | Multi-agent fleets, shared infra |
| MCP Server | 150–400ms per call | MCP protocol logs, tool-level tracing | Low (add to MCP config) | Claude Code, Cursor, agent harnesses |
Library Mode
You import Headroom as a Python package and wrap your agent’s context-building logic.
from headroom import compress_context
def build_agent_context(tool_outputs, logs, rag_chunks):
raw_context = "\n".join([
format_tool_outputs(tool_outputs),
format_logs(logs),
format_rag_chunks(rag_chunks)
])
compressed = compress_context(
raw_context,
preserve_patterns=["ERROR", "FATAL", "def ", "class "],
target_ratio=0.4
)
return compressed
This gives you full control. You can log the before and after token counts, inspect what got dropped, and adjust preserve_patterns based on your domain. The latency overhead is 50–200ms per compression call, depending on context size.
Proxy Mode
Headroom runs as a FastAPI service. Your agent sends requests to the proxy instead of directly to OpenAI or Anthropic. The proxy compresses the context, forwards the request, and returns the response.
# Start the proxy
headroom serve --port 8080 --model kompress-v2-base
# Point your agent at the proxy
export OPENAI_API_BASE=http://localhost:8080/v1
This centralizes compression logic. You get request logs, token savings metrics, and error rates in one place. The latency overhead is 100–300ms, which includes network round-trip and compression time. This mode works well for multi-agent fleets where you want shared observability.
MCP Server Mode
Headroom implements the Model Context Protocol. You add it to your MCP config, and it compresses tool outputs before they reach the agent.
{
"mcpServers": {
"headroom": {
"command": "headroom",
"args": ["mcp"],
"env": {
"HEADROOM_MODEL": "kompress-v2-base",
"HEADROOM_TARGET_RATIO": "0.4"
}
}
}
}
This is the lowest-friction option for Claude Code, Cursor, and other MCP-compatible harnesses. The agent sees Headroom as just another tool. The latency overhead is 150–400ms because MCP adds protocol serialization on top of compression time.
How Compression Handles Structured Formats
The kompress-v2-base model is fine-tuned on JSON, logs, code, and Markdown. It does not use generic text compression (gzip, LZ4). It learns to preserve structure.
For JSON, the model keeps:
- Schema-defining keys (type, id, status)
- Non-null values
- First and last items in large arrays
- Error and warning fields
For logs, it keeps:
- Lines with ERROR, FATAL, WARN
- Timestamps for correlation
- Stack traces
- First occurrence of repeated patterns
For code, it keeps:
- Function and class signatures
- Changed lines in diffs
- Import statements
- Docstrings for public APIs
The model drops:
- Repeated INFO log lines
- Null or default JSON values
- Unchanged code context in diffs
- Boilerplate comments
This is not lossless compression. You lose detail. The bet is that the detail you lose does not change the agent’s answer.
Observability and Failure Modes
Headroom adds a new failure surface. If the compression model drops a critical line, the agent might give the wrong answer. If the proxy goes down, your agent stops working.
Observability hooks you need:
- Token delta logs: Before and after token counts for every compression call.
- Preserved pattern matches: Log which patterns (ERROR, FATAL, def) triggered preservation.
- Compression ratio distribution: Track min, max, and p95 compression ratios across requests.
- Downstream accuracy: Sample agent outputs and compare compressed vs. uncompressed contexts.
Failure modes:
- Over-compression: The model drops a critical error message. The agent misses the root cause and suggests the wrong fix.
- Under-compression: The model preserves too much. You save 10% tokens instead of 60%.
- Proxy downtime: If you run Headroom as a proxy, it becomes a single point of failure. Add health checks and failover.
- Latency spikes: Compression time scales with context size. A 200k-token context might take 1–2 seconds to compress.
When Compression Breaks Semantic Integrity
Headroom works well for verbose, repetitive data. It struggles with dense, information-rich contexts where every token matters.
Good candidates for compression:
- JSON API responses with many null fields
- Log files with repeated INFO lines
- RAG chunks with boilerplate introductions
- Code diffs with large unchanged sections
Bad candidates:
- Mathematical proofs where every step is load-bearing
- Legal documents where exact wording matters
- Short, dense error messages (already minimal)
- Contexts where the agent needs to count occurrences
If your agent needs to count how many times a specific error appears, compression will break that. If your agent needs to verify exact JSON schema compliance, compression might drop the field that violates the schema.
Integration with Agent Frameworks
Headroom supports LangChain, LlamaIndex, and custom agent loops. For LangChain, you wrap the retriever:
from langchain.retrievers import BaseRetriever
from headroom import compress_context
class CompressedRetriever(BaseRetriever):
def __init__(self, base_retriever):
self.base_retriever = base_retriever
def get_relevant_documents(self, query):
docs = self.base_retriever.get_relevant_documents(query)
compressed_docs = [
compress_context(doc.page_content, target_ratio=0.5)
for doc in docs
]
return compressed_docs
For LlamaIndex, you compress the retrieved nodes before they go into the prompt.
For custom agent loops, you compress the accumulated context at each turn:
context_history = []
for turn in agent_loop:
tool_output = run_tool(turn.tool_call)
context_history.append(tool_output)
# Compress the full history before the next LLM call
compressed_history = compress_context(
"\n".join(context_history),
target_ratio=0.4
)
response = llm.generate(
prompt=turn.user_message,
context=compressed_history
)
Cost and Latency Trade-Offs
Compression adds latency but saves token costs. The break-even depends on your LLM pricing and request volume.
Example calculation:
- Agent sends 50k tokens per request
- Headroom compresses to 20k tokens (60% reduction)
- Compression adds 200ms latency
- GPT-4 input tokens cost $0.03 per 1k tokens
Savings per request:
- Token cost saved: (50k - 20k) * $0.03 / 1k = $0.90
- Latency cost: 200ms added to request time
If your agent runs 1,000 requests per day, you save $900 per day. If your SLA requires sub-500ms response times, the 200ms overhead might break your budget.
Technical Verdict
Use Headroom when:
- Your agent workflows routinely exceed 32k tokens per request.
- You have verbose tool outputs (JSON APIs, log files, RAG chunks).
- You can tolerate 100–300ms added latency.
- You have observability to detect when compression drops critical data.
Avoid Headroom when:
- Your contexts are already dense and minimal.
- Your agent needs exact token counts or field-level accuracy.
- You cannot afford the latency overhead.
- Your agent operates in a domain where every token is load-bearing (legal, math, compliance).
Headroom is infrastructure, not magic. It trades latency for cost savings. If your agent is already optimized and you are not hitting context limits, you do not need it. If you are burning $10k per month on tokens because your RAG pipeline dumps 100k tokens into every prompt, Headroom will pay for itself in a week.