---
title: "Caveman Middleware"
description: "Keep your stack. Lose the extra tokens."
canonical: https://caveman.so/products/caveman-middleware
last-updated: 2026-09-21
---

# Caveman Middleware

Keep your stack. Lose the extra tokens.

Your framework, models, tools, provider calls, and original conversation stay in
place. Caveman Middleware projects eligible tool-result text into a smaller copy
of your outbound request. Local compression and original recovery run beside your
application. Your existing SDK still sends inference to your chosen provider.

Available in alpha for TypeScript and Python. Local compression needs no Caveman
account. Client packages are MIT; the engine runtime is Apache-2.0.

## Explore savings by format

The [interactive middleware page](https://caveman.so/products/caveman-middleware#formats)
shows 13 synthetic samples: logs, JSON, CSV, TSV, YAML, TOML, TypeScript source
code, Git diffs, Markdown, HTML, XML, terminal output, and search results.

Every sample includes original and replacement token counts, its estimated net
change after the declared recovery-tool overhead, and excerpts of both payloads.
All 10 replaced originals were recovered byte for byte and checked by SHA-256.
The three unchanged results (Git diffs, HTML, search results) remain visible;
their recovery-tool overhead increases the inferred input count.

These are real local middleware HTTP runtime observations on synthetic data,
using a clean source build at commit
`4971c00dfa488dfac4dbc4f319ca696d8a30eadd`, not a published runtime release.
Token estimates use `o200k_base` and are labeled `inferred`. No provider call,
answer-quality evaluation, complete task, or billed saving was measured.

- [Complete measurements and original content](https://caveman.so/middleware/measurements.json)
- [Method and reproduction instructions](https://caveman.so/middleware/method.txt)
- [Synthetic fixture runner](https://caveman.so/middleware/reproduce.mjs)

## Estimate your input-cost difference

The page calculator uses one measured sample, your monthly number of sends, and
your chosen USD input rate per million tokens. It includes declared recovery-tool
overhead. It excludes output, provider cache pricing, subsequent recovery calls,
retries, runtime costs, and quality changes. The default $3 input rate is an
editable example, not a provider price quotation or a savings promise.

## Keep original detail in reach

Supported tool loops register `caveman_retrieve`, allowing the agent to retrieve
exact original text. Recovery is scoped to application, session, branch, and
cache epoch; handles expire. Your application should keep original history.
The page's recovery control replays captured sample excerpts, not a live runtime.

Images, audio, video, and raw PDF bytes are outside the adapter's text-compression
boundary. Extracted document text can be eligible. User messages, system
instructions, reasoning, error results, and protected content remain outside it.

## Add middleware to your existing agent

TypeScript includes eight adapters: Vercel AI SDK, OpenAI, Anthropic, Google
GenAI, LangChain, Strands, Mastra, and MCP.

Python includes thirteen adapters: LangChain, LiteLLM, OpenAI, Anthropic, Google
GenAI, Strands, Agno, CrewAI, AutoGen, Pydantic AI, LlamaIndex, ASGI, and MCP.
Some names appear in both languages; this is not 21 distinct frameworks.

Start the runtime:

```sh
npm install -g @caveman-ai/cli@1.3.4
caveman setup --install
CAVEMAN_MODE=compress caveman start
```

TypeScript:

```sh
npm install @caveman-ai/sdk@1.1.0 @caveman-ai/middleware@0.1.0-alpha.2
```

Python (3.13+), with LangChain:

```sh
pip install "caveman-sdk==1.1.0" "caveman-middleware[langchain]==0.1.0a1"
```

- [Complete TypeScript quickstart](https://docs.caveman.so/docs/sdk/middleware/typescript)
- [Complete Python quickstart](https://docs.caveman.so/docs/sdk/middleware/python)
- [Compatibility and limitations](https://docs.caveman.so/docs/sdk/middleware/compatibility)
- [Deployment](https://docs.caveman.so/docs/sdk/middleware/deployment)

Use record mode to inspect candidates, compare compress mode on the same tasks,
and retain off mode as a baseline. Compare answer quality, provider costs,
latency, recovery, and retries. Normal runtime unavailability preserves original
inference input; strict mode, startup readiness, cancellation, and failed recovery
have separate error contracts.

## Whole-task evidence is separate

The page also cites the older August 6 wrap-plus-Skill benchmark: 885,793 direct
versus 591,673 Caveman provider-reported input tokens, 33.2% fewer, with 18/18
exact answers in both groups. HTML regressed 9.9%. This older experiment did not
test middleware. Cache buckets were summed without price weighting. Raw harness
and run artifacts are not public. It is not evidence of customer bill savings.

[Read the older report](https://github.com/JuliusBrussee/caveman/blob/8b0c1d3699b8d83e87fe4605b378da20c41555e0/docs/WRAP-BENCHMARK.md)
