---
title: "Caveman guides"
description: "Practical guides to AI agent costs, prompt compression, caching, model routing, and evaluations. Commands, checks, and ways to measure the result."
canonical: https://caveman.so/guides
last-updated: 2026-09-16
---

# Caveman guides

Practical guides to AI agent costs, prompt compression, caching, model routing, and evaluations. Commands, checks, and ways to measure the result.

- [How to evaluate an AI agent optimization before rollout](https://caveman.so/guides/agent-evaluations): Build an acceptance test that a cheaper but subtly wrong answer still fails. Write the rejections first, hold back tuning cases, count retries, read the losing runs.
- [Agent observability: join a trace to money and to finished work](https://caveman.so/guides/agent-observability): Your tracer counts calls; only your app knows the job finished. Stamp a task ID and a verdict on every call, reconcile one hour, then read cost per completed task.
- [Add Caveman Middleware to your existing agent](https://caveman.so/guides/agent-sdk-migration): Set up local tool-result compression and original recovery in a TypeScript or Python application without replacing its framework.
- [How to choose an AI gateway for agents](https://caveman.so/guides/choose-ai-gateway): A decision path instead of a feature matrix. Name the one problem you have now, key sprawl, model access, spend limits, caching or context size, then buy for that.
- [Set up Caveman with Claude Code, Codex, and Gemini CLI](https://caveman.so/guides/coding-agent-setup): The skill and the proxy are two separate installs. Put the skill in for shorter replies, the proxy in for smaller input context, then test both on one real coding task.
- [AI gateway migration: a reversible test plan for agent traffic](https://caveman.so/guides/gateway-migration): Move agent traffic to a new gateway without betting the product on it. Run both paths, diff the request shapes that break, cut over one caller, keep rollback tested.
- [Cut the MCP token overhead you pay on every request](https://caveman.so/guides/mcp-token-overhead): MCP tool definitions ride along on every turn whether you use them or not. Measure that floor with a one word prompt, then cut it without breaking tool selection.
- [How to measure AI agent cost per completed task](https://caveman.so/guides/measure-agent-cost): Count every attempt, including the failed ones, into one number. Set the task boundary, record the right fields per attempt, and divide by the tasks that passed.
- [When routing to a cheaper model actually saves money](https://caveman.so/guides/model-routing): A cheaper model loses money two ways: it breaks a warm cache prefix, and a failed cheap attempt gets retried on the expensive one. Price both before you build a router.
- [Prompt caching and compression: reduce agent cost without losing reuse](https://caveman.so/guides/prompt-caching): Check cache reads before you shorten a prompt. How stable prefixes, cache writes and compression interact, and why cold and warm runs need separate numbers.
- [What prompt compression does to your bytes, and how to check it](https://caveman.so/guides/prompt-compression): Caveman stores the original before it sends a smaller view, and passes the original through whenever anything fails. Verify that yourself with one file and one cmp.
- [How to reduce LLM costs in an agent that already works](https://caveman.so/guides/reduce-llm-costs): Pick the one change worth testing first from your own usage data. A symptom to experiment map for caching, compression, tool output, model choice, and retries.
