Benchmarking Sonnet 5.5: less thinking for the same answers
Measured on our 48-task suite, not on your traffic. At the same effort level, Sonnet 5.5 solved at least as many tasks and usually spent less.
Read07
Research field reports now sit beside letters and release notes. Each piece carries a date and the name of the person who did the work.
Looking for coverage of Caveman?
Visit PressMeasured on our 48-task suite, not on your traffic. At the same effort level, Sonnet 5.5 solved at least as many tasks and usually spent less.
ReadWhat each tool shrinks, how you recover omitted content, telemetry defaults, and what the published numbers measure.
6 min read
Eight TypeScript adapters, thirteen Python adapters, and recoverable tool-result compression inside the agent you already run.
Read articleSeven levers, the measured number behind each, what breaks them, and the cases where they lose money.
8 min readPer-call api_base, fleet-wide config.yaml, Claude Code through both, and a routing callback that never touches inference.
7 min readSame six tasks, 54 runs, an exact-answer check on every one. 33.2 percent against 6.7, and the rows where each lost.
5 min readFour counters, one of them ten times cheaper than the others, and a prefix that decides which one you pay.
4 min read224,655 characters of tool schemas on one first turn. How to measure yours and three ways to bring it down.
5 min read65 percent fewer output tokens on ten prompts, 1,000 to 1,500 input tokens per turn to get it, and plans where it saves nothing.
4 min readThe same two levers as Claude Code, the exact install for each agent, and the meters where token savings do not help.
5 min readPaired tasks, provider counters, a correctness gate, the spread, and four labels that stop an estimate becoming a claim.
5 min readA token counter said 96.2 million saved. Paired trials found a higher bill. The denominator explains both.
7 min readOne real agent loaded 224,655 characters of tool schemas before its first action. Most came from plugins it might never call.
7 min readCompaction can preserve the job while deleting its rules. One benchmark found a 47-token buffer prevented the failure.
7 min readA 998-task replay gained seven net passes while 125 outcomes flipped. Aggregate scores hid both directions.
7 min readNine quality-held pairs produced an attractive diagnostic and three publication blockers. The number stayed out of our claims.
8 min readThe skill now runs at the API level of a green AI platform. No install, nothing to maintain. First platform integration.
2 min readThe price of a token collapsed and the bill went up. What follows from reading that correctly.
15 min readIf we have any of this wrong, that is the more useful reply.
[email protected]