Stay close to the code.
Read, edit, and run commands in your working directory. Keep command sessions between steps.
A coding agent. At home in your terminal.
Big ideas. Small prompt.
Read the repo. Make the change. Check the work.
The retry helper runs one time too many. Fix it.
On it. Let’s make it work.
Retry limit fixed. Boundary cases covered.
export async function retry(run, maxAttempts) { for (let n = 0; n <= maxAttempts; n++) { for (let n = 0; n < maxAttempts; n++) { const result = await run(); if (result.ok) return result; } throw new Error("Retry limit reached");}Built for the way you work
Read, edit, and run commands in your working directory. Keep command sessions between steps.
Connect your provider credentials. Choose the model you want to work with.
Your repository instructions, skills, and supported plugins come with you.
An early signal
Terminal-Bench 2 pilot / 9 September 2026Three tasks. Same model. One attempt each.
estimated API spend
vs. Codex CLI, failures included
tasks passed
Codex CLI passed 2/3
Both agents used GPT-5.6-Sol with medium reasoning. Pebble had programmatic tools and speculation enabled. Routing, skills, plugins, browsing, personal memory, and notebook rollovers were off.
Both passed async cancellation and C++ heap repair. Pebble also passed gRPC; Codex’s server was not running when checked. Runtime was nearly identical; Pebble used more output tokens and API requests.
Costs are estimates, not invoices. Three preselected tasks and one attempt per agent do not establish a general advantage. Raw traces and verifier artifacts are not published here.
| Metric | Codex CLI | Pebble |
|---|---|---|
| Tasks passed | 2/3 | 3/3 |
| Estimated API spend | $0.5916 | $0.4073 |
| Cost per passed task | $0.2958 | $0.1358 |
| Input tokens, including cached | 437,787 | 206,500 |
| Output tokens | 8,917 | 10,112 |
| API requests | 31 | 35 |
| Runtime | 355.7s | 355.0s |
Pebble is taking shape. Be there when it lands.
Join the waitlist. We’ll let you know.