Skip to content
JayPokalePublic

About

Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

633 stars

Watchers

2 watching

Forks

Latest commit

 

History

257 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Chisle

Chisle

Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.

The only tool in this class that publishes the runs where it lost. Here's why.

npm version Works with 12 agents CI Zero deps MIT Star Chisle on GitHub

Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command

Chisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a PostToolUse hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One npx chisle installs it, and the same ruleset ships to Pi, Cursor, Codex, Gemini, Copilot, OpenCode, Antigravity and four more agents. Every number here comes from committed raw transcripts, including the runs where it lost.

Claude Code coding prompts, 14 cells on Haiku 4.5, as percent of the same model with no ruleset. Answer length: caveman 110%, ponytail 110%, Chisle 67%. Billed tokens: caveman 100%, ponytail 120%, Chisle 76%.

Claude Code 2.1.285 on Haiku 4.5, the 14 coding cells of a 26-cell run. Across all 26 prompts Chisle bills 83% where caveman bills 102% and ponytail 105%, the only one under 100%; on short answers it breaks even. Full numbers below.


"Add debounce to a search input that currently fires an API call on every keystroke." Same model, same prompt, one difference: the injected ruleset. Both answers below are the verbatim committed output from benchmarks/results/raw/:

bare agent: 142 lines, 1506 tokensChisle: 35 lines, 602 tokens

Opens with "Let me show you the most common approaches", then ships a reusable generic useDebounce<T> hook in its own file…

// useDebounce.ts
export function useDebounce<T>(
  value: T, delay: number
): T {
  const [debouncedValue, setDebouncedValue]
    = useState<T>(value);
  useEffect(() => { /* … */ }, [value, delay]);
  return debouncedValue;
}

…then Option 2 and Option 3, a comparison table, and a caveats section.

Asks which framework, then answers the question that was actually asked: setTimeout in the effect you already have, no new file, no generic:

useEffect(() => {
  const timer = setTimeout(async () => {
    if (query.trim()) { /* fetch */ }
  }, 300);
  return () => clearTimeout(timer);
}, [query]);

Then two lines on why it works, and "use lodash.debounce if already installed."

Not golfed, boring. Same behaviour, one less abstraction, no second file, and it names the dependency you might already have instead of reinventing it.

What it compresses

axis what how
Output: prose filler, hedging, manufactured structure zero-fluff ruleset, injected per session
Output: code speculative abstractions, unrequested boilerplate YAGNI efficiency ladder
Input: context oversized tool output flooding the window Claude PostToolUse / Pi tool_result: scrub, elide, dedup, plus prevention rules

Where each one attaches to a session:

flowchart LR
    subgraph S["Session start"]
        H1["Ruleset injection<br/>once per active session"]
    end
    subgraph T["Every turn"]
        H2["Mode tracking<br/>Claude + Pi"]
    end
    subgraph L["Every tool call"]
        H3["PostToolUse / tool_result<br/>scrub → elide → dedup"]
    end

    H1 --> M(["Model"])
    H2 --> M
    M -->|writes| O["Output:<br/>terser prose,<br/>YAGNI-first code"]
    M -->|calls a tool| TOOL[["Bash / grep / web / extension tools"]]
    TOOL -->|raw output| H3
    H3 -->|"compressed, rebuilt into<br/>the tool's own shape"| M

    RE["Read / Edit / Write"] -.->|"never touched,<br/>exact bytes feed later edits"| M

    style M fill:#1f2937,stroke:#d78a3c,color:#e6edf3
    style O fill:#14532d,stroke:#2da44e,color:#e6edf3
    style H3 fill:#1f2937,stroke:#2da44e,color:#e6edf3
    style RE fill:#3f1d1d,stroke:#cf3b3b,color:#e6edf3
Loading

The loop on the right is the input axis: tool output is billed again on every later request in the session, so shrinking it once pays repeatedly. Read, Edit, and Write are deliberately outside it.


Install

One command. Auto-detects your agents (Claude Code, Pi, Cursor, Windsurf, Cline, Kiro, Antigravity, Codex, Gemini, Copilot, OpenCode, Hermes) and wires each one. --uninstall puts everything back.

npx chisle
# or via curl
curl -fsSL https://raw.githubusercontent.com/JayPokale/Chisle/main/install.sh | bash
# Windows
irm https://raw.githubusercontent.com/JayPokale/Chisle/main/install.ps1 | iex

Preview first with npx chisle --dry-run, scope with --only claude or --only pi, see everything with npx chisle --help. Remove with npx chisle --uninstall.

Upgrading, per-agent setup, --stats and config: docs/usage.md.


Benchmarks

Claude Code 2.1.285 on Haiku 4.5, 13 live prompts × 2 seeds = 26 cells per arm, billed output tokens vs the same model with no ruleset (writeup + raw cells):

2026-10-01 live rerun: total billed output as percent of the bare model. caveman 102%, ponytail 105%, Chisle 83%.

total bill 95% CI visible answer worst cell backfires
caveman 102% 79–128% 105% 305% 15 / 26
ponytail 105% 86–129% 97% 493% 16 / 26
Chisle 83% 69–95% 76% 170% 11 / 26

Chisle is the only arm below a bare model. It pays on long answers (77%) and coding prompts (76%); on short answers it breaks even (106%).

On Pi + GPT-5.5 Chisle again gives the shortest answers, 37% the length of a bare model's (caveman 56%, ponytail 49%), though on total billed tokens ponytail edges it, 59% to 61%:

Pi benchmark, 6 tasks on GPT-5.5, as percent of the same model with no ruleset. Answer length: caveman 56%, ponytail 49%, Chisle 37%. Billed tokens: caveman 72%, ponytail 59%, Chisle 61%.

Input side: tool output is 67.5% of context in 171 measured Claude Code sessions, and it is re-billed on every later request. The compressor cut ~46% off every eligible output there, and 27.4% of tool output on top of Pi's own truncation (receipts).

prose code judgment input/context worst-case guard publishes failures
caveman ✅ ❌ ❌ ❌ 305% ❌
ponytail ❌ ✅ ❌ ❌ 493% ❌
headroom ❌ ❌ ✅ proxy n/a ❌
Chisle ✅ ✅ ✅ hook 170% ✅

Every table, chart and caveat, including the June suite and the Pi run in full: docs/benchmarks.md. Head to head: docs/comparison.md.


How it works

Before writing code, the agent stops at the first rung that holds:

flowchart TD
    A[Request for code] --> R[Read the problem fully]
    R --> Q1{Does this need<br/>to exist at all?}
    Q1 -->|no| S1[Skip it. Say so in one line]
    Q1 -->|yes| Q2{Already in<br/>this codebase?}
    Q2 -->|yes| S2[Reuse it. Don't rewrite]
    Q2 -->|no| Q3{Stdlib<br/>does it?}
    Q3 -->|yes| S3[Use the stdlib]
    Q3 -->|no| Q4{Native platform<br/>feature covers it?}
    Q4 -->|yes| S4["CSS over JS, DB constraint<br/>over app code"]
    Q4 -->|no| Q5{Already-installed<br/>dependency?}
    Q5 -->|yes| S5[Use it. Never add a new dep<br/>for what a few lines do]
    Q5 -->|no| Q6{Can it be<br/>one line?}
    Q6 -->|yes| S6[One line]
    Q6 -->|no| S7[The minimum code that works]

    S1 & S2 & S3 & S4 & S5 & S6 & S7 --> OUT[Ship it + note what was skipped<br/>and when to add it]

    style Q1 fill:#1f2937,stroke:#d78a3c,color:#e6edf3
    style OUT fill:#14532d,stroke:#2da44e,color:#e6edf3
    style R fill:#1f2937,stroke:#8b949e,color:#e6edf3
Loading

The ladder runs after reading, never instead of it. Note the exit: every rung lands on the same obligation: say what you skipped, so "later" doesn't quietly become "never".

The ladder runs after reading the code, lazy about the solution and never about understanding. Lazy is not negligent: trust-boundary validation, data-loss handling, security, and accessibility are never on the chopping block.

Mark deliberate simplifications so "later" doesn't quietly become "never":

// chisle: global lock, per-account locks if throughput matters
// chisle: O(n) scan, index this when table exceeds ~10k rows

Usage

Command Effect
(nothing) On automatically every session after install
/chisle Re-activate if you'd stopped it
/chisle off Deactivate
stop chisle Deactivate (ruleset and input-side compression)
normal mode Deactivate

Natural language works too: "activate chisle", "chisle mode", "chislify this". Code symbols, function/API names, and error strings stay verbatim, so only the noise around them compresses.


Contributing

See CONTRIBUTING.md. Built by Jay Pokale with Claude, Antigravity, and Codex as co-engineers.

Star History

Star History Chart

License

MIT. The shortest license that works.

Contributors

GitHub contributors

AI co-engineers (pair-work credited in commit trailers and the changelog):

Claude (Anthropic) Codex (OpenAI) Antigravity (Google)


Saved you tokens? ⭐ Star the repo. It costs zero tokens and keeps the benchmarks running.

About

Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

633 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages