38-59% Fewer Output Tokens
680+ real API calls across Haiku, Sonnet, Opus 4.6, and Opus 4.7. A/B tested with automated quality scoring. 38-59% average output reduction per model (up to 90% on verbose prompts). Net savings on your bill: 20-48%.
Important: These results are from our internal testing. Your actual savings will vary based on your usage patterns, prompt types, and workflow. We encourage you to run your own tests to verify results for your specific use case.
Three Integration Modes
Superstack works through three different mechanisms depending on your setup. All modes deliver the same output optimization benefits.
Intercepts your prompts before they reach Claude and injects optimized output rules that reduce response verbosity.
Model Context Protocol server that tools like Cursor can connect to for optimized context injection.
Local proxy that sits between your tools and Claude API. Adds prompt caching for 90% savings on repeated context.
Results by Model
April 2026 benchmark · 47 prompts × 4 models. Prior-generation models (Sonnet 4.5, Opus 4.6/4.7) — current-model July 2026 results below.
59%
average reduction
up to 91% on best prompts
38%
average reduction
up to 89% on best prompts
49%
average reduction
up to 90% on best prompts
28%
average reduction
up to 74% on best prompts
Quality Improved Across All Models
All models tested on single-turn show positive quality deltas (+4.4 to +8.0 points average). Superstack reduces verbosity while improving conciseness, structure, and focus — responses are both shorter and better.
Current-model results
Re-run on the latest models at equal effort (high). Two numbers per model: output reduction, and net savings on your bill (input-inclusive).
73%
fewer output tokens
58%
net savings
59%
fewer output tokens
43%
net savings
59%
fewer output tokens
39%
net savings
40%
fewer output tokens
20%
net savings
Aggregate (bill-weighted) across the 12 prompts every model shares (code-generation-heavy, so these run a little higher than the full 47-prompt set). n=12; a controlled spot-check that keeps effort equal across models. Full detail and confidence intervals in the methodology.
Results by Task Complexity
We grouped our 47 test prompts into 4 difficulty levels. Harder tasks benefit more from Superstack.
Tier 1
Bounded Tasks
46%
output reduction
+6.6 quality
13-15 prompts/model
Code gen, explanations
Tier 2
Project-Aware
47%
output reduction
+8.3 quality
12 prompts/model
Auth, billing, webhooks
Tier 3
Multi-File
35%
output reduction
+6.7 quality
10 prompts/model
Refactoring, pipelines
Tier 4
Debugging
48%
output reduction
+3.3 quality
10-12 prompts/model
Root cause analysis
Why T2 Has the Highest Quality Gain
Project-aware tasks (auth, billing, webhooks) benefit most from Superstack's coding standards injection. Without it, responses include verbose boilerplate and lengthy explanations. With it, Claude stays focused on the project's conventions and produces concise, contextual code.
Results by Task Type
Total token volume saved across Haiku, Sonnet, Opus 4.6, and Opus 4.7 · 188+ A/B comparisons
General Questions
Advisory & best practices
up to 88%
+5.1 quality
Debugging
Root cause focus directives
up to 89%
+4.9 quality
Implementation
Project-aware features
up to 86%
+6.1 quality
Code Generation
Bounded code tasks
up to 90%
+6.6 quality
Refactoring
Already concise output
up to 47%
+0.8 quality
Multi-Turn Session Savings
10 real dev tasks (RBAC, webhooks, caching, etc.) tested with and without Superstack
Haiku 4.5
19%
total token reduction
110K → 89K tokens · 8W/1L/1T
Best: webhooks -46%
Sonnet 4.5
46%
total token reduction
216K → 118K tokens · 7W/2L/1T
Best: webhooks -81%
Opus 4.7
28%
total token reduction
237K → 171K tokens · 6W/4L/0T
Best: webhooks -74%
Savings scale with task complexity — larger tasks like webhooks (-46% to -81%) and RBAC (-16% to -70%) consistently show the strongest reductions across all models. Turn count often drops too (e.g., 6 turns → 3 on webhooks).
MCP Tool Aggregator Savings
Instead of sending all tool definitions every request, Superstack sends one search tool that finds the right tool on demand
4,056
input tokens per request
702
input tokens per request
Haiku 4.5
82.7%
input saved
45% cost saved
Sonnet 4.5
85.5%
input saved
79% cost saved
Opus 4.6
85.4%
input saved
79% cost saved
Opus 4.7
85.2%
input saved
79% cost saved
Consistent Across All Models
MCP aggregator savings are model-agnostic — 82-86% input reduction verified across all 4 Claude models with real API calls (23 downstream tools from filesystem and memory servers). With live tool caching, actual savings reach 97.5% on real MCP servers.
Cache Proxy Savings
Automatically caches repeated context (system prompts, coding rules) so you don't pay for it twice. API users only — subscription users already benefit from hooks.
88%
cache hit rate
100%
cache hit rate
Stacked Savings for API Users
Output optimization (38-59% token reduction), MCP aggregator (82-86% tool token savings), and cache proxy (15-18% input savings) all stack together. API users can see 40-60%+ total cost reduction when using all three.
Code Quality Under Real Conditions
Claude sometimes modifies files it shouldn't, adds features you didn't ask for, or creates unnecessary code. We tested whether Superstack's coding standards injection keeps Claude focused on exactly what you asked.
No Difference
100% discipline — both with and without Superstack
When tests already exist, Claude naturally stays on track. The tests act as guardrails. Superstack doesn't slow it down or get in the way.
83% → 100%
discipline score with Superstack
Without test guardrails, bare Claude occasionally modified files it shouldn't have, added unnecessary files, or created endpoints that weren't asked for.
No-Test Tasks: Bare Claude vs Superstack
Bare Claude (no Superstack)
With Superstack
Why This Matters
Most real projects don't have comprehensive test suites for every feature. When you ask Claude to "add search functionality" or "fix the error handling," there's nothing stopping it from touching files it shouldn't or adding things you didn't ask for. Superstack injects your project's coding standards so Claude respects boundaries — even when there are no tests to enforce them.
8 tasks | 78 total runs | 3 runs per condition | Real Node.js APIs with Express | Post-hoc test verification | Full results at tests/velocity/
What This Means For You
Direct cost savings on every API call. Output tokens cost 5x more than input, so reducing output has outsized impact on your bill.
Monthly savings (1,000 prompts/day)
$69
Haiku
$385
Sonnet
$3,280
Opus
Output optimization only. Add 15-18% more with cache proxy.
Even with flat-rate subscriptions, you have usage limits. Across a working session — where Superstack's context is cached and reused turn after turn — shorter responses stretch your usage further.
With 38-59% output reduction, over a session
Test Methodology
Real API Calls
All tests used actual Claude API calls; token counts come from API response metadata, not estimates. The April 2026 run used temperature=0 for reproducibility. Current models (Sonnet 5, Opus 4.8, Fable 5) do not accept temperature=0, so the July 2026 re-run uses an N-trial paired design with reported confidence intervals instead.
A/B Testing
Each prompt tested twice: once without optimization (baseline) and once with Superstack's output rules injected. Same model, same temperature, same max_tokens.
Automated Quality Scoring
Each response scored on completeness (25%), conciseness (20%), structure (15%), preamble avoidance (15%), efficiency bonus (25% for code tasks). Conciseness uses a flat 100 for any response within the acceptable range — no bias toward shorter responses.
Cross-Model + Multi-Turn + MCP + Cache Proxy
47 prompts tested on Haiku, Sonnet, Opus 4.6, and Opus 4.7. Multi-turn tested on 3 models with 10 real dev tasks. MCP aggregator A/B tested across all 4 models. 870+ total API calls across all dimensions.
Instruction Framing Research
144-call study testing 8 different ways to phrase the same rules. Found that instruction framing alone can shift output by up to 20 percentage points. Results used to continuously improve Superstack's injection quality.
Governance Compliance Scoring
4-category governance compliance: canary token preservation, structural directive adherence, scope boundary respect, and output directive following. Each response independently scored to verify Superstack maintains safety constraints while optimizing output.
680+ API calls | 141 single-turn A/B tests | 4 models | Test date: April 2026 | Full results and benchmark code available at tests/benchmark/ and tests/velocity/