user: "What is the cap... system: "You are a hel... user: "Summarize this... system: "Please note t... user: "Name the Fren... waste: 78% waste: 34% waste: 4% ⚡ cache hit — 93 tokens saved TOKENS SAVED 1.2M this month COST SAVED $847 this month CACHE HIT RATE 43% ↑ from 11%
npm version PyPI version npm downloads VS Code marketplace
LLM token optimization

See inside your
bloated prompts.

Tokoscope audits, compresses, and monitors your LLM token usage across OpenAI, Anthropic, and Gemini so you ship leaner prompts and smaller bills.

Token usage · last 24h live
chat/completions
1.2M
embeddings
640K
after tokoscope
440K
4070%
of tokens in average prompts are waste
$0.003
per 1K tokens adds up fast at scale
3x
typical reduction after prompt compression
Full visibility between
your app and the API.

Drop in one SDK line. Tokoscope sits in the middle, tracks every call, and shows you exactly where money is leaking.

Prompt inspector

Scans your system prompts and inputs for bloat — repeated instructions, redundant context, unnecessary preamble — and scores each one.

HIT
Smart caching

Detects semantically similar requests and serves cached responses. Near-identical prompts stop hitting the API twice.

Auto-compression

Rewrites verbose prompts to their minimum effective form without changing intent. Ships leaner, costs less, still works.

Cost attribution

Break down spend by feature, endpoint, user, or team. Know which part of your product is burning the most — and why.

Budget alerts

Set spend thresholds and get notified via email, Slack, or any webhook before costs spike — not after the invoice lands.

Any LLM, one SDK

Works with OpenAI, Anthropic, and Gemini natively. Any OpenAI-compatible endpoint supported too. One integration, full visibility across all providers.

Works with the tools you already use
OpenAI
Anthropic
Gemini
npm
PyPI
Firebase
VS Code
What developers are saying
"Dropped it into our codebase in 5 minutes. Found out we were wasting 60% of our token budget on one endpoint we'd never thought to optimize."
S
Sean C.
Founder, Lucidity
"The semantic caching alone cut our OpenAI bill by 35% in the first week. Users ask the same questions in different ways constantly."
Y
Yadel F.
AI Operations Founder
"Finally a tool that gives visibility into LLM costs at the feature level. We found our onboarding flow was costing 10x more than our core product."
B
Brad H.
CTO, Power Tech Consulting
Up and running in 2 minutes.
01
Install the SDK
npm install tokoscope
pip install tokoscope
02
Wrap your client
Two lines of code. Works with OpenAI, Anthropic, and Gemini.
03
Watch costs drop
Live dashboard, compression, caching, and alerts — all automatic.
View full docs → Get started free →

Your LLM bill is too high.
Let's fix that.

Start for free — no credit card required. Up and running in 2 minutes.

Get started free → Read the docs

Get token optimization tips and product updates in your inbox.