Topic

Tools for controlling AI token cost

217

we track

214

answered today

10

arrived this week

0

stopped answering

3 we couldn't reach today; we don't claim those are dead. Watching this group for 26 days, checked once a day · how we check

Know your usage pace and resets. Let Codex help you plan.

FreeOpen source
What to know
How it works
İşlemi ayırır (detach), konuşmayı serbest bırakır, sınırlı çıktı kaydeder ve tamamlanınca Codex'e inceleme için döner.
Details →

CodexUsageBar

LiveWatched 13d

A free, open-source macOS menu bar app that tracks your ChatGPT Codex usage limits in real time for developers.

FreeOpen source
What to know
How it works
Lives in the macOS menu bar and shows current Codex usage at a glance.
Pricing
Free and open-source.
Details →

Decentralized on-chain DEX quote aggregator and staking token across several blockchain networks.

What to know
How it works
On-chain DEX aggregator for token swaps and a staking engine for its own BZPX token, across Base, Ethereum, Optimism, and Arbitrum.
Pricing
Free.
Details →

ReserveGauge

LiveWatched 20d

A macOS menu-bar app that watches your AI coding usage on Codex or Claude Code and forecasts whether you'll hit your limit before it resets, all processed locally.

What to know
How it works
A macOS menu-bar app that forecasts your remaining usage for AI coding tools like Codex and Claude Code and flags whether your current pace will last until the next limit reset.
What's different
Processes usage data locally on-device rather than sending it to a server.
Best for
Developers on usage-limited AI coding plans who want to avoid unexpectedly hitting their limit.
Details →

AI Usage Ball

LiveWatched 6d

Live desktop orbs for Claude, Codex & Antigravity

FreeWorks offlineNo subscription
What to know
How it works
Kalan kullanım hakkı ve sıfırlanma zamanını yerelde izler, yalnız sağlayıcıların kendi servisleriyle konuşur; kaynak denetime açık.
Pricing
İlk 100 kullanıcıya ücretsiz ömür boyu kişisel lisans fırsatı.
Details →

LLMSlim

LiveWatched 17d

An open-source Python library that compresses LLM prompts and context to cut token costs while preserving instructions.

Open source
What to know
How it works
Compresses prompts, RAG document contexts, and multi-turn chat logs in one line of code.
Pricing
Open-source Python package.
Best for
Developers who want to cut LLM token costs 40-70% without losing instruction fidelity.
Details →

MCP-Recall

LiveWatched 9d

Compresses MCP tool outputs and stores full results in SQLite so agent sessions fit more tool calls.

What to know
How it works
Compresses MCP tool outputs (claims 94 KB to 3.5 KB, 96%) and stores full results in SQLite for retrieval, enabling up to 30x more tool calls per session.
Details →

Spanlens

LiveWatched 6d

Open-source, self-hostable observability platform for LLM apps: request logging, cost tracking, agent tracing.

Open sourceSelf-hosted
What to know
How it works
Works with OpenAI, Anthropic, and Gemini apps.
Pricing
Open source, MIT license.
Details →

SolReclaim

LiveWatched 15d

Recover your SOL account rental fees with one click.

Watch out · Includes a referral program paying 60% of the service fee in SOL — a financial incentive structure worth noting.

What to know
How it works
Closes unused Solana token accounts and recovers the locked SOL rent deposit (about 0.002 SOL per account) held in each one.
Best for
Solana wallet holders with leftover unused token accounts holding small locked rent deposits.
Watch out
Includes a referral program paying 60% of the service fee in SOL — a financial incentive structure worth noting.
Details →

Pylva

LiveWatched 12d

Tracks the real infrastructure cost of each customer for AI agent builders, so they can set margin caps per customer.

Open source
What to know
How it works
SDK-only telemetry tracks LLM and non-LLM usage per customer
What's different
open source, reactive rules like per-customer cost caps
Pricing
open source
Best for
builders of AI agent products billing by usage
Details →

Codex Quota Panel

LiveWatched 10d

A free open-source Windows dashboard that shows your Codex usage quota, reset timers, and token counts at a glance.

FreeOpen sourceNo subscription
What to know
How it works
Displays weekly quota, reset time, reset-credit cards, and lifetime token usage for Codex
Pricing
Free and open-source
Details →

Shibaita.ai

LiveWatched 7d

Lets Claude Code users compare usage on a public leaderboard instead of only seeing their own local stats.

Watch out · unofficial, not affiliated with Anthropic

Free
What to know
How it works
A CLI reads local logs and submits only aggregate token counts, never prompts or code.
Pricing
free
Watch out
unofficial, not affiliated with Anthropic
Details →

KlaatAI

LiveWatched 9d

Routes prompts to the cheapest capable model tier and offers a terminal coding agent claiming lower cost than Claude Code on a benchmark.

Free
What to know
How it works
Routes every prompt to the cheapest model tier that can handle it, escalating only when needed; its terminal agent Klaat Code claims to solve the same 33-task benchmark as Claude Code at 5.4x lower cost ($0.027 vs $0.146/task).
Pricing
Free beta: 75 requests/day, no card; tool calls are free, only messages count against quota.
Details →

LoRA Speedrun

LiveWatched 12d

Leaderboard site where people compete to fine-tune the same model fastest, with every record re-verified by a referee.

Free
What to know
How it works
same model, same score, same GPU — fastest training run wins, re-run 3x by a referee
What's different
no self-reported times, all records independently verified
Pricing
free to attempt
Details →

An offline iPhone app that generates AI images locally using Stable Diffusion, for people who want image generation without sending prompts to the cloud.

Works offline
What to know
How it works
Runs Stable Diffusion locally on iPhone, fully offline with no internet connection or API costs.
Best for
iPhone users who want AI image generation without sending data to the cloud.
Details →

CC Theme

LiveWatched 11d

An open-source Mac app providing ready-made, reproducible themes for AI desktop apps without repeated prompt generation.

Open source
What to know
How it works
Applies validated, ready-made theme packages to supported AI desktop apps, avoiding repeated image-and-prompt generation.
Details →

nable

LiveWatched 6d

One usage meter for every AI provider you pay for

FreeNo sign-upWorks offline
What to know
How it works
Kullanım yerel oturum kayıtlarından okunur, hesap/anahtar gerekmez; sabit planlarda kullanım, ölçülü planlarda tahmini dolar gösterir.
Pricing
Açık kaynak ve sonsuza dek ücretsiz.
Details →

ProxAI

LiveWatched 9d

Reverse proxy that enforces hard spending caps on LLM API calls across providers by swapping the base URL.

Watch out · AGPL-3.0, no telemetry.

FreeNo trackingSelf-hosted
What to know
How it works
A transparent reverse proxy enforcing hard hourly/daily spending caps per LLM API client via a one-line base_url change, returning 429 once the cap is hit rather than calling the upstream.
What's different
Works with OpenAI, Anthropic, Mistral, Groq, DeepSeek, and 5 more providers.
Pricing
Hosted beta at proxai.eu is free, no credit card; self-hostable as a ~10 MB Docker image.
Watch out
AGPL-3.0, no telemetry.
Details →

Toolport

LiveWatched 26d

A local gateway that lets every AI coding client share one set of MCP server connections instead of configuring each tool separately, cutting how many tools your agent has to load.

Watch out · The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.

FreeNo sign-upWorks offline
What to know
How it works
A local-first, open-source gateway. You set up and authenticate each MCP server once, and every connected AI client — Claude, Cursor, VS Code, Codex — shares that same connection instead of you configuring it per tool.
What's different
Claims up to 91% less tool-token overhead by giving an agent a handful of meta-tools instead of loading hundreds of individual tool definitions; API keys stay in the OS keychain.
Pricing
Free, open source, no sign-up.
Best for
Developers running multiple AI coding clients against the same set of MCP servers who don't want to re-authenticate each one separately.
Watch out
The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.
Details →

Benchmark and gateway that compresses context to cut Codex token usage and session cost.

What to know
How it works
A context-compression gateway for Codex that reduces token usage and boosts cache hit rate.
What's different
Benchmarked at 49.5% fewer input tokens, cache hit rate up from 76.1% to 85.4%, and 35.6% lower total session cost.
Details →

Oriveo

LiveWatched 14d

A synced multi-model AI chat app across iOS, Android, and web using your own API keys.

FreeWorks offline
What to know
How it works
Lets you chat with ChatGPT, Claude, Gemini, and 10+ other AI models in one app across iOS, Android, and web, synced in real time, using your own API keys.
What's different
Includes real-time cost tracking, local-first storage, saved Notes, and cross-checking replies with a second model.
Pricing
Try free, no key needed; no markup or extra subscription with your own keys.
Details →

Tokenez

LiveWatched 6d

Counts your AI usage locally across ChatGPT, Claude, Gemini, Claude Code and Cursor in one dashboard, showing your real total and how close you are to each limit.

FreeNo sign-up
What to know
What's different
Counts locally rather than relying on each provider's own scattered usage view.
Pricing
Free, no signup to try.
Details →

Token Guard

LiveWatched 8d

Tracks and converts LLM API spend across OpenAI, Anthropic and Gemini locally, instead of trusting a cloud dashboard with your usage data.

FreeWorks offlineOpen source
What to know
Pricing
Free and open source.
Best for
Developers juggling multiple LLM provider SDKs and costs.
Details →

Alenia Porter

LiveWatched 15d

A high-performance batch image/video/audio optimizer with GPU acceleration, for developers and creators needing fast local media compression instead of a cloud transcoding service.

What to know
How it works
A universal media optimizer that compresses images, video, and audio, built with Nuitka and Go, with auto-GPU acceleration and a crash-safe fallback.
What's different
Runs with GPU acceleration and a crash-safe fallback instead of relying on a cloud transcoding service.
Best for
Developers and creators who need fast batch compression of images, video, and audio.
Details →

Levitan

LiveWatched 11d

A free, fully offline voice-to-text tool that types what you say into any app on Windows or Mac, running locally with no cloud upload.

FreeWorks offline
What to know
How it works
hold a key, speak, text appears at cursor in any app; runs locally on CPU
Pricing
free forever
Best for
privacy-conscious users who don't want cloud-based voice typing
Details →

Generates a 3D model from a single image entirely on-device using WebGPU, with no server uploads or accounts needed — an open-source tool for developers experimenting with local image-to-3D generation.

Works offlineData stays with youOpen source
What to know
How it works
Generates a 3D model from a single image entirely on-device using WebGPU compute in the browser, with no server uploads or accounts.
What's different
Open source and runs locally via WebGPU rather than requiring cloud GPU processing, also deployable on Vercel for instant testing.
Best for
Developers experimenting with local, private image-to-3D generation.
Details →

Claude Usage Tracker

LiveWatched 10d

A Chrome extension showing your Claude usage limits, reset timers, and a projected time-left-at-this-pace right in the chat interface.

No sign-upWorks offline
What to know
How it works
tracks session/weekly limits locally and projects remaining time at current usage rate
What's different
fully on-device, no account needed
Pricing
free
Details →

MutualGPU

LiveWatched 10d

Lets people with spare GPU capacity share it so others without enough local compute can run demanding AI workloads.

Free
What to know
How it works
community compute-sharing network; first use case runs 3D Gaussian Splat generation via WebGPU
Pricing
free
Details →

Velokey

LiveWatched 14d

A pay-per-token API gateway giving developers access to 100+ AI models without a subscription commitment.

No subscription
What to know
How it works
Pay-per-token API gateway giving access to 100+ AI models.
Pricing
Pay per token, no subscription.
Best for
Developers.
Details →

Avery

LiveWatched 17d

Describes an agent in plain English and it builds a deterministic version that runs locally on your own hardware instead of a metered cloud API, for privacy-conscious teams that can't send sensitive work to the cloud.

Works offline
What to know
How it works
You describe an agent in one plain sentence, and Avery builds a deterministic version of it that runs on your own hardware rather than a metered public cloud, with no code required.
What's different
Runs locally and deterministically instead of routing through a per-token cloud API, aimed at avoiding cost, privacy, and lock-in problems of typical agentic AI projects.
Best for
Teams that want private, auditable automation they own outright, without sending sensitive work to a cloud API.
Details →

AI Usage Tracker

LiveWatched 12d

Local-first dashboard that auto-discovers your Claude and Codex spend across 10+ developer tools, with charts and zero telemetry.

Works offlineData stays with youNo tracking
What to know
How it works
auto-discovers usage across 10+ dev tools locally
What's different
zero telemetry, data stays on your machine
Best for
developers tracking AI tool spend
Details →

Meterbility

LiveWatched 10d

A debugger for AI agent runs that shows exactly what context each model call saw, for developers troubleshooting agents built with Claude Code, Codex, or Cursor.

Works offlineOpen source
What to know
How it works
Captures every step of an agent run, compares two runs, and can fork from any step to test changes
What's different
Local-first and open source, so traces never leave your machine
Best for
Developers debugging or comparing AI agent runs
Details →

A blockchain development agency building smart contracts, enterprise dApps, and token infrastructure.

What to know
How it works
An agency building custom blockchain products — smart contracts, enterprise dApps, token infrastructure and Web3 integrations — tailored to a client's commercial objectives.
Best for
Startups and enterprises wanting a development team to build production-ready blockchain products.
Details →

Codex Limits

LiveWatched 12d

A Mac menu-bar app that tracks your Codex usage limits, token activity and reset times across multiple accounts.

What to know
How it works
Native macOS menu-bar app and widget showing Codex rate limits, credits and reset times
Details →

UltraWork

LiveWatched 12d

Hosted coding environment with flat monthly pricing instead of pay-per-token billing for AI-assisted coding.

What to know
How it works
Flat-rate hosted environment with curated models and API-key access
What's different
Predictable flat pricing instead of per-token billing
Best for
Developers who want to avoid unpredictable AI token costs
Details →

Gives voice-AI developers one account to meter spend across their whole vendor stack (LLM, speech-to-text, text-to-speech, telephony) with spend caps and kill-switches.

Open source
What to know
How it works
Server-side spend caps and per-agent kill-switches across multiple vendor APIs under one key
What's different
Open-source floe-guard component
Best for
Developers building voice AI products across several paid vendor APIs
Details →

ApiCredits

LiveWatched 9d

A prepaid AI API gateway giving one key to access multiple LLM providers at a lower cost than direct billing or aggregators.

No subscription
What to know
How it works
A single key and endpoint gives access to models from OpenAI, Anthropic, Google, xAI, DeepSeek, and Z.ai via OpenAI-, Anthropic-, and Gemini-compatible APIs, with built-in route health, load distribution, and fallback.
What's different
Priced lower than direct provider billing or aggregators like OpenRouter.
Pricing
Prepaid credits starting with a $5 test, no subscription or expiry.
Details →

OSGARD NEW WORLD

LiveWatched 9d

A tool that generates a themed sci-fi/fantasy AI universe with tradable in-game artifacts and a leaderboard.

FreeNo sign-up
What to know
How it works
Generates a themed AI universe (sci-fi, fantasy, cyberpunk or steampunk) with tradable artifacts in about 7 seconds, with a live marketplace, TC token staking and an AI assistant called JARVIS.
Pricing
3 free generations, no signup needed.
Best for
People wanting to generate and trade AI-created world assets.
Details →

GPUCPUCompare

LiveWatched 9d

A free tool that compares GPU/CPU hardware specs and checks compatibility against a large game library.

FreeNo sign-up
What to know
How it works
Compares 733 GPUs and 1,002 CPUs and checks compatibility against 112,113 PC games using public TechPowerUp, Geekbench and Steam data.
What's different
Uses public benchmark data instead of unexplained percentages or simulated benchmarks.
Pricing
Free, no signup.
Best for
PC gamers deciding what hardware to buy.
Details →

PandaCoder

LiveWatched 9d

Free local AI coding assistant positioned as an alternative to token-limited cloud coding tools.

Free
What to know
How it works
Runs locally as a coding assistant with no token limits.
Pricing
Free.
Best for
Developers wanting a private, local alternative to token-limited cloud coding tools.
Details →

Ready-made clone script/template for building a DEX token-tracking platform.

What to know
How it works
A ready-to-develop template for building a platform that tracks token prices, trading pairs, liquidity, and volume across multiple DEXs.
Best for
Businesses wanting to launch a DEX tracking platform faster.
Details →

HygeiaCloud

LiveWatched 9d

Resells certified decommissioned enterprise GPUs (A100/H100/H200) at a discount to AWS/GCP pricing.

What to know
How it works
Intercepts decommissioned enterprise GPUs (A100/H100/H200) before they're shredded, certifies each unit through a 72-hour burn-in, and redeploys them.
What's different
40-70% below AWS and GCP pricing, with zero new hardware manufactured.
Best for
Buyers needing serious GPU compute without new-hardware cost.
Details →

Find AI Credits

LiveWatched 9d

Free directory aggregating AI, cloud, API and GPU credit offers and startup perks.

Free
What to know
How it works
Free directory aggregating free AI, cloud, API, and GPU credit offers and startup perks from AI companies in one place.
Pricing
Free.
Details →

ChorusGraph

LiveWatched 23d

An open-source Python agent runtime, installed with one pip command, with a local cache that skips repeat LLM calls — pitched as a lighter alternative to LangGraph for developers.

Open source
What to know
How it works
An open-source agent graph runtime with a BSP scheduler and a local semantic cache that skips redundant LLM calls when a similar routing request already happened, installed with a single pip command.
What's different
Positioned as a lighter, non-wrapper alternative to LangGraph, with a local cache to cut LLM call costs.
Pricing
Open source (Apache-2.0).
Best for
Developers building LLM agent pipelines who want to reduce redundant model calls.
Details →

NagpurAdsWala

LiveWatched 10d

A performance marketing agency in Nagpur offering Meta Ads, Google Ads, SEO, and lead generation services.

What to know
How it works
A performance marketing agency in Nagpur offering Meta Ads, Google Ads, SEO, landing page optimization, and lead generation.
Details →

Visually shows what a million AI tokens actually amounts to in books, code, or dollars, and compares model pricing.

Free
What to know
How it works
Converts token counts to relatable units and compares input/output pricing across models
Pricing
Free
Details →

Streaklee

LiveWatched 10d

A macOS menu bar app that shows how much of your paid Claude/ChatGPT usage quota you've used before it resets.

Works offline
What to know
How it works
Displays live usage windows, weekly capacity and reset times for AI subscriptions
What's different
Local-first, runs from the menu bar
Best for
People paying for Claude/ChatGPT/Codex who want to use their quota before it resets
Details →

Touki Blockchain

LiveWatched 10d

A no-code platform to create, deploy, verify, and use your own blockchain token.

What to know
How it works
Lets you create, deploy, verify, and use your own token on the Touki blockchain without needing development skills.
Best for
People who want to launch a token without being a developer.
Details →

Generates 3D Gaussian Splats in your browser, for free, using either your device's GPU or a community-shared compute network.

FreeNothing to install
What to know
How it works
runs via WebGPU locally, or offloads to MutualGPU's crowd-sourced GPU network
What's different
no install, and free generation even without a powerful local GPU
Pricing
free
Details →

RamShared

LiveWatched 10d

Lets Linux/WSL2 users use idle GPU VRAM as extra swap memory without freezing the OS.

Watch out · Linux and WSL2 only

What to know
How it works
Cascading tier of zram, GPU VRAM, then NVMe with active latency monitoring and page demotion
What's different
Crash-safe unlike naive GPU swap tools, built with Rust/ublk/NBD drivers
Watch out
Linux and WSL2 only
Details →

AI Karma Tracker

LiveWatched 14d

A browser toolbar popup that shows live usage limits and reset times across the AI tools you use, with peak-hour alerts.

Watch out · No API keys — nothing leaves your browser.

Free
What to know
How it works
A toolbar popup showing live usage, remaining limits, and reset times for the AI tools you're signed into, with peak-hour nudges.
Pricing
Free for one tool, $2/mo for your whole stack.
Watch out
No API keys — nothing leaves your browser.
Details →

Ceiling

LiveWatched 14d

A Windows taskbar utility that tracks remaining usage limits and reset times across Claude, Codex, Cursor, and other AI tools.

Watch out · Windows only; works with Claude, Codex, Cursor, Gemini, and Copilot.

FreeWorks offline
What to know
How it works
Keeps your remaining AI capacity and reset times visible above the Windows taskbar, alerting you when a limit comes back early or partially restores; hovering shows per-model and per-project breakdowns and an equivalent API cost.
Pricing
Free and local-first.
Watch out
Windows only; works with Claude, Codex, Cursor, Gemini, and Copilot.
Details →

Smeltcore

LiveWatched 15d

Benchmark database showing which open-weight AI models run on which consumer GPUs, with install recipes and a GPU advisor.

FreeNo sign-up
What to know
How it works
Provides honest benchmarks for 84 open-weight AI models across 27 consumer GPUs from RTX 3060 to Apple Silicon, each pair rated as runs, tight fit, or won't fit with real speed and VRAM numbers, plus 800+ step-by-step install recipes for llama.cpp, Ollama, ComfyUI, and MLX, and a GPU Advisor.
What's different
No invented numbers.
Pricing
Free under CC BY-SA, no signup.
Details →

An encrypted clipboard manager for developers to stash and quickly paste SSH keys, API tokens and server IPs from the menu bar, instead of keeping them in plaintext notes.

What to know
How it works
A menu-bar clipboard manager that stores SSH keys, API tokens and server IPs, encrypted locally with Argon2id and XChaCha20-Poly1305, with one-click or Enter-to-copy and sync across Mac, iOS, Windows and Web.
What's different
Zero-knowledge encryption applied specifically to developer secrets rather than general clipboard history.
Best for
Developers who need to quickly paste SSH keys, API tokens or server IPs without keeping them in plaintext notes.
Details →

AIVory Smart Inference

LiveWatched 16d

A hosted LLM inference router that's a drop-in replacement for the OpenAI API, automatically sending each request to the cheapest healthy provider. Aimed at developers who want lower inference bills without code changes.

No subscription
What to know
How it works
An LLM inference router and GPU marketplace that acts as a drop-in replacement for the OpenAI API, automatically routing each request in real time to the cheapest healthy provider serving the requested model.
What's different
Continuous automated price discovery across providers, rather than a fixed rate you sign up for once.
Pricing
No subscription.
Best for
Developers who want lower LLM inference bills without changing their SDK or code.
Details →

The Coin Lab

LiveWatched 20d

A no-code platform to launch your own crypto token across 12 blockchains in about a minute, for one payment, built for non-developers rather than blockchain engineers.

Pay once
What to know
How it works
A no-code platform to deploy ERC-20, SPL and BEP-20 crypto tokens live on 12 blockchains in about 60 seconds, without writing Solidity or using a CLI.
What's different
No developer or coding knowledge required to launch a token.
Pricing
One-time payment.
Best for
Non-developers who want to launch their own crypto token quickly.
Details →

RAI

LiveWatched 20d

A self-hostable AI chat platform that routes queries to the cheapest LLM and adds live market data, positioned as an alternative to ChatGPT or Claude you run yourself. For developers and cost-conscious teams.

FreeOpen sourceSelf-hosted
What to know
How it works
A self-hostable, open-source AI chat platform that runs on your own infrastructure, integrates live financial market data into chat, and routes each query to the best/cheapest LLM automatically across OpenAI, Anthropic, SiliconFlow and others.
What's different
Self-hosted with full data ownership and no vendor lock-in, plus intelligent model routing claimed to cut API costs by up to 60% versus using a single provider.
Pricing
Free forever to self-host, with Free / Pro / MAX tiers and per-model quotas.
Best for
Developers and cost-conscious teams who want a ChatGPT/Claude alternative running on their own infrastructure.
Details →

A dashboard for tracking live crypto prices and searching tokens across chains, with no account or subscription required.

Works offline
What to know
How it works
The Basic tier tracks Bitcoin, Solana, and top crypto assets via CoinGecko data on a split-flap terminal dashboard that runs locally; Premium adds DEX Screener search filtered by platform, origin, and blockchain.
Pricing
No subscriptions or setup required for the basic tier.
Details →

agentfdr

LiveWatched 22d

A debugging tool that reads Claude Code's own transcripts and turns them into a flight-recorder view showing loops, errors, and token burn, for developers using coding agents, instead of manually digging through raw logs.

Works offline
What to know
How it works
Reads the transcripts Claude Code writes and turns them into a per-turn timeline with anomaly flags (loops, error streaks, token burns), live watch mode, and cost estimates.
Pricing
Free and open source (MIT license), runs 100% locally with zero instrumentation.
Details →

FastDailyTools

LiveWatched 11d

A growing collection of free, no-signup online calculators covering AI tokens, fees, solar savings, and health tracking.

FreeNo sign-up
What to know
How it works
A collection of free calculators covering AI token estimation, PayPal/Etsy fees, solar savings, water intake, pregnancy due dates, and period tracking.
Pricing
Free, no sign-up or downloads required.
Details →

The ones we've read

Codex Analytics Ledger

LiveWatched 14d

A local-only dashboard for tracking Codex usage and cost, without sending usage data to the cloud.

What to know
How it works
A local-only dashboard for tracking Codex usage and cost.
What's different
Doesn't send usage data to the cloud.
Details →

Token Poker

LiveWatched 10d

Sprint planning poker tool that adds AI-token-usage estimates alongside story points for teams building AI features.

What to know
How it works
Teams vote on story points and forecast expected AI token consumption in the same planning session
What's different
Targets teams building LLM-based features where token cost affects planning
Best for
Software teams building AI-powered products who need to plan both effort and AI cost
Details →

Tokens Forge

LiveWatched 13d

An OpenAI-compatible AI gateway with transparent per-model pricing and usage tracking, paired with a separate AI tool for stock market research.

What to know
How it works
routes requests through a low-cost, OpenAI-compatible gateway with usage tracking, plus an AI stock researcher feature
Best for
developers wanting cheaper, trackable access to AI models
Details →

Clawgate

LiveWatched 18d

A cost-control tool that sets hard-stop budgets and tracks per-project spend for teams using Claude Code and Codex.

What to know
How it works
Applies budget limits and model controls to AI coding tool usage, with built-in prompt compression to cut token spend.
Best for
Teams that want visibility and caps on their Claude Code/Codex spend.
Details →

Search API for AI agents that returns relevant excerpts instead of full pages to cut token usage.

Watch out · Live on every /search call today.

What to know
How it works
A search API that returns the excerpts from each result best matching the query instead of full pages.
What's different
Uses 10x fewer tokens than processing full pages and scores 94.7% on SimpleQA, the highest reported among providers.
Watch out
Live on every /search call today.
Details →

ZooData

LiveWatched 15d

Gives AI agents clean structured data (and pre-built e-commerce intel) instead of raw scraped HTML, for developers building agents.

What to know
How it works
Turns any URL into agent-ready structured JSON, and provides pre-analyzed e-commerce intelligence (competitor, market, traffic, consumer insights) live for Amazon and TikTok. API, CLI, and MCP server included.
What's different
Uses about 75% fewer LLM tokens than working with raw HTML or markdown, and charges only for the fields you use.
Pricing
Starts with 1,000 free credits, no card required.
Best for
Developers building AI agents that need clean structured web or e-commerce data.
Details →

agentrecap

LiveWatched 6d

Simple tool to analyze your local coding-agent sessions

What to know
How it works
Yerel coding-agent oturumlarını okuyup çalışma süresi, yapılan araç çağrıları, değiştirilen kod satırı ve mesaj sayısı gibi veriler çıkarır.
Details →

Public leaderboard of developers' AI tool usage, tracking token spend and surfacing effective AI setups.

What to know
How it works
You connect Claude and Codex to a public board that tracks your token usage and surfaces the skills, tools, and projects you use.
Details →

Just arrived

Tempest

LiveFirst seen today
Details →

Track AI

LiveWatched 1d
Details →

Tokimeter

LiveWatched 1d
No sign-upOpen sourceNo tracking
Details →

97AI.PRO

LiveWatched 2d
Free
Details →
Details →

TokenDam

LiveWatched 2d
FreeData stays with you
Details →

PackMD

LiveWatched 2d
Details →

SendPI

LiveWatched 2d
Details →

Watched the longest, still answering

Auriko

LiveWatched 24d

Trading desk for LLM calls

Watch out · The '30% average cost reduction' figure comes from the product's own published report, not an independent benchmark.

What to know
How it works
Sits between your app and multiple LLM providers, automatically routing each request to the cheapest/fastest provider based on token price, cache behavior, latency and reliability.
What's different
Built by ex-quant traders, applies trading-desk arbitrage logic to LLM API pricing rather than simple static routing rules.
Best for
Teams making heavy, repeated LLM API calls who want automatic cost optimization without manually comparing providers.
Watch out
The '30% average cost reduction' figure comes from the product's own published report, not an independent benchmark.
Details →

Constellation Gate AI

LiveWatched 24d

A desktop proxy that sits in front of your AI agent's calls to block prompt injection, scan for leaked secrets and cut token usage - works alongside your existing Claude or ChatGPT subscription, not a replacement for it.

Watch out · The '#1 in 16 benchmarks' and '20-40% token savings' figures are the vendor's own claims.

What to know
How it works
A desktop app that proxies your AI agent's calls, adding prompt-injection defense, secret scanning, an audit trail, and prompt compression/caching. Works with your existing Claude or ChatGPT subscription or 100+ pay-as-you-go models.
What's different
Claims the #1 rank across 16 public prompt-injection benchmark datasets — a specific, citable security claim.
Best for
Developers running AI agents in production who want prompt-injection defense and token-cost savings without changing models.
Watch out
The '#1 in 16 benchmarks' and '20-40% token savings' figures are the vendor's own claims.
Details →

Dali by Lulu

LiveWatched 24d

An MCP tool that scores your AI image or video prompt before you generate, aiming to catch weak prompts before you spend a paid generation credit.

What to know
How it works
An MCP tool that scores an AI image or video generation prompt before you run it, aiming to flag weak prompts before a paid generation credit is spent.
What's different
Scores prompts pre-generation rather than after, and runs a live dashboard tracking community-wide savings from credits not wasted.
Best for
People generating AI images or video who pay per generation and want to avoid burning credits on bad prompts.
Details →

A lightweight proxy that sits in front of coding-agent tool calls to block repetitive failing calls, mutate strategy when the agent is stuck, and compress bloated tool outputs, working across Claude, OpenAI, DeepSeek, Ollama, and OpenCode.

What to know
How it works
Runs as a background proxy (installed via pip) that detects and blocks repetitive failing tool calls, mutates strategy on loops, and compresses tool outputs, with a real-time dashboard and savings report.
Best for
developers running coding agents who want to cut token waste and loops.
Details →

ManjuAi

LiveWatched 24d

Service that runs deep-research and coding agents using various AI models for a low flat fee.

Free
What to know
How it works
Gives access to deep-research and coding agents with control over report length and full context for coding tasks.
Pricing
Access to most premium models for $5.30.
Details →

SAGE - CLI

LiveWatched 24d

A command-line tool for developers that compresses terminal output to cut token usage and cost when working with AI coding assistants.

What to know
How it works
A command-line tool that compresses terminal output by 85-98% to reduce AI coding assistant token usage, keeping raw logs local and tracking real savings.
Best for
Developers using AI coding assistants who want to cut token costs from verbose terminal output.
Details →

AfterLimit

LiveWatched 24d

Queues a follow-up prompt and auto-sends it to ChatGPT, Claude or Gemini once your usage limit resets, so you don't lose your train of thought waiting around.

What to know
How it works
You type or paste a follow-up prompt, pick a send time (such as when your usage limit resets), and AfterLimit automatically sends it to ChatGPT, Claude, or Gemini at that time.
Best for
People who hit free usage limits on AI chat tools mid-task and don't want to remember to come back.
Details →

ReplyWatcher

LiveWatched 24d

Watches one person's public X posts for signs Codex quota has reset, then pushes ranked alerts (confirmed/possible/noise) to Discord or Lark.

Watch out · Relies entirely on monitoring one individual's public posts as the reset signal — if that account goes quiet or changes behavior, the alerts stop being reliable.

What to know
How it works
Watches one specific person's public X posts, uses AI semantic filtering to separate real Codex quota-reset signals from noise, then pushes ranked alerts (confirmed/possible/noise) to Discord or Lark.
Best for
Codex users who repeatedly hit usage limits and want to know the moment a reset window opens.
Watch out
Relies entirely on monitoring one individual's public posts as the reset signal — if that account goes quiet or changes behavior, the alerts stop being reliable.
Details →

PokeTokenBar

LiveWatched 11d

A menu-bar app that turns your Claude Code/Codex/Gemini CLI token usage into a Pokémon-collecting game while tracking usage limits.

FreeWorks offlineOpen source
What to know
How it works
Reads local usage logs and hatches/evolves a Pokémon as you spend tokens
What's different
Runs entirely on-device, nothing sent anywhere
Pricing
free
Best for
Developers who want a fun way to watch their AI token usage
Details →

TokenPulse

LiveWatched 12d

A free Chrome extension that shows a live token/context bar above Claude, ChatGPT, and Gemini input boxes so you don't get cut off mid-session.

FreeNo sign-up
What to know
How it works
Displays context window usage, rate limits, reset countdowns, and cost live above the input box
Pricing
Free, no account, no API key
Best for
Developers who use multiple AI chat tools daily and hit context limits
Details →

PXINK

LiveWatched 12d

A free browser tool that shrinks big text/code pastes into dense PNGs so AI vision billing costs less than text billing.

FreeData stays with you
What to know
How it works
Wraps text blocks into dense 1928px PNGs the model reads via OCR
What's different
Exploits vision-token vs text-token pricing gap for ~60% savings
Pricing
Free, client-side, no account
Best for
Developers pasting large text/code logs into Claude or Gemini
Details →

Palet

LiveWatched 23d

A free color palette generator and design-token tool for designers and developers, extracting colors from images and exporting to CSS, Tailwind, or design tokens.

FreeCan export my data
What to know
How it works
Generates palettes, extracts colors from images, checks contrast, and exports to CSS, Tailwind, and design tokens.
Pricing
Free.
Details →

Claude session monitor

LiveWatched 19d

A free local VS Code/Cursor extension that turns your Claude Code session files into a browsable dashboard with full transcripts, tool calls, sub-agents and per-response token and cost tracking.

FreeWorks offline
What to know
How it works
A VS Code/Cursor/VSCodium extension that turns local Claude Code session files into a browsable workspace with full transcripts, tool calls, sub-agents, per-response token and cost tracking, an analytics dashboard, and a CLAUDE.md memory view.
What's different
100% local-first — nothing leaves your machine.
Pricing
Free.
Best for
Developers using Claude Code who want to review sessions, costs and tool calls without a CLI-only cost counter.
Details →

Packforai

LiveWatched 19d

Converts PDFs, Word, Excel, PowerPoint, CSV and JSON files into clean Markdown formatted for feeding into ChatGPT or Claude, free to start with no account required.

FreeNo sign-up
What to know
How it works
Converts PDFs, Word, Excel, PowerPoint, CSV, and JSON files into clean Markdown formatted to work well as input for AI models like ChatGPT and Claude.
What's different
Optimizes the output specifically to use fewer tokens and produce better AI answers, rather than just generic format conversion.
Pricing
Free to start, no account needed.
Best for
People who need to feed messy documents into ChatGPT or Claude and want cleaner, more token-efficient input.
Details →

Converly Colors

LiveWatched 19d

Generates AI color palettes as design.md token files for tools like Claude, Cursor, and v0 — a quick color-scale generator for developers and designers doing AI-assisted work.

FreeNo sign-up
What to know
How it works
Generates AI color palettes as design.md token files — the format used by tools like Claude, Cursor and v0 — including 12-step scales, light/dark modes and CSS variables.
What's different
Outputs directly in the design.md token format that AI coding tools expect, rather than a generic palette export.
Pricing
Free, no signup.
Best for
Developers and designers building AI-assisted workflows who need design tokens fast.
Details →

Honey is a free, open-source (MIT) skill that reduces token usage for AI coding agents like Claude Code by enforcing concise code and prose conventions and cheaper bulk reads.

Free
What to know
How it works
enforces YAGNI/stdlib-first code, answer-first prose, compact JSON/ESON handoffs between agents, and renders bulk reads as images to cut token cost
Pricing
free, MIT license
Best for
developers running AI coding agents (Claude Code, GPT, etc.) who want lower token usage
Details →

SuperCompress

LiveWatched 23d

Open-source API that compresses prompts before OpenAI, Claude, Gemini or RAG calls, cutting token costs by about 65% while keeping the important content. For developers wiring it into their own LLM calls.

Open source
What to know
How it works
An open-source API that compresses prompts before they're sent to OpenAI, Claude, Gemini, RAG, chatbot, or agent calls, stripping tokens while preserving the evidence needed to answer correctly.
What's different
Claims to remove more than 65% of tokens on average while preserving over 98% of answer-critical content.
Best for
Developers wiring their own LLM calls who want to cut token costs without integrating a full new pipeline.
Details →

StackSpine

LiveWatched 12d

An open-source control plane that enforces policy, budget limits, and compliance auditing across all the AI providers a company uses in production.

Open source
What to know
How it works
Sits between apps and AI providers to enforce guardrails and log usage for audit.
Pricing
Open-source, Apache 2.0.
Best for
Engineering teams running production AI across multiple providers who need governance.
Details →

burnban

LiveWatched 18d

A local, open-source binary that meters and caps AI agent spend on your own machine for developers who want budget control without sending usage data to a dashboard.

Open source
What to know
How it works
A local binary sits in the request path, meters every token spent by AI agents, and enforces spend caps before they're exceeded; also compares your flat-rate plan against real metered rates.
What's different
Open source (MIT); sync and team budgets are planned but not yet available.
Best for
Developers running AI agents who want local spend control without a hosted dashboard.
Details →

VOLY

LiveWatched 16d

Open-source control plane that routes coding tasks across AI agents like Claude Code, Cursor, and Codex with cost tracking.

Open source
What to know
How it works
An open-source control plane that routes coding tasks across AI coding agents (Claude Code, Cursor, Codex, Wrangler, etc.), runs multi-agent workflows, caches to reduce repeated token spend, falls back on billing/availability issues, and shows cost per task with files touched.
What's different
Gives one measurable layer above multiple separate agent tools instead of running each independently.
Details →

HardwareMon

LiveWatched 18d

Cross-platform hardware monitoring with real-time stats, history and benchmarking for Windows, Linux, macOS and Android, with all data processed locally and no telemetry upload.

Works offlineOpen source
What to know
How it works
Monitors CPU, GPU, memory, disk and network in real time, with historical analytics, process management, local benchmarking, and gaming session detection, across Windows, Linux, macOS (Apple Silicon), and Android.
What's different
All monitoring data is processed on-device with local SQLite storage; no telemetry or benchmark results are uploaded.
Best for
Users who want cross-platform hardware monitoring without sending data to the cloud.
Details →

CodexBar Lite

LiveWatched 7d

Lets developers monitor their OpenAI Codex usage from the macOS menu bar instead of checking a web dashboard, using only their existing CLI session.

Data stays with youOpen source
Details →

ISONGraph

LiveWatched 5d
Open source
Details →

TokenTelemetry

LiveWatched 3d
Works offlineOpen source
Details →

Prompt Compass

LiveWatched 12d

Routes each AI prompt to the cheapest capable model and blocks PII/jailbreak attempts locally, cutting AI API costs for developers.

Works offline
What to know
How it works
On-device ~5ms routing decision per prompt
What's different
Claims up to 96% cost cut by avoiding always-expensive-model routing
Pricing
Free tier, npm SDK and VS Code/Cursor extension
Details →

AgentQuartz

LiveWatched 4d

Claude & Cursor usage in your macOS menu bar

FreeWorks offline
Details →

Hydra

LiveWatched 4d
FreeWorks offline
Details →

LMspend

Watched 17d

A free, local-first CLI that reads logs already on your machine to estimate spend on AI coding tools like Claude Code and Cursor.

FreeWorks offline
What to know
How it works
A local-first CLI that reads logs already on the machine (no proxy, no traffic in the middle) to estimate spend on AI coding tools like Claude Code, Cursor, and Codex from published model pricing, printing monthly reports and shareable cards.
Pricing
Free; install via 'npx lmspend'. Optional dashboard adds history, budgets, alerts, and team roll-ups.
Best for
Individuals and teams wanting to track AI coding tool costs before the invoice arrives.
Details →

Tools for controlling AI token cost — the short answers

Every number here comes from our own daily check — not from a vendor list.

Tools for controlling AI token cost — how many are there?
Tablif is tracking 217 of them. 214 answered our check today, and 3 we couldn't reach — we don't claim those are dead.
Which ones are still maintained?
Tablif knocks on every door once a day and records the answer. 214 of these 217 responded on the latest run, so that number is what "still here" means on this page — not a review score.
Are any of them free?
44 of the live ones say so in their own words, and 14 let you start without making an account. Tablif records the claim the product makes; we don't verify pricing.
Any open-source options?
27 of the live ones on Tablif mention being open source.
What's the newest one?
Tempest — Tablif first saw it today.
Did a person actually look at these?
215 of them Tablif opened and read, then wrote a one-line summary in our own words instead of reusing the founder's tagline. The rest carry keyword labels we haven't confirmed by reading yet — and we say so rather than hiding it.