Topic

Tools for controlling AI token cost · Free

45

match your filters

44

answered today

4

arrived this week

0

stopped answering

Watching this group for 26 days, checked once a day · how we check

Know your usage pace and resets. Let Codex help you plan.

FreeOpen source
What to know
How it works
İşlemi ayırır (detach), konuşmayı serbest bırakır, sınırlı çıktı kaydeder ve tamamlanınca Codex'e inceleme için döner.
Details →

CodexUsageBar

LiveWatched 13d

A free, open-source macOS menu bar app that tracks your ChatGPT Codex usage limits in real time for developers.

FreeOpen source
What to know
How it works
Lives in the macOS menu bar and shows current Codex usage at a glance.
Pricing
Free and open-source.
Details →

AI Usage Ball

LiveWatched 6d

Live desktop orbs for Claude, Codex & Antigravity

FreeWorks offlineNo subscription
What to know
How it works
Kalan kullanım hakkı ve sıfırlanma zamanını yerelde izler, yalnız sağlayıcıların kendi servisleriyle konuşur; kaynak denetime açık.
Pricing
İlk 100 kullanıcıya ücretsiz ömür boyu kişisel lisans fırsatı.
Details →

Codex Quota Panel

LiveWatched 10d

A free open-source Windows dashboard that shows your Codex usage quota, reset timers, and token counts at a glance.

FreeOpen sourceNo subscription
What to know
How it works
Displays weekly quota, reset time, reset-credit cards, and lifetime token usage for Codex
Pricing
Free and open-source
Details →

Shibaita.ai

LiveWatched 7d

Lets Claude Code users compare usage on a public leaderboard instead of only seeing their own local stats.

Watch out · unofficial, not affiliated with Anthropic

Free
What to know
How it works
A CLI reads local logs and submits only aggregate token counts, never prompts or code.
Pricing
free
Watch out
unofficial, not affiliated with Anthropic
Details →

KlaatAI

LiveWatched 9d

Routes prompts to the cheapest capable model tier and offers a terminal coding agent claiming lower cost than Claude Code on a benchmark.

Free
What to know
How it works
Routes every prompt to the cheapest model tier that can handle it, escalating only when needed; its terminal agent Klaat Code claims to solve the same 33-task benchmark as Claude Code at 5.4x lower cost ($0.027 vs $0.146/task).
Pricing
Free beta: 75 requests/day, no card; tool calls are free, only messages count against quota.
Details →

LoRA Speedrun

LiveWatched 12d

Leaderboard site where people compete to fine-tune the same model fastest, with every record re-verified by a referee.

Free
What to know
How it works
same model, same score, same GPU — fastest training run wins, re-run 3x by a referee
What's different
no self-reported times, all records independently verified
Pricing
free to attempt
Details →

nable

LiveWatched 6d

One usage meter for every AI provider you pay for

FreeNo sign-upWorks offline
What to know
How it works
Kullanım yerel oturum kayıtlarından okunur, hesap/anahtar gerekmez; sabit planlarda kullanım, ölçülü planlarda tahmini dolar gösterir.
Pricing
Açık kaynak ve sonsuza dek ücretsiz.
Details →

ProxAI

LiveWatched 9d

Reverse proxy that enforces hard spending caps on LLM API calls across providers by swapping the base URL.

Watch out · AGPL-3.0, no telemetry.

FreeNo trackingSelf-hosted
What to know
How it works
A transparent reverse proxy enforcing hard hourly/daily spending caps per LLM API client via a one-line base_url change, returning 429 once the cap is hit rather than calling the upstream.
What's different
Works with OpenAI, Anthropic, Mistral, Groq, DeepSeek, and 5 more providers.
Pricing
Hosted beta at proxai.eu is free, no credit card; self-hostable as a ~10 MB Docker image.
Watch out
AGPL-3.0, no telemetry.
Details →

Toolport

LiveWatched 26d

A local gateway that lets every AI coding client share one set of MCP server connections instead of configuring each tool separately, cutting how many tools your agent has to load.

Watch out · The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.

FreeNo sign-upWorks offline
What to know
How it works
A local-first, open-source gateway. You set up and authenticate each MCP server once, and every connected AI client — Claude, Cursor, VS Code, Codex — shares that same connection instead of you configuring it per tool.
What's different
Claims up to 91% less tool-token overhead by giving an agent a handful of meta-tools instead of loading hundreds of individual tool definitions; API keys stay in the OS keychain.
Pricing
Free, open source, no sign-up.
Best for
Developers running multiple AI coding clients against the same set of MCP servers who don't want to re-authenticate each one separately.
Watch out
The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.
Details →

Oriveo

LiveWatched 14d

A synced multi-model AI chat app across iOS, Android, and web using your own API keys.

FreeWorks offline
What to know
How it works
Lets you chat with ChatGPT, Claude, Gemini, and 10+ other AI models in one app across iOS, Android, and web, synced in real time, using your own API keys.
What's different
Includes real-time cost tracking, local-first storage, saved Notes, and cross-checking replies with a second model.
Pricing
Try free, no key needed; no markup or extra subscription with your own keys.
Details →

Tokenez

LiveWatched 6d

Counts your AI usage locally across ChatGPT, Claude, Gemini, Claude Code and Cursor in one dashboard, showing your real total and how close you are to each limit.

FreeNo sign-up
What to know
What's different
Counts locally rather than relying on each provider's own scattered usage view.
Pricing
Free, no signup to try.
Details →

Token Guard

LiveWatched 8d

Tracks and converts LLM API spend across OpenAI, Anthropic and Gemini locally, instead of trusting a cloud dashboard with your usage data.

FreeWorks offlineOpen source
What to know
Pricing
Free and open source.
Best for
Developers juggling multiple LLM provider SDKs and costs.
Details →

Levitan

LiveWatched 11d

A free, fully offline voice-to-text tool that types what you say into any app on Windows or Mac, running locally with no cloud upload.

FreeWorks offline
What to know
How it works
hold a key, speak, text appears at cursor in any app; runs locally on CPU
Pricing
free forever
Best for
privacy-conscious users who don't want cloud-based voice typing
Details →

MutualGPU

LiveWatched 10d

Lets people with spare GPU capacity share it so others without enough local compute can run demanding AI workloads.

Free
What to know
How it works
community compute-sharing network; first use case runs 3D Gaussian Splat generation via WebGPU
Pricing
free
Details →

OSGARD NEW WORLD

LiveWatched 9d

A tool that generates a themed sci-fi/fantasy AI universe with tradable in-game artifacts and a leaderboard.

FreeNo sign-up
What to know
How it works
Generates a themed AI universe (sci-fi, fantasy, cyberpunk or steampunk) with tradable artifacts in about 7 seconds, with a live marketplace, TC token staking and an AI assistant called JARVIS.
Pricing
3 free generations, no signup needed.
Best for
People wanting to generate and trade AI-created world assets.
Details →

GPUCPUCompare

LiveWatched 9d

A free tool that compares GPU/CPU hardware specs and checks compatibility against a large game library.

FreeNo sign-up
What to know
How it works
Compares 733 GPUs and 1,002 CPUs and checks compatibility against 112,113 PC games using public TechPowerUp, Geekbench and Steam data.
What's different
Uses public benchmark data instead of unexplained percentages or simulated benchmarks.
Pricing
Free, no signup.
Best for
PC gamers deciding what hardware to buy.
Details →

PandaCoder

LiveWatched 9d

Free local AI coding assistant positioned as an alternative to token-limited cloud coding tools.

Free
What to know
How it works
Runs locally as a coding assistant with no token limits.
Pricing
Free.
Best for
Developers wanting a private, local alternative to token-limited cloud coding tools.
Details →

Find AI Credits

LiveWatched 9d

Free directory aggregating AI, cloud, API and GPU credit offers and startup perks.

Free
What to know
How it works
Free directory aggregating free AI, cloud, API, and GPU credit offers and startup perks from AI companies in one place.
Pricing
Free.
Details →

Visually shows what a million AI tokens actually amounts to in books, code, or dollars, and compares model pricing.

Free
What to know
How it works
Converts token counts to relatable units and compares input/output pricing across models
Pricing
Free
Details →

Generates 3D Gaussian Splats in your browser, for free, using either your device's GPU or a community-shared compute network.

FreeNothing to install
What to know
How it works
runs via WebGPU locally, or offloads to MutualGPU's crowd-sourced GPU network
What's different
no install, and free generation even without a powerful local GPU
Pricing
free
Details →

AI Karma Tracker

LiveWatched 14d

A browser toolbar popup that shows live usage limits and reset times across the AI tools you use, with peak-hour alerts.

Watch out · No API keys — nothing leaves your browser.

Free
What to know
How it works
A toolbar popup showing live usage, remaining limits, and reset times for the AI tools you're signed into, with peak-hour nudges.
Pricing
Free for one tool, $2/mo for your whole stack.
Watch out
No API keys — nothing leaves your browser.
Details →

Ceiling

LiveWatched 14d

A Windows taskbar utility that tracks remaining usage limits and reset times across Claude, Codex, Cursor, and other AI tools.

Watch out · Windows only; works with Claude, Codex, Cursor, Gemini, and Copilot.

FreeWorks offline
What to know
How it works
Keeps your remaining AI capacity and reset times visible above the Windows taskbar, alerting you when a limit comes back early or partially restores; hovering shows per-model and per-project breakdowns and an equivalent API cost.
Pricing
Free and local-first.
Watch out
Windows only; works with Claude, Codex, Cursor, Gemini, and Copilot.
Details →

Smeltcore

LiveWatched 15d

Benchmark database showing which open-weight AI models run on which consumer GPUs, with install recipes and a GPU advisor.

FreeNo sign-up
What to know
How it works
Provides honest benchmarks for 84 open-weight AI models across 27 consumer GPUs from RTX 3060 to Apple Silicon, each pair rated as runs, tight fit, or won't fit with real speed and VRAM numbers, plus 800+ step-by-step install recipes for llama.cpp, Ollama, ComfyUI, and MLX, and a GPU Advisor.
What's different
No invented numbers.
Pricing
Free under CC BY-SA, no signup.
Details →

RAI

LiveWatched 20d

A self-hostable AI chat platform that routes queries to the cheapest LLM and adds live market data, positioned as an alternative to ChatGPT or Claude you run yourself. For developers and cost-conscious teams.

FreeOpen sourceSelf-hosted
What to know
How it works
A self-hostable, open-source AI chat platform that runs on your own infrastructure, integrates live financial market data into chat, and routes each query to the best/cheapest LLM automatically across OpenAI, Anthropic, SiliconFlow and others.
What's different
Self-hosted with full data ownership and no vendor lock-in, plus intelligent model routing claimed to cut API costs by up to 60% versus using a single provider.
Pricing
Free forever to self-host, with Free / Pro / MAX tiers and per-model quotas.
Best for
Developers and cost-conscious teams who want a ChatGPT/Claude alternative running on their own infrastructure.
Details →

FastDailyTools

LiveWatched 11d

A growing collection of free, no-signup online calculators covering AI tokens, fees, solar savings, and health tracking.

FreeNo sign-up
What to know
How it works
A collection of free calculators covering AI token estimation, PayPal/Etsy fees, solar savings, water intake, pregnancy due dates, and period tracking.
Pricing
Free, no sign-up or downloads required.
Details →

PokeTokenBar

LiveWatched 11d

A menu-bar app that turns your Claude Code/Codex/Gemini CLI token usage into a Pokémon-collecting game while tracking usage limits.

FreeWorks offlineOpen source
What to know
How it works
Reads local usage logs and hatches/evolves a Pokémon as you spend tokens
What's different
Runs entirely on-device, nothing sent anywhere
Pricing
free
Best for
Developers who want a fun way to watch their AI token usage
Details →

TokenPulse

LiveWatched 12d

A free Chrome extension that shows a live token/context bar above Claude, ChatGPT, and Gemini input boxes so you don't get cut off mid-session.

FreeNo sign-up
What to know
How it works
Displays context window usage, rate limits, reset countdowns, and cost live above the input box
Pricing
Free, no account, no API key
Best for
Developers who use multiple AI chat tools daily and hit context limits
Details →

PXINK

LiveWatched 12d

A free browser tool that shrinks big text/code pastes into dense PNGs so AI vision billing costs less than text billing.

FreeData stays with you
What to know
How it works
Wraps text blocks into dense 1928px PNGs the model reads via OCR
What's different
Exploits vision-token vs text-token pricing gap for ~60% savings
Pricing
Free, client-side, no account
Best for
Developers pasting large text/code logs into Claude or Gemini
Details →

Palet

LiveWatched 23d

A free color palette generator and design-token tool for designers and developers, extracting colors from images and exporting to CSS, Tailwind, or design tokens.

FreeCan export my data
What to know
How it works
Generates palettes, extracts colors from images, checks contrast, and exports to CSS, Tailwind, and design tokens.
Pricing
Free.
Details →

ManjuAi

LiveWatched 24d

Service that runs deep-research and coding agents using various AI models for a low flat fee.

Free
What to know
How it works
Gives access to deep-research and coding agents with control over report length and full context for coding tasks.
Pricing
Access to most premium models for $5.30.
Details →

Claude session monitor

LiveWatched 19d

A free local VS Code/Cursor extension that turns your Claude Code session files into a browsable dashboard with full transcripts, tool calls, sub-agents and per-response token and cost tracking.

FreeWorks offline
What to know
How it works
A VS Code/Cursor/VSCodium extension that turns local Claude Code session files into a browsable workspace with full transcripts, tool calls, sub-agents, per-response token and cost tracking, an analytics dashboard, and a CLAUDE.md memory view.
What's different
100% local-first — nothing leaves your machine.
Pricing
Free.
Best for
Developers using Claude Code who want to review sessions, costs and tool calls without a CLI-only cost counter.
Details →

Packforai

LiveWatched 19d

Converts PDFs, Word, Excel, PowerPoint, CSV and JSON files into clean Markdown formatted for feeding into ChatGPT or Claude, free to start with no account required.

FreeNo sign-up
What to know
How it works
Converts PDFs, Word, Excel, PowerPoint, CSV, and JSON files into clean Markdown formatted to work well as input for AI models like ChatGPT and Claude.
What's different
Optimizes the output specifically to use fewer tokens and produce better AI answers, rather than just generic format conversion.
Pricing
Free to start, no account needed.
Best for
People who need to feed messy documents into ChatGPT or Claude and want cleaner, more token-efficient input.
Details →

Converly Colors

LiveWatched 19d

Generates AI color palettes as design.md token files for tools like Claude, Cursor, and v0 — a quick color-scale generator for developers and designers doing AI-assisted work.

FreeNo sign-up
What to know
How it works
Generates AI color palettes as design.md token files — the format used by tools like Claude, Cursor and v0 — including 12-step scales, light/dark modes and CSS variables.
What's different
Outputs directly in the design.md token format that AI coding tools expect, rather than a generic palette export.
Pricing
Free, no signup.
Best for
Developers and designers building AI-assisted workflows who need design tokens fast.
Details →

Honey is a free, open-source (MIT) skill that reduces token usage for AI coding agents like Claude Code by enforcing concise code and prose conventions and cheaper bulk reads.

Free
What to know
How it works
enforces YAGNI/stdlib-first code, answer-first prose, compact JSON/ESON handoffs between agents, and renders bulk reads as images to cut token cost
Pricing
free, MIT license
Best for
developers running AI coding agents (Claude Code, GPT, etc.) who want lower token usage
Details →

LLMLite

LiveWatched 23d

Routes each API prompt to the cheapest model that can handle it, cutting AI API costs — a developer tool for teams with real LLM spend, with a free tier.

Free
What to know
How it works
Routes each API prompt through a single endpoint to the cheapest model capable of handling it, cutting AI API costs by 30-80%.
Pricing
Free tier included.
Best for
Developer teams with real LLM API spend looking to cut costs without changing their integration.
Details →

MCP Peek

LiveWatched 7d

Graphical interface for inspecting and managing MCP servers instead of the command line.

Free
Details →

AgentQuartz

LiveWatched 4d
FreeWorks offline
Details →

NavloAI

LiveWatched 5d
FreeNo subscription
Details →

Songvora

LiveWatched 5d
Free
Details →

SolanaForge

LiveWatched 4d
FreeNo sign-upNo subscription
Details →

Hydra

LiveWatched 4d
FreeWorks offline
Details →

97AI.PRO

LiveWatched 2d
Free
Details →

TokenDam

LiveWatched 2d
FreeData stays with you
Details →

LMspend

Watched 17d

A free, local-first CLI that reads logs already on your machine to estimate spend on AI coding tools like Claude Code and Cursor.

FreeWorks offline
What to know
How it works
A local-first CLI that reads logs already on the machine (no proxy, no traffic in the middle) to estimate spend on AI coding tools like Claude Code, Cursor, and Codex from published model pricing, printing monthly reports and shareable cards.
Pricing
Free; install via 'npx lmspend'. Optional dashboard adds history, budgets, alerts, and team roll-ups.
Best for
Individuals and teams wanting to track AI coding tool costs before the invoice arrives.
Details →