Topic

Tools for controlling AI token cost · Open source

27

match your filters

27

answered today

3

arrived this week

0

stopped answering

Watching this group for 26 days, checked once a day · how we check

Know your usage pace and resets. Let Codex help you plan.

FreeOpen source
What to know
How it works
İşlemi ayırır (detach), konuşmayı serbest bırakır, sınırlı çıktı kaydeder ve tamamlanınca Codex'e inceleme için döner.
Details →

CodexUsageBar

LiveWatched 13d

A free, open-source macOS menu bar app that tracks your ChatGPT Codex usage limits in real time for developers.

FreeOpen source
What to know
How it works
Lives in the macOS menu bar and shows current Codex usage at a glance.
Pricing
Free and open-source.
Details →

LLMSlim

LiveWatched 17d

An open-source Python library that compresses LLM prompts and context to cut token costs while preserving instructions.

Open source
What to know
How it works
Compresses prompts, RAG document contexts, and multi-turn chat logs in one line of code.
Pricing
Open-source Python package.
Best for
Developers who want to cut LLM token costs 40-70% without losing instruction fidelity.
Details →

Spanlens

LiveWatched 6d

Open-source, self-hostable observability platform for LLM apps: request logging, cost tracking, agent tracing.

Open sourceSelf-hosted
What to know
How it works
Works with OpenAI, Anthropic, and Gemini apps.
Pricing
Open source, MIT license.
Details →

Pylva

LiveWatched 12d

Tracks the real infrastructure cost of each customer for AI agent builders, so they can set margin caps per customer.

Open source
What to know
How it works
SDK-only telemetry tracks LLM and non-LLM usage per customer
What's different
open source, reactive rules like per-customer cost caps
Pricing
open source
Best for
builders of AI agent products billing by usage
Details →

Codex Quota Panel

LiveWatched 10d

A free open-source Windows dashboard that shows your Codex usage quota, reset timers, and token counts at a glance.

FreeOpen sourceNo subscription
What to know
How it works
Displays weekly quota, reset time, reset-credit cards, and lifetime token usage for Codex
Pricing
Free and open-source
Details →

CC Theme

LiveWatched 11d

An open-source Mac app providing ready-made, reproducible themes for AI desktop apps without repeated prompt generation.

Open source
What to know
How it works
Applies validated, ready-made theme packages to supported AI desktop apps, avoiding repeated image-and-prompt generation.
Details →

nable

LiveWatched 6d

One usage meter for every AI provider you pay for

FreeNo sign-upWorks offline
What to know
How it works
Kullanım yerel oturum kayıtlarından okunur, hesap/anahtar gerekmez; sabit planlarda kullanım, ölçülü planlarda tahmini dolar gösterir.
Pricing
Açık kaynak ve sonsuza dek ücretsiz.
Details →

Toolport

LiveWatched 26d

A local gateway that lets every AI coding client share one set of MCP server connections instead of configuring each tool separately, cutting how many tools your agent has to load.

Watch out · The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.

FreeNo sign-upWorks offline
What to know
How it works
A local-first, open-source gateway. You set up and authenticate each MCP server once, and every connected AI client — Claude, Cursor, VS Code, Codex — shares that same connection instead of you configuring it per tool.
What's different
Claims up to 91% less tool-token overhead by giving an agent a handful of meta-tools instead of loading hundreds of individual tool definitions; API keys stay in the OS keychain.
Pricing
Free, open source, no sign-up.
Best for
Developers running multiple AI coding clients against the same set of MCP servers who don't want to re-authenticate each one separately.
Watch out
The '91% overhead cut' figure is the product's own claim; we haven't independently benchmarked it.
Details →

Token Guard

LiveWatched 8d

Tracks and converts LLM API spend across OpenAI, Anthropic and Gemini locally, instead of trusting a cloud dashboard with your usage data.

FreeWorks offlineOpen source
What to know
Pricing
Free and open source.
Best for
Developers juggling multiple LLM provider SDKs and costs.
Details →

Generates a 3D model from a single image entirely on-device using WebGPU, with no server uploads or accounts needed — an open-source tool for developers experimenting with local image-to-3D generation.

Works offlineData stays with youOpen source
What to know
How it works
Generates a 3D model from a single image entirely on-device using WebGPU compute in the browser, with no server uploads or accounts.
What's different
Open source and runs locally via WebGPU rather than requiring cloud GPU processing, also deployable on Vercel for instant testing.
Best for
Developers experimenting with local, private image-to-3D generation.
Details →

Meterbility

LiveWatched 10d

A debugger for AI agent runs that shows exactly what context each model call saw, for developers troubleshooting agents built with Claude Code, Codex, or Cursor.

Works offlineOpen source
What to know
How it works
Captures every step of an agent run, compares two runs, and can fork from any step to test changes
What's different
Local-first and open source, so traces never leave your machine
Best for
Developers debugging or comparing AI agent runs
Details →

Gives voice-AI developers one account to meter spend across their whole vendor stack (LLM, speech-to-text, text-to-speech, telephony) with spend caps and kill-switches.

Open source
What to know
How it works
Server-side spend caps and per-agent kill-switches across multiple vendor APIs under one key
What's different
Open-source floe-guard component
Best for
Developers building voice AI products across several paid vendor APIs
Details →

ChorusGraph

LiveWatched 23d

An open-source Python agent runtime, installed with one pip command, with a local cache that skips repeat LLM calls — pitched as a lighter alternative to LangGraph for developers.

Open source
What to know
How it works
An open-source agent graph runtime with a BSP scheduler and a local semantic cache that skips redundant LLM calls when a similar routing request already happened, installed with a single pip command.
What's different
Positioned as a lighter, non-wrapper alternative to LangGraph, with a local cache to cut LLM call costs.
Pricing
Open source (Apache-2.0).
Best for
Developers building LLM agent pipelines who want to reduce redundant model calls.
Details →

RAI

LiveWatched 20d

A self-hostable AI chat platform that routes queries to the cheapest LLM and adds live market data, positioned as an alternative to ChatGPT or Claude you run yourself. For developers and cost-conscious teams.

FreeOpen sourceSelf-hosted
What to know
How it works
A self-hostable, open-source AI chat platform that runs on your own infrastructure, integrates live financial market data into chat, and routes each query to the best/cheapest LLM automatically across OpenAI, Anthropic, SiliconFlow and others.
What's different
Self-hosted with full data ownership and no vendor lock-in, plus intelligent model routing claimed to cut API costs by up to 60% versus using a single provider.
Pricing
Free forever to self-host, with Free / Pro / MAX tiers and per-model quotas.
Best for
Developers and cost-conscious teams who want a ChatGPT/Claude alternative running on their own infrastructure.
Details →

PokeTokenBar

LiveWatched 11d

A menu-bar app that turns your Claude Code/Codex/Gemini CLI token usage into a Pokémon-collecting game while tracking usage limits.

FreeWorks offlineOpen source
What to know
How it works
Reads local usage logs and hatches/evolves a Pokémon as you spend tokens
What's different
Runs entirely on-device, nothing sent anywhere
Pricing
free
Best for
Developers who want a fun way to watch their AI token usage
Details →

SuperCompress

LiveWatched 23d

Open-source API that compresses prompts before OpenAI, Claude, Gemini or RAG calls, cutting token costs by about 65% while keeping the important content. For developers wiring it into their own LLM calls.

Open source
What to know
How it works
An open-source API that compresses prompts before they're sent to OpenAI, Claude, Gemini, RAG, chatbot, or agent calls, stripping tokens while preserving the evidence needed to answer correctly.
What's different
Claims to remove more than 65% of tokens on average while preserving over 98% of answer-critical content.
Best for
Developers wiring their own LLM calls who want to cut token costs without integrating a full new pipeline.
Details →

StackSpine

LiveWatched 12d

An open-source control plane that enforces policy, budget limits, and compliance auditing across all the AI providers a company uses in production.

Open source
What to know
How it works
Sits between apps and AI providers to enforce guardrails and log usage for audit.
Pricing
Open-source, Apache 2.0.
Best for
Engineering teams running production AI across multiple providers who need governance.
Details →

burnban

LiveWatched 18d

A local, open-source binary that meters and caps AI agent spend on your own machine for developers who want budget control without sending usage data to a dashboard.

Open source
What to know
How it works
A local binary sits in the request path, meters every token spent by AI agents, and enforces spend caps before they're exceeded; also compares your flat-rate plan against real metered rates.
What's different
Open source (MIT); sync and team budgets are planned but not yet available.
Best for
Developers running AI agents who want local spend control without a hosted dashboard.
Details →

VOLY

LiveWatched 16d

Open-source control plane that routes coding tasks across AI agents like Claude Code, Cursor, and Codex with cost tracking.

Open source
What to know
How it works
An open-source control plane that routes coding tasks across AI coding agents (Claude Code, Cursor, Codex, Wrangler, etc.), runs multi-agent workflows, caches to reduce repeated token spend, falls back on billing/availability issues, and shows cost per task with files touched.
What's different
Gives one measurable layer above multiple separate agent tools instead of running each independently.
Details →

HardwareMon

LiveWatched 18d

Cross-platform hardware monitoring with real-time stats, history and benchmarking for Windows, Linux, macOS and Android, with all data processed locally and no telemetry upload.

Works offlineOpen source
What to know
How it works
Monitors CPU, GPU, memory, disk and network in real time, with historical analytics, process management, local benchmarking, and gaming session detection, across Windows, Linux, macOS (Apple Silicon), and Android.
What's different
All monitoring data is processed on-device with local SQLite storage; no telemetry or benchmark results are uploaded.
Best for
Users who want cross-platform hardware monitoring without sending data to the cloud.
Details →

Tokimeter

LiveWatched 1d
No sign-upOpen sourceNo tracking
Details →

CodexBar Lite

LiveWatched 7d

Lets developers monitor their OpenAI Codex usage from the macOS menu bar instead of checking a web dashboard, using only their existing CLI session.

Data stays with youOpen source
Details →

ISONGraph

LiveWatched 5d
Open source
Details →

TokenTelemetry

LiveWatched 3d
Works offlineOpen source
Details →
Open source
Details →

VoiceGateway

LiveWatched 3d
Open source
Details →