SuperCompress
LiveOpen-source API that compresses prompts before OpenAI, Claude, Gemini or RAG calls, cutting token costs by about 65% while keeping the important content. For developers wiring it into their own LLM calls.
Open source
What to know▼
- How it works
- An open-source API that compresses prompts before they're sent to OpenAI, Claude, Gemini, RAG, chatbot, or agent calls, stripping tokens while preserving the evidence needed to answer correctly.
- What's different
- Claims to remove more than 65% of tokens on average while preserving over 98% of answer-critical content.
- Best for
- Developers wiring their own LLM calls who want to cut token costs without integrating a full new pipeline.
Watched for 17 days · last checked 11d ago





