What Tablif read on Weir's site
Weir describes itself as: “GitHub - yatinannam/weir: A cost-aware gateway for RAG: a semantic cache, a model router and per-request cost accounting. · GitHub”
- Stage: Live, no waitlist
- How it differs: “Weir is not a retriever, vector database or LLM framework. It wraps a RAG service it does not own.”
What it is
Calls itself “gateway”.
- Calls itself: gateway
“Weir is an open-source gateway for RAG that combines semantic caching, cost-aware model routing, and per-request cost and latency accounting.”
Where it puts itself
Rules out “retriever”.
- Says it will not: retriever
“Weir is not a retriever, vector database or LLM framework. It wraps a RAG service it does not own.”
“Weir is not a retriever, vector database or LLM framework.”
Who it sells to
Sells to developers.
- Sells to: developers
“Prerequisites: Docker Desktop (running) and uv . A free Groq API key is optional.”
How it runs
A API. Can be self-hosted. Offers an API. Open source. AI is the core of the product. You can bring your own API key.
- Kind: api
- Runs where: self-hosted
- API: yes
- Open source: yes
- AI role: core to the product
- Bring your own key: yes
“Weir is an open-source gateway for RAG that combines semantic caching, cost-aware model routing, and per-request cost and latency accounting.”
“…docker compose up -d --build # postgres, migrations, hospital-rag, weir”
What it works with
Names GitHub Copilot, Grafana and 2 more on its pages.
- Names: GitHub Copilot · Grafana · Postgres · +1
“GitHub Copilot Write better code with AI”
“Grafana :3000”
What it offers as proof
Live, no waitlist. Claims “40 documents and 158 eval questions”. Publishes a case study. Publishes docs.
- Number claimed: 40 documents and 158 eval questions
- Case study: yes
“You must be signed in to change notification settings”
“40 documents and 158 eval questions”
Sources
Tablif read these pages, last on 10 Oct 2026. Everything above is what the product says about itself; Tablif does not verify every claim.
Similar products in Tablif Town
- EchoCache: Stop paying twice for identical LLM queries.
- Cosmic Engine: RAG strategies that you can find - ALL in one place!
- Throttle: Smart routing for LLM inference, save thousands on API costs
- Swytch: Run a Redis-compatible cache with database-level consistency
- ApexCache: REAL-TIME ACCELERATION FOR YOUR APIs
- Aiondb: Fast multimodal embedded database for RAG apps
- Weave: See output, AI cost and value.
- Cachely: Managed remote build cache for modern build tools