Tools for developersAudio & voice

19 match your filters — 19 of them answered our check today.

Watching this group for 20 days · how we check

This week7 new arrivals

An open-source toolkit for building AI features — chat, RAG, speech, vision — that runs entirely in the browser with no servers and no data leaving the device. Built for developers.

Domain registered 7 months ago

Works offlineData stays with youOpen source
What to know
How it works
An open-source toolkit for running AI entirely in the browser: LLM chat across 76 models, RAG and vector search, Whisper and Kokoro speech, and vision, with no server or API key required. Runs on WebGPU with a WebAssembly fallback.
What's different
Data never leaves the device, and it ships a zero-dependency core plus a shadcn UI registry of 107 components and 36 installable blocks for building on top of it.
Pricing
Open source (MIT license).
Best for
Developers building browser-based AI features without standing up servers or managing API keys.

Watched for 12 days · last checked 1d ago

Details →

Klipzo

Live

Gives students and developers 130+ free browser-based file tools instead of uploading files to third-party servers.

FreeData stays with you

Watched for 1 day · last checked 1d ago

Details →

Single API that gives developers access to multiple AI providers (text, image, TTS, transcription) with one key and one billing model.

Domain registered 1 month ago

What to know
How it works
Normalizes authentication and error responses across providers behind one endpoint
What's different
Credit-based billing replaces per-token pricing for predictable costs
Best for
Developers and small agencies shipping AI features without maintaining multiple integrations

Watched for 4 days · last checked 4d ago

Details →

A browser-based AI audio studio that generates speech, voice clones, music, and sound effects from text, reference audio, or images, for creators, podcasters, and game developers.

Domain registered 1 month ago

Nothing to install
What to know
How it works
Multimodal generation from text prompts, reference audio, or an image, covering speech, multi-character dialogue, voice cloning, music, and ambience.
Best for
Podcasters, filmmakers, game developers, teachers, and marketers building audio scenes.

Watched for 7 days · last checked 1d ago

Details →

A Mac dictation app that turns your voice, screenshots, and screen recordings into a clean AI-ready Markdown prompt for Claude Code, Cursor, or ChatGPT, running fully offline via local Whisper.

Domain registered 2 months ago

Works offline
What to know
How it works
Local Whisper and open models on a Mac turn voice, screenshots, and screen recordings into a clean Markdown prompt.
Best for
Developers using Claude Code, Cursor, or ChatGPT who want offline, private dictation without per-token billing.

Watched for 7 days · last checked 1d ago

Details →

Mac dictation app that turns speech into text system-wide via a hotkey, runs locally for privacy, and is sold as a one-time purchase instead of a subscription.

No subscription
What to know
How it works
System-wide dictation on macOS activated by a customizable push-to-talk key, converting speech to text in any app.
What's different
Runs entirely locally on the Mac for privacy.
Pricing
Lifetime license, one-time purchase.
Watch out
macOS only.

Watched for 8 days · last checked 2d ago

Details →

An API that pulls transcripts and metadata from YouTube, TikTok, Instagram and Facebook into structured data, with AI captions when none exist. For developers building on video content.

Free
What to know
How it works
An API that extracts transcripts (with an AI-generated fallback when captions don't exist), metadata, and prompt-driven structured video analysis from YouTube, TikTok, Instagram, X and Facebook content.
What's different
Covers transcripts, metadata and structured video analysis across five platforms through one API, returning clean JSON.
Pricing
A free YouTube API tier is available alongside the paid endpoints.
Best for
Developers building pipelines on video or social content who need transcripts or structured data as JSON.

Watched for 14 days · last checked 8d ago

Details →

Talk to an AI coding agent from your phone and it writes, tests and deploys code on your own cloud dev box — no laptop needed, for hands-free developers.

Nothing to install
What to know
How it works
A voice-first AI coding agent running on your own persistent cloud dev box — talk to it from your phone and it writes, tests in a real browser, and deploys code to a live URL, with no laptop or setup needed.
What's different
Runs as a persistent machine you own with memory and scheduled tasks, rather than an editor plugin or chatbot session, and auto-routes to the cheapest capable model.
Best for
Developers who want to write and ship code hands-free from a phone, such as during a commute.

Watched for 15 days · last checked 9d ago

Details →

PodKit

Live

A developer API for podcast search, metadata, chapters and transcripts as JSON from $15/month, with an MCP server for AI agents. For developers building on podcast data, not podcasters.

Domain registered 14 days ago

What to know
How it works
A developer API returning podcast search results, episode metadata, chapters, and transcripts as JSON, with a first-party MCP server so AI agents can call the same endpoints as tools.
What's different
Allows caching of responses and requires no attribution logo, and exposes the same data as MCP tools on the same API key and quota.
Pricing
Starts at $15/month.
Best for
Developers building products on top of podcast data, including AI agent builders who want podcast lookup as a tool.

Watched for 11 days · last checked 5d ago

Details →

Bundles AI image, 3D, video, voice and music tools into one subscription for indie game developers and streamers, instead of paying for five separate apps.

What to know
How it works
Bundles AI image generation across 14+ models, image-to-3D mesh export for Unity/Unreal/Godot, image-to-video, voice cloning, music and SFX generation, custom LoRA training, sprite sheet packing and a paint editor under one subscription.
What's different
Replaces stacking five separate AI creator tools with a single subscription, and outputs are commercially usable on every plan.
Pricing
200 free credits at signup, no card required; paid subscription beyond that.
Best for
Indie game developers, streamers and content creators who currently pay for multiple separate AI creation tools.

Watched for 18 days · last checked today

Details →

Real-time talking-avatar API developers plug into an existing STT/LLM/TTS stack to give voice agents a face, for support, sales, or onboarding use cases.

What to know
How it works
An API that streams a real-time talking avatar, connecting to your existing speech-to-text, LLM, and text-to-speech stack.
What's different
Lets teams create custom AI faces in seconds for branded agents.
Pricing
Starting as low as $0.01 per minute at scale.
Best for
Support, sales, tutoring, and onboarding voice agents.

Watched for 8 days · last checked 8d ago

Details →

A developer API that handles the real-time audio streaming, WebSockets and voice-activity detection behind Whisper-grade dictation, so you don't have to build that plumbing yourself.

Domain registered 1 month ago

What to know
How it works
A developer API and UI toolkit that adds Whisper-grade voice dictation to a web app, handling real-time audio chunking, WebSockets, and voice activity detection so users can speak instead of type.
What's different
Provides plug-and-play UI components for React, Next.js, and Vanilla JS, plus pay-as-you-go pricing instead of a flat monthly developer fee.
Best for
Developers who want to add voice dictation to text-heavy forms without building the audio streaming infrastructure themselves.

Watched for 13 days · last checked 7d ago

Details →

A free, on-device tool that lets you draw and talk over your screen so your AI coding agent sees exactly what you mean — built for Claude Code, Cursor and Codex users.

FreeNo sign-upWorks offline
What to know
How it works
Draw directly on your screen and talk while your coding agent watches, then get back annotated screenshots, a voice transcript, and a timeline of what you pointed at, all combined together.
What's different
Runs 100% on-device and works with Claude Code, Codex, and Cursor.
Pricing
Free, no account required.
Best for
Developers who need to show an AI coding agent a visual bug or layout idea instead of describing it in words.

Watched for 13 days · last checked 7d ago

Details →

AI tool for creators and marketers to generate royalty-free songs, instrumentals, and vocals from a text prompt for commercial use.

Domain registered 3 months ago

What to know
How it works
generates tracks from a prompt or purpose-built creation mode, with editing/remix tools
What's different
full commercial rights with no attribution required
Best for
creators, game developers, podcasters, filmmakers, marketers

Watched for 4 days · last checked 2d ago

Details →

API/SDK for embedding audio, video, and screen recording into your own web app without building the infrastructure.

Domain registered 11 years ago

What to know
How it works
Embeddable recording widget plus backend for upload, storage, transcoding, and playback
Best for
Developers who need in-app recording without building media infrastructure

Watched for 6 days · last checked 6d ago

Details →

Prebuilt IFTTT recipes that plug Grok, Gemini and Perplexity into apps you already use — summarize meetings, transcribe audio, or route AI replies into Slack or Drive.

Domain registered 17 years ago

What to know
How it works
Prebuilt IFTTT recipes that plug AI tools like Grok, Gemini, and Perplexity into apps you already use — summarizing meetings when they end, transcribing audio, generating images on a schedule, and routing AI responses into apps like Google Drive, Notion, Slack, or SMS. Set up once and it runs whenever the trigger conditions are met.
Best for
Developers, content creators, or anyone wanting to automate AI-related tasks across their existing apps without building anything custom.

Watched for 12 days · last checked 6d ago

Details →

An API platform giving developers access to over 100 image, video, audio, and chat AI models through one integration.

Domain registered 4 years ago

What to know
How it works
Gives developers access to over 100 image, video, audio, chat, and enhancement AI models through a single API, letting them compare costs and test performance.

Watched for 10 days · last checked 7d ago

Details →

Offline on-device speech to text for Mac and Windows

No sign-upWorks offlineData stays with you

First seen today · last checked today

Details →

A tool that turns spoken descriptions or call transcripts into editable system/data-flow diagrams.

Can export my data
What to know
How it works
Listens to a spoken description of a data flow or a call transcript and builds an editable graph of the actors, systems, decisions and data stores described, then lets you edit by voice, prompt, right-click or inspector.
What's different
Replaces drag-and-box diagramming with verbal description, aimed at not losing the room mid-explanation.
Pricing
Free to start.
Best for
Tech sellers and solution engineers describing architectures live.

Watched for 3 days · last checked 1d ago

Details →