VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation i
Compare 209 quality-filtered Audio Tools tools using public popularity, activity, licensing and source data.
Browse quality-filtered Audio Tools tools ranked by public signals. Filter by source, pricing, or recency to find a fit faster.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation i
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
Open-source framework for conversational voice AI agents
Build voice agents with open-source models
The free and privacy-friendly screen recorder with no limits 🎥
Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Wind
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
A simple, high-quality voice conversion tool focused on ease of use and performance.
UI components and hooks for building video/audio players on the web. Robust, customizable, and accessible. Modern alternative to JW Player and Video.js.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly op
Data manipulation and transformation for audio signal processing, powered by PyTorch
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning
Tools for handling multimodal data in machine learning projects.
The most advanced, fully offline client-side AI suite on Android today.
🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, commun
⏩ Fast-forwards long pauses between sentences — watch lectures ~1.5x faster (browser extension)
🎩 An Alfred 5 Workflow for using OpenAI Chat API to interact with GPT models 🤖💬 It also allows image generation/editing/understanding 🖼️, speech-to-text conv
MCP plugin for Unreal Engine 5.7 & 5.8 — gives AI assistants full read/write access to Blueprints, Materials, Niagara, Animation, Mesh, AI, GAS, Logic Driver, C
Lightning-Fast, On-Device, Multilingual, Accurate TTS