This repo builds an end-to-end deep learning application that supports speech recognition system. It's simple to use and understand 😄
-
Updated
May 23, 2023 - Jupyter Notebook
This repo builds an end-to-end deep learning application that supports speech recognition system. It's simple to use and understand 😄
Offline macOS menu bar app for speech-to-text transcription using AI models (NVIDIA Parakeet, Whisper). 100% local & private.
🎙️ The open-source AI Voice-to-Text & Universal Speech Studio. Real-time dictation, 48kHz Web Audio DSP noise cancellation, 2-Way Babel Live Translator with authentic Urdu/multilingual Neural TTS, Executive MoM PDF generator, Studio EQ mastering, and 100% offline audio tools. Powered by Gemini 2.0, GPT-4o & Claude 3.7.
The Speech Recognition or Speech-to-Text Converter module in Android, implemented using Kotlin, facilitates the conversion of spoken language into written text
Real-time voice to text for Windows. Transcribe microphone input to text in any app. Free offline speech recognition tool for Windows 10/11
CLI that turns a YouTube URL into a text transcript: yt-dlp for the audio, Whisper for the transcription, running 100% locally with GPU support.
Speech to text converter using Google Android
Open-source alternative to Superwhisper and Wispr Flow for Windows. Private, system-wide voice-to-text with local transcription and no subscription.
A speech to text web app for people with speech impairments that has support for Kenyan English & Swahili accents
AI assistant to transcribe, translate and summarize voice messages and audio files
🎙️ Asistente de cues de audio en tiempo real que detecta palabras y frases habladas y activa sonidos configurados para podcasts, radio y televisión en vivo.
Lipi is a lightweight, local-first desktop speech-to-text and voice note application. Built with Tauri v2, Rust, and React 19, it lets you record audio, transcribe speech with OpenAI Whisper or any OpenAI-compatible local/remote endpoint, and manage your notes seamlessly.
Free, private, local speech-to-text app for macOS and Windows — turn audio, video, or YouTube links into .srt/.vtt/.txt/.json subtitles, no cloud, no account.
A lightweight desktop application that captures your voice, converts it to text in real-time, and stores the transcriptions in a local SQLite database — complete with timestamp, language selection, waveform visualizer, and audio playback.
Local offline voice-to-text desktop app for Windows (Whisper, hotkeys, wake word, privacy-first)
**VoxScribe** is a Windows desktop application for speech-to-text transcription using Whisper.cpp. It supports both microphone recording and audio/video file transcription with GPU acceleration (CUDA) and CPU optimizations (AVX2/FMA).
Polyglot Transcribe — near real-time multilingual speech-to-text with AI-generated structured reports (French, Arabic, English)
Sanitized portfolio overview of a university capstone using AI transcription, speaker diarization, image analysis, RAG, and multi-format report generation.
Open-source voice typing for Windows — a Wispr Flow / Superwhisper alternative. Press a shortcut, speak, and AI-polished text lands at your cursor. Local models, your own API keys, or a self-hosted backend.
To associate your repository with the speech-to-text-app topic, visit your repo's landing page and select "manage topics."