AI-powered tool to turn long videos into short, viral-ready clips. Combines transcription, speaker diarization, scene detection & 9:16 resizing — perfect for creators & smart automation.
-
Updated
Apr 2, 2025 - Python
AI-powered tool to turn long videos into short, viral-ready clips. Combines transcription, speaker diarization, scene detection & 9:16 resizing — perfect for creators & smart automation.
speech to text gui for different (e.g. Whisper, Voxtral) models and backends, including whisper.cpp, crispasar, mlx-whisper, faster-whisper, ctranslate2; applies pyannote for diarization
ONNX implementation of Pyannote Speaker Diarization 3.1 pipeline.
A Tool to Transcribe Audio Files with Speaker Diarization
Audio Diarization and Classification to perform Classroom Activity Detection (detect teacher, student, multiple speakers) over an input video.
[Graduation Project] Flow: Enterprise Meeting Content Digitization & Intelligence Platform featuring Modular Monolith, Async AI Worker, Vector Semantic Search (pgvector), and Docker stack.
ONNX implementation of Pyannote Speaker Diarization Community-1 pipeline.
A multimodal speaker diarization system using audio, video, and dialogue cues 🗣️💬
Speaker diarization that exports timeline-aligned tracks for video editing
An intelligent Streamlit application to transcribe and analyze multi-speaker medical consultations. This tool automatically identifies who spoke when (diarization), transcribes their speech (ASR), and assigns their role (Clinician or Patient), even in conversations that mix English and other languages like Hindi or Tamil.
C++ port of the pyannote framework preserving exclusive diarization from community-1
WebSocket based Python implementation that streams live audio to the Deepgram API for real-time transcription and speaker diarization.
FastAPI app for podcast transcription with automatic speaker diarization using pyannote.audio and faster-whisper.
This repository is to experiment the integration with @ggml-org/whisper.cpp for offline STT + pyannote/speaker-diarization-3.1
AI-Powered Speech Recognition & Diarization: A robust Streamlit application leveraging WhisperX and Faster-Whisper for accurate transcription and speaker separation. Features dual-mode processing (Fast/Pro), automatic speaker identification, color-coded Word (.docx) export, and CPU-optimized Docker deployment on AWS EC2.
Event-driven pipeline for evaluating recruiting interviews (FastAPI, Redis Queue, PostgreSQL) with pluggable transcription (WhisperX / OpenAI), speaker diarization (Pyannote), structured LLM scoring via OpenRouter, and human-in-the-loop evidence validation.
Transcription audio locale via whisper.cpp — NestJS + Next.js + pyannote.audio
A simple protocol manager for your audios
Self-hosted meeting transcription with speaker attribution, searchable notes, and live streaming
Local meeting transcription with speaker diarization and known-speaker recognition (Whisper + pyannote). Runs offline; optimized for automatic language detection.
To associate your repository with the pyannote-audio topic, visit your repo's landing page and select "manage topics."