speech-processing

19 projects share this GitHub topic

speech-processing — speechbrain ★11.7kspeech-processingmaths-cs-ai-compendium — ★7.2kawesome-multimodal-ml — ★6.9ktorchscale — ★3.1kwhisper-timestamped — ★2.8kIMS-Toucan — ★2.2kdeepvoice3_pytorch — ★2kawesome-diarization — ★1.9kopen-speech-corpora — ★1.4kStreamSpeech — ★1.3kSincNet — ★1.2kCrisperWhisper — ★1kMultiBench — ★635Speech-Backbones — ★604UniSpeech — ★486nnmnkwii — ★399speechbrain.github.io — ★374VocGAN — ★321Awesome-AVI — ★84maths-cs-ai-compendium★ 7.2kawesome-multimodal-ml★ 6.9ktorchscale★ 3.1kwhisper-timestamped★ 2.8kIMS-Toucan★ 2.2kdeepvoice3_pytorch★ 2kawesome-diarization★ 1.9kopen-speech-corpora★ 1.4kStreamSpeech★ 1.3kSincNet★ 1.2kCrisperWhisper★ 1kMultiBench★ 635Speech-Backbones★ 604UniSpeech★ 486nnmnkwii★ 399speechbrain.github.io★ 374VocGAN★ 321Awesome-AVI★ 84

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
speechbrain
A PyTorch-based Speech Toolkit
★ 11.7k
maths-cs-ai-compendium
Become a cracked AI/ML researcher/engineer with this unconventional textbook covering maths, computing, and…
★ 7.2k
awesome-multimodal-ml
Reading list for research topics in multimodal machine learning
★ 6.9k
torchscale
Foundation Architecture for (M)LLMs
★ 3.1k
whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2.8k
IMS-Toucan
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2.2k
deepvoice3_pytorch
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
★ 2k
awesome-diarization
A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.
★ 1.9k
open-speech-corpora
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1.4k
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech…
★ 1.3k
SincNet
SincNet is a neural architecture for efficiently processing raw audio samples.
★ 1.2k
CrisperWhisper
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
★ 1k
MultiBench
[NeurIPS 2021] Multiscale Benchmarks for Multimodal Representation Learning
★ 635
Speech-Backbones
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
★ 604
UniSpeech
UniSpeech - Large Scale Self-Supervised Learning for Speech
★ 486
nnmnkwii
Library to build speech synthesis systems designed for easy and fast prototyping.
★ 399
speechbrain.github.io
The SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users…
★ 374
VocGAN
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
★ 321
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 84
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.