openspeech
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
File Explorer
Download Latest Version (.zip)- --new-model-addition.md
- bug_report.md
- feature_request.md
- question-help.md
- PULL_REQUEST_TEMPLATE.md
- --new-model-addition.md
- bug_report.md
- feature_request.md
- question-help.md
- pre-commit.yaml
- PULL_REQUEST_TEMPLATE.md
- configuration.html
- cross_entropy.html
- configuration.html
- ctc.html
- configuration.html
- joint_ctc_cross_entropy.html
- configuration.html
- label_smoothed_cross_entropy.html
- configuration.html
- perplexity.html
- configuration.html
- transducer.html
- configuration.html
- filter_bank.html
- configuration.html
- melspectrogram.html
- configuration.html
- mfcc.html
- configuration.html
- spectrogram.html
- augment.html
- data_loader.html
- dataset.html
- load.html
- data_loader.html
- dataset.html
- sampler.html
- lit_data_module.html
- lit_data_module.html
- lit_data_module.html
- lstm_attention_decoder.html
- openspeech_decoder.html
- rnn_transducer_decoder.html
- transformer_decoder.html
- transformer_transducer_decoder.html
- conformer_encoder.html
- contextnet_encoder.html
- convolutional_lstm_encoder.html
- convolutional_transformer_encoder.html
- deepspeech2.html
- jasper.html
- lstm_encoder.html
- openspeech_encoder.html
- quartznet.html
- rnn_transducer_encoder.html
- transformer_encoder.html
- transformer_transducer_encoder.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- configurations.html
- model.html
- openspeech_ctc_model.html
- openspeech_encoder_decoder_model.html
- openspeech_language_model.html
- openspeech_model.html
- openspeech_transducer_model.html
- add_normalization.html
- additive_attention.html
- batchnorm_relu_rnn.html
- conformer_attention_module.html
- conformer_block.html
- conformer_convolution_module.html
- conformer_feed_forward_module.html
- conv2d_extractor.html
- conv2d_subsampling.html
- conv_base.html
- conv_group_shuffle.html
- deepspeech2_extractor.html
- depthwise_conv1d.html
- dot_product_attention.html
- glu.html
- jasper_block.html
- jasper_subblock.html
- location_aware_attention.html
- mask.html
- mask_conv1d.html
- mask_conv2d.html
- multi_head_attention.html
- pointwise_conv1d.html
- positional_encoding.html
- positionwise_feed_forward.html
- quartznet_block.html
- quartznet_subblock.html
- relative_multi_head_attention.html
- residual_connection_module.html
- swish.html
- time_channel_separable_conv1d.html
- transformer_embedding.html
- vgg_extractor.html
- wrapper.html
- lr_scheduler.html
- reduce_lr_on_plateau_scheduler.html
- transformer_lr_scheduler.html
- tri_stage_lr_scheduler.html
- warmup_reduce_lr_on_plateau_scheduler.html
- warmup_scheduler.html
- optimizer.html
- beam_search_base.html
- beam_search_ctc.html
- beam_search_lstm.html
- beam_search_rnn_transducer.html
- beam_search_transformer.html
- beam_search_transformer_transducer.html
- ensemble_search.html
- character.html
- character.html
- grapheme.html
- subword.html
- character.html
- subword.html
- tokenizer.html
- callbacks.html
- metrics.html
- index.html
- Conformer.rst.txt
- ContextNet.rst.txt
- DeepSpeech2.rst.txt
- Jasper.rst.txt
- Listen Attend Spell.rst.txt
- LSTM LM.rst.txt
- QuartzNet.rst.txt
- RNN Transducer.rst.txt
- Transformer LM.rst.txt
- Transformer Transducer.rst.txt
- Transformer.rst.txt
- AISHELL-1.rst.txt
- KsponSpeech.rst.txt
- LibriSpeech.rst.txt
- Openspeech CTC Model.rst.txt
- Openspeech Encoder Decoder Model.rst.txt
- Openspeech Language Model.rst.txt
- Openspeech Model.rst.txt
- Openspeech Transducer Model.rst.txt
- Callback.rst.txt
- Criterion.rst.txt
- Data Augment.rst.txt
- Data Loaders.rst.txt
- Datasets.rst.txt
- Decoders.rst.txt
- Encoders.rst.txt
- Feature Transform.rst.txt
- Metric.rst.txt
- Modules.rst.txt
- Optim.rst.txt
- Search.rst.txt
- Tokenizers.rst.txt
- configs.md.txt
- hydra_configs.md.txt
- intro.md.txt
- index.rst.txt
- fontawesome-webfont.eot
- fontawesome-webfont.svg
- fontawesome-webfont.ttf
- fontawesome-webfont.woff
- fontawesome-webfont.woff2
- lato-bold-italic.woff
- lato-bold-italic.woff2
- lato-bold.woff
- lato-bold.woff2
- lato-normal-italic.woff
- lato-normal-italic.woff2
- lato-normal.woff
- lato-normal.woff2
- Roboto-Slab-Bold.woff
- Roboto-Slab-Bold.woff2
- Roboto-Slab-Regular.woff
- Roboto-Slab-Regular.woff2
- badge_only.css
- theme.css
- badge_only.js
- html5shiv-printshiv.min.js
- html5shiv.min.js
- theme.js
- basic.css
- doctools.js
- documentation_options.js
- file.png
- jquery-3.5.1.js
- jquery.js
- language_data.js
- minus.png
- plus.png
- pygments.css
- searchtools.js
- underscore-1.3.1.js
- underscore.js
- Conformer.html
- ContextNet.html
- DeepSpeech2.html
- Jasper.html
- Listen Attend Spell.html
- LSTM LM.html
- QuartzNet.html
- RNN Transducer.html
- Transformer LM.html
- Transformer Transducer.html
- Transformer.html
- AISHELL-1.html
- KsponSpeech.html
- LibriSpeech.html
- Openspeech CTC Model.html
- Openspeech Encoder Decoder Model.html
- Openspeech Language Model.html
- Openspeech Model.html
- Openspeech Transducer Model.html
- Callback.html
- Criterion.html
- Data Augment.html
- Data Loaders.html
- Datasets.html
- Decoders.html
- Encoders.html
- Feature Transform.html
- Metric.html
- Modules.html
- Optim.html
- Search.html
- Tokenizers.html
- configs.html
- hydra_configs.html
- intro.html
- Conformer.rst
- ContextNet.rst
- DeepSpeech2.rst
- Jasper.rst
- Listen Attend Spell.rst
- LSTM LM.rst
- QuartzNet.rst
- RNN Transducer.rst
- Transformer LM.rst
- Transformer Transducer.rst
- Transformer.rst
- AISHELL-1.rst
- KsponSpeech.rst
- LibriSpeech.rst
- Openspeech CTC Model.rst
- Openspeech Encoder Decoder Model.rst
- Openspeech Language Model.rst
- Openspeech Model.rst
- Openspeech Transducer Model.rst
- Callback.rst
- Criterion.rst
- Data Augment.rst
- Data Loaders.rst
- Datasets.rst
- Decoders.rst
- Encoders.rst
- Feature Transform.rst
- Metric.rst
- Modules.rst
- Optim.rst
- Search.rst
- Tokenizers.rst
- configs.md
- hydra_configs.md
- intro.md
- conf.py
- index.rst
- .buildinfo
- .nojekyll
- genindex.html
- index.html
- make.bat
- Makefile
- objects.inv
- py-modindex.html
- search.html
- searchindex.js
- eval.yaml
- lm_train.yaml
- README.md
- train.yaml
- __init__.py
- configuration.py
- cross_entropy.py
- __init__.py
- configuration.py
- ctc.py
- __init__.py
- configuration.py
- joint_ctc_cross_entropy.py
- __init__.py
- configuration.py
- label_smoothed_cross_entropy.py
- __init__.py
- configuration.py
- perplexity.py
- __init__.py
- configuration.py
- transducer.py
- __init__.py
- __init__.py
- configuration.py
- filter_bank.py
- __init__.py
- configuration.py
- melspectrogram.py
- __init__.py
- configuration.py
- mfcc.py
- __init__.py
- configuration.py
- spectrogram.py
- __init__.py
- augment.py
- data_loader.py
- dataset.py
- load.py
- data_loader.py
- dataset.py
- __init__.py
- sampler.py
- __init__.py
- configurations.py
- initialize.py
- __init__.py
- lit_data_module.py
- preprocess.py
- __init__.py
- character.py
- grapheme.py
- preprocess.py
- subword.py
- __init__.py
- lit_data_module.py
- __init__.py
- lit_data_module.py
- __init__.py
- character.py
- preprocess.py
- subword.py
- __init__.py
- lit_data_module.py
- __init__.py
- README.md
- __init__.py
- lstm_attention_decoder.py
- openspeech_decoder.py
- rnn_transducer_decoder.py
- transformer_decoder.py
- transformer_transducer_decoder.py
- __init__.py
- conformer_encoder.py
- contextnet_encoder.py
- convolutional_lstm_encoder.py
- convolutional_transformer_encoder.py
- deepspeech2.py
- jasper.py
- lstm_encoder.py
- openspeech_encoder.py
- quartznet.py
- rnn_transducer_encoder.py
- squeezeformer_encoder.py
- transformer_encoder.py
- transformer_transducer_encoder.py
- __init__.py
- lstm_lm.py
- openspeech_lm.py
- transformer_lm.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- configurations.py
- model.py
- __init__.py
- openspeech_ctc_model.py
- openspeech_encoder_decoder_model.py
- openspeech_language_model.py
- openspeech_model.py
- openspeech_transducer_model.py
- README.md
- __init__.py
- add_normalization.py
- additive_attention.py
- batchnorm_relu_rnn.py
- conformer_attention_module.py
- conformer_block.py
- conformer_convolution_module.py
- conformer_feed_forward_module.py
- contextnet_block.py
- contextnet_module.py
- conv2d_extractor.py
- conv2d_subsampling.py
- conv_base.py
- conv_group_shuffle.py
- deepspeech2_extractor.py
- depthwise_conv1d.py
- depthwise_conv2d.py
- dot_product_attention.py
- glu.py
- jasper_block.py
- jasper_subblock.py
- location_aware_attention.py
- mask.py
- mask_conv1d.py
- mask_conv2d.py
- multi_head_attention.py
- pointwise_conv1d.py
- positional_encoding.py
- positionwise_feed_forward.py
- quartznet_block.py
- quartznet_subblock.py
- relative_multi_head_attention.py
- residual_connection_module.py
- squeezeformer_attention_module.py
- squeezeformer_block.py
- squeezeformer_module.py
- swish.py
- time_channel_separable_conv1d.py
- transformer_embedding.py
- vgg_extractor.py
- wrapper.py
- __init__.py
- lr_scheduler.py
- reduce_lr_on_plateau_scheduler.py
- transformer_lr_scheduler.py
- tri_stage_lr_scheduler.py
- warmup_reduce_lr_on_plateau_scheduler.py
- warmup_scheduler.py
- __init__.py
- adamp.py
- novograd.py
- optimizer.py
- radam.py
- __init__.py
- beam_search_base.py
- beam_search_ctc.py
- beam_search_lstm.py
- beam_search_rnn_transducer.py
- beam_search_transformer.py
- beam_search_transformer_transducer.py
- ensemble_search.py
- __init__.py
- character.py
- __init__.py
- character.py
- grapheme.py
- subword.py
- __init__.py
- character.py
- subword.py
- __init__.py
- tokenizer.py
- __init__.py
- callbacks.py
- metrics.py
- README.md
- utils.py
- generate_openspeech_configs.py
- hydra_ensemble_eval.py
- hydra_eval.py
- hydra_lm_train.py
- hydra_train.py
- test_conformer.py
- test_conformer_lstm.py
- test_conformer_transducer.py
- test_contextnet.py
- test_contextnet_lstm.py
- test_contextnet_transducer.py
- test_conv2d_subsampling.py
- test_deep_cnn_with_joint_ctc_listen_attend_spell.py
- test_deepspeech2.py
- test_jasper10x5.py
- test_jasper5x3.py
- test_joint_ctc_conformer_lstm.py
- test_joint_ctc_listen_attend_spell.py
- test_joint_ctc_transformer.py
- test_listen_attend_spell.py
- test_listen_attend_spell_with_location_aware.py
- test_listen_attend_spell_with_multi_head.py
- test_lstm_for_causal_lm.py
- test_quartznet10x5.py
- test_quartznet15x5.py
- test_quartznet5x5.py
- test_rnn_transducer.py
- test_squeezeformer.py
- test_squeezeformer_lstm.py
- test_squeezeformer_transducer.py
- test_transformer.py
- test_transformer_transducer.py
- test_transformer_with_ctc.py
- test_vgg_transformer.py
- test_audio_augment.py
- test_lstm_lm.py
- test_transformer_for_causal_lm.py
- test_transformer_lm.py
- labels.csv
- test_lr_scheduler.py
- test_warprnnt_loss.py
- .gitignore
- .pre-commit-config.yaml
- CONTRIBUTING.md
- CONTRIBUTORS.md
- install.sh
- LICENSE
- LICENSE.3rd_party_library
- README.md
- requirements.txt
- setup.cfg
- setup.py
π Installation Guide
1. Get the code
git clone https://github.com/openspeech-team/openspeech
Downloads the entire project code from GitHub to your computer.
cd openspeech
Moves into the project folder you just downloaded.
2. Official Install Script
Easy RecommendedPrerequisites
- Python 3 Python is required to use pip.
pip install openspeech-core
Installs the package published on PyPI directly β no need to clone the source.
After installing, open a new terminal and run the program's version command (e.g. --version) to confirm it worked.
Pulled directly from this repo's README.
3. Python
EasyPrerequisites
pip install openspeech-core
Installs the package published on PyPI directly β no need to clone the source.
$ pip install -v --no-cache-dir --global-option="--cpp_ext" --global-option="--cuda_ext" ./
Installs the Python libraries listed in requirements.txt (or similar).
If it runs without errors and prints output in the terminal, it worked.
Pulled directly from this repo's README.
4. Make
MediumPrerequisites
- Git Needed to download the project code from GitHub.
- Make Usually pre-installed on Linux/macOS. On Windows, install separately (e.g. via MSYS2 or WSL).
cd docs
This project's files live in a subfolder, so move into it first.
make
Compiles the code based on the generated build configuration to produce an executable.
If it finishes without errors, it worked. Try running the generated executable directly.
// repository documentation
Was this content helpful?
(0 ratings)
