Updates to docs and small fixes in preparation for next release [skip ci]

This commit is contained in:
dscripka 2023-06-14 07:58:23 -04:00
parent e05ac019e8
commit 7666ddb257
4 changed files with 54 additions and 16 deletions

17
CHANGELOG.md Normal file
View file

@ -0,0 +1,17 @@
# Change Log
## v0.4.0 - 2023/06/14
### Added
* A new wakeword model, "hey rhasspy"
* Added support for tflite versions of the melspectrogram model, embedding model, and pre-trained wakeword models
* Added an inference framework argument to allow users to select either ONNX or tflite as the inference framework
* The `detect_from_microphone.py` example now supports additional arguments and has improved console formatting
### Changed
* Made tflite the default inference framework for linux platforms due to improved efficiency, with windows still using ONNX as the default given the lack of pre-built Windows WHLs for the tflite runtime (https://pypi.org/project/tflite/)
* Adjusted the default provider arguments for onnx models to avoid warnings (https://github.com/dscripka/openWakeWord/issues/27)
### Removed

View file

@ -1 +1,2 @@
recursive-include openwakeword *.onnx
recursive-include openwakeword *.onnx
recursive-include openwakeword *.tflite

View file

@ -4,13 +4,19 @@
openWakeWord is an open-source wakeword library that can be used to create voice-enabled applications and interfaces. It includes pre-trained models for common words & phrases that work well in real-world environments.
# Updates
**2023/06/14**
- v0.4.0 of openWakeWord released. See the [changelog](CHANGELOG.md) for a full descriptions of new features and changes.
# Demo
You can try an online demo of the included pre-trained models via HuggingFace Spaces [right here!](https://huggingface.co/spaces/davidscripka/openWakeWord).
Note that real-time detection of a microphone stream can occasionally behave strangely in Spaces. For the most reliable testing, perform a local installation as described below.
# Installation & Usage
# Installation
Installing openWakeWord is simple and has minimal dependencies:
@ -18,6 +24,8 @@ Installing openWakeWord is simple and has minimal dependencies:
pip install openwakeword
```
On Linux systems, both the [onnxruntime](https://pypi.org/project/onnxruntime/) package and [tflite-runtime](https://pypi.org/project/tflite-runtime/) packages will be installed as dependencies since both inference frameworks are supported. On Windows, only onnxruntime is installed due to a lack of support for modern versions of tflite.
To (optionally) use [Speex](https://www.speex.org/) noise suppression on Linux systems to improve performance in noisy environments, install the Speex dependencies and then the pre-built Python package (see the assets [here](https://github.com/dscripka/openWakeWord/releases/tag/v0.1.1) for all .whl versions), adjusting for your python version and system architecture as needed.
```
@ -27,6 +35,8 @@ pip install https://github.com/dscripka/openWakeWord/releases/download/v0.1.1/sp
Many thanks to [TeaPoly](https://github.com/TeaPoly/speexdsp-ns-python) for their Python wrapper of the Speex noise suppression libraries.
# Usage
For quick local testing, clone this repository and use the included [example script](examples/detect_from_microphone.py) to try streaming detection from a local microphone. **Important note!** The model files are stored in this repo using [git-lfs](https://git-lfs.com/); make sure it is installed on your system and if needed use `git-lfs fetch --all` to make sure the the models download correctly.
Adding openWakeWord to your own Python code requires just a few lines:
@ -36,7 +46,7 @@ from openwakeword.model import Model
# Instantiate the model
model = Model(
wakeword_model_paths=["path/to/model.onnx"], # can also leave this argument empty to load all of the included pre-trained models
wakeword_models=["path/to/model.onnx"], # can also leave this argument empty to load all of the included pre-trained models
)
# Get audio data containing 16-bit 16khz PCM audio data from a file, microphone, network stream, etc.
@ -48,6 +58,27 @@ frame = my_function_to_get_audio_frame()
prediction = model.predict(frame)
```
Additionally, openWakeWord provides other useful utility functions. For example:
```python
# Get predictions for individual WAV files (16-bit 16khz PCM)
from openwakeword.model import Model
model = Model()
model.predict_clip("path/to/wav/file")
# Get predictions for a large number of files using multiprocessing
from openwakeword.utils import bulk_predict
bulk_predict(
file_paths = ["path/to/wav/file/1", "path/to/wav/file/2"],
wakeword_models = ["hey jarvis"],
ncpu=2
)
```
See `openwakeword/utils.py` and `openwakeword/model.py` for the full specification of class methods and utility functions.
# Recommendations for Usage
## Noise Suppression and Voice Activity Detection (VAD)
@ -91,6 +122,7 @@ The table below lists each model, examples of the word/phrases it is trained to
| alexa | "alexa"| [docs](docs/models/alexa.md) |
| hey mycroft | "hey mycroft" | [docs](docs/models/hey_mycroft.md) |
| hey jarvis | "hey jarvis" | [docs](docs/models/hey_jarvis.md) |
| hey rhasspy | "hey rhasspy" | TBD
| current weather | "what's the weather" | [docs](docs/models/weather.md) |
| timers | "set a 10 minute timer" | [docs](docs/models/timers.md) |
@ -186,4 +218,4 @@ Future release road maps may have non-english support. In particular, [Mycroft.A
# License
All of the code in openWakeWord is licensed under the **Apache 2.0** license. All of the included pre-trained models are licensed under the [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International](https://creativecommons.org/licenses/by-nc-sa/4.0/) license due to the inclusion of datasets with unknown or restrictive licensing as part of the training data. If you are interested in pre-trained models with more permissive licensing, please raise an issue and we will try to add them to a future release.
All of the code in this repository is licensed under the **Apache 2.0** license. All of the included pre-trained models are licensed under the [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International](https://creativecommons.org/licenses/by-nc-sa/4.0/) license due to the inclusion of datasets with unknown or restrictive licensing as part of the training data. If you are interested in pre-trained models with more permissive licensing, please raise an issue and we will try to add them to a future release.

View file

@ -277,14 +277,6 @@ class Model():
if len(x) > 1280:
group_predictions = []
for i in np.arange(len(x)//1280-1, -1, -1):
# group_predictions.extend(
# self.models[mdl].run(
# None,
# {input_name: self.preprocessor.get_features(
# self.model_inputs[mdl],
# start_ndx=-self.model_inputs[mdl] - i
# )}
# )
group_predictions.extend(
self.model_prediction_function[mdl](
self.preprocessor.get_features(
@ -295,10 +287,6 @@ class Model():
)
prediction = np.array(group_predictions).max(axis=0)[None, ]
else:
# prediction = self.models[mdl].run(
# None,
# {input_name: self.preprocessor.get_features(self.model_inputs[mdl])}
# )
prediction = self.model_prediction_function[mdl](
self.preprocessor.get_features(self.model_inputs[mdl])
)