MoshiVis: an open source model for real-time speech dialog and image understanding
General Introduction MoshiVis is an open source project developed by Kyutai Labs and hosted on GitHub. It is based on the Moshi speech-to-text model (7B parameters), with about 206 million new adaptation parameters and frozen Pal...

































































































