PersonaPlex-7B
NVIDIA's full-duplex speech-to-speech model built on Moshi, adding a voice prompt and a text persona prompt so you can set who the agent is and how it sounds. The best open option for persona-driven, interruptible English voice agents.
Overview
Best for: Persona-driven English voice characters and interruptible assistants self-hosted on a single data-centre GPU.
At a glance
NVIDIA Open Model License (card says commercial use ok). FullDuplexBench latency scores 0.170 / 0.240 have no stated units. 24 GB VRAM is a third-party figure; NVIDIA lists A100/H100.
24 kHz
24 kHz
English.
Set by an audio voice prompt (voice conditioning) plus a text persona prompt.
Vendor FullDuplexBench: smooth turn-taking latency 0.170, user-interruption latency 0.240 (units not given on the card).
Self-hosted.
Self-hosted.
NVIDIA lists Ampere (A100) and Hopper (H100), tested on A100 80 GB. Community reports about 19 GB used on a 24 GB A10G; 24 GB is a safe target (third-party).
NVIDIA Open Model License (model card says ready for commercial use); base Moshi weights CC-BY-4.0.
Features
- full duplex
- barge-in and overlapping speech
- persona text prompt
- voice prompt conditioning
Pricing
No public price list.
No licence fee; you pay for GPU time.
n/a (self-hosted)
Free tier: Open weights
Source: huggingface.co
Setup
- Accept the NVIDIA Open Model License on Hugging Face (gated).
- Linux with an Ampere/Hopper GPU (A100 80 GB used in testing).
- Run the Moshi server pointed at the PersonaPlex repo and open the local web UI.
Endpoint
Local: https://localhost:8998
Authentication
Hugging Face token for the gated download; server has no auth by default.
Quick start python
pip install moshi
huggingface-cli login # gated model: accept the licence first
python -m moshi.server --hf-repo "nvidia/personaplex-7b-v1"
# then open https://localhost:8998
Written from the current docs. Check the vendor's SDK version before you ship.
Warnings
Custom NVIDIA licence
Not Apache/MIT: read the NVIDIA Open Model License terms (and keep the CC-BY attribution for Moshi) before shipping.
Gated download
You must accept the licence and share contact details on Hugging Face; automate with a token.
VRAM unclear officially
NVIDIA only lists A100/H100; plan for 24 GB+ and test on smaller cards.
English only, no tools
Same Moshi limits: English speech, no built-in function calling.
Latency claims are benchmark scores
The 0.170 / 0.240 figures are FullDuplexBench scores without stated units; measure on your hardware.
Plus 3 warnings that apply to all open models APIs. See category warnings.
Limits
- English only
- No documented tool calling (community forks exist)
Models and products
| Name | Status |
|---|---|
| nvidia/personaplex-7b-v1 | GA |
| kyutai/personaplex-rl-seamless | Preview |
Docs and sources
Docs
Sources used
- huggingface.co/nvidia/personaplex-7b-v1
- huggingface.co/api/models/nvidia/personaplex-7b-v1
- themenonlab.blog/blog/nvidia-personaplex-full-duplex-voice-ai-how-to-guide
- huggingface.co/abhinavpgagi/personaplex-tool-calling
Minimum VRAM (third-party), latency units, details of Kyutai's RL variant.