GA NVIDIA

PersonaPlex-7B

NVIDIA's full-duplex speech-to-speech model built on Moshi, adding a voice prompt and a text persona prompt so you can set who the agent is and how it sounds. The best open option for persona-driven, interruptible English voice agents.

Est. per minuten/a

Overview

Best for: Persona-driven English voice characters and interruptible assistants self-hosted on a single data-centre GPU.

At a glance

TypeVoice-to-voice
Params B8
VRAM GB24
CPU okNo
LicenceCustom
CommercialYes
Full duplexYes
StreamingYes
Languages1

NVIDIA Open Model License (card says commercial use ok). FullDuplexBench latency scores 0.170 / 0.240 have no stated units. 24 GB VRAM is a third-party figure; NVIDIA lists A100/H100.

Audio in

24 kHz

Audio out

24 kHz

Languages

English.

Voices

Set by an audio voice prompt (voice conditioning) plus a text persona prompt.

Latency

Vendor FullDuplexBench: smooth turn-taking latency 0.170, user-interruption latency 0.240 (units not given on the card).

Regions

Self-hosted.

Compliance

Self-hosted.

Hardware

NVIDIA lists Ampere (A100) and Hopper (H100), tested on A100 80 GB. Community reports about 19 GB used on a 24 GB A10G; 24 GB is a safe target (third-party).

Licence

NVIDIA Open Model License (model card says ready for commercial use); base Moshi weights CC-BY-4.0.

Features

  • full duplex
  • barge-in and overlapping speech
  • persona text prompt
  • voice prompt conditioning

Pricing

No public price list.

How the per-minute estimate was worked out

No licence fee; you pay for GPU time.

Audio token rate

n/a (self-hosted)

Free tier: Open weights

Source: huggingface.co

Setup

  1. Accept the NVIDIA Open Model License on Hugging Face (gated).
  2. Linux with an Ampere/Hopper GPU (A100 80 GB used in testing).
  3. Run the Moshi server pointed at the PersonaPlex repo and open the local web UI.

Endpoint

Local: https://localhost:8998

Authentication

Hugging Face token for the gated download; server has no auth by default.

Quick start python

pip install moshi
huggingface-cli login   # gated model: accept the licence first
python -m moshi.server --hf-repo "nvidia/personaplex-7b-v1"
# then open https://localhost:8998

Written from the current docs. Check the vendor's SDK version before you ship.

Warnings

Custom NVIDIA licence

Not Apache/MIT: read the NVIDIA Open Model License terms (and keep the CC-BY attribution for Moshi) before shipping.

Gated download

You must accept the licence and share contact details on Hugging Face; automate with a token.

VRAM unclear officially

NVIDIA only lists A100/H100; plan for 24 GB+ and test on smaller cards.

English only, no tools

Same Moshi limits: English speech, no built-in function calling.

Latency claims are benchmark scores

The 0.170 / 0.240 figures are FullDuplexBench scores without stated units; measure on your hardware.

Plus 3 warnings that apply to all open models APIs. See category warnings.

Limits

  • English only
  • No documented tool calling (community forks exist)

Models and products

NameStatusNotes
nvidia/personaplex-7b-v1GAAbout 8B params on disk; released January 2026; gated (accept licence).
kyutai/personaplex-rl-seamlessPreviewJune 2026 RL fine-tune published by Kyutai; not verified.

Docs and sources

Docs

Sources used

Not fully verified

Minimum VRAM (third-party), latency units, details of Kyutai's RL variant.

Similar open models APIs

Spotted a wrong price or a dead link?