| Primary product category | AI voice generator (TTS) | AI audio production model |
| One-prompt multi-track generation | - | Yes |
| Dialogue + SFX + music in a single generation | - | Yes |
| Automatic timing arrangement & transitions | - | Yes |
| Output is fully-mixed, broadcast-ready | Single track (mixing done by user) | Fully-mixed master |
| Multi-character dialogue in one generation | Sequential (per-line generation) | Native (single-pass arrangement) |
| Non-verbal expression (laughs, sighs, dialects) | Yes (via tags / SSML) | Yes (embedded automatically) |
| Zero-shot voice cloning | Yes | Yes |
| Multi-modal input - text | Yes | Yes |
| Multi-modal input - reference audio | Yes | Yes |
| Multi-modal input - image (infer voice from a picture) | - | Yes (designed for; rolling out) |
| Long-form voice consistency across hours of content | Stable within tier limits | Native via continuation mode |
| Sound effects generation | Separate module (single-clip generation) | Generated inside the same prompt |
| Background music generation | Separate module | Generated inside the same prompt |
| Public voice library | Large public library (thousands of voices) | Small library (own-upload focused) |
| Languages supported | 30+ languages | English & Mandarin Chinese, expanding |
| Pricing model | Per-character | Per-credit (12 credits per generation) |
| Free tier | Limited monthly characters | 12 credits to try (about 1 generation) |
| Commercial use rights | From paid plans | From paid plans (Basic and up) |
| Post-production required | Yes - for SFX, music, multi-track mixing | No - output is already mixed |