Find the answer.
Keep your work moving.
Search practical answers about voices, transcription, local models, exports, providers, payments, and licenses.
Try “desktop audio”, “multiple voices”, “export”, or “license”.
Help centre
Start with a topic.
Open what you need.
Showing all 24 answers
Getting started
The workspace, platforms, connectivity, and project storage.
What is Voxnara?
Voxnara is a desktop speech workspace for creating audio from text and turning live or recorded audio into editable transcripts. It keeps Text to Speech, Speech to Text, projects, queues, exports, and provider settings in one application.
Which operating systems does Voxnara support?
Voxnara is designed for Windows and macOS. Download availability can differ by release, so use the Download page to see the installers currently published for your platform.
Do I need an internet connection?
Local projects and installed speech features can work without an internet connection. Downloads, updates, license activation, and optional cloud Text to Speech providers require a connection when those services are used.
Where are my projects stored?
Projects and managed media are stored by the desktop application on your computer. Retention settings control when eligible managed files are cleaned up, while project content remains available until you remove it.
Text to Speech
Voices, selection controls, pronunciation, effects, and audio export.
Why do the available voices and languages change?
Voxnara lists only the voices, languages, styles, and controls reported by an installed or connected engine. If a provider is not connected—or does not support a capability—it will not appear as an available choice.
Can one document use multiple voices?
Yes. The Engine settings provide the document defaults. Highlight a passage and use its context menu to assign a different voice or delivery settings without changing the rest of the document.
Can I listen to only part of a document?
Yes. Highlight the words or passage, open the context menu, and choose Play. The selection is prepared and played directly. The bottom transport remains the control for playing the complete document.
What does the Pronunciation dictionary do?
It stores reusable rules for names, abbreviations, dates, specialist terms, or phrases that a voice may otherwise pronounce incorrectly. Rules can be added, edited, and reused across synthesis.
Which controls and effects can I apply to selected text?
A selection can override its voice, speed, pitch, loudness, and emphasis. Local effects include soft voice, breathing, whisper, dynamic-range control, and timbre shaping. Available provider-native options appear only when the selected engine supports them.
Which audio formats can I export?
Voxnara supports WAV, audio-only MP4, OGG, and WebM. When an engine cannot return the chosen format directly, Voxnara creates a WAV master and converts it locally before saving the export.
Speech to Text
Live capture, local models, accuracy, desktop audio, and speaker labels.
How do I start a transcription?
Open Speech to Text and choose Microphone, Desktop audio, or Upload audio. Select the locale and local recognition model, then begin capture or choose a file. New text appears in the transcript canvas and can be saved as a project.
Does transcription appear while recording?
Yes. Supported local recognition displays interim text while you speak and commits stable segments to the canvas during the recording. Final processing can refine the most recent segment after capture stops.
Does local recognition use my GPU?
Compatible Whisper models prefer GPU acceleration. When a supported GPU path is unavailable, Voxnara falls back to bounded CPU processing so recognition still works without consuming every logical processor.
How can I improve transcription accuracy?
Choose a model that supports the spoken language, select the correct locale, reduce background noise, keep the microphone close, and avoid overlapping speakers where possible. Larger compatible local models can improve recognition but need more memory and processing time.
Why is desktop audio capture unavailable?
Desktop capture depends on a supported operating-system loopback device and an active output endpoint. Start audio playback, confirm the intended output device is enabled, then try capture again. Some virtual, exclusive-mode, remote, or disconnected devices may not expose a usable loopback stream.
How does speaker detection work?
Choose the expected number of speakers from 1 to 10. After capture, local diarization compares the audio segments and adds generic speaker labels. You can rename labels afterward, and the transcript remains intact if speaker labeling cannot complete.
License & billing
Plans, checkout, delivery, activation, and moving a license.
What is the difference between Yearly and Lifetime?
Both plans include the same Voxnara features. The RM9.99 Yearly license is valid for one year. The RM29.99 Lifetime license has no commercial expiry.
Does the Yearly plan renew automatically?
No. The current Yearly plan is a one-time payment and does not renew automatically. You decide whether to purchase another period after it expires.
When will I receive my license key?
After Stripe confirms payment, Voxnara generates the license and displays it on the checkout result page. A copy is also sent to the purchase email when email delivery succeeds. Save the key and use the same email during activation.
How do I activate or move my license?
Open Voxnara and enter the purchase email and license key on the activation screen. To move an active license, deactivate it from About → License before activating the new installation. Contact support if the previous device is no longer accessible.
Privacy & troubleshooting
Provider connections, credentials, local transcription, and recovery.
Do I need to connect a cloud speech provider?
No. Provider connections are optional. Connect one only when you want its cloud Text to Speech voices and capabilities. Voxnara hides disconnected providers and unsupported options from the engine controls.
How are provider credentials handled?
Provider credentials are managed by the desktop application through operating-system protected storage. The interface keeps credential values out of project files and displays only the connected profile and supported capabilities.
Is Speech to Text audio uploaded to a cloud service?
Voxnara Speech to Text uses local recognition. If a required local model or locale is unavailable, recognition stops with an error instead of silently sending the recording to a cloud service.
What happens if recording or processing is interrupted?
Committed transcript segments, completed recording chunks, and durable jobs are recoverable. Reopen the project or Queue to continue from saved work, retry a failed job, or remove work you no longer need.
No matching answers
Try a shorter phrase, choose another topic, or ask the support team.
Still need a hand?
Choose the next step.
Read the full workflow guide or contact support with the page, action, and error message you see.