Voxnara

Preparing your workspace

Built for every step
between script and transcript.

Voxnara brings text, voice, and transcription into one desktop workspace—so you can create, refine, and export without breaking your flow.

Windows macOS Version 1.0.0

Voxnara workspace
Voxnara dashboard with recent projects and job queue
Explore the workspace
01

Text to Speech

Direct every word.

Set the voice and delivery for the whole document, then give selected words or passages their own voice, pacing, emphasis, or effect.

Voxnara Text to Speech editor and document engine
Write, direct, preview, and export in one canvas
  1. 01

    Voice & delivery

    Choose from connected or installed voices, then set document-wide rate, pitch, loudness, style, and exact pauses between words when supported.

  2. 02

    Selection controls

    Highlight any passage to override its voice and delivery. Mix multiple voices deliberately without changing the rest of the document.

  3. 03

    Pronunciation

    Build reusable rules for names, abbreviations, dates, specialist terms, and custom phrase replacements.

  4. 04

    Export formats

    Export WAV, audio-only MP4, OGG, or WebM. Voxnara keeps a WAV master and converts locally when the selected engine does not provide the destination format.

Export audio

WAVMP4OGGWebM

Real containers. Validated codecs. No renamed files.

02

Sound Studio

Shape selected passages with local audio processing. Effects are visible on the canvas, saved with the project, and combined only where they are compatible.

  1. Soft voiceGentler presence
  2. BreathingNatural air texture
  3. WhisperUnvoiced character
  4. DRCControlled dynamics
  5. TimbreWarm to bright
  6. EmphasisReduced to strong
03

Speech to Text

Capture speech while it happens.

Turn a microphone, supported desktop output, or imported audio into an editable transcript canvas using local recognition.

  1. 01Capture or importMicrophone · Desktop audio · Audio file
  2. 02Watch live textInterim speech appears before committed segments
  3. 03Edit and exportRefine the transcript, then choose the format
01

Local recognition

STT audio stays on the device. A missing local model or unsupported locale fails closed instead of silently switching to a cloud service.

02

GPU-first processing

Compatible Whisper models prefer GPU acceleration and automatically fall back to bounded CPU processing when needed.

03

Speaker labels

Select an expected count from 1 to 10. Local diarization applies generic speaker labels after capture while preserving the transcript if labeling fails.

04

Crash-safe sessions

Completed recording chunks and committed segments remain recoverable after interruption, device loss, or restart.

05

Editable transcript canvas

Work with timestamps, notes, bookmarks, speaker names, find and replace, revisions, and playback-linked segments.

06

Transcript exports

Create clean TXT or Markdown, structured JSON, and timestamped SRT or WebVTT files.

Voxnara local speech recognition model manager
Choose the local model that fits your device
04

One workspace
for every project.

Create, autosave, search, and return to TTS documents or STT sessions. The queue keeps long-running work visible and recoverable.

  • ProjectsKeep text, transcript segments, engine settings, effects, and export details together.
  • QueueTrack progress, cancel work, retry failures, and continue without losing completed items.
  • Provider-awareOnly connected engines, languages, voices, styles, and controls appear.
Voxnara project and queue dashboard
Projects and jobs stay visible together

Connect only what you choose.

Optional cloud Text to Speech uses your own provider credentials. Voxnara lists only the capabilities reported by each connected service.

  • Microsoft Azure
  • Amazon Polly
  • Google Cloud TTS
  • IBM Watson
  • OpenAI

Put the whole workflow
in one place.

Create voice, capture speech, and keep the project ready for the next edit.