Hermes Agent
A Python package that plugs AverVOX into Hermes through its native speech-provider interface.
Private, system-wide speech for Linux
AverVOX is a local speech layer for Linux. Speak into any focused application, read selected text aloud, hold hands-free conversations with OpenAI-compatible models, and give agents the same private STT/TTS engine.
Turn rough meeting notes into a clear follow-up email, preserving the action items and deadlines.
Choose your workflow
Start with desktop voice, conversational LLMs, or agent infrastructure. The same installation can handle all three.
Press one hotkey, speak naturally, and insert text into the application that already has focus—no browser extension or per-app plugin.
Explore Linux dictationHold streaming spoken conversations with Ollama, LM Studio, vLLM, LocalAI, remote providers, or another OpenAI-compatible endpoint.
Explore LLM voiceLet Hermes, OpenClaw, scripts, and other tools borrow AverVOX through its bridge CLI and warm local daemon.
Explore integrationsPrivate by architecture
Speech recognition, synthesis, voice activity detection, and desktop text insertion run locally. Conversation text goes only to the LLM endpoint you choose.
AverVOX collects no usage analytics, crash reports, or product telemetry. A remote LLM receives conversation text only when you configure one.
Measured through real integrations
The warm daemon keeps speech models loaded, avoiding process startup and model-load cost on every agent response.
Agent-ready
Install host-native packages for Hermes and OpenClaw, or call the bridge CLI directly from your own tool.
A Python package that plugs AverVOX into Hermes through its native speech-provider interface.
A native plugin published for normal OpenClaw discovery and installation, with automatic daemon fallback.
Synthesize, transcribe, discover capabilities, cancel work, and stream PCM without embedding speech models.
Free core, focused upgrade
AverVOX OSS contains the complete core workflow. Pro adds higher-quality voice, hands-free activation, persistence, LAN sharing, and management tools.
| Capability | OSS | Pro |
|---|---|---|
| System-wide dictation, selected-text TTS, and Converse | ✓ | ✓ |
| Faster Whisper, Piper, bridge CLI, warm daemon, streaming, barge-in | ✓ | ✓ |
| Kokoro voices and playback-speed control | — | ✓ |
| Wake word, persistent sessions, and per-profile personas | — | ✓ |
| LAN voice server, dashboard, transcripts, and dictation logs | — | ✓ |
| Price | Free · MIT | $39 one time |
Buy with confidence
The OSS edition is the compatibility check: install it, verify your desktop and audio path, connect your endpoint, and make sure voice fits your workflow.
Use the guided installer or the published Python package.
Dictation, selected-text speech, and Converse use the same underlying desktop and audio integrations as Pro.
X11 and XWayland are supported. Pure Wayland remains best-effort because global hotkeys and text insertion vary by compositor.
Questions before you install
Start with the workflow
Install the MIT-licensed core today. Move to Pro when you want Kokoro, wake-word activation, persistent memory, LAN sharing, and local history.