Building Jarvis: A Local Voice Assistant with Android and Hermes

Jarvis is a voice assistant built around a stock Android phone, a Hermes agent, local speech models, and Homebridge. The phone handles wake detection, recording, playback, and the interface. A separate service host handles transcription, orchestration, and speech synthesis, while a GPU inference server supplies the conversation model.

The design also includes a background development queue: an explicit improvement request can start an isolated coding session, followed by tests, a signed Android build, and wireless installation.

This post describes the implementation snapshot recorded on October 6, 2026. It is based on deployment notes rather than a new end-to-end test. Private addresses, account names, device identifiers, credential locations, and household inventory have been omitted.

A story request in action

Jarvis responds to a request to tell a Peppa Pig story.

Architecture

Jarvis architecture: an Android client communicates with an authenticated voice bridge, which connects local transcription, Hermes orchestration, and speech synthesis. Hermes uses Qwen inference and Homebridge tools. A separate development queue leads to tested, signed updates.

The phone talks to one authenticated bridge. The Hermes gateway and speech worker are loopback-only services behind it; the phone does not need direct access to either.

Component Responsibility
Android client Wake word, voice activity detection, audio capture, interface, playback
Voice bridge Authentication, transcription, request routing, activity events, speech proxy
Hermes conversation profile Sessions, context, memory, model requests, tool execution
Qwen inference server Local conversation inference and structured tool decisions
Persistent speech worker MLX speech models and serialized synthesis
Homebridge connector Device discovery, fresh state, validated writes and readback
Development controller Persistent jobs, worktrees, tests, signed builds, wireless updates

The Qwen deployment is described separately in Deploying Qwen on Intel Arc with vLLM and DFlash2.

One conversation turn

The phone detects “Hey Jarvis” locally, with a button as an alternative entry point. It records a complete utterance as 16 kHz mono PCM16, using local voice activity detection to determine when speech ends. Voice requests are limited to 30 seconds.

The recording goes to /voice, where faster-whisper's small model transcribes it on the service host's CPU. Typed requests enter /text and skip transcription. Ordinary requests then go through Hermes to Qwen and any required tools.

The bridge returns activity events and a final reply. With Accept: application/x-ndjson, the client receives newline-delimited events followed by the result; ordinary JSON remains supported. Activity cues summarize tool work without exposing raw tool arguments or internal reasoning.

Once the reply is complete, the phone splits it into speech segments, requests /speech, and plays the resulting WAV files in order. It downloads one next segment while the current segment plays. After playback, a 30-second follow-up window accepts another utterance without repeating the wake word.

This is a turn-based pipeline. Recording and playback are coordinated to prevent the assistant hearing itself. Full-duplex interruption and speech synthesis from unfinished model tokens are not implemented.

All phone-facing routes require a pairing bearer token. The bridge serializes conversation processing and reports busy requests rather than overlapping turns. A conversation identifier connects requests to server-side history.

Android state ownership

The Android implementation uses native Java and runs without root or a custom ROM. Its interaction logic lives in a pure Java state machine owned by a controller thread. The interface observes immutable snapshots.

The main voice path moves through starting, listening, capturing, processing, speaking, cooldown, and follow-up states. Typed input and recovery have their own paths.

Every transition advances an operation ticket. Audio, HTTP, and speech callbacks must match both the current ticket and expected state. A delayed callback from a cancelled request cannot restart obsolete playback or revive an old recording.

Execution context Ownership
UI thread Controls, transcript display, service snapshots
Controller thread Transitions, operation tickets, timers, recovery, microphone policy
Microphone reader Blocking audio reads and recorder release
Request executor Voice/text HTTP requests and activity events
Speech executor Ordered playback and cancellation
Prefetch executor One next speech segment

The microphone produces 20 ms frames into a bounded queue of 100 frames. Stale frames are rejected, and overflow is reported rather than silently accumulating latency. Delivery epochs discard frames crossing capture-gate changes. Audio collected during processing or playback is suppressed.

Wake detection uses an openWakeWord ONNX pipeline of mel features, embeddings, and a wake classifier. Silero v4 VAD provides speech detection and endpointing. Foreground service behavior, a partial wake lock, and a Wi-Fi lock help maintain listening, while recovery timers handle some failures. Process termination and network outages remain possible.

Conversation context and model routing

Hermes uses separate profiles for conversation and development. The conversation profile calls local Qwen with thinking disabled and a 2,048-token generation cap. Its configured context is 65,536 tokens, with a compression threshold cap of 20,000 tokens.

The conversation profile omits unrelated tool families and the full installed skills index. An earlier short-request comparison reduced input from roughly 17,206 to 8,707 tokens before the home-control addition. That is a historical prompt-size observation, not a fixed current cost or latency guarantee.

The voice run budget is 60 seconds with one API retry. It is cooperative, so it does not impose a hard deadline on every tool or network call. The phone and adapter also handle timeouts.

Development uses a distinct process, session, profile, and Git worktree for each job, with a cloud coding model. Separate profiles prevent ordinary conversation history and coding context from being mixed. They do not provide OS-level isolation: processes still share the host's user permissions and resources. Coding context sent to the cloud leaves the local network; local speech and Qwen inference stay local, while web tools can contact external services.

Speech synthesis that survives long replies

The persistent speech worker uses Qwen3-TTS through MLX. The recorded setup uses an 8-bit 0.6B CustomVoice model for default Chinese and English voices, with an 8-bit 1.7B model loaded for styled voices when needed. Synthesis is serialized to protect model and decoder state.

Each reply captures a voice-settings snapshot. All prefetched segments use that snapshot, so a setting change cannot switch voices halfway through a reply.

Long replies exposed stretching and truncation failures. The implementation addresses them with several controls:

  • Client and server segmentation share a 140-weighted-unit budget, counting CJK characters as four units. Splits prefer sentence and punctuation boundaries while preserving decimals and common abbreviations.
  • Spoken-text cleanup removes display markup without rewriting the visible answer.
  • Each generation receives complete phrase text, with acoustic sampling configured separately from text/code prediction.
  • Decoder streaming state is reset before and after independent generations, including failures.
  • Duration, finite-sample, token-limit, and early-termination checks reject clearly malformed audio. Smaller phrases are retried within bounded limits.
  • The phone validates WAV format and sample rate and monitors actual playback progress.

These checks improve stability but cannot detect every pronunciation error. Prefetching reduces gaps between complete segments; it does not make this token-level audio streaming.

Home control with explicit confirmation

The Homebridge connector is installed only in the conversation profile and authenticates with a dedicated non-admin account. Credentials stay outside Git and tool results, and authentication tokens remain in memory.

Its tools separate discovery, fresh-state reads, and writes. Discovery exposes supported settings without presenting cached values as current state. Reads report when refresh cannot be confirmed. Writes validate an explicit supported change and then perform a separate state readback.

Only matching readback produces confirmed=true. A write timeout is not blindly retried because the device may already have received the command. Ambiguous device names require clarification. The enabled write surface covers lights, switches, outlets, and fans; door, lock, alarm, and camera control are excluded.

An empty inventory can trigger a bounded discovery refresh with a cooldown. That refresh uses the Homebridge Accessories discovery channel without toggling devices.

The deployment notes record end-to-end discovery and fresh-state reads, plus thirteen passing connector tests. Physical device writes were not exercised during setup, so those checks do not establish completed physical-control testing.

Requested improvements and wireless updates

An explicit app-improvement request enters a persistent queue. The assistant acknowledges queuing promptly, and the user can continue talking while coding runs. Acknowledgement means the request was queued; it does not mean an update has been installed.

The controller serializes jobs with a process lock and requires a clean main checkout. Each job starts from the current revision in a separate branch and worktree. The coding agent edits there, while the controller owns signing, commits, and deployment.

Before advancing a job, the controller checks permitted changes, runs server and Android host tests, builds with the existing signing key, and verifies package identity and signing. It advances the main branch only if the original base still matches. A stored APK digest is checked again before installation.

Wireless deployment verifies the intended phone, waits for an allowed idle state, installs an in-place signed update, refuses downgrades, and checks the installed version and running process. The notes record completed jobs through this path. That verifies the tested build and installation, not every possible interaction.

Selected server files have a deployment allowlist, backups, restart checks, and restoration after failed health checks. This is not an atomic rollback covering the APK, Git history, and every service. The idle-state check also has a race if a user begins a turn just before installation.

On the older Android device used for this project, Wi-Fi ADB may require USB re-enablement after reboot. Voice service access and the update channel are independent, so conversation can work while an update waits for the phone.

Interface, history, and operating limits

An animated avatar uses OpenGL ES with a Canvas activity overlay. Visual transitions blend over 800 ms, while capture and playback amplitude drive an audio flame. Tool activity gets brief visual cues for research, calculation, or coding, with listening, speaking, and stopping taking priority. The coding cue is an activity summary rather than a live background-job monitor.

The transcript screen supports selectable message bubbles and a native Markdown subset. User text stays literal. Visible history is currently held in memory and disappears after process death, while the persisted conversation ID can still refer to server-side history. Starting a new conversation clears the local transcript and chooses a new identity without deleting earlier server sessions.

Server-side conversations retain transcribed text and tool results. Coding jobs retain requests, progress, logs, and changes. Ordinary speech processing does not retain payloads as speech logs, though explicit diagnostic artifacts are separate.

The recorded deployment uses authenticated HTTP on a trusted local network. Bearer authentication does not encrypt transport, and remote access was not deployed in this snapshot. Other limits include shared compute during builds and speech processing, variable model and tool latency, imperfect wake detection, and short interruptions during updates.

The most useful design choice has been making ownership explicit: the controller owns interaction state, Hermes owns conversations and tools, the speech worker owns synthesis state, and the release controller owns installation. That makes delayed callbacks, uncertain device writes, malformed audio, and partially completed updates easier to handle and inspect.