Mocchi — a multi-agent chat client for your Hermes workspace

Product Requirements Document & Implementation Plan · Web (PWA) + iOS

Version 1.0 · Prepared 1 Oct 2026 · Author: Mocchi (Hermes Agent) · Reference projects: OpenMausBot (Apache-2.0), Vellum Assistant (MIT) · Backend: Hermes Agent v0.21.3 (NousResearch/hermes-agent)

Thesis. Do not rebuild an agent. You already are one — with skills, memory, cron jobs, browser control, a terminal, ~20 messaging adapters and a workspace full of real projects. What you're missing is a front door: a messaging-app-shaped surface where each agent is a "contact", each project is a "channel", and your whole workspace is reachable from your phone. This PRD builds that front door as a thin client over Hermes' existing JSON-RPC gateway.

1Problem & opportunity

1.1 What exists today

Your Hermes install is a fully-featured agent with a fragmented surface area:

1.2 The gap

There is no place where you can see "these are my agents, these are the threads, this is what ran overnight." Every capability you have is one CLI flag away, and none of it is legible at a glance or reachable from a phone.

1.3 Two reference projects, two lessons

ProjectWhat it provesWhat we take / skip
OpenMausBot
Apache-2.0 · 3.9k★ · TS/Kotlin/Swift · v0.1.91
A real open-source Grok-Bot analogue. Every sidebar entry is a distinct agent with its own personality, model, memory thread and computer. Approval cards for risky actions. Channels per context. Team packages as portable Markdown. BYO engines via ACP and OpenAI-compatible endpoints. iOS + Android apps. Harness server on 127.0.0.1. Take: the information architecture (agents-as-contacts, channels, approval cards, roster metadata), the ACP integration insight, and the licensing model to study.
Skip: re-implementing an agent runtime, Composio, the VM/desktop sandboxing, and the Electron shell.
Vellum Assistant
MIT · 1.4k★ · iOS/Web/desktop
Multi-assistant support ("vellum ps — view running assistants", commands take an assistant ID), device pairing over a tunnel, an SSE event stream as the documented API surface, 8 memory types, SOUL.md identity files, proactivity loop, sandbox + permission tiers. Take: the multi-assistant account model, pairing UX, event-stream-first API design, permission tiers.
Skip: their runtime entirely — notably, Vellum's own README lists Hermes Agent as one of the things it replaces, confirming the space is real.

Verified from the repositories' READMEs and directory listings on 1 Oct 2026. No claim above is inferred from marketing copy alone.

2Product definition

2.1 One-line

A messaging app where every contact is one of your Hermes agents, every channel is a project workspace, and your entire agent fleet is one tap away on your phone.

2.2 Principles

Agent, not app
The client owns zero agent logic. All intelligence, memory and tools stay in Hermes, so the product never goes stale when the agent improves.
Profiles = agents
Multi-agent support is a solved problem upstream. A Hermes profile already is an isolated agent. The client is a profile switcher with a nicer face.
Upgrade-proof
The wire contract is generated and validated upstream. We consume it; we never fork the server.
Workspace-first
The thing that makes this yours is the fleet: skills, memory, cron, browser, files, Telegram — all surfaced, none rebuilt.

2.3 Personas → features

PersonaNeedFeature
Marc, on the couchCheck what ran overnight, steer itActivity feed, push notifications, resume-thread
Marc, in a meetingFire off a task, approve a risky actionApproval cards inline in chat, one-tap Allow/Deny
Marc, at his deskWork a project with a focused agentChannels bound to a profile + working folder + shared instructions
Marc, planningSee the whole fleetAgent roster: model, tools, skills, memory, cost, last-seen

3Architecture — how we build on Hermes without forking it

This is the load-bearing section. The entire plan rests on one verified finding: Hermes already exposes a complete, typed, documented agent protocol over WebSocket, and the same protocol the Electron desktop app uses. We are a third client of that protocol.

3.1 The two integration seams

Seam A — the JSON-RPC gateway (primary)

ui-tui (Ink)  ──stdio JSON-RPC──┐
apps/desktop (Electron) ──WS────┼──►  tui_gateway/server.py  ──►  AIAgent + tools + sessions
Mocchi (our client)      ──WS────┘
                                ws://<host>/api/ws

Seam B — REST (secondary, for non-chat surfaces)

24 routers mounted by hermes_cli/web_server.py, ~200 endpoints, same session token:

PrefixWhat our UI consumes it for
/api/profiles/*Agent roster: create, describe, model, SOUL, export, desktop-overlay
/api/skills, /api/skills/content, /api/skills/toggleSkill library browser + per-agent toggles
/api/cron/jobs, /api/cron/blueprints, /api/cron/fireAutomations screen; run-now; blueprints
/api/eventsActivity feed / system event stream
/api/files/*, /api/fs/*File browser + viewer + upload from phone
/api/tools/toolsets/*, /api/tools/terminal/backendPer-agent tools & terminal backend config
/api/model/*, /api/providers/*Model switcher, provider/OAuth onboarding
/api/audio/* (speak, transcribe, voice-live)Talk mode — see §5.3
/api/pairing/*Device pairing for the phone
/api/ops/doctor, /api/status, /api/system/statsHealth & usage

Seam C — OpenAI-compatible API (for anything non-Hermes)

gateway/platforms/api_server.py serves /v1/chat/completions, /v1/responses, /v1/models, /api/sessions, /api/runs, /api/jobs behind API_SERVER_KEY — and under gateway.multiplex_profiles, secondary profiles live at /p/<profile>/…. That means a single key + URL prefix can address any agent in the fleet from any SDK.

3.2 Auth and transport

3.3 Repo layout

openbot/
├── packages/
│   ├── protocol/        # generated OpenAPI/TS types from the Hermes contract (npm script)
│   ├── client/          # transport-agnostic agent client (WS + reconnect + replay + tickets)
│   └── ui/              # shared design system: roster, thread, approval card, channel
├── apps/
│   ├── web/             # React + Vite PWA  (served at mocchi.yourdomain)
│   └── ios/             # SwiftUI, shares the OpenAPI contract
└── plugins/
    └── mocchi-bridge/   # out-of-tree Hermes plugin: device tokens + optional push

Monorepo, pnpm. One protocol package, two clients — the web app and the iOS app can never disagree about the wire.

4Feature specification

4.1 F1 — Roster (the core screen)

Every profile is an "agent". Data from /api/profiles + agents.list + session.most_recent.

FieldSource
Name, avatar, accent colour, emoji/api/profiles/{name}/description, SOUL.md
Model + provider/api/profiles/{name}/model
Last active, unread, pinnedsession.most_recent + local prefs
Unread badgeDerived from events since last view
Tools/skills chipsagents.list, /api/skills

acceptance Cold start shows the fleet in <1 s from a warm cache, with a skeleton while /api/profiles resolves. Pull-to-refresh. New-agent FAB deep-links to profile creation on the desktop dashboard for v1 (creating a profile from mobile is P2).

4.2 F2 — Thread (chat)

Styled messages (not a raw terminal): markdown, code fences, collapsible tool-call cards, streaming token rendering, and a live activity strip (tool name, elapsed, running/done/failed).

acceptance Reconnect mid-turn: on socket drop and reconnect, session.resume + session.events.since restores the full transcript and any in-flight turn, with no duplicated or lost text.

4.3 F3 — Approval cards highest leverage

The single most valuable feature, and the cheapest: the protocol already pushes approval / sudo / secret / clarify / vault.* requests to the client and blocks the agent thread until answered.

acceptance A tool asks for approval while the app is backgrounded → push arrives → tapping it opens the thread → one tap answers → the agent unblocks.

4.4 F4 — Channels

A channel = profile + working folder + shared instructions + a thread set. Backed by Hermes chat workspaces (/api/chat/workspaces) and session.workspace.move. v1 ships Work / Personal / one channel per real project; the roster inside a channel is a set of profiles.

4.5 F5 — Fleet inspector

Per agent: model, active toolsets (/api/tools/toolsets), skills on/off, memory provider, cron jobs (/api/cron/jobs), token spend (session.usage, /api/analytics/usage), and a "run this agent now" button. Read-mostly in v1; mutations are limited to model switch and skill toggle, both of which are already safe endpoints.

4.6 F6 — Automations

Cron as a first-class screen: list, pause/resume, trigger now (/api/cron/fire), delivery targets (/api/cron/delivery-targets), and blueprint instantiation. Combined with F3, this is the "it worked while I slept" loop.

4.7 F7 — Files from the phone

/api/fs/list + /api/fs/read-text + /api/files/upload + /api/git/status. Read a diff on the train, push a fix. Deliberately scoped: review, not a full IDE.

4.8 F8 — Voice (P2)

/api/audio/transcribe for input and /api/audio/speak / speak-stream for playback of replies. /api/audio/voice-live/session may support a live mode later.

4.9 Non-goals for v1

5Platform specifics

5.1 Web (PWA)

React 19 + Vite, installable, offline shell. Design language: dark-glass panels, subtle borders, pill-shaped bottom dock — matching the Mocchi bookmarks app you already have, so the fleet feels like the same product family.

5.2 iOS

StackSwiftUI + URLSessionWebSocketTask, iOS 17+. No third-party chat SDK.
ProtocolGenerated Swift models from gateway-contract.openrpc.json in the same protocol package pipeline.
Backgroundv1: foreground + push wake. BGAppRefreshTask for polling the activity feed. Live-streamed turns while foregrounded only.
DistributionTestFlight first. No App Store dependency; this is a private tool.

5.3 Auth & pairing (the honest hard part)

Browser WS auth uses a 30-second single-use ticket minted in-page. A native app has no browser page, so it needs a long-lived credential. Rather than touching Hermes core, add an out-of-tree plugin (~/.hermes/plugins/mocchi-bridge) exposing a narrow device-token endpoint — this is rung 4 on Hermes' own Footprint Ladder, and the plugin API is the documented, stable extension point.

Flow: desktop dashboard shows a QR with a short code → phone submits it once → device token stored in the iOS Keychain → all subsequent calls bearer-auth. Revoke via /api/pairing/revoke.

5.4 Push

APNs token registered by the device; a Hermes plugin subscribes to agent events and dispatches. Push is only sent for genuinely interrupt-worthy events: approval requests, completed long runs, and cron results — never for ordinary chatter, or you'll mute the app within a week.

6Implementation plan

Phase 0 — Foundation
Goal: prove the wire. ~3–5 days
  1. Run hermes update first (2106 commits behind) so we build against current contracts.
  2. Scaffold openbot/ pnpm monorepo; add protocol codegen emitting TS + Swift from the live contract file.
  3. Write a terminal harness: connect to ws://localhost:8080/api/ws with a minted ticket, call gateway.ready → session.create → send → stream. If this script works, everything after it is UI.
  4. Commit the harness as an integration test — it's the project's regression net against upstream churn.
Phase 1 — Web MVP
Goal: usable on a laptop. ~2 weeks
  1. packages/client: WS transport, ticket auth, reconnect w/ backoff, session.resume + replay, capability handshake.
  2. Roster screen, thread screen, streaming renderer, tool-call cards.
  3. Approval cards — validate the whole thesis here.
  4. Slash-command composer with complete.slash.
  5. Caddy vhost + TLS; installable PWA.

Exit criterion: from your phone's browser you can message any profile, watch it work, and approve its actions.

Phase 2 — Fleet surfaces
Goal: the workspace, surfaced. ~2 weeks
  1. Channels (chat workspaces), Fleet inspector, Automations/cron screen.
  2. Files browser + upload.
  3. Model switcher, skill toggles.
  4. Full reconnect resilience pass — airplane-mode test.
Phase 3 — Native iOS + push
Goal: the real product. ~3 weeks
  1. SwiftUI client against the same generated protocol.
  2. mocchi-bridge plugin: device tokens, pairing, APNs dispatch.
  3. Push for approvals / long-run completions / cron results.
  4. TestFlight build.
Phase 4 — Polish
~1–2 weeks
  1. Voice (F8), offline transcript cache, per-agent theming.
  2. Optional: agent-to-agent relay surfaced as "teammate" cards (bot_relay.roster / bot_relay.deliver already exist upstream).
  3. Agent packages as portable Markdown — the OpenMausBot team-file idea, backed by hermes profile create --clone.

6.1 Total effort & sequencing

PhaseDurationDepends onValue delivered
0 — Foundation3–5 dayshermes updateWire proven; project scaffolded
1 — Web MVP~2 wksPhase 0The whole product, in a browser
2 — Fleet~2 wksPhase 1Workspace visibility, cron, files
3 — iOS~3 wksPhase 2 + Apple dev accountReal mobile product with push
4 — Polish~1–2 wksPhase 3Voice, teams, theming

Strongly recommended: stop after Phase 1 and use it for two weeks before committing to iOS. Phase 1 delivers ~80% of the value at ~35% of the cost, and the feedback will make Phases 2–3 much sharper. Committing to native before the web client has proven the IA is the classic mistake.

7Risks & mitigations

#RiskSeverity / mitigation
R1Upstream churn. 251 methods and a contract that moves; you're 2106 commits behind today, so this accelerates. High Consume the generated contract, never hand-write types. Keep a Phase-0 integration test that fails loudly on drift. Rebuild contracts as a scheduled job; our breaking-change surface is small because we use ~40 of 251 methods.
R2Native auth. No browser = no ticket mint. Med Out-of-tree plugin for device tokens. Worst case, fall back to the web client in Safari with add-to-home-screen — which is a perfectly good v1 and is exactly what Phase 1 delivers.
R3Scope creep into a second agent runtime. Med Hard non-goal (§4.9). Any PR that reimplements tools, memory or skills is rejected in review. We are a client, full stop.
R4Security exposure. Your terminal, browser and secrets are one tap away. High HTTPS only. Device tokens revocable and rotatable. Every approval card shows the exact command/path. Consider a separate read-mostly profile for phone access — profiles already isolate memory and toolsets, so a "viewer" agent is free.
R5Licensing. OpenMausBot is Apache-2.0 with a source-available enterprise/ carve-out; Vellum is MIT. Low We copy ideas, not code. If any code is ported, Apache-2.0 requires attribution + NOTICE. Hermes itself is the actual dependency — license review happens before any port. This is a feature, not a bug: not forking Hermes is what keeps us safe.
R6Apple distribution. Low TestFlight is sufficient for a private tool. No store dependency. Web/PWA remains the primary surface if this ever bites.
R7Notification fatigue. Low Interrupt-only policy (§5.4), enforced in the bridge plugin with an explicit allowlist.

8Success criteria

9Open questions

  1. Profile-per-role or session-per-role? Profiles are true agents with isolated memory, but they are heavy to create from a phone. Recommend: profiles for genuine roles (research, KDP, ops), sessions for variations within one.
  2. Should the phone get its own restricted profile? Recommendation: yes, as a safety valve — an "operator" profile with a reduced toolset, sharing your memory read-only.
  3. iOS or Android first? iOS assumed. OpenMausBot ships both; you use Telegram on iOS.
  4. Do we want team/agent-relay surfaces (Phase 4)? Upstream bot_relay.roster/deliver already support agent-to-agent messaging; surfacing it is cheap but adds product surface — defer until the core is used daily.
  5. Public or private? Recommendation: private, MIT or Apache-2.0, single-user. Multi-tenant auth is a different and much larger product.

Sources. Hermes integration surfaces verified by direct inspection of the installed source at /home/ubuntu/.hermes/hermes-agent (v0.21.3, commit 4d14aaf) on 1 Oct 2026: tui_gateway/contracts/, apps/shared/src/gateway-contract.openrpc.json, hermes_cli/web_routers/, hermes_cli/web_server*.py, gateway/platforms/api_server.py, tui_gateway/AGENTS.md. Reference projects read from their public repositories on 1 Oct 2026. No unverified claim in this document is presented as fact.