Skip to content

push-to-talk viability at the desk #339

Description

@ojfbot

Question

Is browser-native voice good enough for a hands-free grill at the desk — push-to-talk or continuous, and does the agent speak back?

Type

prototype (HITL, /prototype)

Context

Operator ruling D1: voice-capable at the desk, not in the car. The cockpit stays bound to 127.0.0.1.

Prototype the Web Speech API in the cockpit's browser:

  • push-to-talk vs continuous listening
  • TTS for the agent's turn — and whether it is wanted at all
  • barge-in (interrupting a spoken answer)
  • reduced-motion and a11y; the rail already honours reduced-motion elsewhere

Blocked deliberately on the cadence prototype: voice pushes hard toward one-question-at-a-time and would prejudge that decision if run first.

Note the fallback if voice proves unworkable here: core's queued chat-side skill northstar-voice (rm:rm-l2-ojfbot#S3) was specced for exactly this and never built.

Blocked by

'does a one-thread grill survive a 372px rail' — modality must not prejudge cadence.

Resolution (filled at close)

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions