[APP][Pro] AI Voice Assistant

oke here we go, the hardware device seems to be in a “Beam lock released” loop due to a “Home Assistant event ‘esphome.tts_uri’ dropped; client has not subscribed to actions (yet)”

app logging:

2026-08-06T18:32:01.574Z [log] [APIHELPER] - ApiHelper initialized
2026-08-06T18:32:01.662Z [log] [APIHELPER] - API devices connected
2026-08-06T18:32:01.810Z [log] [APP] - AI voice assistant initialized successfully
2026-08-06T18:32:02.430Z [log] [CONVO][MIC] - Settings saved — context cleared, next turn starts fresh
2026-08-06T18:32:26.341Z [log] [CONVO][MIC] - Turn started (wake word / button / flow), listening…
2026-08-06T18:32:26.374Z [log] [CONVO][MIC] - User started speaking (local VAD)
2026-08-06T18:32:28.561Z [log] [CONVO][MIC] - User stopped speaking (local VAD) — mic closed
2026-08-06T18:32:29.489Z [log] [CONVO][STT] - Heard: "Wat voor weer is vandaag?"
2026-08-06T18:32:33.431Z [log] [CONVO][TTS] - Speaking reply (announce)
2026-08-06T18:32:47.704Z [log] [CONVO][END] - Turn complete — conversation closed
2026-08-06T18:32:48.854Z [log] [CONVO][TTS] - Speaking reply (announce)

esp32 logging:

[20:31:08.294][I][respeaker_xvf3800:528]: Version request successful: 1.0.7
[20:31:58.293][W][api.connection:2532]: ai-voice-assistant (192.168.1.107): Reading failed CONNECTION_CLOSED errno=128
[20:31:58.293][W][component:342]: api set Warning flag: waiting for client connection
[20:32:02.470][W][component:365]: api cleared Warning flag
[20:32:02.475][W][api.connection:1751]: 'ai-voice-assistant' using outdated API 1.6, update to 1.14+
[20:32:02.638][W][micro_wake_word:365]: Wake word detection is already running
[20:32:26.032][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[20:32:26.922][I][esp-idf:000][md_decoder]: D (813246) micro_decoder.decoder_source: Decode finished
[20:32:28.600][I][respeaker_xvf3800:483]: Beam lock released
[20:32:29.509][W][voice_assistant:794]: No text in STT_END event
[20:32:33.458][W][voice_assistant:851]: No text in TTS_START event
[20:32:33.460][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:33.461][I][respeaker_xvf3800:483]: Beam lock released
[20:32:35.469][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:35.470][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:35.472][I][respeaker_xvf3800:483]: Beam lock released
[20:32:37.479][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:37.481][I][respeaker_xvf3800:483]: Beam lock released
[20:32:38.465][I][esp-idf:000][md_reader]: E (824790) esp-tls: [sock=58] select() timeout
[20:32:38.467][I][esp-idf:000][md_reader]: E (824790) transport_base: Failed to open a new connection: 32774
[20:32:38.469][I][esp-idf:000][md_reader]: E (824790) HTTP_CLIENT: Connection failed, sock < 0
[20:32:38.470][I][esp-idf:000][md_reader]: E (824790) micro_decoder.http_client: Failed to open URL: ESP_ERR_HTTP_CONNECT
[20:32:38.473][I][esp-idf:000][md_reader]: E (824790) micro_decoder.audio_reader: Failed to connect to URL: http://172.17.0.2/app/no.arvebjoe.ai-voice-assistant/userdata/audio/tx_4cd9574e-4a59-4a96-9ad9-0b91947bd5c8.flac
[20:32:38.474][I][esp-idf:000][md_reader]: E (824791) micro_decoder.decoder_source: Reader failed to open URL
[20:32:39.496][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:39.498][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:39.499][I][respeaker_xvf3800:483]: Beam lock released
[20:32:41.532][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:41.533][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:41.535][I][respeaker_xvf3800:483]: Beam lock released
[20:32:43.540][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:43.541][I][respeaker_xvf3800:483]: Beam lock released
[20:32:45.610][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:45.612][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:45.613][I][respeaker_xvf3800:483]: Beam lock released
[20:32:47.709][W][voice_assistant:877]: No url in TTS_END event
[20:32:47.711][I][respeaker_xvf3800:483]: Beam lock released
[20:32:48.861][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:48.862][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:48.864][I][respeaker_xvf3800:483]: Beam lock released
[20:32:50.878][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:50.879][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:50.881][I][respeaker_xvf3800:483]: Beam lock released
[20:32:52.951][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:52.953][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:52.955][I][respeaker_xvf3800:483]: Beam lock released
[20:32:55.062][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:55.064][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:55.065][I][respeaker_xvf3800:483]: Beam lock released
[20:32:57.144][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)
[20:32:57.146][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[20:32:57.147][I][respeaker_xvf3800:483]: Beam lock released
[20:32:59.242][W][script:081]: Script 'activate_stop_word_once' is already running! (mode: single)

....here be a logging loop.....

[20:35:00.879][I][respeaker_xvf3800:483]: Beam lock released
[20:35:03.016][W][voice_assistant:877]: No url in TTS_END event
[20:35:03.018][I][respeaker_xvf3800:483]: Beam lock released
[20:35:08.297][I][respeaker_xvf3800:528]: Version request successful: 1.0.7

and my esp config is just the example one with a different device name Respeaker-XVF3800-ESPHome-integration/config/respeaker-xvf-satellite-example.yaml at main · formatBCE/Respeaker-XVF3800-ESPHome-integration · GitHub

oke scratch that, this all seems to be related to using a custom pipeline


oke next question
I now have a hook that outputs the text to my sonos device

But the default voice to sonos is not very nice, but I can also play a url on sonos, would it be possible to have the output available as a url?


and there seems to be a issue where if you ask something and then want to ask another thing the device doesnt react

[21:09:08.305][I][respeaker_xvf3800:528]: Version request successful: 1.0.7
[21:10:27.091][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[21:10:27.956][I][esp-idf:000][md_decoder]: D (3094281) micro_decoder.decoder_source: Decode finished
[21:10:30.509][I][respeaker_xvf3800:483]: Beam lock released
[21:10:30.588][W][voice_assistant:794]: No text in STT_END event
[21:10:33.045][W][voice_assistant:851]: No text in TTS_START event
[21:10:33.047][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[21:10:33.048][I][respeaker_xvf3800:483]: Beam lock released
[21:10:33.119][I][esp-idf:000][md_reader]: D (3099443) micro_decoder.http_client: Connected: status=200 content-type='audio/x-flac'
[21:10:33.120][I][esp-idf:000][md_reader]: I (3099444) micro_decoder.audio_reader: Streaming http://192.168.1.107/app/no.arvebjoe.ai-voice-assistant/userdata/audio/tx_79e63413-01bf-49a9-9783-ba8bdaf0ac1e.flac (FLAC)
[21:10:33.337][I][esp-idf:000][md_reader]: D (3099662) micro_decoder.audio_reader: HTTP read complete
[21:10:33.339][I][esp-idf:000][md_reader]: D (3099664) micro_decoder.decoder_source: Reader finished
[21:10:40.889][I][esp-idf:000][md_decoder]: D (3107215) micro_decoder.decoder_source: Decode finished
[21:10:40.933][W][voice_assistant:877]: No url in TTS_END event
[21:10:41.036][W][api:437]: Home Assistant event 'esphome.tts_uri' dropped; client has not subscribed to actions (yet)
[21:10:41.038][I][respeaker_xvf3800:483]: Beam lock released
[21:10:41.040][I][respeaker_xvf3800:483]: Beam lock released
[21:10:41.375][I][esp-idf:000][md_reader]: D (3107699) micro_decoder.http_client: Connected: status=200 content-type='audio/x-flac'
[21:10:41.378][I][esp-idf:000][md_reader]: I (3107700) micro_decoder.audio_reader: Streaming http://192.168.1.107/app/no.arvebjoe.ai-voice-assistant/userdata/audio/listening_chime.flac (FLAC)
[21:10:41.380][I][esp-idf:000][md_reader]: D (3107700) micro_decoder.audio_reader: HTTP read complete
[21:10:41.380][I][esp-idf:000][md_reader]: D (3107702) micro_decoder.decoder_source: Reader finished
[21:10:41.530][I][esp-idf:000][md_decoder]: D (3107848) micro_decoder.decoder_source: Decode finished
[21:10:45.550][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[21:10:47.657][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[21:10:48.575][I][esp-idf:000][md_decoder]: D (3114900) micro_decoder.decoder_source: Decode finished
[21:10:52.360][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[21:10:58.678][W][api:437]: Home Assistant event 'esphome.wake_word_detected' dropped; client has not subscribed to actions (yet)
[21:10:59.545][I][esp-idf:000][md_decoder]: D (3125870) micro_decoder.decoder_source: Decode finished

I´ve never tested this in HA, so can´t compare. But this app uses an AI large language model, and those are pretty capable out of the box, and you can also ask the AI to do a web search. I´ve had some good results with this so far.

Weather lookup is also implemented, but only for where you Homey is (at home), you can´t ask what the weather is like some other place.

You will need some kind of physical hardware, and i recommend the Home Assistant Voice: Preview Edition (PE) or the ThirdReality Voice & Music Assistant. You can read more about them here:

Thanks — glad you’re enjoying it!

A “Stop processing” card unfortunately can’t work. “Heard something” fires and the pipeline continues in the same breath, so there’s nothing left to brake; and on the cloud engines the model starts answering when you stop talking, not when the transcript lands — the reply is already on its way. Only the custom pipeline is sequential enough for a gate to even be possible.

But what you’re actually after now exists, from a different angle: the custom pipeline’s LLM backend can be set to None. The turn then stops after speech-to-text — what you said goes out on Heard something, and your Flow decides everything from there, including speaking an answer back with the Say card, which still uses your speech backend. No model in the middle, no processing to stop. Its sibling, TTS backend → None, does the mirror image: the assistant still answers but nothing is spoken, so a Flow can send the reply text to a Sonos or wherever you like.

Fair warning on the None LLM route: follow-up questions and every built-in skill (device control, weather, timers, shopping list, music) go with it — your Flow is the assistant now — plus the Flow round trip on top of the latency you already predicted.

Which is why, given your edit, I’d still point you the other way. Set the LLM backend to OpenAI-compatible and aim it at your bridge: your skill becomes the model, and you keep the LED ring, streaming speech that starts before the full answer exists, follow-ups, and my built-in tools alongside your own — one hop less latency than going through Flow. Both paths work; that one costs you nothing.

Both land in the next test build. Do report back on how the latency feels with a heavy skill behind it.

Thank you for the detailed reply. That opens so many possibilities!

Should the app cards be coming one day, I think this app will be the most poweful and versatile Homey AI integration out there! :+1:

New test version (v 1.4.10) is now available with the none-LLM and none-TTS implementation.

Hi Sre,

Your logs cracked it — thanks. Fixed in 1.4.11, plus the feature you asked for.

Test version:

Reply audio as a URL. Device settings now have a Reply audio dropdown — set it to “Send to a Flow as a URL” and a new trigger card, “Reply audio is ready”, fires with Audio URL, Text and Duration tokens. Feed the URL to your Sonos “Play URL” action and you get the assistant’s own voice instead of the Sonos TTS voice. The link expires after a couple of minutes, so play it immediately, and in this mode a follow-up needs the wake word again (the app can’t know when your Sonos finished speaking). The file is a 48 kHz mono FLAC — if Sonos refuses it, tell me and I’ll switch to WAV.

The Beam lock released loop wasn’t your pipeline. The app was advertising audio URLs on 172.17.0.2, a Docker-internal address your device couldn’t reach — hence the ESP_ERR_HTTP_CONNECT timeouts. The firmware then treats the failed announce as finished, the app queues the next segment, and round it goes. Intermittent because it depended on which network interface came up first at app start. It now takes the address from the connection your device makes to it, so it’s right by construction. Affected every device type, not just yours.

Also fixed: the client has not subscribed to actions spam, and No text in STT_END event (that one was already fixed — your build just predated it).

Still open, and I need you here. Your device going deaf to follow-up wake words I have not solved — detection fires, but the device never asks the app to start a turn. It may clear up as a side effect of the URL fix, since that was leaving it in a retry loop. Please check whether it still happens, and if it does, a DEBUG-level device log covering one good turn plus one ignored wake word would let me pin it down rather than guess.

Handy while testing: Settings → Debug now lists every voice device seen on your network and can record what the microphone actually captured, so you can tell a mic problem from a speech-recognition one. Off by default.

If you get a chance: how does the mic behave with Microphone gain at 0, what name does the device show when pairing, does the mute switch work, and does your unit have an API encryption key?

Thanks again.

I just created a PR that fixes two things:

Hi Arve,

Same symptom as reported a few days ago in this thread (Beam lock released loop / esphome.tts_uri event not subscribed).

Setup:

  • Device: Home Assistant Voice PE, firmware 26.6.0 (ESPHome 2026.6.0), no encryption configured
  • Homey Pro (Early 2023), app v1.4.11
  • BLE pairing + WiFi provisioning: successful
  • Network: confirmed reachable (port 6053 open, tested via SSH from another host on the same LAN)
  • mDNS discovery: works, Debug tab shows device as “usable voice satellite”, “paired”, probe “accessible”
  • Device tile in Homey: stuck on “Unavailable”
  • Debug tab detail: “Connected: no”

For comparison, I temporarily ran a local Home Assistant Core instance (Docker) and the same device connected and worked correctly via native ESPHome integration — confirms the device itself, WiFi, and network path are fine. The issue appears isolated to the Homey app’s connection/event-subscription handling for this device/firmware combo.

Happy to provide app logs or test a debug build if useful.

Thanks for the great app, looking forward to a fix!

v1.4.12 — spoken replies on a Sonos, and the app’s own sounds too

This one is mostly about the Reply audio is ready Flow trigger — the option that hands the assistant’s spoken answer to a Flow as an audio link instead of playing it on the device (device settings → Reply
audio → Send to a Flow as a URL). Handy when your voice device has no speaker, or when there’s simply a better speaker in the room.

What’s new:

  • The audio is now MP3 instead of FLAC, so far more speakers will actually play it — a Sonos in particular.
  • The app’s own sounds come through the same trigger now: a short cue when the device starts listening, plus the error sound, the “no API key” sound and the greeting after pairing. On a device without its
    own speaker those were simply silent before.
  • A new “Is a sound effect” tag on the trigger, so a Flow can tell a spoken reply from a cue — play the cue quietly, or ignore effects entirely with a condition card.
  • Fixed: the audio link could point at an address your speaker couldn’t reach, so nothing played.

Also fixed: on newer ESPHome firmware the volume and mute controls, and playing music on the device, could quietly stop working.

All of the reply-audio work above was written by Stefan Renne (@stefanrenne on GitHub) — a properly thorough pull request, tests and documentation included. Thank you! :folded_hands:

Hi Arve,

Following up on my earlier report (Voice PE stuck “Unavailable” despite successful pairing) — after 3+ hours of debugging with the help of Claude AI, we found the root cause.

The bug: In esp-voice-assistant-client.mts, discovery probes intentionally skip SubscribeVoiceAssistantRequest (to avoid stealing the subscription of an already-paired device — the comment in the code explains this well). However, recent firmwares (25.12.4, 26.4.0, 26.6.0 all tested) only reply to VoiceAssistantConfigurationRequest if the client is subscribed. Older firmwares answered without a subscription — which is why pairing works for users on older firmware and silently times out (8s) on recent ones.

How we found it: ran the app locally via homey app run, re-enabled the internal ESP client logger (it’s hardcoded disabled: createLogger('ESP', true)), and the logs showed the exact spot: TCP connect OK → Hello OK → DeviceInfoResponse OK → VoiceAssistantConfigurationRequest sent → total silence → timeout.

Confirmed fix: patching the probe to always subscribe makes pairing succeed instantly. A cleaner fix on your side might be to subscribe temporarily during the probe and unsubscribe right after, or to skip the VA-config check when the device already identified itself via DeviceInfoResponse (projectName already says “Home Assistant Voice PE”).

Happy to test a fix build. Thanks for the great app!

Setup: Voice PE fw 26.4.0, Homey Pro Early 2023, app v1.4.12 (patched from source v1.4.7 on GitHub — note the repo seems behind the store version).


FRANÇAIS

Bonjour Arve,

Suite à mon signalement précédent (Voice PE bloqué “Indisponible” malgré un pairing réussi) — après plus de 3 heures de débogage avec l’aide de Claude AI, nous avons trouvé la cause racine.

Le bug : dans esp-voice-assistant-client.mts, les probes de découverte sautent volontairement SubscribeVoiceAssistantRequest (pour ne pas voler l’abonnement d’un appareil déjà appairé — le commentaire dans le code l’explique bien). Cependant, les firmwares récents (25.12.4, 26.4.0, 26.6.0, tous testés) ne répondent à VoiceAssistantConfigurationRequest que si le client est abonné. Les anciens firmwares répondaient sans abonnement — c’est pourquoi le pairing fonctionne chez les utilisateurs en ancien firmware et expire silencieusement (8s) sur les récents.

Comment on l’a trouvé : appli lancée en local via homey app run, réactivation du logger interne du client ESP (désactivé en dur : createLogger('ESP', true)), et les logs ont montré le point exact : connexion TCP OK → Hello OK → DeviceInfoResponse OK → envoi de VoiceAssistantConfigurationRequest → silence total → timeout.

Correctif confirmé : en patchant le probe pour toujours s’abonner, le pairing réussit instantanément. Un correctif plus propre de votre côté pourrait être de s’abonner temporairement pendant le probe puis se désabonner juste après, ou de sauter la vérification VA-config quand l’appareil s’est déjà identifié via DeviceInfoResponse (le projectName dit déjà “Home Assistant Voice PE”).

Disponible pour tester un build correctif. Merci pour cette super appli !

Config : Voice PE fw 26.4.0, Homey Pro Early 2023, appli v1.4.12 (patchée depuis les sources v1.4.7 sur GitHub — à noter que le dépôt semble en retard sur la version du store).

v1.5.0 is live :tada: — it auto-updates, so you are probably on it already.

New

  • Claude (Anthropic) as a language model in the Custom pipeline, with the model picked from a live dropdown.
  • The cloud services for the pipeline’s speech and text stages (OpenAI, Groq, OpenRouter, DeepSeek…) are now a Server dropdown that fills in the address and a working model for you.
  • The Custom pipeline settings are reorganised into one card per stage.

Fixes

  • Adding a device could reboot your satellite. The pairing probe asked for something that crashes ESPHome 2025.8–2026.5, which looked like pairing timing out or a device stuck Unavailable. Fixed and verified on a real Voice PE across three firmwares.
  • An unavailable device now says which side is down — the device or the voice engine — and the Debug page shows the two separately. A missing API key used to look exactly like a network problem.
  • New Verbose logging switch in Settings → Debug, for when you need to send in a log.
  • Wi-Fi setup over Bluetooth notices a dropped link right away; deleting a device no longer crashes; and the Voice PE now trims the wake word off the start of a command by default.

Please re-check your issue :folded_hands: A lot has changed across 1.4.9 → 1.5.0 in just a few weeks, and I think some of you are sitting on something that has quietly been fixed. Update, try it again, and say either way — I would rather hear “still broken” than leave it untested.

If it is still broken: turn on Verbose logging, reproduce it, send a diagnostic report and post the id here, plus which voice engine you use and what Device connected / Engine connected say in Settings → Debug.

Thank you, Arve! Now it works. Sorry for the late replay! :slightly_smiling_face:

Ho Arve, Can you take a look at this. Set up everthing. No errors but it doesnot work.

87154194-92b4-4345-9a59-b714913be5c8