Speech recognition
- how do I control the stream with my voice
- can I start a poll or a prediction by voice
- the bot does not hear me
- voice commands do nothing
- how do I switch an OBS scene by voice
- the bot answers every phrase I say
- the bot keeps talking and I cannot stop it
What this page is
Voice control for your stream: you say what you want out loud, and the smart AI assistant does it. Starting predictions by voice, running a poll, switching an OBS scene, changing the title and category, cutting a clip, writing a line in chat, flipping a smart plug — all of it without letting go of the mouse mid-game.
Two screens set it up, and they answer different questions:
- Voice → Speech recognition — how the app hears you: microphone, language, when to listen, what you call the bot.
- Chat Bot → Voice control — what the bot is allowed to do once it heard you.
Recognition runs on your own machine, so the sound never leaves the computer. Only the request itself goes to the AI, and only after the app decides you were talking to the bot — and that request is the part you pay for.
How to set it up
1. Pick a recognition model
Voice → Speech recognition, the Recognition model block. Pick one and press Download — it is a one-off download, after which recognition works offline. The row for each model says how much memory it takes and how fast it runs on your graphics card and on your processor: pick the one whose numbers you can live with, not the biggest.
Until a model is downloaded nothing is heard, and the tile says Recognition is not ready — not silence, a missing model.
2. Microphone and language
The Microphone and language block. Find microphones shows their real names; without it the system default is used.
Set Language you speak. Leaving it on Detect automatically makes every phrase go through a second pass, so recognition is about twice as slow.
3. Choose when to listen
The When to listen block, and this is the choice that matters most.
- Hold a key — the microphone opens while the key is held and closes when you let go. Assign the key with Assign; it is not taken away from other programs, so it keeps working inside the game. Everything you say while holding goes to the bot, and no name is needed: the key itself says you are talking to it.
- Always — the microphone stays open for the whole stream. Now the app has to tell your own speech apart from a request, and that is what the bot's name is for.
4. Name the bot, if the microphone stays open
The How you call the bot block. Say the name and the rest of the sentence goes to the bot: "Nova, start a poll".
Leave the name empty and every phrase you say goes to the bot — and every one of them costs credits. The screen says so in red while the microphone is open and the name is missing. With Hold a key an empty name is fine: the key is the trigger.
5. Word that stops the voice over
Say it while the bot is talking and it goes quiet at once. The name is not needed here — by the time you have said it, the bot has finished.
Leave the field empty and the word for the language you speak is used: *stop*, *стоп*, *alto*, *pare*, *stopp*.
6. Allow the actions you want
Chat Bot → Voice control. Every action is a card with its price in credits and a switch. Anything that cannot be undone starts switched off.
There is no list of magic phrases. The bot works out what you meant, so "cut a clip", "clip that" and "save this moment" all reach the same action. The names it can use — OBS scenes, control panel buttons, smart plugs, prediction outcomes — come from what is actually connected, so it cannot invent a scene you do not have.
What it can do
| Action | Say something like |
|---|---|
| Start a poll | "start a poll: which map — Dust or Mirage" |
| Finish the poll | "finish the poll" |
| Start a prediction | "start a prediction: do we win this one, yes or no" |
| Close the betting | "close the betting" |
| Resolve and pay out | "we won, pay it out" |
| Cancel and refund | "cancel the prediction, give the points back" |
| Switch the OBS scene | "switch to the Break scene" |
| Change title and category | "set the title to Chill coding" |
| Save a clip | "clip that" |
| Write in chat | "tell chat the stream ends in ten minutes" |
| Press a control panel button | "press Ad break" |
| Switch a smart plug | "turn the lamp on for a minute" |
Twelve, and that is the whole list. Clips, the title, polls and predictions go through Twitch; on YouTube and Kick the assistant writes in chat, switches scenes and presses panel buttons. One phrase is one action: "lock the betting and switch the scene" does only the first of the two.
Anything that is not an action is an ordinary answer: ask what the weather is, how long the stream has been going, what you played last week, what your chat commands are, what you asked a minute ago.
Ask before anything that cannot be undone
Switched on by default. The bot names the outcome — "Resolving the prediction: 'We win'" — and waits for a "yes". It names the outcome and not the action on purpose: what you are confirming is the thing you cannot take back.
Switch it off and the bot just does it. Quicker, and there is no taking it back.
Where you see what happened
The Stream: Voice tile on the control panel keeps the feed: your phrases and the bot's answers, in order. The colour of the stripe on the left says how it ended — blue for an answer, green for something done, amber for a question back, red for a refusal with the reason next to it.
By default the bot says the short version out loud — "Poll started" — and the details go to that feed. If you have no second screen to read them on, switch on Read the whole result out loud on the Voice control screen.
Three buttons in the tile's header:
- microphone — switches between *hold a key* and *always*, the same setting as on the recognition screen;
- speaker — read answers out loud or leave them as text in the feed;
- power — the outermost one: the bot goes quiet, recognition stops and the model leaves memory. Scenarios triggered by donations and channel points are not affected — they never needed the microphone. Only scenarios triggered by a spoken phrase stop.
What it costs
Voice control is a paid feature. The half that listens runs on your own machine and adds nothing to the bill; the half that understands you and acts is the AI, and that is the half you pay for.
- A voice command is two requests to the AI, not one: a cheap parsing step that repairs what was heard and decides whether the internet is needed, and the answer itself. The card for each action shows the price you will actually be charged.
- Recognition itself is not charged. It runs on your machine, so an open microphone costs nothing until you actually address the bot.
- A web search is ten credits. It only happens when the question needs facts from the internet, and only if Allow web searches is on.
When it does not work
Nothing is heard at all. Check the model is downloaded, then use Check the microphone — say a phrase within four seconds and the app shows what it heard. If it shows nothing, raise Input gain or lower the Speech threshold.
The bot answers every phrase. The microphone is on Always and the name is empty. Give it a name, or switch to Hold a key.
The bot hears you but does nothing. The action is switched off on the Voice control screen, or the platform it needs is not connected — the card says which one.
It says "I did not understand which scene you meant". The name it heard does not match anything connected. Scene names come from OBS, button names from the control panel; say the name as it is written there.
It answers about the wrong thing, though it heard you right. Before answering, the app repairs misheard words against the names it knows — scenes, devices, buttons — and a rare word of your own can be swapped for one of them. Add that word on Voice → Speech recognition, on the Names it should get right tab: the repair step reads that list too, and a word from it is never replaced.
The microphone list says "Previously chosen microphone (not found)". The device is gone from the system — unplugged, renamed, or taken by another program. Recognition does not go quiet: it falls back to the system default microphone. Press Find microphones, pick the one you want and save.
It says "Recognition did not start". The engine could not load, and the microphone has nothing to do with it — the line underneath says what the system reported. Update the app first: builds before 0.14.2 shipped without one of the Windows runtime libraries, and on a clean Windows recognition could not start at all. If the message stays after the update, send the logs to support together with that line.