Speech recognition

People ask this as

What this page is

Voice control for your stream: you say what you want out loud, and the smart AI assistant does it. Starting predictions by voice, running a poll, switching an OBS scene, changing the title and category, cutting a clip, writing a line in chat, flipping a smart plug — all of it without letting go of the mouse mid-game.

Two screens set it up, and they answer different questions:

Recognition runs on your own machine, so the sound never leaves the computer. Only the request itself goes to the AI, and only after the app decides you were talking to the bot — and that request is the part you pay for.

How to set it up

1. Pick a recognition model

Voice → Speech recognition, the Recognition model block. Pick one and press Download — it is a one-off download, after which recognition works offline. The row for each model says how much memory it takes and how fast it runs on your graphics card and on your processor: pick the one whose numbers you can live with, not the biggest.

Until a model is downloaded nothing is heard, and the tile says Recognition is not ready — not silence, a missing model.

2. Microphone and language

The Microphone and language block. Find microphones shows their real names; without it the system default is used.

Set Language you speak. Leaving it on Detect automatically makes every phrase go through a second pass, so recognition is about twice as slow.

3. Choose when to listen

The When to listen block, and this is the choice that matters most.

4. Name the bot, if the microphone stays open

The How you call the bot block. Say the name and the rest of the sentence goes to the bot: "Nova, start a poll".

Leave the name empty and every phrase you say goes to the bot — and every one of them costs credits. The screen says so in red while the microphone is open and the name is missing. With Hold a key an empty name is fine: the key is the trigger.

5. Word that stops the voice over

Say it while the bot is talking and it goes quiet at once. The name is not needed here — by the time you have said it, the bot has finished.

Leave the field empty and the word for the language you speak is used: *stop*, *стоп*, *alto*, *pare*, *stopp*.

6. Allow the actions you want

Chat Bot → Voice control. Every action is a card with its price in credits and a switch. Anything that cannot be undone starts switched off.

There is no list of magic phrases. The bot works out what you meant, so "cut a clip", "clip that" and "save this moment" all reach the same action. The names it can use — OBS scenes, control panel buttons, smart plugs, prediction outcomes — come from what is actually connected, so it cannot invent a scene you do not have.

What it can do

ActionSay something like
Start a poll"start a poll: which map — Dust or Mirage"
Finish the poll"finish the poll"
Start a prediction"start a prediction: do we win this one, yes or no"
Close the betting"close the betting"
Resolve and pay out"we won, pay it out"
Cancel and refund"cancel the prediction, give the points back"
Switch the OBS scene"switch to the Break scene"
Change title and category"set the title to Chill coding"
Save a clip"clip that"
Write in chat"tell chat the stream ends in ten minutes"
Press a control panel button"press Ad break"
Switch a smart plug"turn the lamp on for a minute"

Twelve, and that is the whole list. Clips, the title, polls and predictions go through Twitch; on YouTube and Kick the assistant writes in chat, switches scenes and presses panel buttons. One phrase is one action: "lock the betting and switch the scene" does only the first of the two.

Anything that is not an action is an ordinary answer: ask what the weather is, how long the stream has been going, what you played last week, what your chat commands are, what you asked a minute ago.

Ask before anything that cannot be undone

Switched on by default. The bot names the outcome — "Resolving the prediction: 'We win'" — and waits for a "yes". It names the outcome and not the action on purpose: what you are confirming is the thing you cannot take back.

Switch it off and the bot just does it. Quicker, and there is no taking it back.

Where you see what happened

The Stream: Voice tile on the control panel keeps the feed: your phrases and the bot's answers, in order. The colour of the stripe on the left says how it ended — blue for an answer, green for something done, amber for a question back, red for a refusal with the reason next to it.

By default the bot says the short version out loud — "Poll started" — and the details go to that feed. If you have no second screen to read them on, switch on Read the whole result out loud on the Voice control screen.

Three buttons in the tile's header:

What it costs

Voice control is a paid feature. The half that listens runs on your own machine and adds nothing to the bill; the half that understands you and acts is the AI, and that is the half you pay for.

When it does not work

Nothing is heard at all. Check the model is downloaded, then use Check the microphone — say a phrase within four seconds and the app shows what it heard. If it shows nothing, raise Input gain or lower the Speech threshold.

The bot answers every phrase. The microphone is on Always and the name is empty. Give it a name, or switch to Hold a key.

The bot hears you but does nothing. The action is switched off on the Voice control screen, or the platform it needs is not connected — the card says which one.

It says "I did not understand which scene you meant". The name it heard does not match anything connected. Scene names come from OBS, button names from the control panel; say the name as it is written there.

It answers about the wrong thing, though it heard you right. Before answering, the app repairs misheard words against the names it knows — scenes, devices, buttons — and a rare word of your own can be swapped for one of them. Add that word on Voice → Speech recognition, on the Names it should get right tab: the repair step reads that list too, and a word from it is never replaced.

The microphone list says "Previously chosen microphone (not found)". The device is gone from the system — unplugged, renamed, or taken by another program. Recognition does not go quiet: it falls back to the system default microphone. Press Find microphones, pick the one you want and save.

It says "Recognition did not start". The engine could not load, and the microphone has nothing to do with it — the line underneath says what the system reported. Update the app first: builds before 0.14.2 shipped without one of the Windows runtime libraries, and on a clean Windows recognition could not start at all. If the message stays after the update, send the logs to support together with that line.