Skip to main content

Speech recognition

The Speech recognition tab of the System screen. It decides where the voice messages sent in Qubix conversations are turned into text.

Where recordings are processed

The Voice message recognition switch has three positions:

  • External — the recording is sent to a cloud model. It is billed to your OpenRouter key.
  • Internal — the recording is processed by an engine running on this server and never leaves it.
  • Off — nothing is transcribed; voice messages stay as audio only.

Switching the position recomputes nothing: the text of a voice message travels with the message itself, so everything already transcribed keeps its text, and only new recordings follow the new setting. There is no reindexing and nothing to wait for.

Cloud model

The Cloud model row appears only in the External position — the local engine ships with exactly one model inside it, so there is nothing to choose there.

The list shows, for each model, who receives the recording, the model name and its price. The choice is saved as soon as you pick it.

Engine

The Engine block is read-only. It shows the model that is actually in use and the recognition Language, and below them a line about the current state:

The line saysWhat it means
Voice messages stay without textThe switch is in the Off position.
Transcription runs in the cloudThe cloud position is set up and working.
An OpenRouter key is required in the Connections sectionThe cloud position is selected, but the box has no OpenRouter key. Add it in Connections — until then nothing is transcribed.
Transcription runs on this server; recordings never leave itThe local engine is in place and recordings do not leave the server.
The speech recognition engine is being installed, about 2 GB. This usually takes several tens of minutesThe local position is selected, but the engine is not on this server yet.
If the local engine did not arrive

The engine is downloaded to the server after the local position is chosen, and it is a large download. Picking the same position again asks for it once more — that is exactly how you retry a download that failed (no access to the registry, no free disk space). While you stay on the tab it re-checks on its own, so an engine that arrives soon is picked up without a page reload. That watch is bounded, though — on a slow link the download can outlast it; if the tab still reports the engine as missing, reload the page to start a fresh wait.

If the tab cannot read its own state, it shows Could not load the speech-recognition status. and a Retry button instead of the cards — so a failure is never displayed as an engine that is switched off.

What's next