Engineering

Meet Cue: how we built a real-time question detection model for live calls

By The Assistly team ·

The hardest part of answering a question in a live call is not the answer. It's knowing, in the second after someone stops talking, that a question just happened and that it was meant for you.

That decision runs constantly. A one-hour call produces hundreds of finished lines of speech, and nearly all of them are not questions: "yeah", "mm-hmm", someone thinking out loud, a long explanation that ends in "right?". Somewhere in there is "walk me through how you'd roll that out", and that's the line that matters.

Today we're introducing Cue, the model Assistly now uses to make that call. You can read the short version on the Cue page. This post is the longer one.

What Cue does

Cue reads the newest line of your conversation together with the few lines before it and who said each one. It returns two scores:

  1. Is this your cue? Is there a question, right now, that you should answer?
  2. Does it need your screen? Does answering depend on what you're looking at, like a chart someone is presenting or the code you're writing?

When the first score says yes, Assistly's answer model writes the answer and it streams into your private overlay. When it says no, nothing else happens, and nothing slows down.

Why we built our own

Auto-assist has always worked by asking a capable language model, after every line, whether a question had just been asked. It worked, but it had three costs we could feel:

  • Latency. A large model takes around half a second to answer even a short yes/no question. On a live call, half a second is the gap between help arriving during the pause and help arriving after you've started talking.
  • Waste. Most lines aren't questions, so most of that work was spent confirming that "mm-hmm" is not a question.
  • Over-eagerness. A general model asked "is there a question here?" tends to find one. It would fire on a question asked three lines earlier that you were already halfway through answering.

A dedicated model fixes all three at once. A small encoder trained for exactly one decision is fast enough to run on every line, and because it's trained on what this decision looks like in real conversations, it's stricter about the cases a general model gets wrong.

How Cue works

Cue is a compact multilingual transformer encoder, fine-tuned as a cross-encoder: it reads the recent context and the newest line together, as one input, so it can tell "Toilet." (an echo right after "how does a toilet work?") from "Toilet." in the middle of a story.

We trained it by distillation. Large language models are slow, but they're good judges when you give them time and clear rules. So we had them label tens of thousands of conversation moments using the same rules Assistly's live detector follows, and trained Cue to reproduce those judgments in a fraction of the time. Where the larger models disagreed about a moment, we leaned toward calling it a question. In a live call a missed question is the expensive mistake, and an extra check costs a fraction of a second.

A few details that turned out to matter:

  • Judge the newest line, not the whole window. Asked whether a stretch of conversation "contains a question", models fire on questions that are already being answered. Asking about the line that just finished removes most of those false alarms.
  • A question stays live until someone starts answering. If an interviewer asks something and you say "hmm", the question is still yours. Cue is trained to keep it live through fillers and echoes, and to drop it once an answer begins.
  • Not every question is yours. A question addressed to a colleague by name, a "right?" at the end of someone's sentence, or you asking the other side about their own plans are all questions that shouldn't trigger an answer.

The results

We measured Cue on 634 moments from real calls that it never saw in training, and compared it with the two approaches we had used before: a general-purpose language model asked about every line, and a hosted decision model.

Time to decide (median)Lines sent to the answer model
General-purpose LLM~580 msevery line
Hosted decision model~110 msabout 1 in 4
Cue~30 msabout 1 in 4

Cue makes the call roughly 19 times faster than asking a large model, and three in four lines are settled without the answer model ever running. Because the large model now only sees moments Cue has already flagged, the pipeline raises about 4.5 times fewer false alarms than checking every line with the large model, while catching nearly as many real questions.

Against the hosted decision model, Cue is about as accurate and nearly four times faster, and it runs inside our own backend: no extra network hop and no per-call fee.

Where it runs

Cue runs on Assistly's servers, in the same pipeline as live transcription. There's nothing to install or switch on: it's part of Auto-assist, which is on by default. If Cue is ever unavailable, Assistly falls back to its previous detector automatically, so a question is never dropped on the floor.

What's next

The screen score is the next thing we're improving. Questions that depend on what's on your screen are rarer in real calls, so Cue has seen fewer of them, and for now the answer model still makes the final call there. We're also expanding Cue's training data in more languages and mixed-language calls.

If you want to see it work, start a session and let someone ask you a question. That's your cue.

Try Assistly on your next call

Live guidance in your meetings, calls, and interviews, and clean notes after every session. Free to start, no card required.