Cue · Our model
Cue is Assistly's own real-time question detection model. Every line of your call gets a score in about 30 milliseconds, so the answer starts the moment a question is yours, not half a second later.
Maya · Hiring manager
So the Q3 work is mostly the billing migration, and it's been, uh, a lot.
You
Mm-hmm. Yeah.
Maya · Hiring manager
Walk me through how you'd roll that out without any downtime.
Run both systems side by side: dual-write new invoices, backfill the rest in batches, then flip reads behind a flag you can turn off in seconds.
How Cue works
Each time a line of the conversation finishes, Cue reads it with the few lines before it and who said each one. A question that ends on a fragment or a quick “hmm” still reads as one question.
Is this a question you should answer right now? And does answering it depend on what is on your screen? Two scores, one pass, in tens of milliseconds.
Most of a call is talk, not questions. When the score says it is your cue, Assistly writes the answer. Everything else costs nothing and slows nothing down.
Benchmarks
Measured on 634 moments from real calls that Cue never saw in training, against the two ways we used to detect questions: asking a general-purpose large language model about every line, and a hosted decision model.
Median, per line of conversation.
Out of every 100 lines that could hold a question.
What counts as your cue
“Walk me through how you'd scale that.”
An ask aimed at you, even without a question mark.
“Any other thoughts before we move on?”
An open floor. You're in the room, so it's yours too.
“What's the time complexity of a heap push again?”
You, stuck mid-problem. A real knowledge question.
“Right? You know?”
A tag at the end of someone's sentence. Nothing is asked.
“Priya, are you good with the launch date?”
Addressed to someone else by name. Not your turn.
“What stage is your team at right now?”
You asking them. Their answer, not yours to look up.
Cue is Assistly's own real-time question detection model. It reads a live conversation one line at a time and decides whether the newest line is a question you should answer, and whether answering it needs your screen. When it is, Assistly writes the answer for you.
A general-purpose language model can spot questions, but it takes around half a second each time and has to be asked about every single line. Cue is a small model trained for this one decision. It answers in about 30 milliseconds and lets the large model focus on writing answers.
Cue is built on a multilingual encoder and trained on calls in several languages and mixed-language conversations, so it isn't limited to English. Assistly transcribes dozens of languages and answers in English.
Cue runs on Assistly's servers, in the same pipeline as live transcription. There is nothing to install or turn on: it is part of Auto-assist, which is on by default.
Cue is tuned to lean toward catching a question rather than missing one. A borderline line goes to the answer model, which makes the final call, so an uncertain moment costs a fraction of a second instead of a missed question.
Cue is part of Auto-assist, on by default in Assistly for macOS and Windows. Start a session and the next question you're asked gets answered as it lands.