Cue · Our model

Meet Cue. It knows when it's your turn

Cue is Assistly's own real-time question detection model. Every line of your call gets a score in about 30 milliseconds, so the answer starts the moment a question is yours, not half a second later.

Maya · Hiring manager

So the Q3 work is mostly the billing migration, and it's been, uh, a lot.

0.02

You

Mm-hmm. Yeah.

0.01

Maya · Hiring manager

Walk me through how you'd roll that out without any downtime.

0.97
Your cue · 30 ms
Assistlyonly you see this

Run both systems side by side: dual-write new invoices, backfill the rest in batches, then flip reads behind a flag you can turn off in seconds.

to decide if a line is your cue
30 ms
to decide if a line is your cue
faster than asking a large language model
19×
faster than asking a large language model
of lines settled without the big model
75%
of lines settled without the big model
fewer false alarms than checking every line with an LLM
4.5×
fewer false alarms than checking every line with an LLM

How Cue works

One small model, one decision, every line

  1. 1

    Reads the moment

    Each time a line of the conversation finishes, Cue reads it with the few lines before it and who said each one. A question that ends on a fragment or a quick “hmm” still reads as one question.

  2. 2

    Makes two calls

    Is this a question you should answer right now? And does answering it depend on what is on your screen? Two scores, one pass, in tens of milliseconds.

  3. 3

    Hands off only when it matters

    Most of a call is talk, not questions. When the score says it is your cue, Assistly writes the answer. Everything else costs nothing and slows nothing down.

Benchmarks

Faster than the pause after a question

Measured on 634 moments from real calls that Cue never saw in training, against the two ways we used to detect questions: asking a general-purpose large language model about every line, and a hosted decision model.

Time to decide

Median, per line of conversation.

Cue30 ms
Hosted decision model110 ms
General-purpose LLM580 ms

Lines sent to the big model

Out of every 100 lines that could hold a question.

Cue25 of 100
General-purpose LLM alone100 of 100

What counts as your cue

Not every question mark is a question

Your cue

“Walk me through how you'd scale that.”

An ask aimed at you, even without a question mark.

Your cue

“Any other thoughts before we move on?”

An open floor. You're in the room, so it's yours too.

Your cue

“What's the time complexity of a heap push again?”

You, stuck mid-problem. A real knowledge question.

Not your cue

“Right? You know?”

A tag at the end of someone's sentence. Nothing is asked.

Not your cue

“Priya, are you good with the launch date?”

Addressed to someone else by name. Not your turn.

Not your cue

“What stage is your team at right now?”

You asking them. Their answer, not yours to look up.

Questions about Cue

What is Cue?+

Cue is Assistly's own real-time question detection model. It reads a live conversation one line at a time and decides whether the newest line is a question you should answer, and whether answering it needs your screen. When it is, Assistly writes the answer for you.

How is Cue different from asking a chatbot to spot questions?+

A general-purpose language model can spot questions, but it takes around half a second each time and has to be asked about every single line. Cue is a small model trained for this one decision. It answers in about 30 milliseconds and lets the large model focus on writing answers.

Does Cue work in languages other than English?+

Cue is built on a multilingual encoder and trained on calls in several languages and mixed-language conversations, so it isn't limited to English. Assistly transcribes dozens of languages and answers in English.

Where does Cue run?+

Cue runs on Assistly's servers, in the same pipeline as live transcription. There is nothing to install or turn on: it is part of Auto-assist, which is on by default.

What happens if Cue isn't sure?+

Cue is tuned to lean toward catching a question rather than missing one. A borderline line goes to the answer model, which makes the final call, so an uncertain moment costs a fraction of a second instead of a missed question.

Let Cue catch the next one

Cue is part of Auto-assist, on by default in Assistly for macOS and Windows. Start a session and the next question you're asked gets answered as it lands.