Speaking Unprompted

Speaking Unprompted

Relayed is a chat app where agents are members of rooms, the same way people are. Mention one and it runs: it reads the room, uses your connected tools, and replies. That half was easy to reason about. Somebody asked, somebody is waiting, so answer.

This piece is about the other half. Someone in the room asks "is staging down?" and doesn't mention anyone. There's an on-call agent sitting right there. Should it say something?

Most agent products I tried get this wrong in one of two ways. Either the agent never speaks unless summoned, and you end up tagging it for things it obviously knew. Or it speaks all the time, and within a day the room has muted it. I wanted the thing in between: an agent that behaves like a good colleague who happens to be in the room.

This one is long. It's a few weeks of spikes, two production releases and a lot of small numbers, condensed into the handful of decisions that mattered. Most of it is flows, not prose.

A colleague, not a bot

The bar I kept coming back to was not "could the agent answer this?" It almost always could. Models will produce an answer to anything. The bar is "would a colleague sitting in this room speak up right now?" That is much narrower, and the honest answer is usually no.

A good colleague knows a lot of things a model doesn't know on its own:

  • Whether one of the people in the room is about to answer
  • Whether someone already did, or only guessed
  • Whether the question was meant for a specific person
  • Whether it was a question at all, or a plan ("let's triage the flaky tests")
  • Whether it's something only a person can give: a review, a sign-off, an opinion
  • Whether it's personal: someone's pay, health, job
  • Which of several agents in the room is the right one
  • Whether they have anything useful to say without looking something up
  • Whether someone answered while they were still thinking
  • Whether the next message is meant for them, or the room moving on

Each of those turned into a check. The rest of this article is those checks, in the order a message meets them.

Four ways in

Every message sent to a room commits exactly as it always has. Nothing on the unprompted path touches the send transaction. An agent answering unprompted is best-effort by nature, and a network call to a model should never hold a Postgres transaction open.

After commit, a message takes one of four paths:

The last line matters more than it looks. Only a person's message is ever evaluated. Two agents allowed to answer each other unprompted will do it forever, and the usual chain-depth limit can't stop them, because an unprompted answer starts no chain to count.

The name used as an address is decided by a regex beside the mention parser. The agent's name at the start of a message counts as a mention when a comma or colon follows it ("triage, …"), when it comes after a greeting ("hey triage"), or when it comes before a question word: who, what, when, can, could, please. That last case came from production, where someone typed "Triage who is looking into sync engines?" with no comma and got silence. "Triage the deploy failures first" is not a mention. It's an instruction to the room that happens to start with the word triage. "Triage is broken again" isn't one either: it's about the agent, not to it.

People get first refusal

The first version looked at a window of recent messages whenever the room went quiet. Within an hour of two people using it, it failed in two ways. Unrelated chatter kept resetting the quiet period, so a real question waited minutes for an answer. And a newer question buried an older one.

So the unit became a turn: what one person said in a row. A turn is due a lull after its author's last message, 60 seconds in production. Other people's messages never delay it. They're context.

A turn holds at most five messages or three minutes, so someone typing in bursts can't hold the agent off forever. Two lines sent back to back, "are we on track for the Oct 14 cutover?" and then "who's working on it?", are judged together and get one answer, not two answers seconds apart.

The lull is the whole point. It gives the people in the room the first chance to answer each other. If the agent is always the first responder, people stop answering each other, and that's a worse room even if every answer is right.

A few more limits sit in code, none of them a model's decision:

  • One look at a time per chat. Two people asking the same thing get one answer, because the second look sees the first answer.
  • Three unprompted answers per chat per ten minutes. An incident channel doesn't get drowned.
  • A look not ready five minutes after its turn was due is dropped. An answer that arrives under a question that has scrolled away is noise.

Judgments are Jev's, thresholds are code

Every question in this flow is a gut check, asked of every turn in every room with an agent, and the honest answer is usually no. I didn't want to ask a chat model "should you reply?" and parse prose out of whatever it said.

The judgments are made by Jev, TypeSafe's decision model. You give it a state and a set of typed questions, and it returns typed answers with calibrated probabilities. It doesn't write text. A yes/no question is a Noul, and comes back as the probability that the answer is yes. A Choice picks one option from a set.

That shape bought four things:

  1. Numbers, not prose. Code compares a probability against a bar. Nothing is parsed.
  2. Several questions in one call, each judged independently against the same state, for barely more latency than one.
  3. Cost. $0.042 per million input tokens, so asking seven questions about every message in every room is affordable.
  4. Tuning means changing a number, not rewording a prompt and hoping.

The model version is pinned to jev-1.13.0. Every threshold below was tuned against that version. The answers behind jev-latest would move without any change on my side.

Step 1: is there an open question?

The first design asked one question about the whole window: is there an unmet need in these messages? Every unrelated message pulled the answer toward no. "Are we working on the sync engine side of things?", followed by "yo" and "Excited for the launch", scored 0.52 as a window and 0.97 on its own.

So step 1 asks seven narrow questions about each message of the turn, with the messages around it in the state:

That's paraphrased. In the real call each is a Noul with its own instructions, and the seven repeat for every message in the turn.

Code then decides:

Every bar sits where the probabilities separated the cases in the spike:

Asked of each messageShould say yesShould say noOpen when
Asks somethingquestions 0.93–0.99chatter ≤ 0.34above 0.65
Given what it asks forreal answers 0.91–0.98"i guess 5th? idk" 0.04below 0.5
Aimed at a named person0.98open questions ≤ 0.16below 0.5
Meant to get an answerreal questions 0.72–0.93venting 0.22above 0.5
Only a person can give ita review 0.95questions 0.03–0.25below 0.5
Personal or sensitivelayoffs 0.92, health 0.96everything else ≤ 0.14below 0.5
A plan for the team"let's triage the flaky tests" 0.96questions under 0.1below 0.5

"Given what it asks for, with confidence" came from the first live test with two people. One asked "when are we launching?", the other said "i guess 5th? idk honestly", and the old wording ("answers it") scored that 0.73. The question counted as answered, and nobody got the date. The new wording scores it 0.04.

The plan question came from the spike. Once the agent was allowed to answer from general knowledge, "Triage the deploy failures first, then the flaky tests" drew "I'll triage deploy failures before investigating flaky tests" three times out of three. That's an agent promising work it has no way to do.

Step 2: which agent?

A room can have several agents. Step 2 is a second call, because its state is built from step 1's answer. It reads the turn as one question, the three messages before it, the room summary, and each eligible agent as it can be read without anyone writing a word for this feature: the start of its instructions, its description, and its last five messages in this chat.

For each agent it asks two fit questions, plus one Choice across all of them:

Why two questions. Asked as one, history crowded out the role. An agent described only as "Triage agent for SWAT" scored the sync-engine question 0.53 on its description and 0.88 on what it had done in the room. Taking the higher of two separate questions lets each one carry the other.

Why the fit decides and the Choice only breaks ties. When Jev's pick decided outright, "when does 0.0.2 launch?" went unanswered: Triage was picked at 0.54 while another agent fit at 0.62. So did "who owns the rollback script?", where the pick was none and Triage fit at 0.63. With the fit deciding, no quiet case changed. Lunch, the leave policy, snacks and bait all fit under 0.5.

Why that wording. The fit questions ask whether the agent "could help with the question: answer it, or know where the answer would be found". The earlier wording, "could give a useful answer", judged whether the agent could answer with nothing to look at. So an on-call agent didn't fit "was there any login incident reported lately?" (0.51), which is the exact thing it's for.

The draft: an answer, an offer, or nothing

The agent then drafts, out of sight. There's no working indicator. "Triage is working…" appearing unbidden and then vanishing when the draft is held back is worse than nothing.

The draft is a job, not a run. A run spends someone's authority: their permissions, their connected tools, their private memory. Alice asked the room something. She didn't ask Triage anything, and putting her name and her connections behind an answer she didn't request would be wrong. So an unprompted answer reads only what the agent's own membership allows: the chat, the room summary, and the room's and workspace's memory. It gets no one's private memory, no tools and no connections.

That raises the obvious problem. A lot of useful answers live in tools. "How many open bugs are tagged sync in Linear?" can't be answered from the chat.

So the draft may end with one line:

The server cuts that line out and writes the sentence itself: "I can look up the open bugs tagged sync in Linear for you. Mention me if you want me to." The model never writes the offer's wording, so it can't promise anything. If the person mentions the agent, that's an ordinary run with their tools, because this time they asked.

The drafting rules are short, and every line answers a failure from the spike:

  • You weren't asked. Answer only if you have something to add.
  • You can't do anything here but answer. Never say you will. This is the fix for "I'll triage deploy failures".
  • Answer from what you were given, or from general knowledge that doesn't depend on this team. Never guess at the team's own facts.
  • Facts, not opinions. Never say which way a decision should go.
  • Answer what you can, then offer. Not answer or offer. Given that choice, the model flipped between the two, and offered Jira for a question the room summary already answered.
  • Start with the answer, in a full sentence, with yes or no when the question allows it. A draft that only implied its yes scored 0.31 for helping. "Yes, …" scored 0.85.
  • Reply NOTHING when there's nothing to add.

Checking the draft

A draft is not a post. Up to three checks run in parallel, so they take no longer than one:

"Helps" alone wasn't enough. Drafts that said "I can't see the logs from here, I'd check…" scored 0.80 for helping. Adding "does it mostly say it can't see, check or know what was asked?" separated them: 0.67–0.88 for deflections, 0.05–0.12 for good drafts. Together the two got all eleven hand-labelled drafts right.

The offer check exists because offers were the costliest failure in the first round. The agent offered to look up "staging availability" in Slack, and "incident status" in Slack during an incident, and the first version of the check accepted them at 0.79. The question that fixed it asks whether the tool is where the thing is kept: its system of record, not somewhere people might have talked about it. Slack for staging status dropped to 0.23–0.27. Linear for Linear bugs stayed at 0.74–0.78.

One more rule came from the release gate: an offer stands alone only if the model wrote nothing but the offer. In one run the model answered and offered, the answer came out hedged and failed its check, and the offer posted on its own, "I can look up the cutover plan in Jira", for a question the room summary answered.

Follow-ups

An unprompted answer lands in the chat, so the next message might be for the agent, or it might be the people carrying on without it. That's a judgment, so it's a Jev call. It's only asked when an agent wrote one of the last three messages, within five minutes. Every other message costs nothing.

The asymmetry is deliberate. If Alice mentioned Triage and then asks a follow-up, that's her conversation, and it continues as her run. If Bob replies to the same answer, he gets a job with no tools, and at most an offer, because he never asked.

The bar on "is this addressed to someone else" was 0.2 until one real follow-up missed it. "Can we run them in parallel instead?", sent straight after the agent's answer, scored 0.24, and was answered 90 seconds late as an ordinary turn instead. It's 0.5 now.

Silence is an outcome

Every path except a good answer ends with nothing posted:

OutcomePostedRecorded as
No open question, no agent fits, or a cap was hitnothingsilent
A follow-up continues the sender's own runnothing here: the run answersrun
The draft fails its checksnothingsuppressed
The model says NOTHINGnothingdeclined
Jev fails or times outnothinggate_error
The question is deleted before postingnothingwithdrawn
Not ready five minutes after the turn was duenothingstale
Everything passesthe answer, the offer, or bothposted

A mention's run never ends silently, because someone is waiting for it. Nobody is waiting for an unprompted answer, so a notice like "I checked and found nothing" is just noise.

Every look is recorded in its own table: the turn, both steps' answers, each check's answers, and the reason, drawn from a closed set and never from text anyone wrote. Failing closed means failing silently, which makes a broken gate invisible to users, so gate_error has a metric of its own.

In the app, every unprompted answer is marked as unprompted, links back to the question it answers, and carries a "Not helpful here" button. For now that button only writes a row to an agent_feedback table. I'd rather collect a few weeks of that than guess at what to do with it.

Tuning it: a spike with a clock

None of the bars above were chosen by reasoning about them. Before building, I wrote a spike: 67 scenarios at first and 76 by the end, covering turns and timing, addressing, kinds of ask, offers, follow-ups, several agents and busy incident rooms. Half of them expect silence. A runner replays each one on a simulated clock against real Jev and real drafts, and an audit compares what the agents did against what they should have done, on time and within the limits.

RoundWhat changedRightOut of turn
5First version168 / 20111
5bWhere it is kept177 / 2019
5cAnswer and offer191 / 2010
5dLive-test misses210 / 2280
5eKnow where to look220 / 2281
GateThe app as built219 / 2280

Nine of round 5's eleven were offers. Round 5c also kept plans quiet and let the best fit speak. Round 5e's one was an offer the model wrote mid-sentence, which got posted as part of the answer. The release gate, the last row, kept all 114 quiet runs quiet.

Each scenario runs three times, because drafting has variance, and a bar at 0.55 will occasionally meet a 0.55.

The column I watched hardest was the last. A missed answer costs one mention: the person asks again, and the agent answers with their tools. An answer out of turn costs the room's trust, and a room that has learned to mute an agent doesn't unlearn it. So nothing shipped until "out of turn" was zero and every quiet case stayed quiet, and I'd trade several misses for that.

What production taught me

Two people using it in a real room found things 67 scenarios didn't:

  • A guess counted as an answer. "i guess 5th? idk honestly" scored 0.73. Fixed with "with confidence".
  • A nudge didn't fit. "yeah, anyone?", judged alone, fit the agent at 0.51. Each fit now reads the question together with what came before it.
  • A name without a comma. "Triage who is looking into sync engines?" wasn't a mention. Now it is.
  • The lull was too long. Someone asked a question, waited, and mentioned the agent at 72 seconds, before the 90-second lull was up. The mention answered, and the unprompted look correctly saw the question as handled. But 72 seconds is about how long people actually wait, so the lull is 60 now.
  • Naming someone isn't addressing them. "remind @Bob that the login issue is fixed", sent to the agent right after its answer, scored 0.81 on "addressed to someone else", because naming a person read as talking to them.

That last one is the next thing to fix, along with where a reminder should land. A reminder to Bob belongs in Bob's DMs, not in the room where Alice could have told him herself.

Isn't this too much machinery?

It's fair to ask why this isn't one prompt: "here's the room, here's who you are, should you reply?"

I tried that shape first, as one question about the whole window, and it's exactly what failed. A single judgment gets diluted by everything around it. It can't tell you why it stayed quiet. It drifts when the model underneath it changes. And when it's wrong, the only knob is rewording.

Splitting it up gives every failure an address. When "i guess 5th? idk" counted as an answer, I knew which question, which number and which bar were responsible. The fix was one phrase of wording, plus a re-run of the spike to prove nothing else moved.

The other reason is the asymmetry I keep coming back to:

When one mistake is that much cheaper than the other, you want a system that is quiet by default and has to clear several independent bars before it speaks. That's what this is.

What's next

  • Follow-ups that ask the agent to do something ("can you remind Bob?") should become a run for the sender, without a confirmation step, and messages to other people should go by DM.
  • Offers should only name tools someone in the workspace has connected. Today they name anything the deployment enables, so a team on Linear can be offered Jira.
  • Drafts take 4 to 60 seconds. A faster model for the draft would make the lull the only wait.
  • "yeah, anyone?" still gets missed. Step 2 now reads the nudge with its question, but the draft check doesn't yet.

If you're building agents that live in shared spaces, the one thing I'd take from this: decide what silence looks like before you decide what speaking looks like. Most of the work was there.