Agent proactivity
Agents are good enough that we treat them like coworkers in the apps we work in. One thing still gives them away: you have to tag them. A real coworker doesn't wait to be mentioned. They read the conversation, decide whether something is theirs, and sometimes stay quiet.
Agents have mostly figured out what to say. What's left is when — the gap between a tool in the conversation and someone who feels like part of it.
That habit isn't made up. In a month of messages from our own workspace, people did it again and again, and they noticed every time within half a minute.
No reply
5 seconds later
12 seconds later
14 seconds later
28 seconds later
The obvious approaches
An agent sits in a room for a database cutover. It knows the room's summary: the rollback steps are in the QUARTZ runbook, section 4, and Bob owns the rollback script.
The obvious design is to let it judge every message and answer whenever it can. It almost always can. That's the problem.
Every conversation in this piece comes from our evaluation: scripted rooms replayed on a clock against the agent, or against a simpler rule for comparison. Red marks what went wrong.
Judging every message
Asking "should I answer this?" about every message is expensive, because most messages are chatter. It's also wrong. A message can be answered a few lines later, or be half of a thought.
1 min later
One message at a time
People ask in bursts, and correct themselves. Judged one message at a time, a burst gets several answers, or an answer to the wrong half.
33 s later
Chatter, and questions that aren't the agent's
Most of a working channel isn't asking anything. And some questions are open and answerable, but meant for a person, or about someone's pay or health.
All of these came from one simple rule: answer every question you can. On our conversations it spoke where it shouldn't have about a third of the time.
Episodes
The fix for most of that is to stop judging messages and judge episodes.
An episode is what one person says in a row. It's over when they've been quiet for a while, 60 seconds for us. Only then does the agent look at it as a whole, with everything around it as context.
- Their next message joins it. A clarification, a second question, a correction.
- Other people's messages don't delay it. They're context, and if one answers it, it's closed.
- A reply to an agent is its own episode, and is judged as it arrives.
40 s later
35 s later
35 s later
The decision
Each kind of message runs through a short list of checks, in order. Read down a column to follow one message. Nearly every check can only end in silence.
The rest of this piece walks down those rows. Each one hides cases that aren't obvious. Green marks what the agent did right in a real run. Where it still gets a case wrong, we show only the conversation and the rule. Those are marked as what should happen.
Is it still open?
"Has someone answered?" does the most work, and it's subtler than it looks.
1 min later
The bottom line between a bot and a colleague is often one word. "Answered" has to mean the kind of thing that was asked for: a time for when, a person for who, a reason for why.
A nudge carries its question
"yeah, anyone?" asks nothing on its own. It means the question above it is still open, so it brings that question into the episode.
3 min later
6 min later
Is it the agent's?
An open question still has to belong to the agent. Some questions are about it without being to it. Some sound like they need a person but only need a rule. Some have an answer nobody in the room can see.
An unprompted agent never acts. It answers, or offers to look. Acting spends someone's permissions, and nobody asked it to. If they want it done, they mention it.
Who gets first refusal
The wait gives people the first chance to answer each other. A shorter wait answers Alice sooner. It also talks over more people who were about to answer her.
40 s later
| Wait | Time to an answer | Talked over a person |
|---|---|---|
| 30 s | 32 s | 18 times |
| 60 s | 62 s | 9 times |
| 90 s | 92 s | never |
Our scripts decide how often Bob answers in that window, so we measured real reply times in public chat logs. In a busy channel, a person answers within 30 seconds 13% of the time, within 60 seconds 27%, and within 90 seconds 33%. The step from 30 to 60 matters most. Going to 90 adds little and makes Alice wait longer. We settled on 60.
Deciding to answer isn't the end
Once an agent is picked, it writes its answer out of sight. Then a second decision: should it be posted?
- The agent may write nothing. It's allowed to decide it has nothing to add.
- Does it help the asker? "I can't see that" isn't an answer, however politely it's put.
- For an offer: is that tool where the thing is kept? And would looking it up actually answer the question?
- Did someone answer, or take it on, while it was being written?
- Is it still on time? An answer ready five minutes after it was due is dropped. By then the question has scrolled away.
Follow-ups don't wait
Once the agent has answered, the conversation is partly with it. A reply to it is judged as it arrives, with no wait, the way a person answers someone who just turned to them. It still has to clear three bars: it asks the agent for something more, it isn't sensitive, and the exchange hasn't already had three follow-ups.
1 min later
A busy room
We used to cap unprompted answers at three per room every ten minutes, so an incident channel wouldn't drown. It only ever turned away real questions. Repeated questions are already caught as answered, and venting never gets past the first checks. So there's no cap: every real question gets its own decision.
Harsh's when: the topic case above
Several agents
A room can have several agents. Only one decision runs per room at a time, so a second agent always sees what the first one said. Of the agents that fit, the one whose role fits best answers, and only that one. When two fit almost equally, an overall pick between them breaks the tie. Agents never answer agents, or two of them would talk forever.
The typing indicator
A person who's about to answer shows as typing. That's the natural signal that someone is on it. It isn't a 👀 reaction, and it isn't an "on it" message. People don't react and then also reply, and an "on it" is noise once the answer arrives.
The hard part is when it appears. This is designed, not shipped yet.
Typing stops. The draft was held back
The last one is fine. A person starts typing and thinks better of it all the time.
Where it stands
Ninety conversations we tuned on can't say much on their own. So we wrote 166 new ones, labelled them before running anything, and never tuned on them. Relayed against three simpler options:
| Spoke when it shouldn't | Missed | |
|---|---|---|
| Relayed | 3% | 26% |
| One "should I answer?" prompt to a frontier model | 5% | 19% |
| Answer every question | 33% | 6% |
| Never answer | 0% | 100% |
The restraint held on conversations it had never seen: about 3% wrong posts, the same as on the set we tuned on. The misses didn't hold. Most of them were questions the room's summary could answer, in rooms whose agents weren't set up for that topic. We labelled those as should-answer, and so did a second, independent labeller, who agreed with our labels on 97% of decisions.
That leaves a product question more than a model one: should an agent answer what the room knows, when it's outside what the agent is for? Next is answering that, the edge cases marked should above, the typing indicator, and real rooms.