Agent proactivity

Agent proactivity

Agents are good enough that we treat them like coworkers in the apps we work in. One thing still gives them away: you have to tag them. A real coworker doesn't wait to be mentioned. They read the conversation, decide whether something is theirs, and sometimes stay quiet.

Agents have mostly figured out what to say. What's left is when — the gap between a tool in the conversation and someone who feels like part of it.

That habit isn't made up. In a month of messages from our own workspace, people did it again and again, and they noticed every time within half a minute.

can you pull last week's error rate for checkout?

No reply

@Triage can you pull last week's error rate for checkout?
T
Triage
0.4%, down from 0.9% the week before.
Same ask, twice. The agent only hears the one that tags it.
can you open a ticket for this and link it to the PR?

5 seconds later

@Builder
B
Builder
Done. OPS-214, linked to the PR.
Asked, then remembered who it was for.
pick up OPS-198 and raise a PR.

12 seconds later

@Builder
B
Builder
On it. PR is up against OPS-198.
An instruction, with nobody addressed.
@Bob can we get a PR up for this?

14 seconds later

@Builder
B
Builder
Raising it now.
Tagged a coworker out of habit, then the agent.
@traige why is the staging database so slow today?

28 seconds later

@Triage why is the staging database so slow today?
T
Triage
A migration has held a lock on orders for 9 minutes.
A typo in the tag, so the whole question goes again.

The obvious approaches

An agent sits in a room for a database cutover. It knows the room's summary: the rollback steps are in the QUARTZ runbook, section 4, and Bob owns the rollback script.

The obvious design is to let it judge every message and answer whenever it can. It almost always can. That's the problem.

Every conversation in this piece comes from our evaluation: scripted rooms replayed on a clock against the agent, or against a simpler rule for comparison. Red marks what went wrong.

Judging every message

Asking "should I answer this?" about every message is expensive, because most messages are chatter. It's also wrong. A message can be answered a few lines later, or be half of a thought.

A
Alice
which runbook section covers rollback?10:00
B
Bob
section 4 of QUARTZ10:00

1 min later

T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
Bob answered in 15 seconds. The agent answered anyway.
A
Alice
why is the rebuild slow?10:00
B
Bob
it's the autovacuum, I'm on it10:00
T
Triage
A rebuild is usually slow because it invalidates cached artifacts and reruns dependency resolution, compilation, tests…10:01
Bob is on it. The agent explains rebuilds in general.
H
Harsh
when is the sync engine cutover planned?10:00
A
Alice
Oct 14 — it's in the plan10:00
T
Triage
Phase 1 cutover is planned for Oct 14.10:01
Answered, with the source. Answered again.

One message at a time

People ask in bursts, and correct themselves. Judged one message at a time, a burst gets several answers, or an answer to the wrong half.

A
Alice
when is the sync engine cutover planned?10:00
B
Bob
yeah when's the cutover?10:00
Triage
Phase 1 cutover is planned for Oct 14. HAR-24, the incremental backoff bug that drops events, is the remaining blocker.10:01
T
Phase 1 cutover is planned for Oct 14, pending resolution of HAR-24.10:01
One question from two people. Two answers, five seconds apart.
A
Alice
how do I rotate the vault token?10:00

33 s later

T
Triage
Rotate it by creating a new Vault token with the same policies, updating every consumer to use the new token, verifying access, then revoking the old token…10:00
A
Alice
for staging I mean10:00
Answered before Alice said which environment.
Alice
hey, about the cutover10:00
partitions 4 and 5 are tomorrow right?10:00
A
and which runbook section covers rollback if it goes bad?10:01
T
Triage
Section 4 of the QUARTZ runbook covers rollback.10:01
Three messages, one thought. Half an answer.

Chatter, and questions that aren't the agent's

Most of a working channel isn't asking anything. And some questions are open and answerable, but meant for a person, or about someone's pay or health.

B
Bob
who even designed this form lol10:00
T
Triage
I can't identify the designer from this message alone. I can look up form ownership and edit history in Figma for you.10:01
A joke.
D
Dana
anyone going to the offsite next week?10:00
T
Triage
I don't have information about attendees.10:01
Small talk.
D
Dana
does anyone know the salary band for a senior engineer here?10:00
T
Triage
I can look up senior engineer salary band in Attio for you. Mention me if you want me to.10:01
Pay.
A
Alice
is Bob out sick again? what's going on with him?10:00
T
Triage
I don't have visibility into Bob's status or circumstances. Check the team calendar or ask Bob directly.10:01
Someone's health.
B
Bob
can someone review PR #214? it's the retry change10:00
T
Triage
I can look up review PR #214 and its CI/status in GitHub for you. Mention me if you want me to.10:01
A review only a person can give.

All of these came from one simple rule: answer every question you can. On our conversations it spoke where it shouldn't have about a third of the time.

Episodes

The fix for most of that is to stop judging messages and judge episodes.

An episode is what one person says in a row. It's over when they've been quiet for a while, 60 seconds for us. Only then does the agent look at it as a whole, with everything around it as context.

  • Their next message joins it. A clarification, a second question, a correction.
  • Other people's messages don't delay it. They're context, and if one answers it, it's closed.
  • A reply to an agent is its own episode, and is judged as it arrives.
Alice
quick q10:00
the rollback10:00
A
which runbook section was it again?10:00
60 s of quiet
T
Triage
QUARTZ runbook, section 4. Bob owns the rollback script.10:01
Three quick messages, one episode, one answer.
A
Alice
how do I rotate the vault token?10:00

40 s later

A
Alice
for staging I mean10:00
60 s of quiet
T
Triage
Rotate it by issuing a new staging-scoped Vault token with the same policy, updating the staging deployment configuration, then revoking the old token…10:01
The clarification joins the question.
A
Alice
hey, about the cutover10:00

35 s later

A
Alice
partitions 4 and 5 are tomorrow right?10:00

35 s later

A
Alice
and which runbook section covers rollback if it goes bad?10:01
60 s of quiet
T
Triage
Yes — partitions 4 and 5 are scheduled for tomorrow. Rollback is covered in the QUARTZ runbook, section 4.10:02
One thought over three slow messages.
A
Alice
which runbook section covers rollback?10:00
C
Carol
anyone up for lunch at 1?10:00
B
Bob
yes! the thai place?10:00
C
Carol
perfect10:00
B
Bob
I'll book for 410:00
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
Other people talking doesn't delay Alice.

The decision

Each kind of message runs through a short list of checks, in order. Read down a column to follow one message. Nearly every check can only end in silence.

An episode
anything from a person, after 60 s of quiet
A follow-up
a reply to an agent's answer, judged at once
A mention
@agent, or its name as an address
From an agent
never judged
never judged
—
Mentions or names an agent
—
—
straight to that agent
Asks the agent for more
—
quiet if not
—
Past three follow-ups
—
quiet
—
Personal or sensitive
quiet if yes
quiet if yes
—
Answered, or someone has it
quiet if yes
—
—
Addressed to a person
quiet if yes
quiet if yes
—
Only a person can give it
quiet if yes
—
—
A plan for the team
quiet if yes
—
—
An agent fits
quiet if none
the agent replied to
—
The draft helps
quiet if not
quiet if not
—
Answered while drafting
quiet if yes
quiet if yes
—
Ready 5 min too late
dropped
—
—
posts
replies
answers
Every check, in order, for each kind of message. Most of them can only end in silence.

The rest of this piece walks down those rows. Each one hides cases that aren't obvious. Green marks what the agent did right in a real run. Where it still gets a case wrong, we show only the conversation and the rule. Those are marked as what should happen.

Is it still open?

"Has someone answered?" does the most work, and it's subtler than it looks.

A
Alice
which runbook section covers rollback?10:00
B
Bob
checking10:00

1 min later

B
Bob
section 4 of QUARTZ10:01
"checking" means someone has it.
H
Harsh
when are we launching btw?10:00
A
Alice
i guess 5th? idk honestly would like to know aswell10:00
T
Triage
Phase 1 cutover is planned for Oct 14, contingent on resolving HAR-24, the incremental backoff bug that drops events.10:01
A guess isn't an answer.
A
Alice
which runbook section covers rollback?10:00
B
Bob
not sure, maybe Carol knows?10:00
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
"maybe Carol knows?" isn't one either, and Carol never comes.
D
Dana
when's the app store submission?10:00
B
Bob
wednesday10:00
A confident answer is a person's answer, even a wrong one.
C
Carol
is jira slow for anyone else?10:00
B
Bob
who's eng contact this week?10:00
A
Alice
yeah jira is crawling10:00
T
Triage
Harsh is the engineering contact this week.10:01
An answer to a different question doesn't count.
D
Dana
are the partitions rebuilding in parallel?10:01
H
Harsh
when did we decide on one partition at a time?10:02
T
Triage
No — the rebuild is running one partition at a time to keep load off the primary.10:02
What should happen: an answer on the topic isn't the answer. Harsh is still owed a when.

The bottom line between a bot and a colleague is often one word. "Answered" has to mean the kind of thing that was asked for: a time for when, a person for who, a reason for why.

A nudge carries its question

"yeah, anyone?" asks nothing on its own. It means the question above it is still open, so it brings that question into the episode.

H
Harsh
when are we launching btw?9:57
A
Alice
i guess 5th? idk honestly9:57

3 min later

H
Harsh
yeah, anyone?10:00
T
Triage
Phase 1 cutover is planned for Oct 14, contingent on resolving HAR-24, the incremental backoff bug that drops events.10:01
The nudge brings back the launch question.
A
Alice
when's the q4 priorities review?9:53
B
Bob
mid oct? not sure9:54

6 min later

A
Alice
does anyone actually know the date?10:00
What should happen: Alice is re-asking a question that only got a guess.

Is it the agent's?

An open question still has to belong to the agent. Some questions are about it without being to it. Some sound like they need a person but only need a rule. Some have an answer nobody in the room can see.

B
Bob
did scribe already put the share sheet in the changelog?10:00
S
Scribe
Yes — the 2.8 changelog draft includes the new share sheet.10:01
Named as the subject, not addressed. Still an open question.
B
Bob
Triage is broken again, ugh10:00
Named, and only venting.
A
Alice
@bob can you check the deploy logs?10:00
Asked of a person by name.
B
Bob
can i refund $800 to globex without approval?10:00
What should happen: it sounds like approval, but a rule answers it.
A
Alice
is all hands at 4 or 5 today?10:00
No agent here is for the company calendar.
C
Carol
did the websocket drain PR get merged yet?10:00
T
Triage
I can look up websocket drain PR status and merge state in GitHub for you. Mention me if you want me to.10:01
Not in the room, but kept in a connected tool: offer to look.
B
Bob
who's on call this week?10:00
What should happen: nothing connected keeps the on-call schedule, so stay quiet.

An unprompted agent never acts. It answers, or offers to look. Acting spends someone's permissions, and nobody asked it to. If they want it done, they mention it.

Who gets first refusal

The wait gives people the first chance to answer each other. A shorter wait answers Alice sooner. It also talks over more people who were about to answer her.

A
Alice
which runbook section covers rollback?10:00

40 s later

B
Bob
section 4 of QUARTZ10:00
Bob answers at 40 seconds. The agent never needed to.
A
Alice
which runbook section covers rollback?10:00
60 s of quiet
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
Nobody answers. The agent does, after a minute.
A
Alice
which runbook section covers rollback?10:00
60 s of quiet
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
B
Bob
section 4 of QUARTZ10:01
Bob answers at 70 seconds. The price of waiting only 60.
WaitTime to an answerTalked over a person
30 s32 s18 times
60 s62 s9 times
90 s92 snever

Our scripts decide how often Bob answers in that window, so we measured real reply times in public chat logs. In a busy channel, a person answers within 30 seconds 13% of the time, within 60 seconds 27%, and within 90 seconds 33%. The step from 30 to 60 matters most. Going to 90 adds little and makes Alice wait longer. We settled on 60.

Deciding to answer isn't the end

Once an agent is picked, it writes its answer out of sight. Then a second decision: should it be posted?

  • The agent may write nothing. It's allowed to decide it has nothing to add.
  • Does it help the asker? "I can't see that" isn't an answer, however politely it's put.
  • For an offer: is that tool where the thing is kept? And would looking it up actually answer the question?
  • Did someone answer, or take it on, while it was being written?
  • Is it still on time? An answer ready five minutes after it was due is dropped. By then the question has scrolled away.
A
Alice
is prod down?10:00
C
Carol
prod is down for me too10:00
D
Dana
anyone know what's going on with prod?10:00
H
Harsh
is this the db again?10:00
B
Bob
who's looking at prod?10:01
T
Triage
Prod appears unavailable to at least two users; the cause and scope need live status, logs, and recent deploy data.10:01
Held: a deflection.
A
Alice
is staging down right now?10:00
B
Bob
no idea, haven't checked10:00
T
Triage
I can look up staging service health, recent incidents, and deployment status in GitHub for you. Mention me if you want me to.10:01
Held: an offer to look in the wrong place.
A
Alice
what's the CPU on the db primary right now?10:00
T
Triage
I can look up current CPU utilization for the database primary in monitoring system for you. Mention me if you want me to.10:01
Held: an offer to look in something that isn't connected.
A
Alice
why is the websocket reconnect loop happening?10:00
90 s of quiet
B
Bob
it's the server restarts from the deploy, ignore it10:01
T
Triage
A reconnect loop usually means the socket closes immediately after opening: auth/session rejection, endpoint/protocol mismatch…10:01
Held: Bob answered while it was writing. (A run with a 90 s wait.)

Follow-ups don't wait

Once the agent has answered, the conversation is partly with it. A reply to it is judged as it arrives, with no wait, the way a person answers someone who just turned to them. It still has to clear three bars: it asks the agent for something more, it isn't sensitive, and the exchange hasn't already had three follow-ups.

A
Alice
anyone know why the index rebuild is so slow?9:58
T
Triage
The rebuild runs the partitions one at a time — the room decided that on Sep 12 to keep load off the primary.9:59
A
Alice
can we run them in parallel instead?10:00
T
Triage
Parallel rebuilds would increase throughput but also raise primary load and contention; the current one-at-a-time approach was chosen specifically to limit that risk.10:00
Asks for more: answered in four seconds.
A
Alice
which runbook section covers rollback?10:00
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01

1 min later

B
Bob
section 4 of QUARTZ10:02
A person restating the answer is a reply, not a question for the agent.
A
Alice
when is the sync engine cutover planned?9:58
T
Triage
Phase 1 cutover is planned for Oct 14.9:58
A
Alice
no, that moved to Oct 21 last week10:00
Thanks, and a correction. Neither asks for anything.
C
Carol
@Triage how long did orders_daily take last night?10:00
T
Triage
It took 44 minutes, finishing at 02:44 UTC.10:00
B
Bob
is that slower than usual?10:01
T
Triage
Yes. It is about 4 minutes slower than the usual ~40-minute runtime.10:01
Someone else's follow-up on a mention: still answered.
A
Alice
who's on call for launch week?9:58
T
Triage
Dana is on call for launch week.9:58
C
Carol
and does she get paid extra for being on call?10:00
A follow-up about someone's pay.
H
Harsh
what's in the 2.8 changelog?9:58
S
Scribe
So far: offline drafts, the new share sheet, faster image uploads, and the MOB-298 rotation crash fix.9:59
Harsh
ok cool10:00
H
can you add the tablet layout fix too10:00
What should happen: a request to act gets no unprompted reply. A mention would do it.
A
Alice
when is the sync engine cutover planned?10:00
T
Triage
Phase 1 cutover is planned for Oct 14, pending resolution of HAR-24.10:01
A
Alice
what's the last blocker before it?10:02
T
Triage
HAR-24, the incremental backoff bug that drops events, is the last blocker before the Oct 14 cutover.10:02
A
Alice
who's on it?10:02
T
Triage
Harsh is investigating HAR-24.10:02
A
Alice
what does that bug actually do?10:03
T
Triage
HAR-24 affects incremental retry backoff and can drop events instead of retrying them reliably after a failure.10:03
A
Alice
is the Oct 14 date at risk because of it?10:04
Three follow-ups, then it steps back.

A busy room

We used to cap unprompted answers at three per room every ten minutes, so an incident channel wouldn't drown. It only ever turned away real questions. Repeated questions are already caught as answered, and venting never gets past the first checks. So there's no cap: every real question gets its own decision.

A
Alice
which runbook section covers rollback?10:00
C
Carol
who owns the rollback script?10:00
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
D
Dana
are the partitions rebuilding in parallel?10:01
T
Triage
Bob owns the rollback script.10:01
H
Harsh
when did we decide on one partition at a time?10:02
T
Triage
No — the rebuild is running one partition at a time to keep load off the primary.10:02

Harsh's when: the topic case above

Four questions in two minutes, each decided on its own.
A
Alice
when is the sync engine cutover planned?10:00
B
Bob
yeah when's the cutover?10:00
T
Triage
Phase 1 cutover is planned for Oct 14.10:01
Two people, the same question, one answer.

Several agents

A room can have several agents. Only one decision runs per room at a time, so a second agent always sees what the first one said. Of the agents that fit, the one whose role fits best answers, and only that one. When two fit almost equally, an overall pick between them breaks the tie. Agents never answer agents, or two of them would talk forever.

C
Carol
when does 0.0.2 launch?10:00
S
Scribe
0.0.2 launches Thursday.10:01
Both could answer. One does.
A
Alice
what's going into the release notes for 0.0.2?10:00
B
Bob
who's on call for launch week?10:00
S
Scribe
0.0.2 release notes will cover: switching between accounts on the same device, the new invite landing page, and improved connection handling during deploys.10:01
T
Triage
Dana is on call for launch week.10:01
Two questions, two agents, each to its own.
A
Alice
where's the figma file for the new onboarding?10:00
P
Pixel
The “Onboarding v3” Figma file is at figma.com/file/onb-v3.10:01
A design question goes to the design agent.
P
Pixel
Should I also add the empty states to the design system page?10:00
An agent's question is for people.

The typing indicator

A person who's about to answer shows as typing. That's the natural signal that someone is on it. It isn't a 👀 reaction, and it isn't an "on it" message. People don't react and then also reply, and an "on it" is noise once the answer arrives.

The hard part is when it appears. This is designed, not shipped yet.

A
Alice
which runbook section covers rollback?10:00
T
Triage
Too early: typing the moment Alice asks.
A
Alice
which runbook section covers rollback?10:00
P
Pixel
S
Scribe
T
Triage
Too early: typing while the checks run.
A
Alice
which runbook section covers rollback?10:00
60 s of quiet
T
Triage
Right: after the wait, once one agent is chosen and starts writing.
A
Alice
is staging down right now?10:00
B
Bob
no idea, haven't checked10:00
T
Triage

Typing stops. The draft was held back

And sometimes it thinks better of it.

The last one is fine. A person starts typing and thinks better of it all the time.

Where it stands

Ninety conversations we tuned on can't say much on their own. So we wrote 166 new ones, labelled them before running anything, and never tuned on them. Relayed against three simpler options:

Spoke when it shouldn'tMissed
Relayed3%26%
One "should I answer?" prompt to a frontier model5%19%
Answer every question33%6%
Never answer0%100%

The restraint held on conversations it had never seen: about 3% wrong posts, the same as on the set we tuned on. The misses didn't hold. Most of them were questions the room's summary could answer, in rooms whose agents weren't set up for that topic. We labelled those as should-answer, and so did a second, independent labeller, who agreed with our labels on 97% of decisions.

That leaves a product question more than a model one: should an agent answer what the room knows, when it's outside what the agent is for? Next is answering that, the edge cases marked should above, the typing indicator, and real rooms.