How the agent grades itself
Nobody rates the AI's replies. Instead, every reply is recorded and then judged the next morning by what the customer did โ did they answer within seconds, ask a follow-up, repeat themselves, go quiet, or finish their profile? Those judgements roll up into a short note about each customer that the agent reads before its next reply. LIVE 283 of 283 scored
flowchart TD
T["๐ฌ One conversation turn
customer message + agent reply"] --> REC["๐ผ Recorded as an EPISODE
what was said, which brain answered,
which tools ran, how long it took"]
REC --> STORE[("Stored unjudged
scored: false")]
STORE --> NIGHT["๐ Nightly scorer"]
NIGHT --> LOOK{"What did the customer
do NEXT?"}
LOOK --> SIG["Signals + a score
are attached to the episode"]
SIG --> ROLL["๐ Per-customer rollup:
average score, patterns,
what worked, what blocked them"]
ROLL --> NOTE["โ๏ธ Written into ONE short paragraph
about this customer"]
NOTE --> PROMPT["๐ง Loaded into the agent's prompt
on their next message"]
PROMPT --> T
classDef hi fill:#e7f2ec,stroke:#166b4e,color:#182420
class REC,SIG,NOTE,PROMPT hi
The loop. Nothing is judged in the moment โ the evidence only exists after the customer reacts, which is why scoring happens overnight.
1 ยท What an episode records
Every time the agent replies, it files a record of that single turn: the customer's message and the reply, which "brain" handled it, which tools ran, the agent's own reasoning steps, how many model calls it took and how long, plus a snapshot of who the customer was at that moment โ registered or not, profile status, subscription state, where they were in the conversation.
It is filed unjudged. At the moment of replying there is no evidence about whether the reply was any good.
2 ยท How the score is calculated
This is the heart of it. The scorer finds the customer's next episode and reads the evidence:
flowchart TD
EP["An unscored episode"] --> NEXT{"Did the customer
come back at all?"}
NEXT -->|"NO, and 24h+ have passed"| SILENT["๐ถ went silent
โ1.0"]
NEXT -->|"YES"| SPEED{"How fast?"}
SPEED -->|"under 5 minutes"| Q["โก quick reply
+1.0"]
SPEED -->|"under an hour"| TM["๐ timely reply
+0.5"]
SPEED -->|"slower"| NONE["no timing bonus"]
Q --> WHAT
TM --> WHAT
NONE --> WHAT
WHAT{"What did they SAY?"}
WHAT -->|"thanks / ok"| POS["๐ positive
+0.5"]
WHAT -->|"asked a question"| FU["โ follow-up question
+1.0"]
WHAT -->|"complaint / refusal"| NEG["๐ negative
โ2.0"]
WHAT -->|"repeated themselves"| REP["๐ not understood
โ1.0"]
POS --> REAL
FU --> REAL
NEG --> REAL
REP --> REAL
SILENT --> REAL
REAL{"Did anything REAL happen
in the next 24h?"}
REAL -->|"finished their profile"| DONE["๐ฏ profile completed
+3.0"]
REAL -->|"moved a step forward"| STEP["๐ step advanced
+2.0"]
REAL -->|"nothing"| SUM
DONE --> SUM
STEP --> SUM
SUM["โ Signals are ADDED together
= the episode's score"]
classDef good fill:#e7f2ec,stroke:#166b4e,color:#182420
classDef bad fill:#fdeaea,stroke:#a33,color:#3a1f1f
class Q,TM,POS,FU,DONE,STEP,SUM good
class SILENT,NEG,REP bad
Signals stack. One turn can collect several โ that is how scores range from โ2.5 (slow, annoyed, and repeating themselves) to +4.0 (instant reply, then finished their profile).
| Signal | Weight | What it means |
|---|---|---|
| Profile completed after | +3.0 | The strongest evidence there is โ a real business outcome, not a mood |
| Step advanced | +2.0 | They moved forward in onboarding |
| Quick reply (<5 min) | +1.0 | Engaged and paying attention |
| Follow-up question | +1.0 | The conversation is going somewhere |
| Timely reply (<1 hour) | +0.5 | Still interested |
| Positive acknowledgment | +0.5 | "thanks", "ok" |
| Repeated question | โ1.0 | We did not answer what they asked |
| Went silent (24h+) | โ1.0 | We lost them |
| Negative response | โ2.0 | The reply actively annoyed them |
3 ยท From scores to something the agent can use
Individual scores are not much use to an agent mid-conversation. Each night a second pass gathers everything known about one customer and condenses it:
Counted: how many episodes, their average score, and which signals keep recurring โ a signal seen twice or more becomes either a "strategy that works" or a "blocker".
Noticed: when they tend to engage, what they tend to ask about, and whether their recent scores are improving, steady, or declining.
Written: all of that becomes one short paragraph in plain language โ and that paragraph, not the numbers, is what the agent actually reads.
The agent never sees a score, a signal name, or an episode. It sees something closer to what a colleague would tell you before handing over a call.
4 ยท How it reaches the reply
flowchart LR
EPS[("๐ผ Episodes
one per turn")] --> SC["๐ Nightly scorer"]
SC --> EPS
EPS --> RU["๐ Nightly rollup
per customer"]
RU --> UI[("๐ One paragraph
per customer")]
UI --> IC["Context builder
loads it with 11 other sources"]
IC --> TH["Rendered into the prompt as
'Behavioral Insights'"]
TH --> RP["โ๏ธ The reply"]
RP -.->|"becomes the next episode"| EPS
classDef hi fill:#e7f2ec,stroke:#166b4e,color:#182420
class UI,IC,TH hi
Two nightly passes, one paragraph, then straight into the prompt. Episodes themselves never enter a prompt โ there are far too many.
Where this stood, and where it stands
The limits, stated plainly
- It measures responsiveness more than quality. A customer who replies fast to a mediocre answer scores better than a thoughtful one who takes an hour. See the caveat above.
- It is a next-day judgement. Nothing is graded live, and nothing changes the reply that is happening right now.
- "Repeated themselves" is a guess. It compares message text, so a customer who rephrases will not always be caught.
- The agent reads a paragraph, not the data. If that paragraph is wrong, the numbers behind it being right does not help.