learning ยท pipeline deep-dive

How the agent grades itself

Nobody rates the AI's replies. Instead, every reply is recorded and then judged the next morning by what the customer did โ€” did they answer within seconds, ask a follow-up, repeat themselves, go quiet, or finish their profile? Those judgements roll up into a short note about each customer that the agent reads before its next reply. LIVE 283 of 283 scored

flowchart TD
    T["๐Ÿ’ฌ One conversation turn
customer message + agent reply"] --> REC["๐Ÿ“ผ Recorded as an EPISODE
what was said, which brain answered,
which tools ran, how long it took"] REC --> STORE[("Stored unjudged
scored: false")] STORE --> NIGHT["๐ŸŒ™ Nightly scorer"] NIGHT --> LOOK{"What did the customer
do NEXT?"} LOOK --> SIG["Signals + a score
are attached to the episode"] SIG --> ROLL["๐Ÿ“Š Per-customer rollup:
average score, patterns,
what worked, what blocked them"] ROLL --> NOTE["โœ๏ธ Written into ONE short paragraph
about this customer"] NOTE --> PROMPT["๐Ÿง  Loaded into the agent's prompt
on their next message"] PROMPT --> T classDef hi fill:#e7f2ec,stroke:#166b4e,color:#182420 class REC,SIG,NOTE,PROMPT hi

The loop. Nothing is judged in the moment โ€” the evidence only exists after the customer reacts, which is why scoring happens overnight.

1 ยท What an episode records

Every time the agent replies, it files a record of that single turn: the customer's message and the reply, which "brain" handled it, which tools ran, the agent's own reasoning steps, how many model calls it took and how long, plus a snapshot of who the customer was at that moment โ€” registered or not, profile status, subscription state, where they were in the conversation.

It is filed unjudged. At the moment of replying there is no evidence about whether the reply was any good.

2 ยท How the score is calculated

This is the heart of it. The scorer finds the customer's next episode and reads the evidence:

flowchart TD
    EP["An unscored episode"] --> NEXT{"Did the customer
come back at all?"} NEXT -->|"NO, and 24h+ have passed"| SILENT["๐Ÿ˜ถ went silent
โˆ’1.0"] NEXT -->|"YES"| SPEED{"How fast?"} SPEED -->|"under 5 minutes"| Q["โšก quick reply
+1.0"] SPEED -->|"under an hour"| TM["๐Ÿ• timely reply
+0.5"] SPEED -->|"slower"| NONE["no timing bonus"] Q --> WHAT TM --> WHAT NONE --> WHAT WHAT{"What did they SAY?"} WHAT -->|"thanks / ok"| POS["๐Ÿ™‚ positive
+0.5"] WHAT -->|"asked a question"| FU["โ“ follow-up question
+1.0"] WHAT -->|"complaint / refusal"| NEG["๐Ÿ˜  negative
โˆ’2.0"] WHAT -->|"repeated themselves"| REP["๐Ÿ” not understood
โˆ’1.0"] POS --> REAL FU --> REAL NEG --> REAL REP --> REAL SILENT --> REAL REAL{"Did anything REAL happen
in the next 24h?"} REAL -->|"finished their profile"| DONE["๐ŸŽฏ profile completed
+3.0"] REAL -->|"moved a step forward"| STEP["๐Ÿ“ˆ step advanced
+2.0"] REAL -->|"nothing"| SUM DONE --> SUM STEP --> SUM SUM["โž• Signals are ADDED together
= the episode's score"] classDef good fill:#e7f2ec,stroke:#166b4e,color:#182420 classDef bad fill:#fdeaea,stroke:#a33,color:#3a1f1f class Q,TM,POS,FU,DONE,STEP,SUM good class SILENT,NEG,REP bad

Signals stack. One turn can collect several โ€” that is how scores range from โˆ’2.5 (slow, annoyed, and repeating themselves) to +4.0 (instant reply, then finished their profile).

SignalWeightWhat it means
Profile completed after+3.0The strongest evidence there is โ€” a real business outcome, not a mood
Step advanced+2.0They moved forward in onboarding
Quick reply (<5 min)+1.0Engaged and paying attention
Follow-up question+1.0The conversation is going somewhere
Timely reply (<1 hour)+0.5Still interested
Positive acknowledgment+0.5"thanks", "ok"
Repeated questionโˆ’1.0We did not answer what they asked
Went silent (24h+)โˆ’1.0We lost them
Negative responseโˆ’2.0The reply actively annoyed them
๐Ÿ“‹
An honest caveat about what this measures. Of the 283 scored episodes, 82 scored purely on "they replied fast" and 46 collected no signal at all. Only a handful carry a real business outcome. So today the score leans more toward how responsive the customer is than how good the reply was. It is a useful signal, not a verdict.

3 ยท From scores to something the agent can use

Individual scores are not much use to an agent mid-conversation. Each night a second pass gathers everything known about one customer and condenses it:

Counted: how many episodes, their average score, and which signals keep recurring โ€” a signal seen twice or more becomes either a "strategy that works" or a "blocker".

Noticed: when they tend to engage, what they tend to ask about, and whether their recent scores are improving, steady, or declining.

Written: all of that becomes one short paragraph in plain language โ€” and that paragraph, not the numbers, is what the agent actually reads.

The agent never sees a score, a signal name, or an episode. It sees something closer to what a colleague would tell you before handing over a call.

4 ยท How it reaches the reply

flowchart LR
    EPS[("๐Ÿ“ผ Episodes
one per turn")] --> SC["๐ŸŒ™ Nightly scorer"] SC --> EPS EPS --> RU["๐ŸŒ™ Nightly rollup
per customer"] RU --> UI[("๐Ÿ“Š One paragraph
per customer")] UI --> IC["Context builder
loads it with 11 other sources"] IC --> TH["Rendered into the prompt as
'Behavioral Insights'"] TH --> RP["โœ๏ธ The reply"] RP -.->|"becomes the next episode"| EPS classDef hi fill:#e7f2ec,stroke:#166b4e,color:#182420 class UI,IC,TH hi

Two nightly passes, one paragraph, then straight into the prompt. Episodes themselves never enter a prompt โ€” there are far too many.

Where this stood, and where it stands

โœ…
It was silently dead for months. The nightly scorer compared two timestamps, one of which carried a timezone and one of which did not. That threw an error on every single episode, the error was caught and discarded per-episode, and the job reported success while scoring zero. Fixed on 6 August: a backlog stuck since March drained in one run, and two signals โ€” "went silent" and "completed their profile after" โ€” fired for the very first time. All 283 episodes are now scored, and every customer profile carries a written note.

The limits, stated plainly