RILLA · MODEL BEHAVIOR · CASE STUDY · 2026

Rick Copilot & Autopilot

I designed Rilla’s AI coaching system, first helping managers write feedback, then testing what happens when the model acts without them.

The core design question: when has a model earned the right to speak in someone else’s voice?

Role
Product Designer
Team
1 designer, 2 engineers
Timeline
Copilot, Feb–Mar 2026. Autopilot, Mar 2026 to present
Shipped
Copilot launched on stage at Rilla Masters 2026. Autopilot is in a live experiment.
WHO WRITES THE FEEDBACK BEFORE The manager, alone Writes every comment by hand. About 10% of calls get one. RICK COPILOT Rick drafts, manager sends Suggestions appear in the workflow. The manager reads, edits, and sends. RICK AUTOPILOT Rick sends on its own Only when a moment clears the bar, and only with the manager’s consent. manager in control model acts on its own
The problem

Great coaching doesn’t scale

Managers have dozens of reps and limited time. Only about 10% of recorded calls ever receive coaching.

The opportunity was to use AI to close that gap without losing what makes human feedback valuable.

RECORDED CALLS THAT GET COACHING ~10% of recorded calls ever get a coaching comment coached never reviewed
Rick Copilot

The obvious AI interface was the wrong one

Copilot started as a chat sidebar, modeled after tools like Cursor.

Sales managers didn’t want to prompt an AI. They wanted to see good feedback, verify it quickly, and move on.

So I cut the chat.

WHAT THE MANAGER HAS TO DO FIRST FIRST CONCEPT · CHAT SIDEBAR ask Rick… A thread and a prompt box. The manager has to ask first. SHIPPED · THE COMMENT IS ALREADY THERE The comment is attached to the moment. Read, edit, send.

Put the AI directly in the workflow

Copilot surfaces high-value coaching moments with a comment ready to send.

The manager reads, edits if needed, and sends. The interaction itself becomes the prompt.

AI proposes. The manager decides.

One comment, three layers of context

A good coaching comment needs to know three things: what good coaching looks like, how this company sells, and how this manager speaks.

THREE LAYERS, ONE COMMENT GENERAL What good coaching is Ships with the product by default, trained on sales-coaching books. COMPANY This org’s sales process Generated from materials the org uploads: scripts, pricing, process. MANAGER This manager’s voice Learned from the comments a manager writes, and their Rick Copilot edits. GENERATION REPORTS BACK gaps conflicts sections cost sounds like a model sounds like your manager

Correct wasn’t enough. It had to sound human.

In ~42 manager interviews, the recurring problem wasn’t substance. It was voice.

Managers shortened comments, removed formality, added warmth, and rewrote openings. So we built those edits back into the manager knowledge base.

SAME FEEDBACK, TWO VOICES GENERAL BASELINE ONLY “I noticed that during this portion of the conversation there may have been an opportunity to more effectively establish rapport prior to discussing pricing. Consider leading with discovery questions.” + MANAGER KNOWLEDGE BASE “You went to price too fast. Ask what’s driving the decision first. You’ll close more.” Shorter. No hedging. Says the thing. It goes out under a name reps already trust. Both carry the same judgment. Only one of them gets opened.

Behavior became the feedback loop

Thumbs up and down were used by less than 1% of managers, so I stopped asking for feedback explicitly.

Instead, every interaction became a signal.

WHAT THE MANAGER DOES IS THE SIGNAL suggestion streams in 3 fields, in parallel manager reads it 4s dwell before it counts VIEWED_AND_SKIPPED weak negative signal SUBMITTED good output SUBMITTED_EDITED learns the manager’s voice copilot/engagements deduped per moment trains the knowledge base that writes the next one thumbs up / down used by <1%, dropped

The manager trains the model simply by doing their job.

Copilot moved most coaching to AI

40% → 70–80%
Share of coaching comments sent that were AI-written
~157K
AI-written comments per month
Rick Autopilot

Then we removed the human approval step

Copilot showed that the model could write useful feedback with very little editing.

Autopilot asked the harder question: should it be allowed to send that feedback on its own?

THE STEP AUTOPILOT REPLACES COPILOT coaching moment found in the call Rick drafts in the manager’s voice manager approves reads, edits, sends comment sent by the manager AUTOPILOT coaching moment found in the call Rick drafts in the manager’s voice confidence gate the model decides comment sent automatically replaced by
The trust experiment

Whose name should the AI speak under?

Before Autopilot, about 86% of manager comments got read, but only about 10% of calls received one.

I tested whether AI could preserve that trust while dramatically increasing coverage.

THREE COHORTS, ONE VARIABLE: WHOSE NAME IS ON IT COHORT A The manager’s name No disclosure. Reads like a comment they wrote. READ MOST COHORT B Labeled as Rick AI The same comment, sent openly as the model. READ LESS COHORT C Nothing sent Control. Whatever their manager wrote, if anything. BASELINE WHAT ACTUALLY HAPPENED Cohort A led, then declined. Managers told reps about the feature, and Autopilot sent more often than a manager would. Once reps realized the comments were automated, the advantage disappeared.

Borrowed trust worked, until people noticed

Manager-named comments were read more at first.

Then reps realized the comments were automated, and the advantage declined.

So we stopped borrowing the name and ran it again, with every AI comment labeled as Rick AI.

READ RATE BY COMMENT SOURCE · DAILY, LABELED COHORT 0% 20% 40% 60% 80% 100% 06/21 06/27 07/03 07/09 07/15 07/20 COMMENT SOURCE MEAN READ RATE Rick AI Autopilot 73.3% Manager Autopilot 72.2% Human 62.4%

The two AI lines were nearly identical: 73% read when labeled as Rick AI, 72% under the manager’s name. Labeling the comment as AI cost nothing in read rate.

Manager-written comments averaged 62% in the same cohort. That isn’t a verdict on human coaching. Those reps were receiving AI comments too, and attention per comment drops as volume rises.

The model could imitate a manager’s voice. It turned out not to need to.

Designing for trust

More feedback wasn’t always better

A bad automated comment could damage trust in the entire coaching channel.

So I designed Autopilot around restraint, not volume.

HOW AUTOPILOT DECIDES call recorded MANAGER CONSENT manager opted in? yes no nothing sent HUMAN FIRST manager hasn’t commented? yes no manager’s stands CONFIDENCE GATED a moment clears the bar? yes no silence PERSONALIZED match what the rep is ready for ONE COMMENT MAX send one comment

Sometimes the right AI behavior is silence

Autopilot only acts when confidence is high enough.

No qualifying moment means no comment.

SAME BAR, TWO CALLS · SIMULATED SCORES CALL A · 5 CANDIDATE MOMENTS the bar · 0.80 0.22 0.41 0.58 0.70 0.88 0.00 1.00 confidence → 0.88 clears the bar. One comment is sent. CALL B · 5 CANDIDATE MOMENTS the bar · 0.80 0.18 0.33 0.49 0.62 0.74 0.00 1.00 confidence → the best is 0.74. Nothing is sent.
Coverage matters. Trust matters more.
Outcome

From AI-assisted to autonomous coaching

Copilot is now live for all managers, with AI writing the majority of coaching feedback. That adoption gave us the foundation to test Autopilot with select organizations.

40% → 70–80%
Share of coaching comments sent that were AI-written
~157K / month
AI-written comments

Autopilot is still a live experiment. So far, it changed our original hypothesis: the model didn’t need to pretend to be the manager for people to read its feedback.

The system is now moving toward Rick AI speaking as Rick AI, labeled every time, with manager oversight over what it sends and where it is allowed to act.

The question shifted from:

Can AI coach like a manager?

to:

How do we let AI coach without replacing the manager?