Case 02 · EdTech learning marketplace

Diagnosing and reducing recurring customer dissatisfaction

EdTech · Reducing recurring customer dissatisfaction through frontline capability · 2021 — 2025 · via Hugo

100+
Agents trained across ~20 cohorts
8
QA analysts managed
97%
CSAT sustained through growth
84%
Drop in issue-recognition DSATs (27 → 4)
The business problem

Quality was declining at scale while volume grew — CSAT slipped from 92% to 90% to 89% across 2024 with no clear cause.

The diagnosis

62% of dissatisfaction was policy-related and driven by incomplete resolution and missed issue recognition, not agent effort or tone.

The intervention

Rebuilt the quality system around issue recognition and resolution completeness, with calibration, onboarding redesign, targeted coaching and workforce planning behind it.

The business impact

84% drop in issue-recognition dissatisfaction, 97% CSAT, 35% fewer backlog occurrences while volume grew 19%, and roughly 25% higher productivity.

Overview. The client is an online marketplace of live classes for kids and teens. I led quality assurance, training, recruiting and workforce planning for Hugo's support team across multiple years — from training the first cohort through building the QA rubric, calibration process and a library of recurring performance experiments, and later into leadership development, a tiered support structure (T1/T2/T3), a high-touch Concierge program, and the account's early adoption of AI-assisted workflows.

The team grew from ~48 agents to 100+ trained across roughly 20 cohorts. This is the longest-running and broadest-scope engagement in my portfolio, and performance over that span wasn't a straight line up — it included a real quality dip in 2024 that I helped diagnose and reverse heading into 2025.

My role
  • Built training decks and curricula from scratch, edited Lessonly course content, and drove Knowledge Base and policy improvements.
  • Trained 100+ agents across ~20 cohorts using a classroom-style, case-study-driven onboarding model.
  • Managed 8 QA analysts — owning a six-dimension rubric, scoring standards, QA sheet and calibration end to end.
  • Built personality-based recruiting: Big Five data used to predetermine the profile a project needed, then screening against it.
  • Presented Weekly Business Reviews to the client and led the design of recurring performance experiments.
  • Owned capacity projection and monthly schedule-building for a rotating, three-shift team.
  • Created and launched mentorship programs for Team Leads and for QAs/Coaches.
The challenge

The recurring quality issues were consistent across quarters: agents sacrificing accuracy for speed, missing the deeper issue behind a surface-level question, inconsistent brand voice, and at times misalignment between Hugo's internal QA scoring and the client's own reviews.

That infrastructure was tested in 2024. CSAT opened the year at target (92%), drifted to 90% through Q2–Q3, then hit 89% in Q4 — a 4-point decline versus 2023, driven mainly by policy-related DSATs (62% of the total, concentrated in refunds) and a stubborn 'Incomplete Resolution' category as agents struggled to adapt to a wave of product and policy changes. That dip triggered the account's most data-driven intervention yet.

What I did

Building the QA framework, rubric and calibration process

I built the QA rubric and rating system from scratch, scoring every reviewed ticket across six dimensions — Accuracy of Response, Due Diligence, Brand Customization, Personalization, Thoroughness and Empathy — averaged into a single composite Klaus score.

To diagnose why a score moved, I maintained two companion taxonomies: a DSAT taxonomy categorising every 1–3-star rating (Incomplete Resolution, OS Policy, Issue Recognition, Lack of Empathy, User Error, Delayed Response, System Error) and a CSAT taxonomy doing the same for 4–5-star ratings. Together they turned “quality went down” into “quality went down because of X, in Y% of tickets.”

Calibration ran weekly or bi-weekly: a small set of representative tickets scored independently, then discussed until the group reached consensus, with deviations and rubric ambiguities logged to feed the next revision. Alignment with the client's own QA went from 16% at the first session to 50% by the second, and weekly calibration later pushed policy-ticket rating accuracy to a sustained 98–100%.

Training and onboarding 100+ agents

I trained the first batch of 40+ agents using a classroom-style model where every ticket was treated as a case study: debated by the group, then approved by a QA or coach. New agents spent their first 3–4 weeks on email-only queues so they could focus on issue recognition and product knowledge before being measured on speed, and worked through knowledge base articles with written essay assignments to build real comprehension.

Data showed that agents who tested strongest on logical and rapid problem-solving — two of seven screening parameters — were the best predictors of on-the-job success, and that agents with a naturally skeptical disposition outperformed natural optimists on issue recognition. Once onboarding coaches reviewed roughly 75% of a new agent's responses before send, new-agent CSAT in the first two weeks jumped to 97–100%, up from 91–95%, with no negative impact on the wider team.

The experiment library

  • 01Stroop Experiment — diagnostic and conditioning method to build cognitive flexibility and root-cause issue recognition.
  • 02Peer Reviews — weekly cross-agent ticket review against the QA rubric to surface blind spots.
  • 03Workgroups — groups of 5 agents with a peer moderator plus a dedicated QA/TL for real-time shift support.
  • 04Service Mechanics Blueprint — visual decision trees for common issue types.
  • 05Habit Forming Principles — a cue-routine-reward loop for brand voice and conversation mechanics.
  • 06Improv Games — role-play exercises run in daily stand-ups to build empathy and connection.
  • 07Floating Notes — shortcut-based glossary of on-brand personalisation language.
  • 08Hall of Fame — recognition track spotlighting the behaviours the rubric rewards.

Personality-based recruiting and project fit

Instead of screening candidates against one generic bar, I used Big Five assessment data to predetermine the personality profile a given project or bucket actually needed — a high-touch Concierge role calls for a different trait mix than a high-volume T1 queue — then screened and placed candidates against that specific profile. Placement decisions were grounded in measured traits, not gut feel.

Capacity projection and workforce management

Each month I pulled four weeks of ticket data from Intercom across every bucket, then broke it down by hour and weekday with pivot tables to find the actual demand curve across chat, email and Facebook channels.

That curve fed a headcount model across a three-shift structure (roughly 7am–4pm, 3pm–midnight, 11pm–8am) with explicit assumptions for shrinkage (20%) and productivity (80%), so numbers reflected real available capacity. I built the monthly shift schedule by hand against those targets while balancing leave, swaps and overtime — and documented the whole process plus its governance rules into a reusable build guide.

Results — audit summary
MeasureResult
Agents scoring ≥50% on Klaus, Q4 202113 agents, up 44% vs prior quarter
QA calibration accuracy (internal ↔ client QA)16% alignment → 50% by second session
Concierge program, early results85% of conversations rated 5-star post-launch
Ticket handling efficiency, Q1 202535% fewer backlog occurrences while volume grew 19%; productivity up ~25% to >2 tickets/hour
T3 knowledge & escalations, Q1 202528 workflows handed off, 36 docs completed, 413 miscategorised tickets restored, 96% peak weekly resolution
Stroop Experiment, 2025 relaunch84% drop in issue-recognition DSATs (27 → 4); 83% drop in incomplete-resolution DSATs

Leadership development

I created and launched two parallel mentorship tracks — one for Team Leads, one for QAs/Coaches — where each group designed its own approach rather than following a prescribed format. A mid-program survey showed mentees energised about learning from mentors and each other, with the gap between sessions flagged as too long. I separately surveyed the teams that mentees themselves managed, to see whether mentorship was changing how they led; results were largely positive, with delegation and time management the clearest areas still to develop. Both rounds of feedback fed a redesign and a reusable playbook.

Results — audit summary
  • Issue-recognition DSATs after Stroop relaunch27 → 4
  • New-agent CSAT in first two weeks91–95% → 97–100%
  • Internal ↔ client QA alignment16% → 50%
  • Q1 2025 backlog occurrences−35% on +19% volume
  • Sustained CSAT around 97% in the account's early years while repeatedly onboarding new cohorts — new agents averaged 96–100% CSAT in their first quarter.
  • Diagnosed and reversed a real 2024 quality dip with data rather than generic retraining.
  • Built infrastructure — rubric, taxonomies, calibration, experiment library, workforce build guide — that outlived any single cohort.

More case studies

Full index