Building a quality operating system for a 500+ agent healthcare operation
Healthcare / Pharmacy · Diagnosing the root causes of customer and compliance failures, separating agent performance from systemic issues, and building the governance to improve both safely at scale · Apr 2025 — present · via Hugo
A regulated healthcare operation was scaling fast, and quality results could not distinguish individual agent performance from systemic process and policy failures.
Attributing every failure by domain showed 42.5% came from ownership gaps rather than knowledge — so more generic training would not have moved the number.
Built the QA function from zero: a 100-point risk-based rubric with patient-safety auto-zeros, a DSAT attribution model, root-cause categories, severity tiers, coaching governance and escalation thresholds.
500+ agents supported, 2,500+ reviews logged and attributed across 9 sites, and 93% CSAT sustained against a 90% target every period.
Overview. The client is a US digital-first pharmacy service that streamlines prescription fulfilment and delivery across the United States. Hugo supports patients and prescribers across voice, chat and internal task channels — insurance processing, payments and complaints.
The engagement began as a referral-sourced pilot in April 2025, scaling from a 13-person Wave 1 to a 20-person Wave 2 within three weeks, and has since grown into a 500+ agent operation. I set up the QA function from scratch and now manage 3 QA Team Leads and 15 QA analysts across 9 global sites, with more than 2,500 logged quality reviews. CSAT has held at 93% against a 90% target every period since program inception in January 2026.
The framework has gone through four versions as the account matured — weekly-rolling DSAT evaluation and formal DSAT-to-score integration in v2, a fuller DSAT classification model in v3, and refined zero-tolerance and compliance categories in v4. Each version answered something real the operation ran into, not a scheduled rewrite.
- Set up the QA function from scratch for a 500+ agent operation; now manage 3 QA Team Leads and 15 QA analysts, and review calls and chats as part of the rotation.
- Designed the 100-point, 5-dimension QA rubric, including automatic-zero triggers for patient-safety violations.
- Built the domain-based DSAT attribution framework separating true agent error from process/policy context and ambiguous cases.
- Designed the root-cause hypothesis framework classifying failures as knowledge, ownership or interpretation gaps.
- Built the 3-tier (Minor / Major / Critical) consequence ladder tying response severity to patient-safety risk.
- Designed the backoffice QA and coaching-coordination process, including escalation with external vendor Team Leads.
- Ran the weekly QA reporting cadence and the weekly cross-team Roundtable review.
The operation was scaling faster than its quality infrastructure. A 13-person pilot grew to 20+ people within three weeks and ultimately to 500+ agents across nine global sites — and in a regulated pharmacy environment that created two problems at once.
Risk: patient-safety failures could not be treated like ordinary service defects — wrong dosage guidance, a missed adverse-event escalation or mishandled PII is an incident, not a low score. Attribution: a poor customer outcome was not necessarily an agent failure, because policy, process and tooling could produce the same symptom.
The challenge was therefore not simply to improve QA. It was to build a quality system that could distinguish risk from routine defects, individual performance from systemic problems, and symptoms from root causes — while staying consistent at scale.
Transformation approach
- 01Diagnose — built a DSAT attribution model to distinguish agent error from process, policy and tooling issues.
- 02Design — created a 100-point QA framework with zero-tolerance patient-safety triggers and risk-based severity tiers.
- 03Implement — established QA and coaching governance, escalation thresholds, sampling rules and weekly operational reporting.
- 04Measure — introduced post-coaching verification and root-cause tracking to determine whether interventions actually changed behaviour.
The governance principles underneath the framework
- 01Patient safety overrides performance metrics.
- 02Quality is observable and measurable.
- 03Feedback and coaching are Team-Lead-owned responsibilities.
- 04QA reviews and verifies; it does not coach.
- 05Systemic issues require training, not repeated coaching.
- 06Management decisions are based on documented trends, not anecdotes.
Designing the 100-point, 5-dimension rubric
Rather than treating all quality failures equally, I weighted the framework according to customer, compliance and operational risk.
Every reviewed ticket is scored across five weighted dimensions, with defined criteria at four performance levels (0 / 25 / 50 / 100) for each sub-category.
Certain failures — giving medical or dosage advice, missing an AE/PQC escalation, mishandling PII — trigger an automatic zero regardless of how the rest of the interaction scored, so patient-safety risk can never be averaged away by an otherwise-good conversation.
| Dimension | Points | What it covers |
|---|---|---|
| Accuracy & Compliance | 35 | Policy application, safety compliance (AE/PQC/verification), correct workflow execution |
| Detection & Risk Prevention | 20 | Spotting red flags and contradictions; proactive prevention guidance |
| Communication & CX | 20 | Greeting and rapport, structure and clarity, emotional intelligence |
| System Use & Documentation | 15 | Correct use of LDZ/Zendesk, ticket hygiene and notes |
| Professional Conduct & Alignment | 10 | Process adherence, escalation discipline, internal standards |
Key diagnostic insight — not every poor outcome is a frontline problem
I introduced a validation gate that separated agent error from policy, process and tooling failures before those outcomes entered performance scoring. That prevented the QA system from penalising agents for problems they could not control, while creating a separate route for systemic issues to reach policy owners.
Not every DSAT is the agent's fault, and treating them all as if they were erodes trust in the QA process. I built a domain-based attribution model that routes each DSAT to its real cause before it ever touches an agent's record. Only DSATs classified as agent error after this validation gate reach scoring or escalation logic; everything else is logged for systemic trend analysis.
Each agent-error domain carries example-backed sub-drivers so classification stays consistent across analysts rather than becoming a judgment call. A separate, multi-select layer of compliance flags (verification/HIPAA, consent, PII handling, AE/PQC documentation, closing protocol) is tracked independently of the DSAT driver — ticked even on calls the customer was satisfied with — so a compliance gap can never hide behind a good CSAT.
| Domain | Counts against agent? |
|---|---|
| Resolution & Outcome, Issue Recognition, Communication Quality, Tone & Emotional Experience, Expectation & Guidance | Yes — agent error |
| Process / Policy / Tools (cause outside the agent's control) | No — routed to policy owners as context |
| Ambiguous / Unclear (not enough signal) | Excluded until classifiable |
| Root cause category | Share of failures |
|---|---|
| Ownership gap (“didn't want to look”) | 42.5% |
| Knowledge gap (“didn't know where to look”) | 29.5% |
| Interpretation gap (“confused”) | 28.1% |
Diagnosing causes, not symptoms
Ownership gaps — agents closing tickets without investigating, or not wanting to dig further — were the single largest driver of failures, ahead of both knowledge and comprehension issues. That reframed coaching priorities: more training wouldn't have fixed the biggest problem. Accountability coaching would.
| Agent-error DSATs / week | DSAT score (of 15) | Pattern classification |
|---|---|---|
| 0 | 15 | — |
| 1 | 10 | Signal — score impact only |
| 2 | 5 | Pattern — At-Risk status + mandatory coaching |
| 3 | 0 | Escalated pattern — deep dive & performance recommendation |
| 4+ (within 2 weeks) | 0 | Sustained failure — flagged, consequences determined |
Thresholds grounded in real volume
The thresholds aren't arbitrary. Agents typically handle 150–200 tickets a week, so 1–2 agent-error DSATs (roughly 1%) is treated as normal variation, 3 or more (roughly 2%+) as elevated risk, and 4–6 (3–4%) as sustained failure requiring intervention — fair to agents working real volume, while still protecting patients.
| Severity | Example triggers | Consequence path |
|---|---|---|
| Minor | Documentation gaps, unclear wording, mild tone issues | Coaching → retraining → PIP → formal warning |
| Major | Incorrect policy explanation, poor empathy in sensitive cases, skipped workflow steps | Written warning → final warning → PIP → suspension |
| Critical (zero tolerance) | Medical advice given, missed AE/PQC escalation, wrong refill or storage guidance, PII mishandling | Immediate removal from inbox; suspension/termination after investigation |
Separating QA from coaching — and verifying it worked
QA reviews interactions, identifies defects, flags when coaching is required, and verifies outcomes afterward — it never delivers coaching directly. That's Team-Lead-owned: review the findings, deliver standards-based feedback, confirm the agent understood it, set expectations, and log the session. Uncompleted coaching is logged as not done; there's no partial credit for intent.
Verification is what most QA systems skip. After coaching, QA runs a 2-week follow-up window to check one thing only: did behaviour actually change on subsequent tickets. An apology in the moment isn't evidence. If the same defect appears in three or more agents, or spans teams, it stops being an individual coaching problem and escalates to systemic training, measured by defect reduction rather than attendance.
A tiered performance-escalation model
Built a tiered performance-escalation model linking repeated defects to documented root-cause analysis, coaching, improvement plans and formal consequences — with reset rules for sustained improvement, so flags clear after four consecutive weeks at target and the line is visible well before anyone reaches it.
Rolling out the Resolution Checklist and the Roundtable report
Two of the most-used tools I introduced were deliberately simple. The Resolution Checklist targets the framework's biggest DSAT driver — Resolution & Outcome — with a concrete pre-close checklist, turning “did you actually finish the job” from a judgment call into something checkable before a ticket leaves the queue.
The Roundtable document is the weekly operational report structuring the cross-team review: QA score trends, DSAT pattern classifications, confirmed critical events, and coaching/training status in one recurring format, so leadership sees the same shape of data every week. Both are still in active use.
Running the review operation at scale
Sampling scales review intensity to risk: new agents get 10–15 tickets reviewed per week, established agents 5–7, and high-risk workflows are sampled weekly regardless of tenure. Since launch that has produced over 2,500 DSAT reviews and 215 CSAT reviews across 9 sites — Lagos, Cape Town (x2), Johannesburg and Missouri — spanning intake, refills, checkout, shipping and specialty workflows like the Medicare Bridge Program.
- CSAT, every period since inception93% vs 90% target
- Of identified failures traced to ownership gaps42.5%
- Quality reviews logged and attributed2,500+
- Agents supported across 9 global sites500+
- Pilot scale-up without loosening standards13 → 20+ in 3 weeks
- Framework iterations driven by operational learning4
- The outcome wasn't simply a higher QA score. The organisation gained a repeatable mechanism for identifying whether a problem belonged to the agent, the process, the policy or the system — and responding accordingly.
- Root-cause data identified ownership gaps — not knowledge gaps — as the leading driver of failures, redirecting coaching investment to where it mattered.
- Built a defensible, healthcare-appropriate safety net: zero-tolerance triggers isolate true patient-safety risk from the much larger pool of ordinary, coachable mistakes.
- Escalation thresholds designed around real ticket volume (150–200/week), with a genuine path back to a clean record for agents who improve.
More case studies
Full index- 02EdTech learning marketplaceReducing recurring customer dissatisfaction through frontline capability84% drop in issue-recognition DSATs
- 03Conversational commerce SaaS platformDesigning and scaling a new customer-support channel90% onboarding pass rate vs 60% target
- 04Clinical decision-support platformRedesigning support operations to achieve 100% SLA compliance100% SLA compliance, Q1–Q3