How StatQuestions works

From raw customer text to a prioritized fix list - what happens at each step

Your data in → prioritized answers out

Raw Data InputSupport emailsVoice recordingsSurvey responsesChat transcriptsUnstructured textAI ProcessingNLP extractionIssue detectionSentiment analysisConfidence scoringStatistical proofActionable InsightsCategorized issuesRoot causes rankedImpact by departmentχ² statistical proofIntervention guidance
1

Export from your source system

10–30 min

Pull a CSV or XLSX from your helpdesk, contact center, or shared inbox. Zendesk, ServiceNow, Freshdesk, Salesforce Service Cloud, Intercom, Genesys, Five9, Outlook, Front - anything you can export works. You need ticket ID, the text body, a date field, and ideally a case owner or business unit column. Minimum: 5,000 rows. Ideal: 20,000+.

2

Upload and map your columns

2 min

Drop the file in. StatQuestions detects your column structure and asks you to confirm the mapping: which column is the text body, which is the date, which is the business unit. No template required.

3

Classification runs automatically

5–20 min depending on volume

Issues are classified against a taxonomy aligned to your industry. Each row gets an issue type, a business unit assignment, and a sentiment score. Duplicate contacts are identified and collapsed - a pattern that appeared 47 times counts as 47 cases, not one. This is what makes the rework math accurate.

4

Review the issue dashboard

15 min

Issues are ranked by frequency and by which business unit generates the highest rate. Each issue type shows repeat contact count, escalation rate, and sentiment trend. Click any issue to drill into verbatim examples.

5

Run the waste audit (Tier 2)

5 min

Input your average handle time and fully-loaded labor rate. StatQuestions calculates rework hours and cost per issue type. Chronic patterns (issues that recur every week) are flagged separately from transient spikes. This is the number you bring to the budget conversation.

6

Drill into root cause (Tier 2–3)

10 min

Select any high-cost issue. The 5 Whys runs automatically against the text evidence, grounded-ds-control in the most statistically frequent patterns - not a random sample. Each why level shows supporting verbatims. You get a causal chain, not an AI guess.

7

Connect outcome data for causal ranking (Tier 3)

15 min

Upload a churn file, refund export, or CSAT time series. StatQuestions joins it to the ticket data and identifies which failure patterns statistically predict bad outcomes. Interventions are then ranked by downstream dollar impact - fix the thing that reduces churn the most, not the thing that generates the most tickets.

8

Export and charter the project

5 min

Export the waste audit, root cause analysis, or full DMAIC report. Or click "Start Project" on any issue to create a tracked improvement action in the Action Backlog with the impact estimate pre-filled.

Common questions before starting

  • What if my data has PII? Remove or anonymize customer name and email before uploading. The text body and metadata (date, BU, case ID) are what the analysis runs on - personal identifiers add no analytical value.
  • How much history should I export? 90 days minimum. 12 months ideal - it's what separates chronic patterns (every week for a year) from seasonal spikes (Q4 only).
  • Do I need a business unit column? Strongly recommended. Without BU attribution, you get total company patterns. With BU, you get which team is causing which failure - that's what makes the root cause actionable.
  • What's the minimum viable dataset? 5,000 rows with at least a text body and date field. Below that, patterns aren't statistically separable from noise. Above 20,000 is where the analysis becomes highly reliable.
  • Can I combine data from multiple systems? Yes. Export from each system separately (e.g., Zendesk tickets + Salesforce cases), add a "source" column to each, then upload as one combined file. The issue taxonomy normalizes across sources.
  • What structured data makes the analysis causal? Handle time per ticket, refund/credit amounts, account tier/ARR, SLA breach flag, agent ID, or churn date. Text alone gives you thematic patterns. Text plus any of these gives you cost and causation.