Qualitative data is free text: survey answers, support tickets, marketplace reviews, customer messages. A large language model (LLM) can read thousands of such passages a month, tag topics, score sentiment, extract specific requests and produce a summary that used to take several staff days of manual reading. The conditions for trustworthy results: a topic set defined by people, a human-labelled test sample to measure accuracy against, and a process that runs every month rather than once.
Why customer feedback usually goes unread
A 30-person retail business receives roughly 400 support tickets, 200 marketplace reviews and 150 post-purchase survey responses a month. That is 750 passages. Reading and classifying them by hand takes about 25 person-hours, and the result depends on how tired the reader is that day.
The familiar outcome: feedback gets skimmed during a crisis and ignored the rest of the time. The business knows "customers complain about delivery" but not whether that is 12 percent or 40 percent, rising or falling, or which carrier is responsible.
What LLMs can do with qualitative data
Four tasks, in decreasing order of reliability:
- Classification against a fixed topic set: each comment is tagged with one or more topics (delivery, product quality, staff attitude, price, returns). Accuracy is high when the topics are clear and have examples.
- Sentiment scoring: positive, neutral, negative with intensity. Good enough for trend tracking, not for handling individual cases.
- Extracting specifics: the product that failed, an order number, a feature request, a competitor mentioned by name. Works well on structured feedback, less well on abbreviations and slang.
- Summarising and discovering new topics: "what did customers say this month that the topic set does not cover". Useful but needs a human re-read, because the model can merge wrongly or amplify something rare.
The underlying concepts of probability, accuracy and drift apply in full; see machine learning fundamentals for business leaders.
The 5-step process
Step 1: gather the data in one place
Tickets from the helpdesk, reviews from marketplaces, surveys from forms, messages from chat channels. Each row needs at least: text, date, source, and a customer or order id if available. This is usually the slowest step the first time, and the reason from spreadsheets to smart systems recommends standardising data first.
Step 2: define the topic set by hand
Someone who knows the business lists 8 to 15 topics, each with a one-sentence definition and 3 real examples. Do not let the model propose the topic set at this stage; the topic set is the shared language for month-to-month comparison, and changing it constantly destroys comparability.
Step 3: hand-label a test sample
Pick 150 to 200 comments at random, have two people label them independently, and reconcile the differences. This sample is the yardstick: every change to instructions or model version is rerun against it to see whether accuracy went up or down. Skipping this step is the most common reason feedback-analysis projects lose credibility within three months.
Step 4: write instructions, run, measure, revise
The model's instructions include the topic set with definitions and examples, a fixed output format, and a rule for unclear cases (tag "other" rather than guess). Run on the test sample, compare with human labels, read the errors to refine definitions or add examples. With a clear topic set, agreement with humans typically reaches 85 to 92 percent after two or three rounds, higher than two humans agree with each other in many cases.
Step 5: run monthly, publish a report, keep a human sampling
Each month the system tags all new feedback and produces a table: counts and shares by topic, change from the previous month, sentiment by topic, and the most negative comments listed for direct reading. The owner re-reads 30 random tagged comments to catch drift early.
Real costs
For 750 comments a month averaging 80 words:
- Model fees: under 100,000 dong a month with a mid-tier commercial model. Model cost is not the issue at this scale.
- Initial build: data gathering, topic set, test sample, instructions, reporting: 2 to 4 weeks, a few tens of millions of dong depending on how many data sources need connecting.
- Operations: 2 to 4 person-hours a month reading the sample and the report, plus AI system management fees if outsourced.
Against 25 hours of manual reading a month, payback is usually within the first quarter, before counting the value of actually knowing what customers are saying.
5 common mistakes
- Too many or overlapping topics. Thirty topics with blurry boundaries make both humans and the model inconsistent. Start with 10 and merge aggressively.
- No test sample. Without measurement you cannot tell when the system is wrong, and it will be wrong when the provider changes the model version.
- Trusting sentiment scores per case. Sentiment is for group trends; handling an individual customer still needs a human reader.
- Sending unnecessary identifying data to the model. Names, phone numbers and addresses should be masked before sending; the model does not need them to classify.
- Doing it once. The value is in month-to-month comparison. The first month's report is only a baseline.
From analysis to action
A feedback report only matters when each topic has an owner: delivery belongs to operations, product to purchasing, staff attitude to the shift lead. Every month each owner receives their slice plus the 10 most negative verbatim comments. This is where the analysis connects to operations, and the point at which to consider letting an AI agent turn urgent feedback into a ticket for the right person the same day instead of waiting for the monthly report.
Siri9 builds this under AI integration and keeps accuracy stable under AI system management.
Frequently asked questions
How much feedback a month makes this worthwhile?
From about 200 comments upward. Below that, one person reading by hand for a few hours is still more efficient and understands more deeply.
Does the model handle Vietnamese without diacritics, abbreviations and mixed English?
Current commercial models handle unaccented Vietnamese and mixed English reasonably well; industry-specific abbreviations need explaining in the instructions. The test sample will tell you the real accuracy on your data.
Do we need to train a custom model?
No, for most small and mid-sized businesses. Good instructions with examples reach 85 to 92 percent. Custom training is only worth considering at very high volumes; see fine-tuning vs RAG.
How do we start?
Send Siri9 an export of your 100 most recent comments (identifying data masked) and a list of data sources; we run a trial classification and return the results with a scope estimate within 3 working days.
