
BackContact
Classifying feedback
wwiitthh AAII
wwiitthh AAII
Classifying feedback
wwiitthh AAII
wwiitthh AAII
Client
Self-initiated case study
UX research with AI
Role
Researcher & Designer
created
2026
scope
R&D, Prompt design, Evaluation
Client
This is a self-initiated case study
Role
Researcher & Designer
created
2026
scope
R&D, Prompt design, Evaluation
Client
Self-initiated case study
UX research with AI
Role
Researcher & Designer
created
2026
scope
R&D, Prompt design, Evaluation
Classifying feedback
wwiitthh AAII
wwiitthh AAII
context
The problem
Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.
I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.
context
The problem
Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.
I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.
context
The problem
Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.
I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.
outcomes
consistently
Classifies
I tested several prompting methods. Few-shot prompting produced the most accurate and decisive results.
Adding nuanced examples helped the model distinguish similar cases and apply the categories more consistently.
ready to work with
Feedback
Each message comes back with a category, priority, and suggested specialist - a structured feedback that is easier to filter, group, and compare.
I tested this in a real workflow, by connecting Zapier and Gemini.
what matters
Designed to push
Simple complaints can receive an automatic acknowledgment without human review. Critical signals are escalated immediately.
how I worked
How I designed the logic
I needed to decide what "good" means for the business.
Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.
I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.
Then I tested three strategies against those criteria.
Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.
how I worked
How I designed the logic
I needed to decide what "good" means for the business.
Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.
I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.
Then I tested three strategies against those criteria.
Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.
how I worked
How I designed the logic
I needed to decide what "good" means for the business.
Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.
I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.
Then I tested three strategies against those criteria.
Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.


Promp testing results and flow design for the system
highlights
Product & design considerations
highlights
Product & design considerations
highlights
Product & design considerations
Afterthoughts
What I learned
Afterthoughts
What I learned
Afterthoughts
What I learned
What AI
shouldn't do
The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.
What AI
shouldn't do
The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.
What AI
shouldn't do
The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.
BackContact
