Client

Self-initiated case study

UX research with AI

Role

Researcher & Designer

created

2026

scope

R&D, Prompt design, Evaluation

Client

This is a self-initiated case study

Role

Researcher & Designer

created

2026

scope

R&D, Prompt design, Evaluation

Client

Self-initiated case study

UX research with AI

Role

Researcher & Designer

created

2026

scope

R&D, Prompt design, Evaluation

context

The problem

Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.

I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.

context

The problem

Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.

I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.

context

The problem

Food delivery platforms (like Uber Eats) deal with a high volume of unstructured, subjective feedback. For CX teams, manual analysis is a bottleneck that slows down rapid response to critical signals, e.g., food safety or logistics failures.

I wanted to understand whether an AI classification system could turn that noise into structured, routing-ready tickets - and what it would take to design the logic well.

outcomes

consistently

Classifies

I tested several prompting methods. Few-shot prompting produced the most accurate and decisive results.

Adding nuanced examples helped the model distinguish similar cases and apply the categories more consistently.

ready to work with

Feedback

Each message comes back with a category, priority, and suggested specialist - a structured feedback that is easier to filter, group, and compare.

I tested this in a real workflow, by connecting Zapier and Gemini.

what matters

Designed to push

Simple complaints can receive an automatic acknowledgment without human review. Critical signals are escalated immediately.

how I worked

How I designed the logic

I needed to decide what "good" means for the business.

Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.

I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.

Then I tested three strategies against those criteria.

Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.

how I worked

How I designed the logic

I needed to decide what "good" means for the business.

Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.

I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.

Then I tested three strategies against those criteria.

Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.

how I worked

How I designed the logic

I needed to decide what "good" means for the business.

Accurate + indecisive system = useless.
A system that escalates everything = breaks the capacity.

I started by defining the evaluation criteria: routing accuracy, decisiveness, calibrated escalation, and resistance to over-thinking simple cases.

Then I tested three strategies against those criteria.

Once the prompting strategy was selected, I designed the classification taxonomy and business rules: what counts as low-risk, what triggers escalation, what signals food safety vs logistics vs service quality. Then I connected it to Zapier to simulate end-to-end workflow routing.

Gallery item 1
Gallery item 2
Promp testing results and flow design for the system

highlights

Product & design considerations

highlights

Product & design considerations

highlights

Product & design considerations

Afterthoughts

What I learned

Afterthoughts

What I learned

Afterthoughts

What I learned

What AI
shouldn't do

The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.


What AI
shouldn't do

The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.


What AI
shouldn't do

The prompt is only one part.
Most of the design work is deciding what the model should handle, where it should stop,
what should be escalated, and how to test those decisions.


Create a free website with Framer, the website builder loved by startups, designers and agencies.