Testing for Emotion and How Desirability Studies Work

There's a reason some products feel instantly trustworthy and others don't. It's not an accident, and it's testable.

Testing for Emotion and How Desirability Studies Work - Clay

A product can work perfectly and still lose to a worse one that simply feels better. Users forgive sluggish load times and clunky navigation far more often than they forgive a design that leaves them cold. That gap between "functions correctly" and "people actually want to use it" is exactly what desirability testing is built to measure.

Most teams run usability tests as a matter of course. Fewer run desirability tests, and the ones who skip them tend to assume a product's emotional pull will sort itself out once the functional problems are fixed. It rarely does. Desirability testing gives you a structured way to find out how a product makes people feel: whether it reads as trustworthy, cheap, modern, or forgettable, and which specific design choices drive that impression.

Key Takeaways

  • Desirability testing measures emotional response to design (how a product looks and feels), while usability testing measures whether the product actually works.
  • The Microsoft Desirability Toolkit, a set of 118 product reaction words developed at Microsoft in 2002, remains the most widely used instrument for this kind of research.
  • Visual appeal registers almost instantly. Research from Carleton University found people form an opinion about a web page's visual appeal in as little as 50 milliseconds.
  • Desirability, feasibility, and viability are three separate lenses, and a product idea needs all three to succeed, not just one.
  • The biggest risk in desirability testing isn't running it badly. It's letting a polished, well-liked design distract from a product that doesn't actually solve the user's problem.
  • Desirability testing works best early and often: at concept validation, through design iterations, and again before launch.

What Is Desirability Testing?

Desirability testing evaluates how emotionally appealing a product is to the people who use it. Rather than asking, "Does this work?" it asks, "Does this feel right?" Researchers gather that data through qualitative methods like interviews, surveys, and focus groups, along with a handful of methods built specifically for measuring emotional reaction, which we'll get into below.

What Is Desirability Testing?

What Is Desirability Testing?

The output isn't a pass/fail score. It's a picture of which design elements build trust or excitement, and which ones quietly push people away before they've even tried the product.

Why Desirability Testing Matters

Usability testing checks whether a product functions: can someone complete a task, find a button, finish a checkout flow.

Desirability testing checks something adjacent but different: how the interface's look and feel (its color, typography, imagery, and overall polish) make someone feel while they're doing it.

The distinction matters because first impressions form faster than most teams assume. In a widely cited study out of Carleton University, researchers found that people can judge a web page's visual appeal in as little as 50 milliseconds, and that snap judgment correlates closely with ratings made after much longer exposure. A visitor has often decided whether your product looks credible before they've read a single word of copy.

That's the gap desirability testing closes. It won't tell you whether your checkout flow is broken. It will tell you whether your product looks like something worth trusting with a credit card in the first place, and both questions need answering before launch.

Desirability, Feasibility, and Viability

Product teams often use a three-lens framework to evaluate whether an idea is worth building: desirability, feasibility, and viability. Each answers a different question, and a strong idea needs to hold up under all three.

Desirability asks whether the product meets what users actually want, measured through research, testing, and direct feedback. It's the lens most people mean when they talk about "good UX."

Feasibility asks whether the product is buildable with the technology, time, and team you actually have. A brilliant concept that requires infrastructure you don't have access to isn't feasible yet, whatever its emotional appeal.

Rinsing vs Desirability vs Feasibility vs Viability

Rinsing vs Desirability vs Feasibility vs Viability

Viability asks whether the product can sustain itself: enough demand, a workable cost structure, room to compete. An idea people love, and your team can build, still needs to make business sense over time.

These three don't compete with each other. A product only clears the bar when it's desirable, buildable, and financially sound at once, which is why desirability testing should sit alongside technical and business validation rather than after it.

Preference Testing and A/B Testing

Two methods get confused often enough that they're worth separating clearly.

Preference testing shows people two or more design options, side by side, and asks which they prefer and why. It's a direct way to gauge reaction to colors, layouts, or branding before anything gets built, and it's where most desirability insight comes from early in a project.

A/B testing works differently. It ships two live variations of a real feature to real users and measures what they do: click-through rate, conversion, time on task. It answers "which version performs better," not "which version do people like more," and the two answers don't always agree. A design can win on stated preference and lose on behavior, or the reverse.

A/B Testing Example by Clay Global

A/B testing example

Used together, they cover both sides of the picture. Preference testing shapes the design direction before development starts. A/B testing confirms or corrects that direction once it's live and being used for real.

Writing Better Preference Questions

Specific questions produce more useful answers than vague ones. Instead of asking "which icon do you prefer," ask "which icon best represents an incoming message" or "which of these feels more reliable." The added context gives people a real basis for choosing instead of a coin flip.

Neutral wording matters just as much. A leading question nudges people toward the answer they think you want to hear, which defeats the point of testing in the first place. A few examples that hold up well in practice:

  • Which homepage version feels more trustworthy?
  • Which font is easiest to read at a glance?
  • Which notification design grabs your attention first?
  • Would you rather access this feature from the homepage or a dropdown menu?

Well-built questions remove guesswork on both ends: the participant knows exactly what they're evaluating, and the researcher gets a clean signal instead of noise.

The Psychology Behind Desirability

Desirability is, at root, an emotional connection. People gravitate toward products that reflect something about their identity, values, or aspirations, shaped by personal history, culture, and the story a brand tells about itself.

Visual design does real work here. Color, shape, texture, and a smooth overall experience combine to make a product feel more memorable and more aligned with what a user already believes about themselves. That combination of emotional resonance and design craft is what drives preference and, over time, loyalty.

Color Wheel Illustration by Clay Global

Color Wheel Illustration

We've designed for 100M+ Speedtest users, Discover cardholders, and Snapchat's AR commerce experience. Scale and complexity don't intimidate us. Tell us about your product.

Methods for Desirability Testing

Emotional response doesn't show up on its own. Each method below captures it differently: some ask people to put a feeling into words, some ask them to rate it on a scale, and newer ones read it straight off the body before a person has consciously registered it.

Picking the right one comes down to how much nuance you need and how much time and budget you have to get it.

The Microsoft Desirability Toolkit

The most widely used instrument for this kind of research is the Microsoft Desirability Toolkit, also known as Product Reaction Cards, developed by researchers Joey Benedek and Trish Miner in 2002.

The method is simple to run: after a usability session, participants are handed a set of 118 words, each printed on its own card, and asked to pick the five that best describe the product. The set is deliberately skewed, 60% positive terms and 40% negative, so participants have a genuine, unbiased set of options rather than a leading list.

Product Reaction Cards Example

Product Reaction Cards Example

What makes the method valuable isn't the word count itself. It's what happens next: the moderator asks participants to explain why they picked each word, which turns a simple sorting exercise into a rich qualitative conversation about what the product actually communicates.

Semantic Differential Scales

Semantic differential scales ask participants to rate a product along a spectrum between two opposite adjectives, such as "attractive" and "unattractive" or "modern" and "dated."

Because the scale is bounded and consistent across participants, it surfaces subtle differences in perception that open-ended feedback often misses, and the results are easy to aggregate across a larger sample.

Card Sorting

Card sorting asks participants to organize a set of concepts or features into categories that make sense to them. It's typically associated with information architecture, but applied to desirability research, it reveals which qualities or features people mentally group together, and which ones don't fit the story a brand is trying to tell.

Interviews and Focus Groups

Structured conversations, one-on-one or in a group, give researchers room to dig past a surface reaction and into the motivations behind it. They're slower and more resource-intensive than card-based methods, but they're often the only way to surface the "why" behind an emotional response that a rating scale can only measure the size of.

Focus Group Steps

Steps to conduct effective focus group research

A/B Testing

Covered above as a comparison point to preference testing, A/B testing remains a legitimate desirability tool in its own right when the variable being tested is aesthetic rather than functional, comparing two visual treatments of the same feature to see which one users engage with more.

Running a Desirability Test

A useful desirability test starts with a specific question, not a general one. "Do people like our new dashboard?" is too broad to act on. "Does the new color palette read as more trustworthy than the old one to our target demographic?" gives you something you can actually test and answer.

From there, the mechanics matter less than the sample. Recruit participants who reflect the real range of your target audience, not just in age or location but in income, digital comfort, and lifestyle. A narrow sample will hand you a confident, well-documented answer to the wrong question.

Target Audience Key Factors

key factors for identifying a target audience

Real-world scenarios beat abstract ones: show people the design in a context close to how they'd actually encounter it, and watch for hesitation or delight as closely as you listen to their stated opinions, since the two don't always match.

Run the session, then treat the raw output, word selections, scale ratings, transcripts, as the start of the analysis rather than the end. Look for patterns across participants rather than reacting to any single strong opinion, and cross-check what people say against what they actually chose when given a real decision to make.

Case Study

When Clay rebranded Streetbeat, an AI-driven investment platform, the goal wasn't just a new logo. The team needed the identity to read as precise and trustworthy to an audience evaluating whether to hand over their portfolio to an algorithm. That meant validating specific emotional targets - confidence, clarity, technical credibility - rather than general aesthetic approval, before the direction was locked in.

Streetbeat Case Study by Clay

The resulting identity used a rotating geometric mark, a green-teal gradient signaling the platform's dual fiat-and-crypto positioning, and a typographic system chosen specifically for precision over decoration. The lesson generalizes well beyond fintech: desirability testing works best when it's validating a specific emotional target the brand needs to hit, not a vague sense of "does this look good."

Common Pitfalls in Desirability Testing

  1. 1.

    The most common failure is letting a well-liked design distract from a product that still doesn't work. A polished, emotionally appealing interface sitting on top of a broken checkout flow or a confusing task won't save the product, and desirability testing can't tell you that on its own. Run it alongside usability testing, never as a substitute for it.
  2. 2.

    Sample bias is the second-biggest risk. Recruiting only convenient participants, coworkers, existing power users, people from a single demographic, produces confident results that don't hold up once the product reaches a broader audience. Deliberately recruiting across income, lifestyle, and digital comfort level catches blind spots a narrow sample never surfaces.
  3. 3.

    Leading questions are the third. Any question that hints at the "right" answer, or a moderator's tone that signals approval or disapproval, quietly corrupts the data. Neutral wording and a moderator who stays genuinely curious rather than persuasive protect the results.
  4. 4.

    Finally, budget and timeline constraints push teams toward skipping desirability testing entirely, treating it as a nice-to-have layered on top of usability work. Lighter-weight versions, a short reaction-card session added to the end of an existing usability test, or a remote unmoderated preference survey, capture most of the value at a fraction of the cost of a standalone study.

Where Desirability Testing Fits in Product Development

Desirability testing pays off most when it runs early and keeps running. The three moments that matter most are concept validation, before real investment goes into a direction, design iterations (while changes are still cheap to make), and pre-launch, as a final check before the product meets a wider audience.

It also depends on design and marketing working from the same brief. A product can nail its visual desirability testing and still feel disjointed if the brand voice, marketing imagery, and in-product experience are being validated separately, so cross-team alignment on what "desirable" means for this specific product matters as much as the testing method itself.

Treat it as ongoing rather than a one-time gate. User expectations shift, competitors ship new patterns that reset the baseline, and a product that felt fresh eighteen months ago can start to feel dated without anyone on the team consciously deciding to let it.

What's Next for Desirability Testing

Eye-tracking and facial-expression analysis are moving from research labs into more accessible, lower-cost tools, giving teams a behavioral layer (where attention actually goes) alongside the self-reported data that reaction cards and interviews provide.

Neither replaces the other. Behavioral signals show what draws the eye. Self-reported reactions explain why it matters to the person looking. The teams getting the most out of desirability testing right now are the ones pairing both rather than betting on one replacing the other.

Ready to build a web layout that performs across every device and keeps users engaged? Let our team take it from wireframe to launch. Get in touch.

Read More

FAQ

What is the difference between desirability testing and usability testing?

Usability testing measures whether a product functions correctly, whether someone can complete a task without getting stuck. Desirability testing measures the emotional response to the design itself: whether it feels trustworthy, modern, or appealing. A product can pass one and fail the other.

When should I run a desirability test?

Ideally more than once: during concept validation before real development investment happens, again during design iterations while changes are still cheap, and once more before launch as a final check.

How many participants do I need for a desirability test?

Qualitative methods like reaction cards and interviews tend to surface most major patterns with 5 to 8 participants per user segment. Quantitative methods like semantic differential scales benefit from larger samples, generally 20 or more, to produce statistically meaningful comparisons.

What are the Microsoft Product Reaction Cards?

A set of 118 words, 60% positive and 40% negative, developed at Microsoft in 2002 to give usability test participants a structured vocabulary for describing their emotional response to a product. Participants pick the five words that fit best and explain their choices.

Can desirability testing be done remotely?

Yes. Reaction-card sorting, semantic differential surveys, and preference tests all translate well to remote, unmoderated formats using standard survey tools. Interviews and focus groups lose some of the in-person nuance but remain workable over video.

Does a beautiful design guarantee a successful product?

No. A visually appealing product built on a broken or confusing experience will still fail. Desirability and usability need to be validated together, not treated as substitutes for one another.

How is desirability testing different from A/B testing?

A/B testing compares two live variations based on real user behavior, like clicks or conversions. Desirability testing (including preference testing) asks directly how something makes people feel, often before it's built. They answer different questions and work well in sequence.

What's the difference between desirability, feasibility, and viability?

Desirability asks whether users want the product. Feasibility asks whether it's technically buildable with the resources available. Viability asks whether it can sustain itself financially. All three need to hold up for a product to succeed long-term.

Do small teams need formal desirability testing?

Rarely in the full sense. A minimum viable version works for most small teams: five representative participants, a short reaction-card exercise added to a session you're already running, and a few open-ended questions about how the design makes them feel. That's enough for a qualitative signal on direction. Save the larger sample sizes and standalone sessions for decisions with real budget or reputational risk attached, like a full rebrand or a pricing page redesign.

What's a common mistake teams make with desirability testing?

Testing a polished visual direction while ignoring whether the underlying product actually works. A well-liked design on top of a confusing or broken flow won't save the product, so desirability testing should run alongside usability testing, not instead of it.

How do semantic differential scales work?

Participants rate a product along a spectrum between two opposite adjectives, such as "modern" versus "dated" or "trustworthy" versus "unreliable." The bounded scale makes it easy to compare results across a large sample and spot which specific qualities are weakest.

Should desirability testing happen before or after usability testing?

Both have a place at every stage, but a common and efficient approach is to add a short reaction-card exercise to the end of an existing usability session, capturing emotional response without running a separate study from scratch.

Bring the Data Into the Room

Desirability testing turns a subjective question, "do people like this," into something a team can actually act on: specific words, specific ratings, specific design choices tied to specific emotional reactions. That specificity is what separates a real signal from a hunch in a design review.

The products people stay loyal to are rarely the ones that merely function. They're the ones that made someone feel something the first time they opened them, and kept doing it. Testing for that feeling, deliberately and often, is how you build it on purpose instead of hoping it shows up.

Clay's Team

About Clay

Clay is a UI/UX design & branding agency in San Francisco. We team up with startups and leading brands to create transformative digital experience. Clients: Facebook, Slack, Google, Amazon, Credit Karma, Zenefits, etc.

Learn more

Share this article

Clay's Team

About Clay

Clay is a UI/UX design & branding agency in San Francisco. We team up with startups and leading brands to create transformative digital experience. Clients: Facebook, Slack, Google, Amazon, Credit Karma, Zenefits, etc.

Learn more

Share this article

Link copied

Thank you for subscribing!

We'll send you a subscription every couple of weeks.